ComputerVision (857) — 4/5
PixVerse R2: Scaling Real-Time Omni World Models, PixVerse, 2026.09
Paper/Blog Link My Issue
#Article #Blog #read-later #WorldModels #Scalability #Realtime Issue Date: 2026-09-23 Comment
元ポスト:
cua-s1-forms, cua-ai, 2026.09
Paper/Blog Link My Issue
#Article #NLP #Dataset #Transformer #SyntheticData #OpenWeight #ComputerUse #Encoder #One-Line Notes #Author Thread-Post Issue Date: 2026-09-19 Comment
github:
https://github.com/trycua/cua/tree/main/libs/cua-s1
dataset:
https://huggingface.co/datasets/cua-ai/cua-s1-forms
元ポスト:
フォーム入力に特化したCUAを実施できるモデルのようで、これもいわゆるSystem One Modelとして紹介されている。Transformer Encoderをバックボーンとする。
まあしかしこれは言ってしまえばただの合成データでFinetuningをしたBERTに見える。Jevのような様々なタスクに対してzero-shotの汎化をしめすものではないであろう点に注意。Jevがどれだけ汎用なのかはわからんが...
【ノーカット】AIに関わる米中戦争と日本の現状 国立情報研 佐藤教授, 時事通信映像センター, 2026.09
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Alignment #AIAgents #GenerativeAI #FoundationModel #Safety #Video #Selected Papers/Blogs #VisionLanguageModel #Robotics #WorldModels #VisionLanguageActionModel #EmbodiedAI #Reading Reflections #WorldActionModel Issue Date: 2026-09-18 Comment
非常におもしろかった...色々と考えさせられるなあ。
特にフィジカルAIにおいて、中国は人海戦術で、(簡単な動作を学習できたら複雑な動作を実現することにつながるという仮説の元)非専門家の簡単な動作に関するデータを大量に収集し、最近では仮想環境でデータを合成し活用する動きがある(Sim-to-Real)に対して、日本の場合は、特定の企業の現場のデータを収集しようとする動きがある。(このような文脈において、日本の政策として考えた場合に)特定の企業が有する現場のデータは、その現場特有の特徴が入り込むから実はロボットの学習データに向いていない、という話は、なるほど、と思うなどした。
たとえば、
- [Paper Note] Open X-Embodiment: Robotic Learning Datasets and RT-X Models, Open X-Embodiment Collaboration+, arXiv'23, 2023.10
- [Paper Note] LeRobot: An Open-Source Library for End-to-End Robot Learning, Remi Cadene+, ICLR'26, 2026.02
のように、多様なデータを統一された枠組みで学習をすることで汎化することを狙うという戦略の場合は、現場データの現場特有のバイアス問題はどの程度緩和されるのだろうか。
- Superposition, Memorization, and Double Descent, Transformer Circuits Thread, 2023.01
のDouble Decentのような議論を考えると、同じ種別のデータで、きちんと多様なデータが集まっていれば、現場固有のバイアスが含まれていても汎用的な特徴量としては学習されずらい、という現象は起こるように思える。
また、仮想空間と実世界のgapもあるようで、視覚的なgapは合成データ等で埋め合わせがある程度できそうな一方で、触覚や力感のような繊細な部分や、物理的なセンサーのノイズや、実際のデバイスの物理的な個体差のようなものは、仮想空間上では再現しづらいという話もあるようである。
ただ、なんとなーく、まず汎用的なモデルを学習して基礎的な動作(LLMで言うところのatomic skill)を満足にできる基盤モデルを用意し(これはおそらく仮想空間や非専門家データで足りる)、その基盤モデルはin-context learning能力を備えていて、現場のロボットに適用する際にはfew-shotのdemonstrationや、ルーブリックのようなものを与えて制御する、というのでうまくいくシナリオは起きそうな気がしており、そうなると大量の現場データは必ずしも必要ではなさそうだよね、という気はする。
ただし、日本が人海戦術をとれないのであれば、汎用基盤モデル→現場での適用、というレールの上に乗るのであれば、直接的に汎化させるには専門的すぎる現場のデータが与えられたときに、そこからどのようにして汎用的な基盤モデルを学習できるのか、というところがうまくいけば、ある程度形になるのではなかろうか、と思うなどした。
Introducing ChatGPT Images 2.5, OpenAI, 2026.09
Paper/Blog Link My Issue
#Article #NLP #TextToImageGeneration #Blog #Proprietary #Editing #ImageSynthesis #Author Thread-Post Issue Date: 2026-09-09 Comment
元ポスト:
スケッチを入力することが可能になったようだ
LLM-jp-4-VL 9Bリリース, LLM-jp, 2026.09
Paper/Blog Link My Issue
#Article #Pretraining #NLP #Dataset #OpenWeight #Japanese #OpenSource #Selected Papers/Blogs #VisionLanguageModel #One-Line Notes Issue Date: 2026-09-06 Comment
元ポスト:
モデル: https://huggingface.co/llm-jp/llm-jp-4-vl-9b
RefinedVision:
https://huggingface.co/datasets/llm-jp/RefinedVision
OpenSource AIの定義を満たすためにFineVisionデータセットを精査(e.g. GPT-4oの出力等は利用規約上問題がある)した上で修正(FineVisionは185サブセット, 24M VQAで構成されている)したとのこと。すごい。。。
ライセンス・利用規約上の問題を精査するだけでなく、品質(画像の品質、QAの品質)の観点でも精査の上修正しているようである。
関連:
- [Paper Note] FineVision: Open Data Is All You Need, Luis Wiedmann+, NeurIPS'26, 2025.09
Atlas: A World Model for Spatial Intelligence, World Labs, 2026.09
Paper/Blog Link My Issue
#Article #Transformer #DiffusionModel #Blog #Proprietary #read-later #Selected Papers/Blogs #3D Reconstruction #WorldModels #SpatialUnderstanding #Author Thread-Post #4D(Scene + Time) Issue Date: 2026-09-06 Comment
元ポスト:
解説:
Orbis: Our foundation model that creates living worlds and streams them in real time, Visko, 2026.08
Paper/Blog Link My Issue
#Article #Blog #Proprietary #VideoGeneration/Understandings #interactive #TextToVideoGeneration #Realtime #Initial Impression Notes #Author Thread-Post Issue Date: 2026-09-06 Comment
元ポスト:
リアルタイムに、かつテキストでのインタラクティブな指示を与えつつ動画を生成できるモデルに見える。
How Far Are We from a Native Autoregressive Video Model?, Weiyang Jin, 2026.08
Paper/Blog Link My Issue
#Article #Blog #read-later #VideoGeneration/Understandings #Author Thread-Post Issue Date: 2026-09-06 Comment
元ポスト:
Introducing H3 Max by fal, fal, 2026.08
Paper/Blog Link My Issue
#Article #NLP #Blog #Proprietary #VideoGeneration/Understandings #Initial Impression Notes #Author Thread-Post Issue Date: 2026-09-04 Comment
元ポスト:
Minimax H3を事後学習(指示追従とvisual qualityを改善)することで、Artificial Analysisによる評価でSoTAを達成した動画生成モデルとのこと。
Claude: Fable 5.1 and Mythos 5.1, Anthropic, 2026.09
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Proprietary #Selected Papers/Blogs #VisionLanguageModel #Author Thread-Post Issue Date: 2026-09-04 Comment
元ポスト:
所見:
GPT-6 Astra: A new generation of intelligence, 2026.09
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Blog #Proprietary #Selected Papers/Blogs #VisionLanguageModel #Initial Impression Notes #Author Thread-Post Issue Date: 2026-09-04 Comment
元ポスト:
GPT-5.6 Solから大幅に性能向上。scaling lawはいつまで続くのだろうか
ベンチマーク上は
- Claude: Fable 5.1 and Mythos 5.1, Anthropic, 2026.09
を超えている。
所見:
アーキテクチャに関する所見:
- [Paper Note] Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model, Nanbeige Lab+, arXiv'26, 2026.07
- [Paper Note] Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation, Sangmin Bae+, NeurIPS'25
より注意深く指示に従い、CUAの能力も大幅に向上しているようである:
CUAによってBlenderで作成されたhouseの例がなかなかすごい
Artificial Analysisによる評価によると、Artificial Analysis Indexでは5.6-Solと同等、Coding IndexでもFable 5と同等であり、OpenAIによるベンチマークスコアとかなり乖離があるように見える。パブリックなベンチマークに対して過剰に適合しているのか。それともArtificial Analysisのベンチマーク群とその重みづけにより算出されるスコアの相性が悪いのか。(まあこれを考えても不毛だけど)
SRE-Benchと呼ばれるモデルがコンパイル済みのバイナリからソフトウェアをリバースエンジニアリングできるかを測定するベンチマークが飽和したとのこと:
SRE-Bench:
https://benchlm.ai/benchmarks/srebench
所見:
Artificial Analysis Indexのスコアが更新されてFable 5.1とAstraのスコアが同じになったようである😅:
Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency, Qwen Team, 2026.08
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #OpenWeight #Selected Papers/Blogs #VisionLanguageModel #Initial Impression Notes #Author Thread-Post Issue Date: 2026-08-30 Comment
ベンチマークスコアはOpus 4.6超え。パラメータは125B-A6B、51Bの N-gram Embeddings。Gated Delta Net layerとQwen Sparse Attention と呼ばれる Layerを3:1の比率で積み上げていっているようである。N-gram Embedding、MTPも導入されている。また、Gated Residualと呼ばれる技術により、Residual Streamからの読み込み/書き込みを制御している。
- 関連:
- [Paper Note] Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models, Xin Cheng+, arXiv'26, 2026.01
HF:
https://huggingface.co/Qwen/Qwen3.8-Flash-Next?spm=a2ty_o06.30285417.0.0.1d73c921FsyOPe&file=Qwen3.8-Flash-Next
technical report:
https://github.com/QwenLM/Qwen3.8-Flash-Next/blob/main/tech_report.pdf
解説:
所見:
EgoSuite-Open100K, LightwheelAI, 2026.08
Paper/Blog Link My Issue
#Article #Dataset #Robotics #3D (Video) #EmbodiedAI #EgocentricView #Author Thread-Post Issue Date: 2026-08-22 Comment
pj page: https://egosuite100k.lightwheel.ai/
元ポスト:
LeRobotフォーマットとのこと
- [Paper Note] LeRobot: An Open-Source Library for End-to-End Robot Learning, Remi Cadene+, ICLR'26, 2026.02
- LeRobot Humanoid: An Open, Low-Cost, 3D-Printed Humanoid for Robot Learning, LeRobot, 2026.05
dots3-note Preview: A Small but Mighty Step Toward Long-Horizon Agency in Real Life, dots studio, 2026.08
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Blog #OpenWeight #VisionLanguageModel Issue Date: 2026-08-22 Comment
HF: https://huggingface.co/dots-studio/dots3-note-prev
元ポスト:
SenseNova-U1.5-8B-MoT, SenseNova, 2026.08
Paper/Blog Link My Issue
#Article #NLP #TextToImageGeneration #OpenWeight #Editing #UMM #ImageSynthesis #Author Thread-Post Issue Date: 2026-08-22 Comment
元ポスト:
WildArtifactBench, Meta, 2026.08
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Evaluation #MultiModal #LLM-as-a-Judge #VisionLanguageModel #Initial Impression Notes #Author Thread-Post Issue Date: 2026-08-22 Comment
元ポスト:
AI Agentによるpreference judgeに基づく、win rate/Elo Scoreによるマルチモーダルエージェントのベンチマークのようである
text-2-image-human-preferences-2m, Datapoint AI, 2026.08
Paper/Blog Link My Issue
#Article #NLP #Dataset #TextToImageGeneration #Blog #DPO #One-Line Notes #ImageSynthesis Issue Date: 2026-08-21 Comment
21kのimageペアに対して2Mのpair-wiseでの人間のpreferenceが付与されたデータとのこと。500 promptに対して、30個のT2Iモデルが利用されている。
Cua: Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation, trycua, 2026.08
Paper/Blog Link My Issue
#Article #Tools #NLP #Infrastructure #AIAgents #ComputerUse #Selected Papers/Blogs #Author Thread-Post #SandboxEnvironment Issue Date: 2026-08-20 Comment
元ポスト:
Putting sign language AI into users’ hands, Google Deepmind, 2026.08
Paper/Blog Link My Issue
#Article #MachineTranslation #NLP #LanguageModel #Blog #MultiLingual #One-Line Notes #Author Thread-Post Issue Date: 2026-08-14 Comment
元ポスト:
Sign-Language to Text (SL2T) Model
50種類以上の手話で学習された手話をテキストに翻訳するモデル。スマホのカメラを通じてリアルタイムにstreaming textとして翻訳される。
Meet North Micro Vision: A 2.4B Native-Resolution Vision-Language Model, Cohere, 2026.08
Paper/Blog Link My Issue
#Article #NLP #SmallModel #OpenWeight #VisionLanguageModel #Author Thread-Post Issue Date: 2026-08-14 Comment
HF:
https://huggingface.co/CohereLabs/North-Micro-Vision-Instruct
Apache 2.0
元ポスト:
Reflections on Video DeltaNet, Haoyi Zhu, 2026.08
Paper/Blog Link My Issue
#Article #Blog #read-later #VideoGeneration/Understandings #LinearAttention #Author Thread-Post Issue Date: 2026-08-14 Comment
元ポスト:
videoに対してdelta netのようなlinear attentionモデルを適用することに関する考察のようである。
LTX-2.5, Lightricks, 2026.08
Paper/Blog Link My Issue
#Article #DiffusionModel #OpenWeight #VideoGeneration/Understandings #TextToVideoGeneration #ImageToVideoGeneration #Author Thread-Post Issue Date: 2026-08-14 Comment
元ポスト:
VISTA: A Visual Harness for Reasoning in an Interactive World, Han+, 2026.08
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Reasoning #read-later #Selected Papers/Blogs #2D (Image) #3D (Scene) #memory #LongHorizon #Author Thread-Post #AgentHarness Issue Date: 2026-08-10 Comment
元ポスト:
2025年度GENIAC採択事業において収集・開発したロボット動作データセットおよびロボット基盤モデルを一般公開, AIRoA, 2026.08
Paper/Blog Link My Issue
#Article #Dataset #Blog #OpenWeight #Robotics #VisionLanguageActionModel #EmbodiedAI #Author Thread-Post Issue Date: 2026-08-10 Comment
Dataset:
https://huggingface.co/datasets/airoa-org/airoa-moma-5k
HF:
https://huggingface.co/airoa-org/airoa-pi05-hsr-base
元ポスト:
MiniMax-H3, MiniMaxAI, 2026.08
Paper/Blog Link My Issue
#Article #NLP #Transformer #TextToAudio #SpeechProcessing #OpenWeight #VideoGeneration/Understandings #UMM #Omni #TextToVideoGeneration #Author Thread-Post Issue Date: 2026-08-09 Comment
元ポスト:
Video Arenaと呼ばれるベンチマーク (Image-to-Video, Image-to-Video) でOpenWeightモデルでSoTA、全体で2位:
[Paper Note] Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents, MAI-UI Team, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #ComputerUse #VisionLanguageModel #Initial Impression Notes #GUI Issue Date: 2026-08-09 Comment
pj page: https://tongyi-mai.github.io/Qwen-UI-Agent/
多くのベンチマークでFrontier Modelを上回る性能を示すGUI Agent。デモを見るとなかなかインパクトがある。
元ポスト:
Introducing Inkling-Small, THINKING MACHINES, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #OpenWeight #VisionLanguageModel #UMM #One-Line Notes #Author Thread-Post Issue Date: 2026-08-09 Comment
HF: https://huggingface.co/thinkingmachines/Inkling-Small
関連:
- Inkling: Our open-weights model, THINKING MACHINES, 2026.07
- 276B-A12B
- NVFP4
- 小規模なモデルだが、Inklingと同等以上の性能を達成(ただし、一般的な知識や事実性に関してはInklingの方が上)
- 事前学習データのDataMixtureや学習レシピにいくつかの改良
- Inklingを教師モデルとしたOPD
- codingに特化したagentic RLをスケールアップ
元ポスト:
Introducing MazeBench, MazeBench, 2026.07
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Planning #VisionLanguageModel #3D (Scene) #SpatialUnderstanding #One-Line Notes #LongHorizon Issue Date: 2026-08-05 Comment
元ポスト:
3D空間において、カメラアングルの移動とブロックの移動をすることで実施可能なパズルゲームによる新たなベンチマークで、視覚的・空間的な推論能力と、長期のplanning能力を評価する。
Introducing PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models, Kimi Team, 2026.07
Paper/Blog Link My Issue
#Article #NLP #Dataset #Evaluation #Selected Papers/Blogs #VisionLanguageModel #One-Line Notes Issue Date: 2026-08-05 Comment
元ポスト:
Vision系の新たなベンチマーク。42種類のベンチマークのフロンティアモデルが失敗する事例に基づいて評価対象とするatomicな能力を定義し、それらのatomicなスキルを独立して評価可能なサンプルを作成し評価するベンチマーク。サンプルは与えられた情報のみから、評価対象とするatomicな能力を駆使すれば成功できるverifiableなタスクとして定義されているようである。能力ごとにモデルのスコアが算出可能。
実例としては、たとえば、魚が泳いでいる複数の池のイラストが与えられて、真ん中の池には何匹の魚がいる?といったクエリに応答するなどである。これには localization の能力が求められる(Countingも必要な気がするが)。
Scaling Video Pretraining with Imagination Models, induction labs, 2026.07
Paper/Blog Link My Issue
#Article #Pretraining #FoundationModel #Blog #VideoGeneration/Understandings #Reading Reflections #Author Thread-Post Issue Date: 2026-08-04 Comment
元ポスト:
動画で事前学習することでvision系タスクの基盤モデルとして有効に機能する、という話が最近増えてきたように感じる。
SIGReg from First Principles, Reza Bayat, 2026.07
Paper/Blog Link My Issue
#Article #Tutorial #RepresentationLearning #Self-SupervisedLearning #read-later #WorldModels #Author Thread-Post #LatentRepresentation Issue Date: 2026-08-04 Comment
元ポスト:
JEPA関連のチュートリアル
関連:
- [Paper Note] Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture, Mahmoud Assran+, CVPR'23, 2023.01
- [Paper Note] LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics, Randall Balestriero+, arXiv'25, 2025.11
- [Paper Note] LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels, Lucas Maes+, arXiv'26, 2026.03
- JEPAwiki, mishig, 2026.04
著者ポスト:
Introducing FLUX-mimic: Scaling Video-Action Models for General Purpose Dexterity, mimic robotics, 2026.07
Paper/Blog Link My Issue
#Article #NLP #VideoGeneration/Understandings #Robotics #EmbodiedAI #WorldActionModel Issue Date: 2026-08-03 Comment
元ポスト:
FLUX3をベースに構築されたVision Action Model
- 大規模な動画生成モデルをバックボーンとすることで、ロボットのActionの予測をVLAと比較して最大10倍のデータ効率で実現
- 動画生成バックボーンの性能が向上すれば、Video Action Modelの性能もそれに伴い向上することが示唆される
Introducing MAI-Image-2.5- Pro and MAI-Voice-2-Flash, MAI, 2026.07
Paper/Blog Link My Issue
#Article #NLP #SpeechProcessing #TextToImageGeneration #Proprietary #TTS #ImageSynthesis #Author Thread-Post Issue Date: 2026-08-03 Comment
元ポスト:
関連:
- Building a hill-climbing machine: Launching seven new MAI models, Mustafa Suleyman, MAI, 2026.06
Introducing Cosmos 3 Edge, Nvidia, 2026.07
Paper/Blog Link My Issue
#Article #NLP #Blog #SmallModel #OpenWeight #VisionLanguageModel #Robotics #WorldModels #EmbodiedAI #Author Thread-Post Issue Date: 2026-07-31 Comment
元ポスト:
HF: https://huggingface.co/nvidia/Cosmos3-Edge?linkId=100000431533162
autoregressive -> diffusion の two-tower モデル
Cosmos3 テクニカルペーパー:
- [Paper Note] Cosmos 3: Omnimodal World Models for Physical AI, NVIDIA+, arXiv'26, 2026.06
Introducing Composite-Bench: the strongest open-weights model isn't Kimi K3, composite, 2026.07
Paper/Blog Link My Issue
#Article #NLP #Evaluation #Blog #ComputerUse #VisionLanguageModel #LongHorizon #Initial Impression Notes #Author Thread-Post Issue Date: 2026-07-24 Comment
元ポスト:
エンタープライズ向けのサービスに基づく専門家の操作に基づくbrowser-useベンチマークのようである。ざっくりブログを読んだが実際にどのようなサービスなのか知らないためあまりよくわからなかったので、公開されているベンチマークを見る方が直感的な理解は捗りそう。
本ベンチマーク上では、GLM-5.2がOpenWeightモデルの中では最も性能が高いようである。
FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence, Black Forest Labs, 2026.07
Paper/Blog Link My Issue
#Article #TextToAudio #MultiModal #TextToImageGeneration #Blog #Proprietary #VideoGeneration/Understandings #Editing #UMM #One-Line Notes #ImageSynthesis #TextToVideoGeneration #WorldActionModel #Author Thread-Post Issue Date: 2026-07-24 Comment
元ポスト:
モデルは将来的にオープンになるようである
Announcing Bonsai 27B: The First 27B-Class Model to Run on a Phone, PrismRL, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MultiModal #Blog #OpenWeight #Selected Papers/Blogs #VisionLanguageModel #Initial Impression Notes #Author Thread-Post Issue Date: 2026-07-19 Comment
元ポスト:
HF: https://huggingface.co/collections/prism-ml/bonsai-27b
Bonsaiシリーズ:
- Announcing 1-bit Bonsai: The First Commercially Viable 1-bit LLMs, 2026.03
- Introducing Ternary Bonsai: Top Intelligence at 1.58 Bits, PrismML, 2026.04
- Introducing 1-bit and Ternary Bonsai Image 4B: Image Generation for Local Devices, PrismML, 2026.05
1-bit, ternary weightによって、27B級モデルがエッジデバイス上で動作する。
How We Serve Ideogram V4 Lightning Fast on fal, fal, 2026.07
Paper/Blog Link My Issue
#Article #NLP #TextToImageGeneration #Blog #OpenWeight #Initial Impression Notes Issue Date: 2026-07-19 Comment
元ポスト:
関連:
- Ideogram 4: Open image model at the forefront of design, Ideogram, 2026.06
Ideogramに基づいた、高速なTextToImage Modelのようである。
Introducing Muse Spark 1.1, Meta, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #MultiModal #Proprietary #VisionLanguageModel #One-Line Notes #Author Thread-Post Issue Date: 2026-07-19 Comment
元ポスト:
Muse Spark:
- Introducing Muse Spark: Scaling Towards Personal Superintelligence, Meta, 2026.04
大幅にAgenticな能力やコーディング能力が向上。Agenticな能力では多くのベンチマークでOpus4.8超え、CodingはGPT-5.5にベンチマーク上では及ばず(SWE Bench Proでは勝っているが、OpenAIから30%以上のpromptが評価で不適切との報告があった Separating signal from noise in coding evaluations, OpenAI, 2026.07 )。
Alexandr Wang氏によるポスト:
VQAからDocument Parsingへ:Nemotron-3-Nano-Omniに対する日本語文書の構造化出力Post-training, Stockmark, 2026.07
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #Japanese #VisionLanguageModel #reading Issue Date: 2026-07-16 Comment
元ポスト:
Inkling: Our open-weights model, THINKING MACHINES, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MultiModal #SpeechProcessing #Blog #Reasoning #OpenWeight #MoE(Mixture-of-Experts) #Selected Papers/Blogs #VisionLanguageModel #UMM #KeyPoint Notes #Reference Collection #AudioLanguageModel #Author Thread-Post Issue Date: 2026-07-16 Comment
HF: https://huggingface.co/thinkingmachines/Inkling
元ポスト:
THINNING MACHINESによる最初のOpenWeightモデル。975B-41B
- フルスクラッチで学習したReasoningモデル
- 1M context window
- text, image, audio, video 45Tトークンで学習
- バランスよくさまざまな領域で性能を発揮するように訓練し、Tinker上でのfinetuningやカスタマイゼーションの良い基盤として機能することを目指した
- vision, audioドメインに関してはencoder freeアーキテクチャを採用
- audio signalの入力は dMel spectrogramsと呼ばれる手法を採用
- [Paper Note] dMel: Speech Tokenization made Simple, Richard He Bai+, arXiv'24, 2024.07
- 画像は40x40のパッチに分割され、4つのhMLPと呼ばれるレイヤーでエンコーディング
- [Paper Note] Three things everyone should know about Vision Transformers, Hugo Touvron+, ECCV'22, 2022.03
- IFの訓練には、Rubric basedなgraderとclaims graderの2種類のgraderを用いることで、helpfulnessとhallucinationを同時に低減
- Rubric basedなgraderはチェックリストに基づいてスコアリング
- claims graderはfactualな主張をagentic web searchを通じてverificationする
- アーキテクチャはDeepSeek V3を踏襲し、256のexpertsと2つのshared expertを採用し、6つのexpertsがactivateされる
- SWAとGlobal attentionの比率は5:1でKV Headは8つ (GQA)
- **相対位置エンコーディングを用いることでRoPEよりもlong contextに対してより高い外挿性能を示した**
- [Paper Note] Self-Attention with Relative Position Representations, Peter Shaw+, NAACL'18
- K, Vのprojection後、**およびattentionとMLPが残差ストリームに合流する前にshort convolutionを導入**
- QKV projection後に convolution を導入するアーキテクチャは下記研究で提案
- [Paper Note] Primer: Searching for Efficient Transformers for Language Modeling, David R. So+, NIPS'21, 2021.09
- optimiserはMuonとAdamWのハイブリッドで、前者は巨大な行列の重みに対して適用し、そのほかは後者を利用
- 重み自体が多様体上に存在するように制約することが有効であったことに着想を得て、weight decayを学習率の2乗に連動させることでモデルの重みの大きさを安定させたとのこと
- Modular Manifolds, Jeremy Bernstein+, THINKING MACHINES, 2025.09
- [Paper Note] Why Gradients Rapidly Increase Near the End of Training, Aaron Defazio, arXiv'25, 2025.06
- 事後学習としては、まずKimi K2.5を含むOpenWeightモデルで合成データ上でSFTをし、その後合成、あるいは人間が作成したenvironment上で大規模なRLを実施
- RLは非同期RLを異様し、ロールアウト数に対して推論能力が対数線形にスケールした。
- 最終的に30Mロールアウト以上のRLを実施した
- システムメッセージを変更し、トークン単位のコストを調整することでreasoning effortを調整
- Controlling Reasoning Effort in LLMs, Sebastian Raschka, 2026.07
- **276B-A12BのInkling-Smallもプレビュー段階にあり、agenticな能力においてInklingに近い性能を達成しており、今後公開予定**
1M context windowを実現した工夫は特に書かれていなかったように感じるが、よく使われるのはKV CacheをコンパクトにするためのSparse Attention、あるいはLinear Attentionなのに対し、本モデルでは使われていないように見える。linear attentionでなくとも、SWAとGlobal Attentionのハイブリッド、かつGQAによって、KV Cacheの肥大化を実用レベルで抑制できるのだろうか。また、外挿性能に関してはRoPEは採用せず、相対位置エンコーディングのを採用したという点が効いているのだろうか。1Mレベルでのcontextでの評価がなさそうに見えるので性能はよくわからない。
RoPEのlong contextでの限界を示した以下の研究を思い出した:
- [Paper Note] RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably, Yufeng Du+, arXiv'26, 2026.05
Artificial Analysisによる評価:
アーキテクチャでの目新しい点:
所見:
公式:
所見:
LingBot-VLA-V2, Robbyant, 2026.07
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #Robotics #VisionLanguageActionModel #EmbodiedAI #Author Thread-Post Issue Date: 2026-07-12 Comment
元ポスト:
[Paper Note] Waypoint-1.5: A Real-Time Video World Model for Consumer Hardware, Rajpal+, 2026.07
Paper/Blog Link My Issue
#Article #DiffusionModel #WorldModels #Game #Realtime Issue Date: 2026-07-08 Comment
HF: https://huggingface.co/Overworld/Waypoint-1.5-1B
関連:
- Waypoint-1: Real-time Interactive Video Diffusion from Overworld, Overworld, 2026.01
元ポスト:
pj page: https://over.world/waypoint-1.5
CS2-10k: A Large-Scale Egocentric Counter-Strike 2 Dataset, Reka Labs, 2026.06
Paper/Blog Link My Issue
#Article #Dataset #WorldModels #Initial Impression Notes Issue Date: 2026-07-07 Comment
Counter-Strikeのプロレベルの試合に基づいた、1万時間以上一人称視点の映像と、フレームごとのannotation(キーボード操作状態、マウス移動の軌跡、プレイヤーの移動経路等)を含むデータセットのようである。
映像だけでなく、annotationが付与されている点が有用。
元ポスト:
Introducing computer use in Gemini 3.5 Flash, Google, 2026.06
Paper/Blog Link My Issue
#Article #NLP #Blog #Proprietary #ComputerUse #VisionLanguageModel #Initial Impression Notes Issue Date: 2026-07-05 Comment
元ポスト:
quick start guide:
computer useがだいぶ使いやすくなってきたなあ(小並感)
MolmoMotion: Language-guided 3D motion forecasting, Ai2, 2026.06
Paper/Blog Link My Issue
#Article #DiffusionModel #VideoGeneration/Understandings #3D (Scene) #FlowMatching #Robotics #Author Thread-Post Issue Date: 2026-06-18 Comment
元ポスト:
視覚基盤モデルの構築, 産総研, 片岡裕雄, 2026.06
Paper/Blog Link My Issue
#Article #FoundationModel #Slide #read-later #Robotics Issue Date: 2026-06-09 Comment
元ポスト:
Ideogram 4: Open image model at the forefront of design, Ideogram, 2026.06
Paper/Blog Link My Issue
#Article #NLP #TextToImageGeneration #OpenWeight #ImageSynthesis #Author Thread-Post Issue Date: 2026-06-05 Comment
元ポスト:
HF: https://huggingface.co/collections/ideogram-ai/ideogram-4
Introducing Gemma 4 12B: a unified, encoder-free multimodal model, Google, 2026.06
Paper/Blog Link My Issue
#Article #NLP #MultiModal #OpenWeight #VisionLanguageModel #2D (Image) #UMM #SpatialUnderstanding #One-Line Notes #Reference Collection #AudioLanguageModel #audio #Author Thread-Post Issue Date: 2026-06-04 Comment
元ポスト:
vision/audioエンコーダーを無くしたvision/audio nativeなマルチモーダルLLM
HF: https://huggingface.co/google/gemma-4-12B
アーキテクチャ図:
A Functional Taxonomy of World Models, Fei-Fei Li, 2026.06
Paper/Blog Link My Issue
#Article #Tutorial #NLP #Post #Selected Papers/Blogs #VideoGeneration/Understandings #Robotics #WorldModels #VisionLanguageActionModel #KeyPoint Notes #TextToVideoGeneration #Reading Reflections #WorldActionModel #Author Thread-Post Issue Date: 2026-06-04 Comment
元ポスト:
以下ポストの内容の要約(と意訳、間違ってたらごめんなさい)
- 世界モデルは現在最も重要だが、最も多義的な概念の一つになっている。
- 様々な分野がWorld Modelを構築していると主張するが、意味するところが実際には大きく異なる
- (実際 [Paper Note] Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond, Meng Chu+, arXiv'26, 2026.04
のような研究も存在し似たような問題意識のもと様々な分野での統一的な分類体系が提案されている)
- 世界モデルという用語のもともとの枠組みは「部分観測マルコフ決定過程 (POMDP)」であり、
- エージェントは行動を実行し、行動は世界の状態に影響を与え、エージェントは観測データを受け取り(≠状態を認識する)、新たな観測データに基づいてアクションが実行される、といったループが繰り返される枠組みである
- ここで、「状態」とは、ある時点における世界で何が起きているかに関する完全なdescriptionであり、エージェントは状態自体を認識することはできず、行動と状態から生じた部分的な観測データのみである。
- 現在様々な世界モデルと呼ばれるものが存在するが、構造としては上記のループを持っており、それらの切り口が異なっているにすぎない。
- 世界モデルのカテゴリ1: Renderer
- Rendererは人間の目に見えるピクセルで「観測」を出力する。
- たとえば、テキストのプロンプトを映像に変換するText-To-Videoモデル、ユーザの入力に応じてリアルタイムにフレームを生成するシステムはレンダラーに相当する。
- これらモデルは観測者にとって「見えるもの」を生成しているにすぎず、実際の3次元構造を明示的に理解しているわけではない(i.e., 見えるもの≠実在するもの)。
- ビジネスとして最も成長(してきており、学習データもインターネット上の動画が活用できるため他の2カテゴリと比べて多い)
- 世界モデルのカテゴリ2: Simulator
- Simulatorは「状態」を出力する。これは実際に人間やコンピュータが相互作用可能な世界の表現である。
- Rendererは単に視覚的なものであるが、Simulatorは実世界の幾何学的・物理的・動的なダイナミクスを理解することが求められる。
- Simulatorは建築家やゲーム開発者などの視覚を超えた(たとえば構造・物理的な)正確性を必要とする職種や、RLの学習の環境として利用できる。
- Simulator は Rendererと次のPlannerの土台となる技術(Simulatorは RendererとPlannnerの双方をバイパスできる)であるが、学習データが最も不足
- 世界モデルのカテゴリ3: Planner
- Plannerは「行動」を出力する。観測と目標が与えられた時に「次に何をすべきか」を出力する。
- Vision Language Action Model / World Action Model は Planner に該当し、これらはロボットが次に何をすべきかを決定できる。
- 現在研究初期段階で、研究所内での閉じられた環境でのデモ中心で、実世界で活用するためにはまだまだ多くの課題が残る。
- これら3つのカテゴリは現在世に出ているWorld Modelの多くを説明しており、区別をする際に役に立つ。
- が、これらカテゴリは独立したものではなく、これらは世界の機能に関する基本的な知識(幾何学、物理学、ダイナミクス)の上に成り立つ。
- これら3つのカテゴリは最近は互いが融合してくる流れにあり、たとえば事前学習された Renderer は、次に何が起こるか・何をすべきか(=Planner)を予測するためのバックボーンとして利用できることが示されてきており、これは Renderer と Plannerが 融合した例と言える。
- (この辺の話はBackboneとしてVision Encoderを持つVLA系全般の研究と、事前学習済みのVision Encoderを用いずに事前学習の方法をそもそも改善するような方向などだろうか)
上記の話に基づくと、たとえばターミナルでのWorld Modelに相当すると考えられる
- [Paper Note] ECHO: Terminal Agents Learn World Models for Free, Vaishnavi Shrivastava+, arXiv'26, 2026.05
は3つのカテゴリのうちにどれに該当するだろうか。
次のアクションを予測できるので、まずPlannerには該当すると思われる。また、ある時点においてターミナル上で何が起きているかの記述(ターミナルの出力)を予測しているので、Simulatorの役割を果たしていると思われる(ただ、ターミナルの出力だけがターミナルの状態を完全に記述した情報なの?定義としてそれでいいの?という疑問はあるのが)。このため、Planner と Simulator が融合した研究と言えるのではなかろうか。
Holo3.1: Fast & Local Computer Use Agents, H Company, 2026.06
Paper/Blog Link My Issue
#Article #NLP #AIAgents #OpenWeight #ComputerUse #VisionLanguageModel #Author Thread-Post Issue Date: 2026-06-03 Comment
HF: https://huggingface.co/collections/Hcompany/holo31
元ポスト:
関連:
- Holo3: Breaking the Computer Use Frontier, H Company, 2026.03
Develop Physical AI Reasoning, World, and Action Models with NVIDIA Cosmos 3, nvidia, 2026.05
Paper/Blog Link My Issue
#Article #NLP #MultiModal #OpenWeight #Selected Papers/Blogs #VideoGeneration/Understandings #Robotics #WorldModels #UMM #reading #Omni #One-Line Notes #WorldActionModel #Author Thread-Post Issue Date: 2026-06-02 Comment
元ポスト:
公式:
encoder-freeなOmniモダリティモデルで、かつ将来の世界の状態、およびactionを予測可能なWorldActionModel
Step-3.7-Flash, stepfun-ai, 2026.05
Paper/Blog Link My Issue
#Article #NLP #AIAgents #MultiModal #OpenWeight #MoE(Mixture-of-Experts) #read-later #VisionLanguageModel #Author Thread-Post Issue Date: 2026-05-31 Comment
元ポスト:
公式:
[Paper Note] LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding, Wang+, 2026.05
Paper/Blog Link My Issue
#Article #NLP #CVPR #read-later #Selected Papers/Blogs #ObjectLocalization #VisionLanguageModel #2D (Image) #UMM #3D (Video) #text #ObjectDetection #GUI #Author Thread-Post Issue Date: 2026-05-30 Comment
元ポスト:
Introducing 1-bit and Ternary Bonsai Image 4B: Image Generation for Local Devices, PrismML, 2026.05
Paper/Blog Link My Issue
#Article #DiffusionModel #TextToImageGeneration #SmallModel #Selected Papers/Blogs #One-Line Notes #ImageSynthesis #LowPrecision #Author Thread-Post Issue Date: 2026-05-27 Comment
元ポスト:
HF: https://huggingface.co/collections/prism-ml/bonsai-image
Ternary Weight {-1, 0, 1}による画像生成モデル
PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects, Ropedia, 2026.05
Paper/Blog Link My Issue
#Article #NLP #VisionLanguageModel #Robotics #Simulation #3D Object Generation #Author Thread-Post Issue Date: 2026-05-27 Comment
元ポスト:
MagicLite, Microsoft, 2026.05
Paper/Blog Link My Issue
#Article #Tools #NLP #SmallModel #ComputerUse #One-Line Notes #Author Thread-Post Issue Date: 2026-05-27 Comment
元ポスト:
Faraを用いたブラウザベースのChatUIを備えたCUAで、ローカルストレージへのアクセスも可能な模様
Marlin-2B, NemoStation, 2026.05
Paper/Blog Link My Issue
#Article #Temporal #VideoGeneration/Understandings #VisionLanguageModel #3D (Video) #reading #Grounding #Author Thread-Post Issue Date: 2026-05-27 Comment
元ポスト:
何が、いつ起きたかに答えるVideo VLMで、イベントごとのキャプションとtimestampのspanを出力してくれるようである。2Bモデルなので軽量である。
例は以下:
より安全で透明性の高い AI エコシステムに向けて、コンテンツ来歴の取り組みを前進, OpenAI, 2026.05
Paper/Blog Link My Issue
#Article #TextToImageGeneration #Proprietary #2D (Image) #One-Line Notes #ImageSynthesis #AI Detector #Author Thread-Post Issue Date: 2026-05-27 Comment
元ポスト:
画像生成にSynthID追加、また、画像がChatGPT, Codex, OpenAI APIから生成されたものかを判定するツールの一般向けプレビューを開始
https://openai.com/ja-JP/research/verify/
OlmoEarth v1.1: A more efficient family of models, Ai2, 2026.05
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #FoundationModel #OpenWeight #2D (Image) #Author Thread-Post Issue Date: 2026-05-27 Comment
元ポスト:
関連:
- OlmoEarth-v1-Large, Ai2, 2025.11
- Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis, Ai2, 202604
Vision in the Age of LLMs, Lucas Beyer, 2026.05
Paper/Blog Link My Issue
#Article #Tutorial #Pretraining #Transformer #MultiModal #ContrastiveLearning #Video #read-later #VisionLanguageModel #Backbone Issue Date: 2026-05-21 Comment
関連:
- [Paper Note] An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy+, ICLR'21
- [Paper Note] Sigmoid Loss for Language Image Pre-Training, Xiaohua Zhai+, ICCV'23
- [Paper Note] PaliGemma: A versatile 3B VLM for transfer, Lucas Beyer+, arXiv'24, 2024.07
元ポスト:
Gemini Omni を発表, Google Japan Blog, 2026.05
Paper/Blog Link My Issue
#Article #NLP #MultiModal #SpeechProcessing #Blog #Proprietary #VideoGeneration/Understandings #Omni Issue Date: 2026-05-20 Comment
元ポスト:
ZAYA1-VL-8B: Efficient Open Visual Intelligence, ZYPHRA, 2026.05
Paper/Blog Link My Issue
#Article #NLP #SmallModel #OpenWeight #MoE(Mixture-of-Experts) #VisionLanguageModel #One-Line Notes Issue Date: 2026-05-12 Comment
HF: https://huggingface.co/Zyphra/ZAYA1-VL-8B
元ポスト:
画像トークンには双方向のattentionを適用できるようなアーキテクチャを採用
MolmoAct 2: An open foundation for robots that work in the real world, Ai2, 2026.05
Paper/Blog Link My Issue
#Article #NLP #MultiModal #Reasoning #OpenWeight #OpenSource #read-later #Robotics #VisionLanguageActionModel #Author Thread-Post Issue Date: 2026-05-08 Comment
元ポスト:
関連:
- [Paper Note] MolmoAct: Action Reasoning Models that can Reason in Space, Jason Lee+, arXiv'25
dataset:
https://huggingface.co/collections/allenai/molmoact2-datasets
models:
https://huggingface.co/collections/allenai/molmoact2-models
著者ポスト:
著者ポスト2:
著者ポスト3:
Introducing Moonlake's 3D Agent: Computer Use Capabilities For World Modeling, The Moonlake Team, 2026.04
Paper/Blog Link My Issue
#Article #AIAgents #ComputerUse #VisionLanguageModel #2D (Image) #3D (Scene) #Initial Impression Notes #3D Object Generation #ImageTo3D Issue Date: 2026-04-30 Comment
元ポスト:
Blenderを操作可能なComputer Use Agentのようである
Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis, Ai2, 202604
Paper/Blog Link My Issue
#Article #Embeddings #NLP #Proprietary #read-later Issue Date: 2026-04-25 Comment
元ポスト:
関連:
- OlmoEarth-v1-Large, Ai2, 2025.11
ちょっとこれはしっかり読まないと具体的にはわからないかもしれない
Building a Fast Multilingual OCR Model with Synthetic Data, Nvidia, 2026.04
Paper/Blog Link My Issue
#Article #NLP #Dataset #SyntheticData #MultiLingual #OCR #Initial Impression Notes Issue Date: 2026-04-25 Comment
元ポスト:
日本語サンプルも全6ヶ国語中全体の17%含まれておりかなり含まれているOCRモデル学習用の合成データとモデル
model:
https://huggingface.co/nvidia/nemotron-ocr-v2
data:
https://huggingface.co/datasets/nvidia/OCR-Synthetic-Multilingual-v1
スループットが非常に高い
Flipbook is an infinite visual browser generated entirely on demand in real time, Shah+, 2026.04
Paper/Blog Link My Issue
#Article #Blog #VideoGeneration/Understandings #interactive #Realtime #Initial Impression Notes #GUI Issue Date: 2026-04-25 Comment
元ポスト:
画面上のピクセルを全てVideo Generationによってinteractiveに描画するGUIのデモのようである
Gemini Embedding 2: Our first natively multimodal embedding model, Google, 2026.03
Paper/Blog Link My Issue
#Article #Embeddings #NLP #MultiModal #Blog #Proprietary #read-later #Selected Papers/Blogs #KeyPoint Notes #Author Thread-Post Issue Date: 2026-04-25 Comment
元ポスト:
単一のモデルで、マルチモーダルな情報を統合されたembedding空間で表現し、マトリョーシカ表現によって3種類の次元で取得でき、100+言語をサポートしかつcontext windowは8192。オーディオをわざわざ書き起こしてテキストモダリティに変換する必要もなく直接unifiedなembeddingを取得可能というなかなか便利そうな代物。
(以前のIssueを誤って削除したため再掲)
Generally Availableになったとのこと:
vismatch (formerly Image Matching Models), gmberton, 2026.04
Paper/Blog Link My Issue
#Article #Library #2D (Image) #One-Line Notes #needs-revision #Author Thread-Post Issue Date: 2026-04-25 Comment
元ポスト:
50種類以上のimage matchingモデルを統一的なinterfaceでシームレスに利用可能なライブラリとのこと
Hunyuan3D-2, tencent, 2026.04
Paper/Blog Link My Issue
#Article #NLP #Transformer #DiffusionModel #OpenWeight #TextTo3D #ImageTo3D Issue Date: 2026-04-25 Comment
元ポスト:
PaddleOCR, PaddlePaddle
Paper/Blog Link My Issue
#Article #Tools #NLP #DocParser #VisionLanguageModel #OCR #Initial Impression Notes #Author Thread-Post Issue Date: 2026-04-23 Comment
元ポスト:
ブラウザ上でも動作可能らしい
Introducing ChatGPT Images 2.0: A new era of image generation, OpenAI, 2026.04
Paper/Blog Link My Issue
#Article #NLP #ChatGPT #TextToImageGeneration #Proprietary #Selected Papers/Blogs #ImageSynthesis #Initial Impression Notes #Author Thread-Post Issue Date: 2026-04-22 Comment
元ポスト:
めとゃめちゃ良くなってそう
関連:
関連:
Artificial Analysisによる評価(SoTA):
Introducing Claude Design by Anthropic Labs, Anthropic, 2026.04
Paper/Blog Link My Issue
#Article #NLP #Selected Papers/Blogs #VisionLanguageModel #Design Issue Date: 2026-04-18 Comment
元ポスト:
HY-World-2.0, Tencent, 2026.04
Paper/Blog Link My Issue
#Article #Transformer #DiffusionModel #OpenWeight #WorldModels #Author Thread-Post Issue Date: 2026-04-16 Comment
元ポスト:
テクニカルレポート: https://3d-models.hunyuan.tencent.com/world/world2_0/HY_World_2_0.pdf
Introducing Claude Opus 4.7, Anthropic, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MultiModal #Proprietary #Selected Papers/Blogs #VisionLanguageModel #Reference Collection #Author Thread-Post Issue Date: 2026-04-16 Comment
元ポスト:
Artificial Analysisによる評価:
GDPval-AAでGPT-5.4超えのSoTA
IntelligenceでもSoTA(同等)
所見:
所見:
新たなtokenizerを用いている。knowledge cutoffも更新されている。すなわち、新たなベースモデルが事前学習された可能性が高い
tokenizerが更新された=必ずしもベースモデルも新しいということではないよねという指摘:
デグレしたベンチマークがある模様:
所見:
Introducing ERNIE‑Image, Baidu, 2026.04
Paper/Blog Link My Issue
#Article #NLP #Transformer #DiffusionModel #TextToImageGeneration #OpenWeight #Selected Papers/Blogs #2D (Image) #One-Line Notes #ImageSynthesis #Author Thread-Post Issue Date: 2026-04-15 Comment
HF: https://huggingface.co/baidu/ERNIE-Image
ERNIEからtext-to-imageモデルがOpenWeightモデルとしてリリース。ベンチマークとしては公式ブログ上ではOpenWeightモデルの中でトップで、nano banana 2.0に匹敵するようなスコアが出ているように見える
Gemini Robotics-ER 1.6: Powering real-world robotics tasks through enhanced embodied reasoning, Google Deepmind, 2026.04
Paper/Blog Link My Issue
#Article #NLP #Reasoning #Proprietary #Robotics #VisionLanguageActionModel #SpatialUnderstanding #Reference Collection #Initial Impression Notes #Author Thread-Post #MultiView Issue Date: 2026-04-15 Comment
元ポスト:
おー、とうとうDeepmindからVLAがでた。プロプライエタリモデル
私が知らなかっただけで、以前からリリースされていたようだ:
- Building the Next Generation of Physical Agents with Gemini Robotics-ER 1.5, Google, 2025.09
-
https://developers.googleblog.com/en/building-the-next-generation-of-physical-agents-with-gemini-robotics-er-15/
ポイント解説:
LLM-jp-4-VL 9B betaリリース, LLM-jp, 2026.04
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #Japanese #OpenSource #VisionLanguageModel #Author Thread-Post Issue Date: 2026-04-14 Comment
元ポスト:
Introducing Muse Spark: Scaling Towards Personal Superintelligence, Meta, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MultiModal #Proprietary #read-later #Selected Papers/Blogs #VisionLanguageModel #One-Line Notes #Reference Collection #Author Thread-Post Issue Date: 2026-04-11 Comment
元ポスト:
-
-
元ポストのベンチマークスコアを見るとマルチモーダルの性能はフロンティアモデル(gpt5.4, Opus 4.6, Gemini 3.1 Pro)と同等、text/reasoningはフロンティアモデルより少しスコアが低く、特に抽象的な思考が苦手(ARC-AGI-2)。HEALTH分野はhealthは高スコアだがmedicalは少し低めのスコア、Agenticな分野では、SWE Bench Verified/Proよスコアは少し低め、terminal useは明確にスコアが低くtool useは少しスコアが低い、という感じにみえる。
codingとlong horizon taskに継続的に投資するとのこと。
中の人による解説:
全てをフルスクラッチから作り直したっぽい。
Artificial Analysisによる解説:
一気にOpenWeight最強のGLM-5.1超え
所見:
所見:
所見:
第三者によるおそらく独自のベンチマークによる評価の結果、(おそらく101モデルのうち)全体で3位となっているらしい(つまり、既存ベンチマークにoverfittingしているわけではないという考えがある)。
Unfolding Robotics: The Open-Source Recipe for Teaching a Robot to Fold Your Clothes, Hugging Face, 2026.04
Paper/Blog Link My Issue
#Article #Tutorial #NLP #OpenWeight #OpenSource #read-later #Selected Papers/Blogs #Robotics #VisionLanguageActionModel Issue Date: 2026-04-07 Comment
元ポスト:
Introducing WildDet3D: Open-world 3D detection from a single image, Ai2, 2026.04
Paper/Blog Link My Issue
#Article #Dataset #OpenWeight #OpenSource #read-later #Selected Papers/Blogs #3D (Video) #ObjectDetection #Initial Impression Notes Issue Date: 2026-04-07 Comment
元ポスト:
wildな環境においてzero shot(click, text, bounding boxで対象を指定)で動作する単眼の3D Object Detectionモデルとのこと。データセットもコードも公開
Vision Language Models (Better, faster, stronger), merve+, 2025.03
Paper/Blog Link My Issue
#Article #Survey #NLP #Blog #VisionLanguageModel #Initial Impression Notes Issue Date: 2026-04-07 Comment
元ポスト:
1年前のVLMに関するトレンドをまとめた記事のようだが、その後も同トレンドが継続している模様
JEPAwiki, mishig, 2026.04
Paper/Blog Link My Issue
#Article #Blog #read-later #WorldModels #LatentRepresentation Issue Date: 2026-04-07 Comment
元ポスト:
Moonlake: Causal World Models should be Multimodal, Interactive, and Efficient — with Chris Manning and Fan-yun Sun, LatentSpace, 2026.04
Paper/Blog Link My Issue
#Article #Tutorial #read-later #WorldModels Issue Date: 2026-04-05 Comment
元ポスト:
ibm-granite_granite-4.0-3b-vision, ibm-granite, 2026.03
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #VisionLanguageModel Issue Date: 2026-04-04 Comment
元ポスト:
Build with Veo 3.1 Lite, our most cost-effective video generation model, Google, 2026.03
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #Proprietary #VideoGeneration/Understandings #TextToVideoGeneration #ImageToVideoGeneration Issue Date: 2026-04-04 Comment
元ポスト:
Qwen3.5-Omni: Scaling Up, Toward Native Omni-Modal AGI, Qwen Team, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #SpeechProcessing #Proprietary #VisionLanguageModel #2D (Image) #3D (Video) #Omni #AudioLanguageModel #audio #text Issue Date: 2026-04-04 Comment
元ポスト:
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory, Skywork AI, 2026.04
Paper/Blog Link My Issue
#Article #Transformer #SyntheticData #DiffusionModel #OpenWeight #VideoGeneration/Understandings #WorldModels #interactive #Game #3D (Video) #LongHorizon #Realtime #Initial Impression Notes Issue Date: 2026-04-02 Comment
元ポスト:
Unreal Engineで合成されたデータに基づいて学習されたDiTベースのWorld Modelらしい。
Acknowleagementから察するに、Wan2.2がベースモデルで、self-forcingが学習に用いられている。
- Wan2.2, Alibaba Wan, 2025.07
- [Paper Note] Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion, Xun Huang+, NeurIPS'25
また、action control moduleをアーキテクチャに導入することで、汎用的な動画生成モデルにキーボード、マウス等のアクションによるコントロールを実現している模様。
- [Paper Note] GameFactory: Creating New Games with Generative Interactive Videos, Jiwen Yu+, arXiv'25, 2025.01
デコードの高速化には量子化を利用しているとのこと。
Gemma 4: Byte for byte, the most capable open models, Google, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #MultiModal #SpeechProcessing #Reasoning #OpenWeight #MoE(Mixture-of-Experts) #Selected Papers/Blogs #VisionLanguageModel #2D (Image) #3D (Video) #One-Line Notes #Reference Collection #AudioLanguageModel #audio #text #Initial Impression Notes Issue Date: 2026-04-02 Comment
元ポスト:
2B, 4B, 26BのMoEモデルと31BのDenseモデルの4種類のモデルファミリーで、マルチモーダル(vision)対応。2B, 4Bはaudioも入力として扱える。
edgeデバイス向けのモデルは128k, 他は256kのコンテキストウィンドウ。140+の多言語サポート。
Apache 2.0ライセンス
arenaで同サイズのモデル群でSoTAといった話がブログ中に記述されている。
モデルカードには一般的なベンチマーク群とのスコアも記載されている。
https://ai.google.dev/gemma/docs/core/model_card_4?hl=ja
(そもそも既存のベンチマークにもコンタミネーションがあると思われるが、)arenaに関しては特定の企業に対してデータを提供し、複数のモデルの亜種をテストできるという慣行があり、リーダーボードにバイアスがあるであろう点には注意:
- [Paper Note] The Leaderboard Illusion, Shivalika Singh+, NeurIPS'25
artificial analysisによる評価:
Qwenがproprietaryになったことから、ライセンス的に使いやすく、日本語に強そうなモデルとしては筆頭ではなかろうか。日本語性能が気になる。
アーキテクチャ解説:
ポイント解説:
所見:
attentionのscaleをsqrt(d)でスケールさせる代わりに、QK-norm, V normを適用するなど。
NvidiaによるNVFP4へのpost-trainingによる量子化:
https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4
量子化後の性能も比較されており、知識、数学、コーディング、terminac useなど6種類のベンチマークでオリジナルのモデルと遜色ない性能が出ている旨記載されている。
解説:
https://newsletter.maartengrootendorst.com/p/a-visual-guide-to-gemma-4
所見(encoder-freeにした裏側でパッチ化→projection + x/y軸のpositional embeddingを実施している話):
Holo3: Breaking the Computer Use Frontier, H Company, 2026.03
Paper/Blog Link My Issue
#Article #NLP #AIAgents #MultiModal #OpenWeight #MoE(Mixture-of-Experts) #ComputerUse #read-later #VisionLanguageModel #One-Line Notes #GUI #Environment Issue Date: 2026-04-02 Comment
元ポスト:
HF: https://huggingface.co/Hcompany/Holo3-35B-A3B
関連:
- Holo2: Cost-Efficient Models for Cross-Platform Computer-Use Agents, H Company, 2025.11
Qwen3.5をファインチューニングすることで実現。以前のシリーズもQwenベースだったが、新たなQwenのリリースに伴いより強力なベースモデルを得て、かつシナリオをベースにして自動でwebsiteを構築しverifiableが可能な独自のEnvironmentを保持しており、多様な合成データの活用とRLを実現することで、性能が向上していると思われる。
sarashina2.2-ocr, SBIntuitions, 2026.03
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #Japanese #Selected Papers/Blogs #DocParser #OCR #Initial Impression Notes Issue Date: 2026-03-31 Comment
元ポスト:
縦書き文書に強いのは大変ありがたい
dots.ocrよりも日本語文書に対するCERとBLEUのスコアが良い。素晴らしい
chandra-ocr-2, datalab-to, 2026.03
Paper/Blog Link My Issue
#Article #MultiLingual #OpenWeight #Selected Papers/Blogs #VisionLanguageModel #OCR #One-Line Notes Issue Date: 2026-03-21 Comment
元ポスト:
日本語の認識性能がGemini-2.5-Flashよりも高い。マルチリンガルでの認識性能がこらほど網羅的に列挙されているのはありがたい。
OpenClaw — Personal AI Assistant, openclaw, 2026.03
Paper/Blog Link My Issue
#Article #Tools #NLP #AIAgents #Repository #ComputerUse #Selected Papers/Blogs #WorkspaceAgents Issue Date: 2026-03-19 Comment
2026.04.07:
xperience-10m, ropedia-ai, 2026.03
Paper/Blog Link My Issue
#Article #Dataset #Selected Papers/Blogs #Robotics #3D (Video) #EmbodiedAI Issue Date: 2026-03-17 Comment
元ポスト:
Ropediaとは:
- Interactive Intelligence from Human Xperience, Ropedia, 2025.12
アナウンス:
5日で1.66M downloadsとのこと:
最近の主要なembodied AIのための基盤モデルに使われているとのこと:
computer-use-large, markov-ai, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Dataset #AIAgents #ComputerUse #Selected Papers/Blogs #VisionLanguageModel #3D (Video) #One-Line Notes Issue Date: 2026-03-15 Comment
元ポスト:
12,300時間程度の、プロフェッショナルなソフトウェア(AutoCAD, Blender, Excel, Photoshop, Salesforce VSCode)利用しているスクリーンのレコーディングデータとのこと。
CC-BY-4.0!?
FLUX.2-klein-9B, black-forest-labs, 2026.01
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #Transformer #TextToImageGeneration #SmallModel #Selected Papers/Blogs #2D (Image) #Editing #One-Line Notes Issue Date: 2026-03-15 Comment
元ポスト:
github: https://github.com/black-forest-labs/flux2
そもそも2025年11月にリリースされているFLUX.2は結構色々なところで名前を見かけるのでおさえておいたほうが良いかもしれない
https://bfl.ai/blog/flux-2
kleinはFLUX.2シリーズの中で最も軽量なモデルとのこと。2ヶ月程度で既に110k DLされている。
Reka Edge: Frontier-Level Edge Intelligence for Physical AI, Reka, 2026.03
Paper/Blog Link My Issue
#Article #NLP #MultiModal #OpenWeight #VisionLanguageModel Issue Date: 2026-03-14 Comment
元ポスト:
Moondream 3 Preview: Frontier-level reasoning at a blazing speed, Moondream, 2025.09
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #FoundationModel #Reasoning #SmallModel #OpenWeight #Selected Papers/Blogs #VisionLanguageModel #KeyPoint Notes Issue Date: 2026-03-12 Comment
HF: https://huggingface.co/moondream/moondream3-preview
9B-A2Bの小規模なVLMで、
- visual reasoning: 小規模だが実タスクに適用可能なvisual reasoning性能
- trainable: Visual系のタスクは人間でもzero shotではできないことが多く、簡単にfinetuningできることが重要で
- fast: vision系のアプリケーションはリアルタイムのlavencyが求められることが多く
- inexpensive: 安くスケーラブルでなければならない
をテーマにしたモデルのようである。
object detection, pointing, 構造化された出力(犬の群の個々の犬の毛と首輪の色動画)、OCRなどの様々なタスクが実行可能で、GPT5, Gemini 2.5 Flash, Claude 4 Sonnetをこの規模感のモデルで、objec' detection, counting, document understanding, hallucinationに関するベンチマークで上回る。
前身のモデルであるmoondream2は、5Mダウンロードを達成したようだ
vikhyatk/moondream2
Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis, Black Forest Labs, 2026.03
Paper/Blog Link My Issue
#Article #Pretraining #NLP #MultiModal #SpeechProcessing #Self-SupervisedLearning #read-later #2D (Image) #FlowMatching #3D (Video) #Omni #RectifiedFlow #audio Issue Date: 2026-03-10 Comment
backbone modelは下記のFLUX.2と呼ばれるモデル:
FLUX Commercial Licensing:
https://bfl.ai/licensing
先行研究:
- The Simulation Company, Simile, 2026.02
先行研究から読みたい
元ポスト:
Awesome World Models, knightnemo,
Paper/Blog Link My Issue
#Article #Survey #Robotics #WorldModels Issue Date: 2026-03-08
Awesome World Models for Robotics, leofan90,
Paper/Blog Link My Issue
#Article #Survey #Robotics #WorldModels Issue Date: 2026-03-08
Awesome From Video Generation to World Model, ziqihuangg, 2026.03
Paper/Blog Link My Issue
#Article #Survey #Robotics #WorldModels Issue Date: 2026-03-08 Comment
元ポスト:
Yuan3.0-Ultra, YuanLabAI, 2026.03
Paper/Blog Link My Issue
#Article #NLP #MultiModal #OpenWeight #MoE(Mixture-of-Experts) #VisionLanguageModel #UMM #One-Line Notes #Initial Impression Notes Issue Date: 2026-03-07 Comment
元ポスト:
MoEのwarmupが終わり安定してきたタイミングでルーティングがされにくいExpertを枝刈りし、残ったexpertに対してバランスよくルーティングがされるようなrearrangeをするアルゴリズム Layer-Adaptive Expert Pruning (LAEP)によって、パラメータサイズを1515Bから1010Bまで削減し、49%程度事前学習の効率を改善したとのこと。
RAG, multimodal document understanding, tabular data analysis, content summarizationにおいて、非常に高い性能を獲得している。tool useに関してはGPT-5.2(effort不明)以外には負けているので、優秀ではあるが特に秀でているというわけではないよつに見える(BFCVv3)。
しかし他のベンチマークでこれらフロンティアモデル群をここまでPass@1やAccで抜くのは、驚きではあるが、実際にどのような評価をしているのかはテクニカルレポートを見た方が良いと思われる。
ocr-bench, davanstrien, 2026.03
Paper/Blog Link My Issue
#Article #Tools #NLP #Evaluation #Repository #LLM-as-a-Judge #OCR #One-Line Notes #Initial Impression Notes Issue Date: 2026-03-06 Comment
元ポスト:
自分が試したいドキュメントのコレクションに対して、5つほどのOpenなOCRで実際に書き起こしを行い、VLM-as-a-JudgeでスコアリングしELOでの当該ドキュメントセットに対するスコアボードを作成するツール
非常に興味深く実用的だが、個人的にOlmOCRもサポートして欲しいなぁと思うなど。あと、機密性の高い文書などを扱う場面では、セキュリティ面にどれだけ配慮されているのかが気になってしまう。
HY-WU (Part I): An Extensible Functional Neural Memory Framework and An Instantiation in Text-Guided Image Editing, Tencent HY Team, 2026.03
Paper/Blog Link My Issue
#Article #Personalization #PEFT(Adaptor/LoRA) #2D (Image) #memory #Editing #One-Line Notes #ImageSynthesis #Adaptive Issue Date: 2026-03-06 Comment
元ポスト:
source imageとpromptから、frozenされたモデルに対するadapter weightを(finetuningなしで)動的に生成し、インスタンス固有のパラメータを用いることでinstance specificな演算を実現する
関連:
- [Paper Note] Doc-to-LoRA: Learning to Instantly Internalize Contexts, Rujikorn Charakorn+, arXiv'26, 2026.02
- [Paper Note] Text-to-LoRA: Instant Transformer Adaption, Rujikorn Charakorn+, ICML'25, 2025.06
NEO-unify: Building Native Multimodal Unified Models End to End, SenseTime, 2026.03
Paper/Blog Link My Issue
#Article #NLP #MultiModal #Post #Architecture #VisionLanguageModel #UMM #One-Line Notes #Pixel-based Issue Date: 2026-03-06 Comment
Vision EncoderやVAEを用いずに、pixel,wordの入力でnativeなunified modelを構築する。
takeawayとしては
- エンコーダーフリーなアーキテクチャでも、意味とピクセルの表現の両方を保持できる
- image reconstruction, image editingの両者において高い性能を獲得
- understandingとgenerationのtransformerを別々に事前学習し、その後両者を組み合わせて(Mixture of Transformer)追加のSFTをしているようだが、その際に両者のtransformerがconflictすることなく、understandingタスクは安定したままgenerationタスクは素早く収束するといった挙動を示した
- mid-training後により大規模なweb-scaleでの事前学習をするようだが、その際に競合モデルよりもよりデータ効率良く学習ができた
という感じらしい
Build with Nano Banana 2, our best image generation and editing model, Google, 2026.02
Paper/Blog Link My Issue
#Article #NLP #TextToImageGeneration #Proprietary #Editing #ImageSynthesis Issue Date: 2026-02-28 Comment
元ポスト:
NDLOCR-Liteの公開について, NDL Lab, 2026.02
Paper/Blog Link My Issue
#Article #NeuralNetwork #NLP #Blog #Repository #Japanese #Selected Papers/Blogs #Encoder-Decoder #OCR #One-Line Notes Issue Date: 2026-02-28 Comment
元ポスト:
江戸期以前の和古書、清代以前の漢籍といった古典籍資料のデジタル化画像からテキストデータを作成するOCRとのこと。以前はGPUで動作していたが、CPUで動作するようにした軽量版とのこと。すごい。
The First Fully General Computer Action Model, Standard Intelligence Team, 2026.02
Paper/Blog Link My Issue
#Article #Pretraining #FoundationModel #DiffusionModel #ComputerUse #3D (Video) #One-Line Notes #WorldActionModel Issue Date: 2026-02-27 Comment
元ポスト:
公式ポスト:
関連:
- [Paper Note] Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos, Bowen Baker+, arXiv'22, 2022.06
Training Recipeの部分を読むと、上記研究で提案されているVideo PreTrainingと同じ手法を用いているように見える。
つまり、Inverse Dynamics Modelを学習し、大量のvideoデータに対してアクションラベルを付与し、付与されたアクションラベルを用いて半教師あり学習によるnext action predictionを実施することによって基盤モデルを学習する、というアプローチ。
この基盤モデルによってたとえば1時間のサンフランシスコをdrivingしている動画によってfinetuningすることで、自動運転をするようなモデルが学習できる、といったことが実現可能な模様。
[Paper Note] Preconditioned inexact stochastic ADMM for deep models, Nature Machine Intelligence 2026, 2026.02
Paper/Blog Link My Issue
#Article #NeuralNetwork #MachineLearning #NLP #LanguageModel #Optimizer #Initial Impression Notes #Nature Machine Intelligence Issue Date: 2026-02-24 Comment
元ポスト:
パラメータサイズが大きい場合にMuon超え...?
所見:
Qwen3.5: Towards Native Multimodal Agents, Qwen Team, 2026.02
Paper/Blog Link My Issue
#Article #NLP #ReinforcementLearning #MultiModal #MultiLingual #OpenWeight #MoE(Mixture-of-Experts) #read-later #Selected Papers/Blogs #VisionLanguageModel #UMM #KeyPoint Notes #Scalability #Environment Issue Date: 2026-02-17 Comment
元ポスト:
最新のQwenがリリース・・・!!
- Vision+TextのUMMを採用。
- real-world agentsのために訓練
- hybrid linear attention + sparse MoE + 環境スケーリングに基づくlarge scale RLを実施
- decodingのスループットがQwen3-Maxと比較して8.6--19.0倍
- 201の言語と方言をサポート
- 397B-A17B
- Gated DeltaNet
- Gated Attention
- context length: 262k
- Multi token prediction
- 言語系タスクではGPT5.2と比較して少し劣る程度、agenticなベンチマークでは大きく上回るものも存在(ただし、Claude 4.5 Opusには届いていないベンチマークが多いように見える)
- Vision系タスクでは全体的にGPT5.2, Opus 4.5よりも優秀に見え、Gemini 3 Proと同等か少し劣る程度に見える。
世はlinear attention時代
所見:
INT4モデル:
dots.ocr-1.5, rednote-hilab, 2026.02
Paper/Blog Link My Issue
#Article #NLP #MultiModal #StructuredData #SmallModel #OpenWeight #DocParser #OCR Issue Date: 2026-02-16 Comment
元ポスト:
[Paper Notes] Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity, Bytedance Seed, 2026.02
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Reasoning #Proprietary #VisionLanguageModel Issue Date: 2026-02-16 Comment
元ポスト:
所見:
Ming-flash-omni-2.0, inclusionAI, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Transformer #MultiModal #SpeechProcessing #DiffusionModel #Speech #OpenWeight #MoE(Mixture-of-Experts) #2D (Image) #Omni #text Issue Date: 2026-02-12 Comment
元ポスト:
関連:
- Ming-flash-omni-Preview, inclusionAI, 2025.10
- [Paper Note] Ming-Omni: A Unified Multimodal Model for Perception and Generation, Inclusion AI+, arXiv'25, 2025.06
公式ポスト:
Introducing Lab: The Full-Stack Platform for Training your Own Models, Prime Intellect, 2026.02
Paper/Blog Link My Issue
#Article #MachineLearning #NLP #LanguageModel #Infrastructure #ReinforcementLearning #AIAgents #Blog #ScientificDiscovery #PostTraining #Selected Papers/Blogs #One-Line Notes #Reference Collection #Environment Issue Date: 2026-02-11 Comment
元ポスト:
事後学習、特にAgenticな研究の民主化のためのプラットフォームの提供
所見:
利用例 (Environment Hub):
new-datasets-in-machine-learning, librarian-bots, 2026.02
Paper/Blog Link My Issue
#Article #RecommenderSystems #Survey #MachineLearning #InformationRetrieval #NLP #Dataset #Evaluation #SpeechProcessing #Robotics #Live Issue Date: 2026-02-09 Comment
元ポスト:
ModernBERTをFinetuningした分類器を用いてデータセットやベンチマークを提案している研究を自動分類して検索できるようにしている。有用
Intern-S1-Pro, internlm, 2026.02
Paper/Blog Link My Issue
#Article #NLP #MultiModal #Reasoning #PositionalEncoding #OpenWeight #MoE(Mixture-of-Experts) #VisionLanguageModel #Science Issue Date: 2026-02-05 Comment
元ポスト:
ポイント解説:
関連:
- [Paper Note] Intern-S1: A Scientific Multimodal Foundation Model, Lei Bai+, arXiv'25, 2025.08
Fourier Position Encoding (FoPE) + upgraded time-series modeling
MiniCPM-o-4_5, OpenBMB, 2026.02
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #SpeechProcessing #DiffusionModel #OpenWeight #AutomaticSpeechRecognition(ASR) #VisionLanguageModel #TTS #Omni #AudioLanguageModel Issue Date: 2026-02-05 Comment
元ポスト:
The Second Pre-training Paradigm, Jim Fan, X, 2026.02
Paper/Blog Link My Issue
#Article #Pretraining #NLP #LanguageModel #MultiModal #Post #Robotics #WorldModels #One-Line Notes Issue Date: 2026-02-05 Comment
事前学習がnext word predictionから過去の行動と状態によって条件付けられ次の(ある期間の)世界の状態を予測するワールドモデリング(next physical state prediction)へのパラダイムシフトの予想(というよりこのパラダイムシフトの真っ只中にいる)。人間の脳が処理する情報の多くは視覚であり、言語的な領域は部分的なことであることや、猿は言語的な能力が低くても視覚や運動、触覚などの感覚的情報から世界の物理法則を理解し知的なアクションをとるメンタルモデルを確立していることなどを引き合いに説明している。
New Holo2 model takes the lead in UI Localization, H Company, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Blog #OpenWeight #ComputerUse #Selected Papers/Blogs #VisionLanguageModel #Grounding #GUI Issue Date: 2026-02-05 Comment
HF: https://huggingface.co/Hcompany/Holo2-235B-A22B
元ポスト:
関連:
- Holo1.5 - Open Foundation Models for Computer Use Agents, H Company, 2025.09
Project Genie: Experimenting with infinite, interactive worlds, Google Deepmind, 2026.01
Paper/Blog Link My Issue
#Article #NLP #GenerativeAI #Proprietary #WorldModels #interactive Issue Date: 2026-01-30 Comment
元ポスト:
Googleからのworld model
Accelerating Diffusion Models with an Open, Plug-and-Play Offering, Nvidia, 2026.01
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #Tools #NLP #DiffusionModel #TextToImageGeneration #Distillation #PostTraining #2D (Image) #Editing #3D (Video) #TextToVideoGeneration #ImageToTextGeneration #TrainingFramework Issue Date: 2026-01-29 Comment
元ポスト:
self forcingも実装されている
- [Paper Note] Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion, Xun Huang+, NeurIPS'25
Introducing Agentic Vision in Gemini 3 Flash, Google Deepmind, 2026.01
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Proprietary #VisionLanguageModel #One-Line Notes Issue Date: 2026-01-29 Comment
元ポスト:
visual reasoningとコード実行の融合
DeepSeek-OCR-2, DeepSeek-AI, 2026.01
Paper/Blog Link My Issue
#Article #NLP #OCR #Compression Issue Date: 2026-01-27 Comment
元ポスト:
関連:
- DeepSeek-OCR: Contexts Optical Compression, DeepSeek, 2025.10
Minimax Agent, Minimax, 2026.01
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #GenerativeAI #ComputerUse Issue Date: 2026-01-27 Comment
code: https://github.com/MiniMax-AI/Mini-Agent
元ポスト:
Composing Weight and Data Sparsity in MoE: Improving compute efficiency through varying compute per token, Perceptron, 2026.01
Paper/Blog Link My Issue
#Article #Pretraining #NLP #MultiModal #MoE(Mixture-of-Experts) #read-later #VisionLanguageModel #Routing #Sparse #Initial Impression Notes Issue Date: 2026-01-23 Comment
元ポスト:
MoEがトークン単位でactivateするweightをサブセットにするweight sparcityによって効率化を実現する手法とみなしたときに、それぞれのinputに情報量の濃淡があることから現在のトークンごとにweightを割り当てるのではなく、weightごとにトークンを割り当てるというもう一つの軸を考えることができ(=Data Sparcity)、これをweightごとにトークンのsubsetしか持たないような実現方法をとるとcontextが損なわれauto-regressiveの前提が崩れるためtrain-inference-mismatchが生じるので、null experts(受け取ったトークンに対して何もしない)を実装して実現するみたいな話のように見えるが全くまだ読めていない。
Waypoint-1: Real-time Interactive Video Diffusion from Overworld, Overworld, 2026.01
Paper/Blog Link My Issue
#Article #Controllable #NLP #Transformer #MultiModal #DiffusionModel #OpenWeight #WorldModels #interactive #3D (Video) #One-Line Notes #RectifiedFlow #Realtime Issue Date: 2026-01-22 Comment
blog:
https://over.world/blog/the-path-to-real-time-worlds-and-why-it-matters
pj page:
https://over.world/
元ポスト:
リアルタイムにzero latencyでマウス(カメラも自由に動かせる)、キーボード、テキストでinteraction可能なworld model
ICLR 2026 Acceptance Prediction: Benchmarking Decision Process with A Multi-Agent System, Zhang+, 2026.01
Paper/Blog Link My Issue
#Article #NLP #Dataset #LanguageModel #AIAgents #Evaluation #MultiModal #ScientificDiscovery #VisionLanguageModel #AcademicWriting #Live #One-Line Notes Issue Date: 2026-01-20 Comment
元ポスト:
conference paperのpeer reviewに関するベンチマーク。accept/rejectを予測する。papers, reviews, rebuttalsそしてfinal decisionsが紐づけられている。
A Visual Introduction to Rectified Flows, Alec Helbling, 2026.01
Paper/Blog Link My Issue
#Article #Tutorial #MachineLearning #Blog #read-later #FlowMatching #RectifiedFlow Issue Date: 2026-01-19 Comment
元ポスト:
action100m-preview, Meta, 2026.01
Paper/Blog Link My Issue
#Article #Dataset #Robotics #VisionLanguageActionModel #3D (Video) Issue Date: 2026-01-16 Comment
元ポスト:
Next generation medical image interpretation with MedGemma 1.5 and medical speech to text with MedASR, Google Research, 2026.01
Paper/Blog Link My Issue
#Article #NLP #MultiModal #SpeechProcessing #Blog #OpenWeight #AutomaticSpeechRecognition(ASR) #VisionLanguageModel #Medical Issue Date: 2026-01-14 Comment
元ポスト:
ポイント解説:
GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image Generation, Z.ai, 2026.01
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MultiModal #DiffusionModel #TextToImageGeneration #OpenWeight #Editing Issue Date: 2026-01-14 Comment
元ポスト:
🍫 Local Cocoa: Your Personal AI Assistant, Fully Local 💻, synvo-ai, 2026.01
Paper/Blog Link My Issue
#Article #Tools #NLP #LanguageModel #AIAgents #MultiModal #Selected Papers/Blogs #ContextEngineering #memory Issue Date: 2026-01-09 Comment
元ポスト:
NVIDIA Cosmos Reason 2 Brings Advanced Reasoning To Physical AI, Nvidia, 2026.01
Paper/Blog Link My Issue
#Article #NLP #Reasoning #LongContext #SmallModel #OpenWeight #ObjectLocalization #VisionLanguageModel #Robotics #SpatialUnderstanding #EmbodiedAI #Physics Issue Date: 2026-01-06 Comment
HF: https://huggingface.co/nvidia/Cosmos-Reason2-8B?linkId=100000401175768
元ポスト:
LightX2V: Light Video Generation Inference Framework, ModelTC, 2025.12
Paper/Blog Link My Issue
#Article #Library #LLMServing #VideoGeneration/Understandings #3D (Video) Issue Date: 2025-12-24 Comment
元ポスト:
A2UI: A Protocol for Agent-Driven Interfaces, Google, 2025
Paper/Blog Link My Issue
#Article #Tools #NLP #AIAgents #SoftwareEngineering #VisionLanguageModel #One-Line Notes Issue Date: 2025-12-22 Comment
AI Agent (Gemini)を用いてUIを自動生成できるツールらしい
元ポスト:
Introducing Mistral OCR 3, Mistral AI, 2025.12
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Proprietary #VisionLanguageModel #OCR #One-Line Notes Issue Date: 2025-12-19 Comment
元ポスト:
MistralによるOCR。他のOCRに比べてmulti-lingual, 様々なデータセットで高い性能を発揮。APIでのみ提供されている模様。
[Paper Note] Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning, Meta, 2025.12
Paper/Blog Link My Issue
#Article #Library #MultiModal #SpeechProcessing #python #Encoder #2D (Image) #3D (Video) #audio Issue Date: 2025-12-19 Comment
元ポスト:
様々なモダリティ(画像・動画・音声等)をエンコードできるPerception Encoderに最近リリースされたSAM Audio (Audio-Visual / Audio-frame) も組み込まれた模様
code:
https://github.com/facebookresearch/perception_models
Seed1.8, ByteDance Seed, 2025.12
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Proprietary #ComputerUse #VisionLanguageModel Issue Date: 2025-12-18 Comment
元ポスト:
GUI Agentとして性能はトップレベル(Opusが比較対象に入っていないが)で、
テキスト、画像モダリティでの検索でもトップレベル、codingやツール利用などは少し劣るように見える。
LLM系、VideoUnderstanding系ののベンチマークではフロンティアモデル群と同等、VLM系のタスクではフロンティアモデル群と同等以上の性能に見える。
が、一方のモダリティはGPT5で比較しているのに対し、他方はGPT5.1であったりしており、比較対象が少し恣意的にピックされているのでは?という気もする。
LongCat-Video-Avatar, meituan-longcat, 2025.12
Paper/Blog Link My Issue
#Article #Transformer #DiffusionModel #VariationalAutoEncoder #OpenWeight #VideoGeneration/Understandings #3D (Scene) #One-Line Notes #Audio-Text-to-Video #Audio-Text-Image-to-Video #Video Continuation Issue Date: 2025-12-17 Comment
元ポスト:
アーキテクチャはDiTベースのDiffusion Modelで、3D Variational AutoencoderによってEncode/Decodeされ、3D RoPEによって位置情報が埋め込まれる。DiT Blockでは、テキストとaudio用のcross attentionが用いられてこれらのモーダルに関する情報が組み込まれる。audioはWav2Vecでエンコードされ、テキストはUMT5[^1]によってエンコードされる。
[^1]: multilingualなT5で100言語以上がサポートされている模様
bu-30b-a3b-preview: Meet BU-30B-A3B-Preview — bringing SoTA Browser Use capabilities in a small model that can be hosted on a single GPU., browser-use, 2025.12
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #ComputerUse #VisionLanguageModel Issue Date: 2025-12-17 Comment
元ポスト:
Molmo 2: State-of-the-art video understanding, pointing, and tracking, Ai2, 2025.12
Paper/Blog Link My Issue
#Article #NLP #MultiModal #SmallModel #OpenWeight #OpenSource #Selected Papers/Blogs #VideoGeneration/Understandings #VisionLanguageModel #2D (Image) #3D (Video) #KeyPoint Notes Issue Date: 2025-12-17 Comment
テクニカルレポート:
https://www.datocms-assets.com/64837/1765901660-molmo_v2_2026-techreport-3.pdf
HF:
https://huggingface.co/collections/allenai/molmo2
Qwen3とOlmoをベースにしたvariantsが存在し、Olmoの方はバックボーンのLLMも含めて全てがオープンになっている。MetaのPerceptionLMと比較して1/8の動画データ量で高い性能を達成できており、データのcurationの品質と、grounding basedな目的関数の工夫によって実現されているとのこと。
proprietaryなモデル群と比較すると、trackingは圧勝、そのほかはGPT5-miniと同様なものが多い。モデルによってタスクの優劣が結構分かれており、Video関連タスクをタスクをまたいで汎化させることにはclosedでも苦戦しているように見える。
オープンモデルとの比較で言うと圧勝で、LongVideoのQAに関してだけは、Eagle2.5-8Bと呼ばれるモデルが勝っている。
あとは全体を通じてLLMのバックボーンがQwen3の場合の性能が良いことが興味深い。バックボーンに採用するLLMに応じて性能が結構変わる。これはアーキテクチャがそもそもConnectorを利用するタイプのもので、Unifiedなアーキテクチャではないことが要因としては考えられる。
元ポスト:
demo:
コードベースが公開:
https://github.com/allenai/molmo2
AutoGLM-Phone-9B, Zhipu AI, 2025.12
Paper/Blog Link My Issue
#Article #NLP #SmallModel #OpenWeight #VisionLanguageModel Issue Date: 2025-12-10 Comment
元ポスト:
GLM-4.6V, Zhipu AI, 2025.12
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #VisionLanguageModel Issue Date: 2025-12-10 Comment
元ポスト:
Introducing Mistral 3 The next generation of open multimodal and multilingual AI, Mistral AI, 2025.12
Paper/Blog Link My Issue
#Article #NLP #MultiModal #Blog #MultiLingual #OpenWeight #VisionLanguageModel #One-Line Notes Issue Date: 2025-12-03 Comment
元ポスト:
マルチモーダルなベンチマークがほとんどないように見えるMM-MT-Benchというもののみ?
[Paper Notes] Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem, Longpre+, 2025.11
Paper/Blog Link My Issue
#Article #Analysis #NLP #LanguageModel #OpenWeight #VisionLanguageModel Issue Date: 2025-11-30 Comment
元ポスト:
MITとHuggingFaceの調査によると、open weightモデルのDLにおいて、米国のAI産業における中国のモデルDL数が米国のモデルを初めて抜いた模様。
ダッシュボード: https://huggingface.co/spaces/economies-open-ai/open-model-evolution
生成AI革命の最前線:拡散を超える「流れ」の思想とMambaの台頭, laughman-ai, 2025.10
Paper/Blog Link My Issue
#Article #Blog #FlowMatching #reading #RectifiedFlow #FlowMaps Issue Date: 2025-11-28
Flow With What You Know, Scott H. Hawley, 2024.11
Paper/Blog Link My Issue
#Article #Blog #read-later #FlowMatching #RectifiedFlow #Physics Issue Date: 2025-11-28
GPT-4V-Act, ddupont808, 2023.10
Paper/Blog Link My Issue
#Article #NLP #Repository #ComputerUse #VisionLanguageModel #One-Line Notes #Grounding Issue Date: 2025-11-25 Comment
GPT4V(VLM)と、SoMを用いてVLMによってWebUIとClick/Keyboard操作を通じてinteractできる実装
- [Paper Note] Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V, Jianwei Yang+, arXiv'23, 2023.10
OCR Arena, extend.ai, 2025.11
Paper/Blog Link My Issue
#Article #NLP #Evaluation #VisionLanguageModel #OCR #One-Line Notes Issue Date: 2025-11-25 Comment
元ポスト:
OCRのアリーナ(=ユーザがPDFをアップロードし2モデルでOCRし優劣をユーザが判定しその結果からElo Rateを算出する)。
言語間の性能差はわからないので参考程度にすると良いと思われる。
Introducing Nano Banana Pro, Google, 2025.11
Paper/Blog Link My Issue
#Article #GenerativeAI #Proprietary #Selected Papers/Blogs #2D (Image) Issue Date: 2025-11-21 Comment
元ポスト:
所見:
所見:
Hunyuan Video 1.5 Technical Report, Tencent, 2025.11
Paper/Blog Link My Issue
#Article #Transformer #DiffusionModel #OpenWeight #VideoGeneration/Understandings Issue Date: 2025-11-21 Comment
pj page:
https://hunyuan.tencent.com/video/zh?tabIndex=0
HF:
https://huggingface.co/tencent/HunyuanVideo-1.5
元ポスト:
TAURO Project, note, 2024.10
Paper/Blog Link My Issue
#Article #Tutorial #NLP #Blog #ScientificDiscovery #Japanese #Robotics Issue Date: 2025-11-20 Comment
元ポスト:
👀👀👀
NVIDIA-Nemotron-Parse-v1.1, NVIDIA, 2025.11
Paper/Blog Link My Issue
#Article #NLP #TabularData #OpenWeight #read-later #DocParser #VisionLanguageModel #OCR Issue Date: 2025-11-20 Comment
元ポスト:
olmocr2と比較して性能はどうだろうか、特に日本語
- olmOCR 2: Unit test rewards for document OCR, Ai2, 2025.10
Introducing SAM 3D: Powerful 3D Reconstruction for Physical World Images, Meta, 2025.11
Paper/Blog Link My Issue
#Article #FoundationModel #Blog #read-later #Selected Papers/Blogs #3D Reconstruction #3D (Scene) Issue Date: 2025-11-20 Comment
元ポスト:
解説:
Introducing Meta Segment Anything Model 3 and Segment Anything Playground, Meta, 2025.11
Paper/Blog Link My Issue
#Article #ImageSegmentation #FoundationModel #Blog #read-later #Selected Papers/Blogs #2D (Image) #3D (Video) Issue Date: 2025-11-20 Comment
元ポスト:
今度はSAM3、最近毎日なんか新しいの出てるな
SAM 3.1:
https://huggingface.co/facebook/sam3.1
元ポスト:
Awesome Spatial Intelligence in VLMs, mll-lab-nu, 2025.11
Paper/Blog Link My Issue
#Article #Survey #NLP #MultiModal #Repository #VisionLanguageModel #SpatialUnderstanding Issue Date: 2025-11-18 Comment
元ポスト:
VLM, マルチモーダルなLLMにおけるSpatial Intelligenceに関する論文リスト
How to Train a State-of-the-Art Pathology Foundation Model with $1.6k, Kaplan+, 2025.11
Paper/Blog Link My Issue
#Article #Transformer #FoundationModel #Medical Issue Date: 2025-11-15 GPT Summary- OpenMidnightは、Midnight病理基盤モデルを再現・改善したもので、12,000枚の全スライド画像を用いて$1.6Kでトレーニングし、複数のベンチマークで最先端の性能を達成。大規模データなしでもトップパフォーマンスが可能であり、トレーニングパイプライン、コード、モデルの重みを公開して研究を促進する。 Comment
HF: https://huggingface.co/SophontAI/OpenMidnight
元ポストより
> The surprising performance of our model points to the challenges of the pathology FM space.
> Performance doesn't seem to scale with compute or dataset size, and for some benchmarks, really simple baselines perform shockingly well.
> In our mind, this indicates both that current models aren't being trained efficiently, and that the current benchmarks are poor.
まだデータセットサイズや計算量に応じてスケールしているようには見えず、現在のモデルが効率的に学習ができてとらず、かつ現在のベンチマークがモデルの性能を適切に測れていないのでは、といった話が記述されている。興味深い。
SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds, Google DeepMind, 2025.11
Paper/Blog Link My Issue
#Article #NLP #Blog #Reasoning #ComputerUse #VisionLanguageModel #3D (Scene) #Game Issue Date: 2025-11-14 Comment
元ポスト:
もはやAIがゲームをできるのは当たり前の時代だが、どのくらいOODに汎化するのかは気になる。
Holo2: Cost-Efficient Models for Cross-Platform Computer-Use Agents, H Company, 2025.11
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Blog #OpenWeight #ComputerUse #Selected Papers/Blogs #VisionLanguageModel #Grounding #GUI Issue Date: 2025-11-14 Comment
HF: https://huggingface.co/collections/Hcompany/holo2
元ポスト:
関連:
- Holo1.5 - Open Foundation Models for Computer Use Agents, H Company, 2025.09
OlmoEarth-v1-Large, Ai2, 2025.11
Paper/Blog Link My Issue
#Article #NLP #FoundationModel #OpenWeight #2D (Image) Issue Date: 2025-11-06 Comment
元ポスト:
衛星画像で学習されたモデルらしい
Do we still need geometry for Visual Localization and Mapping?, Paul-Edouard Sarlin, 50th Pattern Recognition and Computer Vision Colloquium - CVUT, 2025.10
Paper/Blog Link My Issue
#Article #Tutorial #Slide #ObjectLocalization #Geometric #Mapping Issue Date: 2025-11-04 Comment
元ポスト:
ICCV 2025 Report, Kataoka+, LIMIT.Lab, cvpaper.challenge, Visual Geometry Group (VGG), 2025.10
Paper/Blog Link My Issue
#Article #Survey #Slide #read-later #ICCV Issue Date: 2025-11-01 Comment
元ポスト:
Awesome World Models, Siqiao Huang, 2025.10
Paper/Blog Link My Issue
#Article #Survey #WorldModels Issue Date: 2025-11-01 Comment
元ポスト:
LongCat-Flash-Omni Technical Report, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #SpeechProcessing #OpenWeight #MoE(Mixture-of-Experts) #2D (Image) #UMM #3D (Video) #Omni #audio #text Issue Date: 2025-11-01 Comment
元ポスト:
HF: https://huggingface.co/meituan-longcat/LongCat-Flash-Omni
text, image/video, audioをinputし、audioを生成するomniモデル
Nemotron-VLM-Dataset-v2, Nvidia, 2025.10
Paper/Blog Link My Issue
#Article #NLP #Dataset #VisionLanguageModel Issue Date: 2025-10-29 Comment
元ポスト:
From Egocentric Perception to Embodied Intelligence: Building the World in First Person, Ziwei Liu, 2025.10
Paper/Blog Link My Issue
#Article #Tutorial #ICCV Issue Date: 2025-10-29 Comment
元ポスト:
Multimodal Reasoning for Human-Centric Generative Models, Ziwei Liu, 2025.10
Paper/Blog Link My Issue
#Article #Tutorial #ICCV Issue Date: 2025-10-29 Comment
元ポスト:
Native Multimodal Models: Architecture, Post-Training, and Evaluation, Ziwei Liu, 2025.10
Paper/Blog Link My Issue
#Article #Tutorial #MultiModal #ICCV Issue Date: 2025-10-29 Comment
元ポスト:
Ming-flash-omni-Preview, inclusionAI, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MultiModal #SpeechProcessing #TextToImageGeneration #OpenWeight #AutomaticSpeechRecognition(ASR) #Architecture #MoE(Mixture-of-Experts) #Selected Papers/Blogs #VideoGeneration/Understandings #Editing #TTS #Routing #UMM #Omni #Sparse #ImageSynthesis #Initial Impression Notes Issue Date: 2025-10-28 Comment
元ポスト:
過去一番多くのタグを付与した気がするが、果たして大規模、Omniモーダルかつ、UMMにしたことによる恩恵(=様々なモダリティを統一された空間上に学習させる恩恵)はどの程度あるのだろうか?
アーキテクチャを見ると、モダリティごとに(モダリティ単位でのバイアスがかかった)Routerが用意されexpertにルーティングされるような構造になっている。
OmniモーダルでUMMを大規模にスクラッチから事前学習:
- [Paper Note] ERNIE 5.0 Technical Report, Haifeng Wang+, arXiv'26, 2026.02
LMMs Engine, EvolvingLMMs-Lab, 2025.10
Paper/Blog Link My Issue
#Article #MachineLearning #NLP #MultiModal #Repository #PostTraining #Selected Papers/Blogs #UMM #One-Line Notes Issue Date: 2025-10-27 Comment
元ポスト:
事前学習済みのLLM, VLM, dLM, DiffusionModelなどからUMMを学習できる事後学習フレームワーク。
LigerKernelでメモリ使用量を30%削減し、SparseAttentionもサポートし、Muon Optimizerもサポートしている。
LongCat-Video Techcal Report, Meituan LongCat Team, 2025.10
Paper/Blog Link My Issue
#Article #Transformer #DiffusionModel #TextToImageGeneration #LongContext #VariationalAutoEncoder #OpenWeight #VideoGeneration/Understandings Issue Date: 2025-10-26 Comment
元ポスト:
HF: https://huggingface.co/meituan-longcat/LongCat-Video
公式ポスト:
Supercharge your OCR Pipelines with Open Models, merve+, 2025.10
Paper/Blog Link My Issue
#Article #Survey #NLP #OCR Issue Date: 2025-10-24 Comment
元ポスト:
LightOnOCR-1B: The Case for End-to-End and Efficient Domain-Specific Vision-Language Models for OCR, Taghadouini+, 2025.10
Paper/Blog Link My Issue
#Article #NLP #DocParser #VisionLanguageModel #OCR Issue Date: 2025-10-24 Comment
元ポスト:
olmOCR 2: Unit test rewards for document OCR, Ai2, 2025.10
Paper/Blog Link My Issue
#Article #NLP #Supervised-FineTuning (SFT) #ReinforcementLearning #MultiLingual #Japanese #GRPO #Selected Papers/Blogs #DocParser #VisionLanguageModel #OCR #One-Line Notes Issue Date: 2025-10-23 Comment
元ポスト:
モデル: https://huggingface.co/allenai/olmOCR-2-7B-1025-FP8
Apache2.0ライセンスでSoTA更新。そしてさすがの学習データとコードも公開
テクニカルレポート: https://github.com/allenai/olmocr/blob/main/olmOCR-2-Unit-Test-Rewards-for-Document-OCR.pdf
果たして日本語は…SFT Datasetのtop5にjaはなかったように見える
所見:
demoを試した見たが日本語スライドでも非常に性能が良い
DeepSeekOCRとの比較:
LFM2-VL-3B: A New Efficient Vision-Language for the Edge, LiquidAI, 2025.10
Paper/Blog Link My Issue
#Article #NLP #SmallModel #MultiLingual #OpenWeight #VisionLanguageModel Issue Date: 2025-10-22 Comment
元ポスト:
HF: https://huggingface.co/LiquidAI/LFM2-VL-3B
SigLIP2とLFM2がバックボーン
- Introducing LFM2: The Fastest On-Device Foundation Models on the Market, LiquidAI, 2025.07
dots.ocr, rednote-hilab, 2025.07
Paper/Blog Link My Issue
#Article #NLP #SmallModel #MultiLingual #OpenWeight #DocParser #VisionLanguageModel #OCR Issue Date: 2025-10-22 Comment
100+言語のdots.ocr benchと呼ばれるものでの性能も報告されているが、日本語性能はどのくらいなのだろうか
MIT Licence
参考:VLMを使った多言語ドキュメントパーサ「dots.ocr」を試す, kun432, Zenn
https://zenn.dev/kun432/scraps/b91fce6fbeb30c
日本語もかなりいけてそう
Chandra, datalab-to, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MultiLingual #OpenWeight #DocParser #OCR Issue Date: 2025-10-22 Comment
元ポスト:
SoTA.だったdots.ocrというモデルをoutperformしている模様
40+ languagesをサポート
AI PUBS OpenRAIL-M Modifiedライセンス🤔
https://huggingface.co/datalab-to/chandra/blob/main/LICENSE
dots.ocrはMIT Licence
- dots.ocr, rednote-hilab, 2025.07
DeepSeek-OCR: Contexts Optical Compression, DeepSeek, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MultiLingual #read-later #Selected Papers/Blogs #DocParser #Encoder-Decoder #OCR #Reference Collection #Compression Issue Date: 2025-10-20 Comment
元ポスト:
英語と中国語では使えそうだが、日本語では使えるのだろうか?p.17 Figure11を見ると100言語に対して学習したと書かれているように見える。
所見:
所見:
OCRベンチマーク:
- [Paper Note] OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations, Linke Ouyang+, CVPR'25, 2024.12
(DeepSeek-OCRの主題はOCRの性能向上というわけではないようだが)
所見:
所見+ポイント解説:
所見:
textxをimageとしてエンコードする話は以下の2023年のICLRの研究でもやられているよというポスト:
- [Paper Note] Language Modelling with Pixels, Phillip Rust+, ICLR'23, 2022.07
関連:
- [Paper Note] Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text
Inputs in Multimodal LLMs, Yanhong Li+, arXiv'25, 2025.10
- [Paper Note] PixelWorld: Towards Perceiving Everything as Pixels, Zhiheng Lyu+, arXiv'25, 2025.01
関連:
literature:
上記ポストでは本研究はこれらliteratureを完全に無視し “an initial investigation into the feasibility of compressing long contexts via optical 2D mapping.” と主張しているので、先行研究を認識し引用すべきだと述べられているようだ。
karpathy氏のポスト:
Generative Modeling by Estimating Gradients of the Data Distribution, Yang Song, 2021.05
Paper/Blog Link My Issue
#Article #Tutorial #MachineLearning #DiffusionModel #read-later #ScoreMatching Issue Date: 2025-10-20 Comment
元ポスト:
Find3D: Localizing Semantic Concepts in the 3D Space , Ziqi Ma, 2025.10
Paper/Blog Link My Issue
#Article #Blog #ObjectLocalization #3D (Scene) Issue Date: 2025-10-20 Comment
元ポスト:
画像生成AIにおけるEulerサンプラーの詳細解説, あらもり, 2024.07
Paper/Blog Link My Issue
#Article #DiffusionModel #Blog #Samplers Issue Date: 2025-10-10
Stable Diffusionにおけるサンプラーの役割を理解する, moykeen, 2024.01
Paper/Blog Link My Issue
#Article #DiffusionModel #Blog #Samplers Issue Date: 2025-10-10
Introducing Stable Diffusion 3.5, StabilityAI, 2024.10
Paper/Blog Link My Issue
#Article #Transformer #DiffusionModel #TextToImageGeneration #Blog #OpenWeight #Selected Papers/Blogs Issue Date: 2025-10-10 Comment
SD3.5
Ming-UniVision: Joint Image Understanding and Generation via a Unified Continuous Tokenizer, inclusionAI, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #OpenWeight #UMM Issue Date: 2025-10-03 Comment
HF: https://huggingface.co/inclusionAI/Ming-UniVision-16B-A3B
元ポスト:
Apriel-1.5-15b-Thinker, ServiceNow-AI, 2025.09
Paper/Blog Link My Issue
#Article #NLP #MultiModal #Reasoning #SmallModel #OpenWeight #VisionLanguageModel Issue Date: 2025-10-01 Comment
元ポスト:
Artificial Analysisによるベンチマーキングでは現状<20BでSoTAなReasoningモデルな模様。
MIT License
公式ポスト:
Nvidiaによるポスト:
GLM-4.6: Advanced Agentic, Reasoning and Coding Capabilies, Zhipu AI, 2025.09
Paper/Blog Link My Issue
#Article #NLP #MultiModal #OpenWeight #MoE(Mixture-of-Experts) #read-later #VisionLanguageModel #One-Line Notes Issue Date: 2025-09-30 Comment
元ポスト:
続報:
Artificial Intelligenceによる評価:
OpenWeightモデルの中でトップレベルのベンチスコア
HFにてモデルが公開された模様。ベンチマークのスコアを見て思ったが、106BA12Bのモデルと9Bモデルのスコア差がベンチマークによっては小さいので、場合によってはSLMの方でtest time scacingを効かせた方が、時間的な制約がきつい場合は現実的には高い性能が出るのでは?
InternVL3.5-Flash, OpenGVLab, 2025.09
Paper/Blog Link My Issue
#Article #Reasoning #OpenWeight #VisionLanguageModel Issue Date: 2025-09-29 Comment
元ポスト:
HunyuanImage-3.0, Tencent, 2025.09
Paper/Blog Link My Issue
#Article #NLP #MultiModal #OpenWeight #UMM #One-Line Notes Issue Date: 2025-09-29 Comment
元ポスト:
所見:
テキスト生成+画像理解・生成が可能なUnified Multimodal Models (UMMs)。テキストはtokenizer、画像は生成用エンコーダ、理解用エンコーダを用意してエンコードしDecoder-Only Tranformerに入力。auto-regressiveに生成し、テキストはDe-Tokenizerでテキスト化、画像の場合は専用のDecoderでデコードする。
Qwen-Image-Edit-2509, Qwen Team, 2025.09
Paper/Blog Link My Issue
#Article #NLP #DiffusionModel #VisionLanguageModel #Encoder #Editing Issue Date: 2025-09-24 Comment
テクニカルレポート: https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-Image/Qwen_Image.pdf
Qwen3-VL, Qwen Team, 2025.09
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #VisionLanguageModel Issue Date: 2025-09-23 Comment
元ポスト:
DocVQAのオラクルはラベルノイズと曖昧性の観点から94--95という主張:
Qwen3 VL cookbook:
https://github.com/QwenLM/Qwen3-VL/tree/main/cookbooks
元ポスト:
続報:
MagicBench, ByteDance-Seed, 2025.09
Paper/Blog Link My Issue
#Article #NLP #Dataset #LanguageModel #Evaluation #TextToImageGeneration #UMM Issue Date: 2025-09-19 Comment
元ポスト:
英文と中文両方存在する
Magistral-Small-2509, MistralAI, 2025.09
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MultiModal #Reasoning #OpenWeight #VisionLanguageModel Issue Date: 2025-09-18 Comment
元ポスト:
granite-docling-258M, IBM, 2025.09
Paper/Blog Link My Issue
#Article #NLP #MultiModal #OpenWeight #DocParser #VisionLanguageModel Issue Date: 2025-09-18 Comment
元ポスト:
Apache 2.0, 言語は英語のみ
Holo1.5 - Open Foundation Models for Computer Use Agents, H Company, 2025.09
Paper/Blog Link My Issue
#Article #NLP #Supervised-FineTuning (SFT) #ReinforcementLearning #OpenWeight #ComputerUse #GRPO #VisionLanguageModel #GUI Issue Date: 2025-09-16 Comment
7BのみApache 2.0ライセンス。3BはQwenのライセンスを継承し、72Bはnon-commercialライセンスらしい
モデルカードとブログによると下記モデル群とSonnet 4 よりもComputer Use関連ベンチマーク(GUI上での位置を特定するUI LocalizationとScreen Contentの理解およびQA関連のベンチマーク)で高性能とのこと:
- [Paper Note] UI-Venus Technical Report: Building High-performance UI Agents with RFT, Zhangxuan Gu+, arXiv'25
- [Paper Note] UI-TARS: Pioneering Automated GUI Interaction with Native Agents, Yujia Qin+, arXiv'25, 2025.01
- Qwen2.5-VL-32B-Instruct, Qwen Team, 2025.03
モデルカードによるとopen sourceデータのmixと、合成データ、人手でアノテーションされたデータを用いて、SFT->GRPOによって学習されたとだけ書かれている。
画像モデルのバックボーンとして最初に何を選ぶべきか?, ちくわぶ, 2025.09
Paper/Blog Link My Issue
#Article #Analysis #Blog #Backbone Issue Date: 2025-09-13 Comment
こちらの論文を参考にしている:
- [Paper Note] Battle of the Backbones: A Large-Scale Comparison of Pretrained Models across Computer Vision Tasks, Micah Goldblum+, NeurIPS'23
Backbone選定の際は参照のこと。2024年以後のモデルは含まれていない点に注意。
CLOCKBENCH: VISUAL TIME BENCHMARK WHERE HUMANS BEAT THE CLOCK, LLMS DON’T ALEK SAFAR (OLEG CHICHIGIN), 2025.09
Paper/Blog Link My Issue
#Article #NLP #Dataset #LanguageModel #Evaluation #Contamination-free #VisionLanguageModel Issue Date: 2025-09-07 Comment
リーダーボード: https://clockbench.ai
元ポスト:
様々な種類の時計(e.g., 反転、フォントの違い, invalidな時刻の存在, 大きさ, フォーマットなど; p.2参照のこと)の時刻を読み取り(あるいはvalidな時刻か否かを判定し)、読み取った時刻に対してQA(e.g., X時間Y分Z秒進める、戻した時刻は?長針を30/60/90度動かした時刻は?この時刻がニューヨークの時間だとしたらロンドンの時刻は?)を実施するベンチマーク。人間の正解率は89.1%に対してSoTAモデルでも13.3%程度。contaminationに配慮して全てスクラッチから作成され、全体の評価データはprivateなままにしているとのこと。
続報:
Qwen3-VL-235B-InstructがGPT-5 Chat超え
FineVision: Open Data Is All You Need, Wiedmann+, Hugging Face, 2025.09
Paper/Blog Link My Issue
#Article #Pretraining #NLP #Dataset #Blog #Selected Papers/Blogs #VisionLanguageModel Issue Date: 2025-09-05 Comment
HF: https://huggingface.co/datasets/HuggingFaceM4/FineVision
元ポスト:
【論文解説】高速・高品質な生成を実現するFlow Map Models(Part 1: 概要編), Masato Ishii (Sony AI), 2025.09
Paper/Blog Link My Issue
#Article #Tutorial #MachineLearning #Video #read-later Issue Date: 2025-09-04
