NLP (4152) — 18/21
Ling-2.5-1T, inclusionAI, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #Reasoning #OpenWeight #MoE(Mixture-of-Experts) Issue Date: 2026-02-16 Comment
Ringに続いてLingもリリース
関連:
- Ring-1T-2.5-FP8, inclusionAI, 2026.02
元ポスト:
AI 101: "On-Policy Distillation Zeitgeist", Turing Post, 2026.02
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel #ReinforcementLearning #Blog #PostTraining #On-Policy #One-Line Notes #SelfDistillation Issue Date: 2026-02-16 Comment
元ポスト:
最近よくみかける on-policy self-distillationに関する解説
QED-Nano: Teaching a Tiny Model to Prove Hard Theorems, LM Provers Team, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #Supervised-FineTuning (SFT) #ReinforcementLearning #Blog #Mathematics #SmallModel #PostTraining #Proofs #Rubric-based #Initial Impression Notes Issue Date: 2026-02-16 Comment
元ポスト:
ポイント解説:
早くもReasoning Cacheが利用されている:
- [Paper Note] Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL, Ian Wu+, arXiv'26, 2026.02
4B級のモデルで特定タスクに特化したモデルを作りたい場合に非常に役立ちそうなレシピ
Building Olmo in the Era of Agents, Nathan Lambert, LTI Colloquim, 2026.02
Paper/Blog Link My Issue
#Article #Tutorial #Survey #LanguageModel #AIAgents #Reasoning #Slide #OpenSource #read-later #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-02-16 Comment
元ポスト:
うーんこれは時間をとってしっかり読んで色々まとめたい・・・
[Paper Notes] Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity, Bytedance Seed, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #LanguageModel #AIAgents #Reasoning #Proprietary #VisionLanguageModel Issue Date: 2026-02-16 Comment
元ポスト:
所見:
GPT‑5.2 derives a new result in theoretical physics, OpenAI, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Blog #ScientificDiscovery #Physics #Human-in-the-Loop Issue Date: 2026-02-14 Comment
元ポスト:
The Simulation Company, Simile, 2026.02
Paper/Blog Link My Issue
#Article #MachineLearning #FoundationModel #Post #WorldModels #Initial Impression Notes Issue Date: 2026-02-13 Comment
やはり次のFoundation Modelsの軸としてWorld Modelsやシミュレーションが注目されているように感じる。実際、シミュレーションによって様々なデータが合成できれば現在の基盤モデルをさらに引き上げると思われる。
関連:
Karpathy氏のポスト:
続報:
Introducing GPT‑5.3‑Codex‑Spark: An ultra-fast model for real-time coding in Codex, OpenAI, 2026.02
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #LanguageModel #AIAgents #Blog #Coding #SoftwareEngineering Issue Date: 2026-02-13 Comment
元ポスト:
所見:
Gemini 3 Deep Think: Advancing science, research and engineering, Google, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Coding #Reasoning #Mathematics #Proprietary #SoftwareEngineering #VisionLanguageModel #Science Issue Date: 2026-02-13 Comment
まずはUltra Subscriberに公開し、その後徐々にAPIアクセスを解禁していくとのこと。
LiveCodeBench:
MiniMax M2.5: SOTA in Coding and Agent, designed for Agent Universe, MiniMax, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Blog #Coding #OpenWeight #SoftwareEngineering #Selected Papers/Blogs Issue Date: 2026-02-13 Comment
元ポスト:
OsenHands IndexでClaude Sonnet 4.5超えの初めてのOpenWeightモデル:
コストパフォーマンスにおいては、低コストなモデル群の中では抜きん出た性能
まだHF上にWeightは公開されていないようだが後ほど公開されると思われる。
所見:
weightが公開:
https://huggingface.co/MiniMaxAI/MiniMax-M2.5
元ポスト:
UnslothがGGUF版を公開:
A2A: The Agent2Agent Protocol, DeepLearning.AI, 2026.02
Paper/Blog Link My Issue
#Article #Multi #Tutorial #LanguageModel #AIAgents #Video #SoftwareEngineering #A2A Issue Date: 2026-02-13 Comment
元ポスト:
元ポスト:
Ring-1T-2.5-FP8, inclusionAI, 2026.02
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #LanguageModel #AIAgents #Attention #Reasoning #LongContext #OpenWeight #LongHorizon #LinearAttention Issue Date: 2026-02-12 Comment
元ポスト:
関連:
- Ring-1T, inclusionAI, 2025.10
MLA + lightning linear attentionのハイブリッド
- MHA vs MQA vs GQA vs MLA, Zain ul Abideen, 2024.07
- [Paper Note] Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention, Zhen Qin+, ICML'24, 2024.05
Harness engineering: leveraging Codex in an agent-first world, Ryan Lopopolo, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #GenerativeAI #Blog #Coding #SoftwareEngineering #One-Line Notes Issue Date: 2026-02-12 Comment
OpenAI社内でのコードを1行も人間が書かないで製品をリリースする取り組みに関する詳細なレポートのようである。初期の設計などで想像以上に時間がかかってしまった点(これはCodexの能力の問題ではない)や、実装を続ける中で品質に責任を持つ人間の能力(というより時間)がボトルネックになっていったため、極力Codexが自律的に品質管理ができるような実行・検証環境を用意することで負担を低減した話や、Codexに膨大なマニュアルを読ませて処理をさせるのではなく、どこにどのような情報が格納されているのかといったマップ(目次)を与えることがコンテキストエンジニアリング上重要だったことなどを通じてエージェントにとってリポジトリ全体の可読性を高めることが重要だったといった話や、プロジェクトの期間が長引くにつれて、リポジトリ内に共有されていないcontextが増大していき、それらをリポジトリに統合する作業が生じるなどの課題も生じたといったような話など色々と書かれている。
microgpt.py, Andrej Karpathy, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #python #Selected Papers/Blogs #MinimalCode Issue Date: 2026-02-12 Comment
元ポスト:
[Paper Note] Accelerating Mathematical and Scientific Discovery with Gemini Deep Think, Google DeepMin, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Blog #Mathematics #ScientificDiscovery #Test-Time Scaling #read-later #KeyPoint Notes #Physics #Human-in-the-Loop Issue Date: 2026-02-12 Comment
元ポスト:
- 数学について
- verifierを通じて解の修正と再生成を繰り返すが、問題が解けないことを認めることで(無駄な修正・再生成を減らすことで)効率を大幅に改善
- 博士課程レベル・オリンピックレベルを超えてもtest-time scalingが継続する
- 検索を融合することで既存文献を取り入れ正確性向上
- 完全自動で出版できるレベルの研究を実施可能なところまできている(level0--5のlevel2)
- コンピュータサイエンス・物理学について
- ネットワーク側で広範な解空間を探索してlong-trailな解も捉え推論に組み込むことが可能で、自動的なverificationと人間によるverificationを通じてoutputを生成する
- たとえば10年間未解決だったオンライン列モジュラ最適化と呼ばれる問題や、モデル学習時のノイズ除去による理論的な証明などを実施できている
論文:
- [Paper Note] Towards Autonomous Mathematics Research, Tony Feng+, arXiv'26, 2026.02
[Paper Note] Position: Humans are Missing from AI Coding Agent Research, Wang+, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #UserBased #AIAgents #Coding #read-later #Selected Papers/Blogs #interactive #One-Line Notes #Initial Impression Notes Issue Date: 2026-02-12 Comment
# Authors
Zora Zhiruo Wang, John Yang, Kilian Lieret, Alexa Tartaglini, Valerie Chen, Yuxiang Wei,
Zijian Wang, Lingming Zhang, Karthik Narasimhan, Ludwig Schmidt, Graham Neubig, Daniel Fried, Diyi Yang
元ポスト:
現在のコーディングエージェントは自動的にタスクを完了させ、難易度の高いベンチマークを解けることが実用的な価値とみなされているが、今後より実用的な価値を高めプロダクト化するためには単独でタスクをこなすのではなく、人間開発者やユーザとの相互作用をするような枠組みが次のブレイクスルーとなりうるというposition。非常に共感できる。
Ming-flash-omni-2.0, inclusionAI, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #Transformer #MultiModal #SpeechProcessing #DiffusionModel #Speech #OpenWeight #MoE(Mixture-of-Experts) #2D (Image) #Omni #text Issue Date: 2026-02-12 Comment
元ポスト:
関連:
- Ming-flash-omni-Preview, inclusionAI, 2025.10
- [Paper Note] Ming-Omni: A Unified Multimodal Model for Perception and Generation, Inclusion AI+, arXiv'25, 2025.06
公式ポスト:
GLM-5: From Vibe Coding to Agentic Engineering, Z.ai, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #OpenWeight #MoE(Mixture-of-Experts) #Selected Papers/Blogs #KeyPoint Notes #Reference Collection #LongHorizon #SparseAttention Issue Date: 2026-02-12 Comment
関連:
- GLM-4.7: Advancing the Coding Capability, Z.ai, 2025.12
GLMシリーズの最新モデルGLM-5がリリースされた
元ポスト:
- DeepSeek Sparse Attentionを採用:
- DeepSeek-V3.2-Exp: Boosting Long-Context Efficiency with DeepSeek Sparse Attention, DeepSeek-AI, 2025.09
- [Paper Note] DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models, DeepSeek-AI+, arXiv'25, 2025.12
- 事前学習データを23Tから28.5Tトークンへ
- パラメータ数は4.5の355B-A32から744B-A40Bへ
- RLのインフラとして4.5から引き続きSlimeを採用
- slime, THUDM & Zhihu, 2025.09
- long-horizonなタスクに秀でており、reasoning, coding, agenticタスクにおける各種ベンチマークでOpus 4.5, GPT-5.2, Gemini 3 Proと同等程度の性能
FP8版も公開されている模様(Hopper以後のアーキテクチャでないとサポートされていない点に注意
所見:
元ポスト:
unslothがGGUF版をすでにリリースしている模様。早い:
https://unsloth.ai/docs/models/glm-5
アーキテクチャ解説:
アーキテクチャ解説:
所見:
ENGRAM, EvolvingLMMs-Lab, 2026.02
Paper/Blog Link My Issue
#Article #Tools #LanguageModel #AIAgents #Privacy #MCP #memory #One-Line Notes Issue Date: 2026-02-12 Comment
元ポスト:
MCPに対応しているAI Agentであれば互換性がある暗号化されたストレージの実装なようで、サードパーティのストレージにデータを預けなくてもローカルのストレージでLLMに対して知識を提供可能な模様。
最近DeepSeekが提案したEngramとは異なるので注意:
- [Paper Note] Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models, Xin Cheng+, arXiv'26, 2026.01
Introducing Lab: The Full-Stack Platform for Training your Own Models, Prime Intellect, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #MachineLearning #LanguageModel #Infrastructure #ReinforcementLearning #AIAgents #Blog #ScientificDiscovery #PostTraining #Selected Papers/Blogs #One-Line Notes #Reference Collection #Environment Issue Date: 2026-02-11 Comment
元ポスト:
事後学習、特にAgenticな研究の民主化のためのプラットフォームの提供
所見:
利用例 (Environment Hub):
Sabotage Risk Report: Claude Opus 4.6, Anthropic, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Proprietary #Safety #read-later #Sabotage Issue Date: 2026-02-11 Comment
元ポスト:
[Paper Note] OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis, Li+, 2026.02
Paper/Blog Link My Issue
#Article #InformationRetrieval #Search #LanguageModel #Supervised-FineTuning (SFT) #AIAgents #SyntheticData #OpenSource #Selected Papers/Blogs #Reproducibility #DeepResearch #One-Line Notes #LongHorizon #Initial Impression Notes #Environment Issue Date: 2026-02-10 Comment
元ポスト:
APIに依存せずオフラインコーパスと検索を利用し、高品質なDeepResearchのlong horizonなtrajectoryを合成可能な環境を構築。合成したtrajectoryでNemotron-3-nano-30B-A3B-BaseをSFTすることで、Kimi-K2, GLM-4.6などの10倍以上大きいサイズのモデルよりもBrowseCompで高い性能を獲得。同サイズのTongyiDeepResearchもoutperform。
Deterministicなプロセスで、オフラインコーパスからデータを合成し外部APIに依存しないため完全に再現性があり、かつAPIのコストやrate limitにも引っかからないという利点がある。検索エンジン、コード、データ、合成データ、モデル、全てを公開。
完全に再現性のある研究は素晴らしい。
Composer 1.5 のご紹介, Cursor Team, 2026.02
Paper/Blog Link My Issue
#Article #ReinforcementLearning #AIAgents #GenerativeAI #Blog #Coding #SoftwareEngineering #PostTraining #One-Line Notes #Scalability Issue Date: 2026-02-10 Comment
事前学習モデルに対して、RLをさらにスケールさせることで性能が継続的に向上し、自己要約能力も備えさせることでcontext windowの問題に対処しているとのこと。
(関連)Composer: 強化学習で構築する高速フロンティアモデル:
https://cursor.com/ja/blog/composer
new-datasets-in-machine-learning, librarian-bots, 2026.02
Paper/Blog Link My Issue
#Article #RecommenderSystems #Survey #ComputerVision #MachineLearning #InformationRetrieval #Dataset #Evaluation #SpeechProcessing #Robotics #Live Issue Date: 2026-02-09 Comment
元ポスト:
ModernBERTをFinetuningした分類器を用いてデータセットやベンチマークを提案している研究を自動分類して検索できるようにしている。有用
Context-Bench: A benchmark for agentic context engineering, Letta Research, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Evaluation #Blog #ContextEngineering Issue Date: 2026-02-09 Comment
元ポスト:
Knowledge Editing for LLMs Papers, zjunlp, 2024.07
Paper/Blog Link My Issue
#Article #Survey #LanguageModel #KnowledgeEditing Issue Date: 2026-02-08
Introducing GPT-5.3-Codex: Expanding Codex across the full spectrum of professional work on a computer, OpenAI, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Coding #Proprietary #SoftwareEngineering #Selected Papers/Blogs #Reference Collection Issue Date: 2026-02-06 Comment
元ポスト:
terminal bench 2.0でOpus 4.6超え:
所見:
Advancing finance with Claude Opus 4.6, Anthropic, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Financial #Proprietary #SoftwareEngineering #Selected Papers/Blogs #One-Line Notes #Reference Collection Issue Date: 2026-02-06 Comment
元ポスト:
全体的に能力が向上しているが、ターミナルでのコーディング、BrowseComp(Agentic search), HLE, Financial Analysis, GDPValにおけるOffice Task, Novel Problem Solvingの能力が大きく向上しているように見える。
Context Windowが1Mとのことで素晴らしい
OpenHands Indexでトップとのことだが、Codex 5.3との比較はまだの模様:
50% time horizonが脅威の14.5時間:
Intern-S1-Pro, internlm, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #MultiModal #Reasoning #PositionalEncoding #OpenWeight #MoE(Mixture-of-Experts) #VisionLanguageModel #Science Issue Date: 2026-02-05 Comment
元ポスト:
ポイント解説:
関連:
- [Paper Note] Intern-S1: A Scientific Multimodal Foundation Model, Lei Bai+, arXiv'25, 2025.08
Fourier Position Encoding (FoPE) + upgraded time-series modeling
MiniCPM-o-4_5, OpenBMB, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #LanguageModel #SpeechProcessing #DiffusionModel #OpenWeight #AutomaticSpeechRecognition(ASR) #VisionLanguageModel #TTS #Omni #AudioLanguageModel Issue Date: 2026-02-05 Comment
元ポスト:
[Paper Note] THE MILLION-LABEL NER: BREAKING SCALE BARRIERS WITH GLINER BI-ENCODER, Stepanov+, 2026.02
Paper/Blog Link My Issue
#Article #NeuralNetwork #EfficiencyImprovement #Encoder #NER #EntityLinking Issue Date: 2026-02-05 Comment
元ポスト:
The Second Pre-training Paradigm, Jim Fan, X, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #Pretraining #LanguageModel #MultiModal #Post #Robotics #WorldModels #One-Line Notes Issue Date: 2026-02-05 Comment
事前学習がnext word predictionから過去の行動と状態によって条件付けられ次の(ある期間の)世界の状態を予測するワールドモデリング(next physical state prediction)へのパラダイムシフトの予想(というよりこのパラダイムシフトの真っ只中にいる)。人間の脳が処理する情報の多くは視覚であり、言語的な領域は部分的なことであることや、猿は言語的な能力が低くても視覚や運動、触覚などの感覚的情報から世界の物理法則を理解し知的なアクションをとるメンタルモデルを確立していることなどを引き合いに説明している。
Time Horizon 1.1, METR, 2026.01
Paper/Blog Link My Issue
#Article #Metrics #LanguageModel #AIAgents #Evaluation #Scaling Laws #Selected Papers/Blogs Issue Date: 2026-02-05 Comment
元ポスト:
続報:
関連:
- [Paper Note] Measuring AI Ability to Complete Long Tasks, Thomas Kwa+, arXiv'25, 2025.03
Fine-tuning open LLM judges to outperform GPT-5.2, together.ai, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #Evaluation #Blog #LLM-as-a-Judge #DPO #RewardModel #One-Line Notes #Initial Impression Notes Issue Date: 2026-02-05 Comment
元ポスト:
Reward Bench 2:
- [Paper Note] RewardBench 2: Advancing Reward Model Evaluation, Saumya Malik+, arXiv'25, 2025.06
LLMでLLMを評価するというパラドックスに違和感はあるが、一般論として、「生成」するよりも「検証」することがモデルにとって簡単なタスクであるためうまくいきます(LLM-as-a-Judge)、といった説明が書いてあり、数千程度のサンプルでOpenLLMをDPOすることによって、GPT-5.2のようなFrontierモデルをReward Benchで上回ることができた、といった話が書かれている。
ただし、上記Reward Bench 2研究で示されている通り、**Reward Benchでの性能が高いReward Modelだからといって、必ずしもRLによって下流タスクの性能が向上するとは限らない点には注意**であり、元論文に従うとBest-of-Nサンプリングのようなtest-time-scalingのパラダイムとして利用するのが現在の実務上は良さそうである。
Together Evaluations now supports comparing top commercial APIs vs. open source models, together.ai, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #Evaluation #Blog #PEFT(Adaptor/LoRA) #PostTraining #One-Line Notes Issue Date: 2026-02-05 Comment
元ポスト:
OpenLLMのFinetuningをサポートしているプラットフォームにおいて、データセットをアップロードすると
- Prompt optimization (GEPA)
- Fine-tuning (PEFT + full finetuning)
の両方を実施し、コスト-性能のパレート最適なポイントを評価し、かつGPT等とのProprietaryモデルとの比較もした評価もできるようになりました、といった話の紹介。
GEPA:
- [Paper Note] GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning, Lakshya A Agrawal+, ICLR'26, 2025.07
Finetuningがサポートされているモデル群:
-
https://docs.together.ai/docs/fine-tuning-models
New Holo2 model takes the lead in UI Localization, H Company, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #AIAgents #Blog #OpenWeight #ComputerUse #Selected Papers/Blogs #VisionLanguageModel #Grounding #GUI Issue Date: 2026-02-05 Comment
HF: https://huggingface.co/Hcompany/Holo2-235B-A22B
元ポスト:
関連:
- Holo1.5 - Open Foundation Models for Computer Use Agents, H Company, 2025.09
Qwen3-Coder-Next: Pushing Small Hybrid Models on Agentic Coding, QwenTeam, 2026.02
Paper/Blog Link My Issue
#Article #LanguageModel #Attention #Blog #Coding #LongContext #SmallModel #MoE(Mixture-of-Experts) #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-02-04 Comment
HF: https://huggingface.co/collections/Qwen/qwen3-coder-next?spm=a2ty_o06.30285417.0.0.3bdec921Ja5TZI
元ポスト:
A3BでSWE Bench ProにおいてClaude Sonnet 4.5超え
関連:
- [Paper Note] Gated Delta Networks: Improving Mamba2 with Delta Rule, Songlin Yang+, ICLR'25, 2024.12
開発者の方のポスト:
int4 model from Cerebras:
https://huggingface.co/Intel/Qwen3-Coder-Next-int4-AutoRound
元ポスト:
Latest open artifacts (#18): Arcee's 400B MoE, LiquidAI's underrated 1B model, new Kimi, and anticipation of a busy month, Interconnects, 2026.02
Paper/Blog Link My Issue
#Article #Analysis #LanguageModel #Blog #OpenWeight Issue Date: 2026-02-03 Comment
paid userしか全文は閲覧できない
元ポスト:
Moltbook is the most interesting place on the internet right now, Simon Willisons's blog, 2026.01
Paper/Blog Link My Issue
#Article #Multi #LanguageModel #AIAgents #GenerativeAI #Blog #Conversation #Selected Papers/Blogs #Reference Collection Issue Date: 2026-02-01 Comment
元ポスト:
興味深い:
話したことのないhumanとの会話をあたかもあったことのように話し始める:
所見:
Andrej Karpathy氏もエージェントを参加させたようである:
所見:
Introducing the OpenHands Index, OpenHands, 2026.01
Paper/Blog Link My Issue
#Article #Analysis #LanguageModel #AIAgents #Evaluation #Blog #SoftwareEngineering #Selected Papers/Blogs #KeyPoint Notes Issue Date: 2026-01-30 Comment
元ポスト:
SWE Bench(pythonプログラムリポジトリに対するissueを解決するタスク)がSWE関連の代表的なベンチマークだがこれらはソフトウェアエンジニアリングのサブタスクの一つしか反映しておらず、より多くのタスクの解決能力でSWE Agentの能力を評価し、かつコストの軸でも評価をしてどのモデルがパレート最適なものなのかを見つけられるようなindexを作って評価しました、という話に見える。
タスクとしては以下の5つをピックしているとのこと:
> 1. Issue Resolution
> 2. Frontend Development
> 3. Greenfield Development
> 4. Software Testing
> 5. Information Gathering
これらのタスクを総合的に評価するとClaude 4.5 Opusが最も性能が高くコストも高い。次点でGPT-5.2-Codexという結果。またコストが最も安く平均的な性能が高いモデルとしてはDeepSeekV3.2-Reasonerとなった。また、特定のタスク、たとえばGreenfield developmentではGPT-5.2-Codexの性能が抜きん出ているなど、個別のタスクで見るとモデル間の優劣がはっきりと見えるような結果になっている。
以下のモデルが追加:
Claude 4.6 Opus
GPT 5.2 Codex
Kimi K2.5
GLM-4.7
MiniMax M2.5
Project Genie: Experimenting with infinite, interactive worlds, Google Deepmind, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #GenerativeAI #Proprietary #WorldModels #interactive Issue Date: 2026-01-30 Comment
元ポスト:
Googleからのworld model
PLaMo 2.2 Primeをリリースしました, PFN, 2026.01
Paper/Blog Link My Issue
#Article #Multi #LanguageModel #Supervised-FineTuning (SFT) #Proprietary #Japanese #DPO #PostTraining #InstructionFollowingCapability #Medical #RolePlaying Issue Date: 2026-01-29 Comment
関連:
- [Paper Note] Generalizing Verifiable Instruction Following, Valentina Pyatkin+, NeurIPS'25, 2025.07
- JFBench: 実務レベルの日本語指示追従性能を備えた生成AIを目指して, PFN, 2026.01
non-thinkingモデルである点に注意
JFBench: 実務レベルの日本語指示追従性能を備えた生成AIを目指して, PFN, 2026.01
Paper/Blog Link My Issue
#Article #Dataset #LanguageModel #InstructionTuning #Evaluation #Japanese #InstructionFollowingCapability Issue Date: 2026-01-29 Comment
元ポスト:
Trinity Large, Arcee, 2026.01
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #Pretraining #LanguageModel #OpenWeight #MoE(Mixture-of-Experts) #read-later #Selected Papers/Blogs #Stability #One-Line Notes #Reference Collection #Sparse #Initial Impression Notes Issue Date: 2026-01-29 Comment
テクニカルレポート:
https://github.com/arcee-ai/trinity-large-tech-report/
HF:
https://huggingface.co/arcee-ai
GLM4.7やDeepSeekV3と比較してスループットやTTFTが二倍以上。
非常にsparseなMoE(400B-A13B, 4/256のexpertsにルーティング)であるため学習を安定させるためにDense layerを増やし、モメンタムを考慮したexpertのバランシングや、z-lossと呼ばれるlogitのスケールをコントロールするような手法を導入することで安定した学習を実現。2048 Nvidia B300 GPUsで、17Tトークンの事前学習33日で完了
元ポスト:
これほどsparseなMoEをここまで安定させて学習できるのは非常に興味深いと思われる。
インタビュー:
やると決めてチームビルディングも含めて非常に短期間(6ヶ月)で達成したとのことだが、気になる。
解説:
所見(風刺):
ポイント解説:
アーキテクチャ解説:
Accelerating Diffusion Models with an Open, Plug-and-Play Offering, Nvidia, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #EfficiencyImprovement #Tools #DiffusionModel #TextToImageGeneration #Distillation #PostTraining #2D (Image) #Editing #3D (Video) #TextToVideoGeneration #ImageToTextGeneration #TrainingFramework Issue Date: 2026-01-29 Comment
元ポスト:
self forcingも実装されている
- [Paper Note] Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion, Xun Huang+, NeurIPS'25
Introducing Agentic Vision in Gemini 3 Flash, Google Deepmind, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #AIAgents #Proprietary #VisionLanguageModel #One-Line Notes Issue Date: 2026-01-29 Comment
元ポスト:
visual reasoningとコード実行の融合
Introducing Prism, OpenAI, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #ChatGPT #GenerativeAI #MultiModal #AcademicWriting #DeepResearch #One-Line Notes Issue Date: 2026-01-29 Comment
デモを見るとdraftをベースに関連研究をdeepresearchしてワンクリックでbibtexにexport, ホワイトボードに描いた図をドラッグ&ドロップして論文に反映などしている。Overleafの競合。
元ポスト:
所見:
Open Coding Agents: Fast, accessible coding agents that adapt to any repo, Ai2, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Coding #SoftwareEngineering #read-later Issue Date: 2026-01-29 Comment
開発者の方のブログ:
https://timdettmers.com/2026/01/27/building-open-coding-agent-sera/
HF:
https://huggingface.co/collections/allenai/open-coding-agents
14Bモデルリリース:
DeepSeek-OCR-2, DeepSeek-AI, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #OCR #Compression Issue Date: 2026-01-27 Comment
元ポスト:
関連:
- DeepSeek-OCR: Contexts Optical Compression, DeepSeek, 2025.10
A few random notes from claude coding quite a bit last few weeks., Andrej Karpathy, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Coding #Post #SoftwareEngineering Issue Date: 2026-01-27
Minimax Agent, Minimax, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #LanguageModel #AIAgents #GenerativeAI #ComputerUse Issue Date: 2026-01-27 Comment
code: https://github.com/MiniMax-AI/Mini-Agent
元ポスト:
Continual Learning with RL for LLMs, CAMERON R. WOLFE, PH.D., 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #ReinforcementLearning #Blog #PostTraining Issue Date: 2026-01-26 Comment
元ポスト:
RLHF Book - Code Examples, Nathan Lambert, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #ReinforcementLearning #Repository #PostTraining #Selected Papers/Blogs #MinimalCode #Initial Impression Notes Issue Date: 2026-01-26 Comment
元ポスト:
Qwen 1.7Bモデルでの様々なRLアルゴリズムでのミニマルコード集。学習曲線つきで非常に実用的
A well known important feature to stabilize RL training is implementing the LM head in fp32 precision to help with gradients ... , Nathan Lambert, X, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #ReinforcementLearning #Post #PostTraining #Stability #One-Line Notes Issue Date: 2026-01-24 Comment
関連:
- MiniMax-M1, MiniMax, 2025.06
- [Paper Note] MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning
Attention, MiniMax+, arXiv'25, 2025.06
RLを安定化するためのtipsとそれによりMiniMax M1のplotが再現できたという話な模様。RLはこういった細かいテクニックが大事だと思うので、共有して頂けるのは大変ありがたい。
関連:
- [Paper Note] Defeating the Training-Inference Mismatch via FP16, Penghui Qi+, arXiv'25, 2025.10
- train-inference-gap && ReinforcementLearning ラベルが紐づいたissueも参照のこと
Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigations, Anthropic, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #Alignment #Evaluation #Blog #read-later Issue Date: 2026-01-23 Comment
元ポスト:
eval awareness mitigation
Composing Weight and Data Sparsity in MoE: Improving compute efficiency through varying compute per token, Perceptron, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #Pretraining #MultiModal #MoE(Mixture-of-Experts) #read-later #VisionLanguageModel #Routing #Sparse #Initial Impression Notes Issue Date: 2026-01-23 Comment
元ポスト:
MoEがトークン単位でactivateするweightをサブセットにするweight sparcityによって効率化を実現する手法とみなしたときに、それぞれのinputに情報量の濃淡があることから現在のトークンごとにweightを割り当てるのではなく、weightごとにトークンを割り当てるというもう一つの軸を考えることができ(=Data Sparcity)、これをweightごとにトークンのsubsetしか持たないような実現方法をとるとcontextが損なわれauto-regressiveの前提が崩れるためtrain-inference-mismatchが生じるので、null experts(受け取ったトークンに対して何もしない)を実装して実現するみたいな話のように見えるが全くまだ読めていない。
Waypoint-1: Real-time Interactive Video Diffusion from Overworld, Overworld, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #Controllable #Transformer #MultiModal #DiffusionModel #OpenWeight #WorldModels #interactive #3D (Video) #One-Line Notes #RectifiedFlow #Realtime Issue Date: 2026-01-22 Comment
blog:
https://over.world/blog/the-path-to-real-time-worlds-and-why-it-matters
pj page:
https://over.world/
元ポスト:
リアルタイムにzero latencyでマウス(カメラも自由に動かせる)、キーボード、テキストでinteraction可能なworld model
Claude's new constitution, Anthropic, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #Blog #Safety #One-Line Notes Issue Date: 2026-01-22 Comment
ClaudeのAI Modelで利用される新たなConstitution
関連:
- [Paper Note] Constitutional AI: Harmlessness from AI Feedback, Yuntao Bai+, arXiv'22
元ポスト:
IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMs, Cheng+, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #ReinforcementLearning #Blog #PostTraining #KeyPoint Notes #Scalability Issue Date: 2026-01-22 Comment
元ポスト:
RLにおけるロールアウト数nのスケーリングは、シグモイド関数のような形状になりどこかのポイントで明確にサチるポイントが存在し、それ以上増やしても少量のゲインしか得られないポイントが存在する。これらのトレンドはeasy/hardな問題の双方で共通して見出されるが、原因は大きく異なっており、nを大きくするとeasyな問題ではworst@kが改善し、hardな問題ではbest@kが改善することで性能が向上する。つまり、簡単な問題に対してはより安定して正解できてミスが減り、困難な問題に対しては探索空間が広がり1回でも正解できる可能性が高まる。また、また、ハードウェア制約によりバッチサイズは基本的に固定されるので、ロールアウト数nと1バッチあたりに含められる問題数はトレードオフの関係となる。
このロールアウト数nに関する性質は、異なるベースモデル間で共通して生じるが、サチるポイントが異なる。問題セットのサイズで見ると、サイズが小さいと早々にoverfitするためサチるnのポイントも早くなる。問題難易度の分布がmixしているものであればnによるスケーリングのトレンドは維持されるが、評価する際のmetricsによってサチるぽいんとが左右される。nのスケーリングはdownstreamタスクの性能も向上させる。
と言った話らしい。
Fantastic Pretraining Optimizers and Where to Find Them 2.1: Hyperball Optimization, Wen+, 2026.01
Paper/Blog Link My Issue
#Article #NeuralNetwork #EfficiencyImprovement #Pretraining #LanguageModel #Optimizer #read-later #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-01-22 Comment
元ポスト:
シンプルな手法で、先行研究によってモデルのパラメータサイズやデータのスケールが大きくなるとMuonのような行列ベースのoptimiserの高速化の恩恵が小さくなる現象を改善しているとのこと。
具体的には、重みを更新する際にweight decayのようなソフトにweightのノルムをコントロールするような仕組みを入れるのではなく、optimiserの重みに対する更新量と、更新後のネットワークの重みをフロベニウスノルムで正規化し、最適化の軌跡を半径Rの超球面の表面上に位置するように明示的に制約する(ここで、Rは最初の重み行列のフロベニウスノルム)。Muonを含む様々なoptimiserでも機能して学習効率を高めるため、インパクトの大きな重要研究に見える。
関連(concurrent works):
- [Paper Note] Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models, Yonggan Fu+, arXiv'25, 2025.11
- [Paper Note] Controlled LLM Training on Spectral Sphere, Tian Xie+, arXiv'26, 2026.01
関連:
- [Paper Note] Fantastic Pretraining Optimizers and Where to Find Them, Kaiyue Wen+, ICLR'26, 2025.09
ICLR 2026 Acceptance Prediction: Benchmarking Decision Process with A Multi-Agent System, Zhang+, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #Dataset #LanguageModel #AIAgents #Evaluation #MultiModal #ScientificDiscovery #VisionLanguageModel #AcademicWriting #Live #One-Line Notes Issue Date: 2026-01-20 Comment
元ポスト:
conference paperのpeer reviewに関するベンチマーク。accept/rejectを予測する。papers, reviews, rebuttalsそしてfinal decisionsが紐づけられている。
GLM-4.7-Flash, Z.ai, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Coding #OpenWeight #MoE(Mixture-of-Experts) #One-Line Notes Issue Date: 2026-01-20 Comment
元ポスト:
関連:
- GLM-4.7: Advancing the Coding Capability, Z.ai, 2025.12
30B-A3BのMoEモデルで、gpt-oss-20B, Qwen3-30B-A3B-Thinking-2507を、SWE Bench Verified, tau2_bench, BrowseComp(SWEタスク, tooluse, 検索)等で大幅にoutperform。AIME, GPQA, HLEなどの推論系のベンチマークも同等以上。つまり、agenticなタスクに適した能力を有することが示唆される。
ポイント解説:
10,924x: The Instability Bomb at 1.7B Scale, TayKolasinski, 2026.01
Paper/Blog Link My Issue
#Article #Tutorial #MachineLearning #LanguageModel #Blog #Selected Papers/Blogs #Reproducibility #ResidualStream Issue Date: 2026-01-19 Comment
元ポスト:
関連:
- [Paper Note] mHC: Manifold-Constrained Hyper-Connections, Zhenda Xie+, arXiv'25, 2025.12
- [Paper Note] Hyper-Connections, Defa Zhu+, ICLR'25, 2024.09
part1:
https://taylorkolasinski.com/notes/mhc-reproduction/
HC, mHCの説明が美しい図解と数式で説明されている。分かりやすい!
HCの課題とmHCがどのように解決したかを数式的、直感的に理解でき非常に有用
Pocket Flow: 100-line LLM framework. Let Agents build Agents, The-Rocket, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #Library #AIAgents #python #SoftwareEngineering #read-later #Selected Papers/Blogs #MinimalCode #Initial Impression Notes Issue Date: 2026-01-19 Comment
元ポスト:
たったの100行で実現されるミニマルなAI Agent/LLMフレームワークで、9種類の抽象化(Node, Flow, Shared, ...)でchat, agent, workflow, RAG, MCP, A2Aなどの様々なLLMをベースとした機能を実装できるフレームワークな模様。コード読みたい
Context Rot: How Increasing Input Tokens Impacts LLM Performance, CHROMA TECHNICAL REPORT, 2025.07
Paper/Blog Link My Issue
#Article #LanguageModel #Blog #LongContext #read-later #ContextEngineering #ContextRot Issue Date: 2026-01-17
OctoCodingBench, MiniMaxAI, 2026.01
Paper/Blog Link My Issue
#Article #Dataset #AIAgents #Evaluation #Coding #SoftwareEngineering Issue Date: 2026-01-16 Comment
元ポスト:
FrogMini-14B-2510, Microsoft, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #Supervised-FineTuning (SFT) #AIAgents #Coding #OpenWeight #SoftwareEngineering #One-Line Notes Issue Date: 2026-01-16 Comment
元ポスト:
strong modelから合成されたbug fixのtrajectoryでSFTすることで小規模モデルでSWE Benchの性能改善
Narrow Misalignment is Hard, Emergent Misalignment is Easy, Turner+, 2025.07
Paper/Blog Link My Issue
#Article #Analysis #LanguageModel #Alignment #PEFT(Adaptor/LoRA) #PostTraining #One-Line Notes #EmergentMisalignment Issue Date: 2026-01-15 Comment
openreview: https://openreview.net/forum?id=q5AawZ5UuQ
一般的にevilになることを学習することが、狭義にevilになるよりも簡単だ、という知見を示した研究とのこと。
LongCat-Flash-Thinking-2601, Meituan, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #OpenWeight #MoE(Mixture-of-Experts) #Selected Papers/Blogs Issue Date: 2026-01-15 Comment
元ポスト:
解説:
coding, agentiaなベンチでTopTierを獲得した560B-27BのMoEモデル。MIT Licence
1MコンテキストウィンドウのZigzag attentionのモデルもcoming soon...だと...!?
Zigzag attentionはおそらく以下だろうか:
- [Paper Note] Efficient Context Scaling with LongCat ZigZag Attention, Chen Zhang+, arXiv'25, 2025.12
[Paper Note] Training large language models on narrow tasks can lead to broad misalignment, Nature 649, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #Alignment #Safety #read-later #Selected Papers/Blogs #Nature #EmergentMisalignment Issue Date: 2026-01-15 Comment
元ポスト:
元ポストによると、以下のような時系列でEmergent Misalignmentのliteratureは形成されていったらしい:
- [Paper Note] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs, Jan Betley+, arXiv'25, 2025.02
- [Paper Note] Persona Features Control Emergent Misalignment, Miles Wang+, arXiv'25, 2025.06
- [Paper Note] Model Organisms for Emergent Misalignment, Edward Turner+, arXiv'25, 2025.06
- [Paper Note] Convergent Linear Representations of Emergent Misalignment, Anna Soligo+, arXiv'25, 2025.06
- Narrow Misalignment is Hard, Emergent Misalignment is Easy, Turner+, 2025.07
- [Paper Note] School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs, Mia Taylor+, arXiv'25, 2025.08
- [Paper Note] Natural Emergent Misalignment from Reward Hacking in Production RL, Monte MacDiarmid+, arXiv'25, 2025.11
- [Paper Note] Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs, Jan Betley+, arXiv'25, 2025.12
Next generation medical image interpretation with MedGemma 1.5 and medical speech to text with MedASR, Google Research, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #MultiModal #SpeechProcessing #Blog #OpenWeight #AutomaticSpeechRecognition(ASR) #VisionLanguageModel #Medical Issue Date: 2026-01-14 Comment
元ポスト:
ポイント解説:
GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image Generation, Z.ai, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #LanguageModel #MultiModal #DiffusionModel #TextToImageGeneration #OpenWeight #Editing Issue Date: 2026-01-14 Comment
元ポスト:
Cowork: Claude Code for the rest of your work, Anthropic, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #GenerativeAI #Blog #WorkspaceAgents Issue Date: 2026-01-13 Comment
元ポスト:
競合(こちらは完全にオフラインで動作する):
- 🍫 Local Cocoa: Your Personal AI Assistant, Fully Local 💻, synvo-ai, 2026.01
MedReason-Stenographic, openmed-community, 2026.01
Paper/Blog Link My Issue
#Article #Dataset #LanguageModel #QuestionAnswering #Chain-of-Thought #SyntheticData #Evaluation #Reasoning #Medical #KeyPoint Notes Issue Date: 2026-01-12 Comment
元ポスト:
MiniMax M2.1を用いてMedical QAに対してreasoning traceを生成。生成されたreasoning traceをstenographic formatと呼ばれる自然言語からフィラーを排除し、論理の流れのみをsymbolicな表現に変換することで合成されたデータセットとのこと。
ユースケースとしては下記とのこと:
> 1. Train reasoning models with symbolic compression
> 2. Fine-tune for medical QA
> 3. Research reasoning compression techniques
> 4. Benchmark reasoning trace quality
個人的には1,3が興味深く、symbolを用いてreasoning traceを圧縮することで、LLMの推論時のトークン効率を改善できる可能性がある。
が、surfaceがシンボルを用いた論理の流れとなると、汎化性能を損なわないためにはLLMが内部でシンボルに対する何らかの強固な解釈が別途必要になるし、それが多様なドメインで機能するような柔軟性を持っていなければならない気もする。
AI Safetyの観点でいうと、論理の流れでCoTが表現されるため、CoTを監視する際には異常なパターンがとりうる空間がshrinkし監視しやすくなる一方で、surfaceの空間がshrinkする代わりに内部のブラックボックス化された表現の自由度が高まり抜け道が増える可能性もある気がする。結局、自然言語もLLMから見たらトークンの羅列なので、本質的な課題は変わらない気はする。
SETA: Scaling Environments for Terminal Agents, CAMEL-AI, 2026.01
Paper/Blog Link My Issue
#Article #Tools #LanguageModel #ReinforcementLearning #AIAgents #SyntheticData #Evaluation #Blog #Repository #SoftwareEngineering #PostTraining Issue Date: 2026-01-12 Comment
元ポスト:
HF: https://huggingface.co/datasets/camel-ai/seta-env
GitHubのreadmeに日本語がある!?
FineTranslations, Penedo+, 2026.01
Paper/Blog Link My Issue
#Article #MachineTranslation #Pretraining #Dataset #LanguageModel #SyntheticData #mid-training #One-Line Notes Issue Date: 2026-01-10 Comment
元ポスト:
FineWeb2のテキストを英訳することで合成されたパラレルコーパスらしい
Demystifying evals for AI agents, Anthropic, 2026.01
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel #AIAgents #Evaluation #Blog #Selected Papers/Blogs Issue Date: 2026-01-10 Comment
元ポスト:
NousCoder-14B: A Competitive Olympiad Programming Model, Joe Li, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #ReinforcementLearning #Blog #Coding #OpenWeight #PostTraining #read-later Issue Date: 2026-01-09 Comment
元ポスト:
HF:
https://huggingface.co/NousResearch/NousCoder-14B
Apache 2.0
PipelineRLを採用している模様。興味深い。
Introducing LFM2.5: The Next Generation of On-Device AI, LiquidAI, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #ReinforcementLearning #Blog #SmallModel #OpenWeight #Japanese #PostTraining #Selected Papers/Blogs #VisionLanguageModel #One-Line Notes #AudioLanguageModel Issue Date: 2026-01-09 Comment
元ポスト:
日本語に特化した言語モデルも存在し、Sarashina2.2-1b-instruct-v0.1, TinySwallow-1.5B-InstructよりもJMMLU, M-IFEval (ja), GSM8K (ja)においてより高い性能を発揮している。
LFM2.5-1.2B-Base: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-1.2B-Base)
LFM2.5-1.2B-Instruct: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct),
[LEAP](
https://leap.liquid.ai/models?model=lfm2.5-1.2b-instruct),
[Playground](
https://playground.liquid.ai/chat?model=cmk1jyp8f000204i56yy76uwh)
LFM2.5-1.2B-JP: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-1.2B-JP),
[LEAP](
https://leap.liquid.ai/models?model=lfm2.5-1.2b-jp)
LFM2.5-VL-1.6B: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-VL-1.6B),
[LEAP](
https://leap.liquid.ai/models?model=lfm2.5-vl-1.6b),
[Playground](
https://playground.liquid.ai/chat?model=cmk0wefde000204jp2knb2qr8),
[Demo](
https://huggingface.co/spaces/LiquidAI/LFM2.5-VL-1.6B-WebGPU)
LFM2.5-Audio-1.5B: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-Audio-1.5B),
[LEAP](
https://leap.liquid.ai/models?model=lfm2.5-audio-1.5b),
[Playground](
http://playground.liquid.ai/talk)
LiquidAIのモデルは日本語に特化したモデルが多く存在するのが特徴的に感じる。
LFM2-2.6B-Transcript, LiquidAI, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #OpenWeight #RecurrentModels #Transcript Issue Date: 2026-01-09 Comment
関連:
- Introducing LFM2: The Fastest On-Device Foundation Models on the Market, LiquidAI, 2025.07
[Paper Note] On the Slow Death of Scaling, Hooker+, 2026.01
Paper/Blog Link My Issue
#Article #NeuralNetwork #EfficiencyImprovement #LanguageModel #Scaling Laws #Author Thread-Post Issue Date: 2026-01-09 Comment
元ポスト:
著者ポスト:
Qwen3-VL-Embedding and Qwen3-VL-Reranker: For the Next Generation of Multimodal Retrieval, Qwen Team, 2026.1
Paper/Blog Link My Issue
#Article #Embeddings #RepresentationLearning #MultiModal #MultiLingual #read-later #Reranking Issue Date: 2026-01-09 Comment
元ポスト:
technical report: https://github.com/QwenLM/Qwen3-VL-Embedding/blob/main/assets/qwen3vlembedding_technical_report.pdf
ポイント解説:
🍫 Local Cocoa: Your Personal AI Assistant, Fully Local 💻, synvo-ai, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #Tools #LanguageModel #AIAgents #MultiModal #Selected Papers/Blogs #ContextEngineering #memory Issue Date: 2026-01-09 Comment
元ポスト:
The next equalizer is not model architecture, but mastery over data behavior, gm8xx8, 2025.12
Paper/Blog Link My Issue
#Article #Pretraining #LanguageModel #SyntheticData #Post #Selected Papers/Blogs #DataMixture #PhaseTransition Issue Date: 2026-01-07 Comment
関連(4-epochまで再利用するのがコスパが良いことを示した研究):
- [Paper Note] Scaling Data-Constrained Language Models, Niklas Muennighoff+, NeurIPS'23
関連(合成データの比率によるPhaseTransition):
- [Paper Note] Data Mixing Can Induce Phase Transitions in Knowledge Acquisition, Xinran Gu+, NeurIPS'25 Spotlight, 2025.05
- [Paper Note] Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls, Feiyang Kang+, EMNLP'25, 2025.10
- [Paper Note] Why Less is More (Sometimes): A Theory of Data Curation, Elvis Dohmatob+, arXiv'25, 2025.11
NVIDIA Cosmos Reason 2 Brings Advanced Reasoning To Physical AI, Nvidia, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #Reasoning #LongContext #SmallModel #OpenWeight #ObjectLocalization #VisionLanguageModel #Robotics #SpatialUnderstanding #EmbodiedAI #Physics Issue Date: 2026-01-06 Comment
HF: https://huggingface.co/nvidia/Cosmos-Reason2-8B?linkId=100000401175768
元ポスト:
VAETKI, NC-AI-consortium, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #Reasoning #MultiLingual #OpenWeight #MoE(Mixture-of-Experts) Issue Date: 2026-01-03 Comment
元ポスト:
Solar-Open-100B, upstage, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #Reasoning #OpenWeight #MoE(Mixture-of-Experts) #Korean Issue Date: 2026-01-03 Comment
元ポスト:
ポイント解説:
K-EXAONE-236B-A23B, LG AI Research, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #Reasoning #MultiLingual #OpenWeight #MoE(Mixture-of-Experts) Issue Date: 2026-01-03 Comment
関連:
- EXAONE-Deep-32B, LG AI Research, 2025.03
Multi Token Prediction
Sliding Window Attention
256k context length
MoE
元ポスト:
A.X-K1, SK Telecom, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #Reasoning #OpenWeight #MoE(Mixture-of-Experts) #Korean Issue Date: 2026-01-03 Comment
元ポスト:
Production-Grade Agentic AI System, FareedKhan-dev, 2025.12
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel #AIAgents #SoftwareEngineering #read-later Issue Date: 2026-01-03 Comment
元ポスト:
Recursive Language Models: the paradigm of 2026, PRIME Intellect, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #Blog #LongContext #read-later #Selected Papers/Blogs #LatentReasoning #reading #RecursiveModels #ContextRot Issue Date: 2026-01-02 Comment
関連研究:
- [Paper Note] Recursive Language Models, Alex L. Zhang+, arXiv'25, 2025.12
- Context Rot: How Increasing Input Tokens Impacts LLM Performance, CHROMA TECHNICAL REPORT, 2025.07
- [Paper Note] Scaling Long-Horizon LLM Agent via Context-Folding, Weiwei Sun+, arXiv'25, 2025.10
- [Paper Note] AgentFold: Long-Horizon Web Agents with Proactive Context Management, Rui Ye+, arXiv'25, 2025.10
- [Paper Note] Agentic Context Engineering: Evolving Contexts for Self-Improving
Language Models, Qizheng Zhang+, arXiv'25, 2025.10
IQuest-Coder, IQuestLab, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #Coding #OpenWeight #SoftwareEngineering Issue Date: 2026-01-01 Comment
元ポスト:
Deriving the DPO Loss from First Principles, aayush garg, 2025.12
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel #ReinforcementLearning #Blog #DPO #PostTraining #read-later Issue Date: 2025-12-31 Comment
元ポスト:
関連:
- Deriving the PPO Loss from First Principles, aayush garg, 2025.12
Today's conversations about AI-assisted programming are strikingly similar to those from decades ago about the choice between low-level languages like C versus high-level languages like Python, Arvind Narayanan, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Coding #Post #SoftwareEngineering Issue Date: 2025-12-31
LLMRouter: An Open-Source Library for LLM Routing, Feng+, 2025.12
Paper/Blog Link My Issue
#Article #Tools #LanguageModel #python #SoftwareEngineering #Routing #Orchestration Issue Date: 2025-12-30 Comment
元ポスト:
SpecBundle & SpecForge v0.2: Production-Ready Speculative Decoding Models and Framework, Spec Forge Team+, lmsys org, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #Blog #LLMServing #SpeculativeDecoding Issue Date: 2025-12-28 Comment
元ポスト:
Reverse Engineering a Phase Change in GPT's Training Data... with the Seahorse Emoji 🌊🐴, PRATYUSH MAINI, 2025.12
Paper/Blog Link My Issue
#Article #Analysis #LanguageModel #ChatGPT #Reasoning #SelfCorrection #mid-training #One-Line Notes Issue Date: 2025-12-28 Comment
元ポスト:
Is there seahorse emoji?という質問に対するLLMのreasoning trajectoryと、self correctionの挙動が、OpenAIのどの時点のモデルで出現するか、しないかを線引くことで、mid-trainingにself correction形式のデータが追加されたのがいつ頃なのかを考察している。
mini-sglang: A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems, sgl-project, 2025
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #LanguageModel #python #Repository #LLMServing #SoftwareEngineering #read-later #Selected Papers/Blogs #MinimalCode Issue Date: 2025-12-28 Comment
元ポスト:
めっちゃ勉強したい
ノーコードで言語モデルの「学習」を体験できるMN-Core Playground _ SLM Customizeの遊び方, PFN, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #Blog #SmallModel #Japanese #PostTraining Issue Date: 2025-12-27 Comment
元ポスト:
Aligning to What? Rethinking Agent Generalization in MiniMax M2, MiniMaxAI, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #Alignment #AIAgents #Blog #Reasoning #read-later Issue Date: 2025-12-27 Comment
元ポスト:
Deriving the PPO Loss from First Principles, aayush garg, 2025.12
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel #ReinforcementLearning #Blog #PostTraining #read-later Issue Date: 2025-12-27 Comment
元ポスト:
LFM2-2.6B-Exp, LiquidAI, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #SmallModel #OpenWeight #RecurrentModels Issue Date: 2025-12-25 Comment
元ポスト:
関連:
- Introducing LFM2: The Fastest On-Device Foundation Models on the Market, LiquidAI, 2025.07
ポイント解説:
LFM2にRLによるpost trainingを実施し、指示追従、知識、数学を伸ばしているとのこと。(ドキュメントにもこれは書かれている)
日本語もサポートされている。2.6Bモデルは、22 conv+8 attnと書かれている。
アーキテクチャは下記で、LIV Operatorは入力に応じて異なる線形変換をするオペレータだが、学習された結果convolutionするのが最適ということになったのだろうか?よくわからない。
>Architecture: Hybrid model with multiplicative gates and short convolutions: 10 double-gated short-range LIV convolution blocks and 6 grouped query attention (GQA) blocks.
GLM-4.7: Advancing the Coding Capability, Z.ai, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Coding #Reasoning #OpenWeight #SoftwareEngineering #One-Line Notes #Reference Collection Issue Date: 2025-12-25 Comment
元ポスト:
HF: https://huggingface.co/zai-org/GLM-4.7
デザインアリーナでtop2:
Artificial Intelligence Indexにおいて、OpenModelの中でトップ:
GLM-4.6と比較して、コーディング/SWE, reasoning, tooluseなどの能力が大幅に向上
Interleaved Thinking, Preserved Thinking, Turn-level Thinkingの3つの特性がある。
Interleaved Thinkingは全てのレスポンスとtool callingの前にreasoningを挟むことで、IFや生成品質を向上。
Preserved Thinkingは過去のターンの全てのthinking blockのトークンを保持し、再計算もしないのでマルチターンでの一貫性が増す。
Turn-level Thinkingはターンごとにreasoningを実施するか否かをコントロールでき、latency/costを重視するか、品質を重視するかを選択できる、といった特徴がある模様。
モデルサイズは358B
【LLM強化学習④】強化学習のコツ(後編), Yuu Jinnai, JSAI公式チャンネル
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel #ReinforcementLearning #Video #PostTraining #read-later Issue Date: 2025-12-25 Comment
元ポスト:
MiniMax M2.1: Significantly Enhanced Multi-Language Programming, Built for Real-World Complex Tasks, MiniMax, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #Blog #Coding #Reasoning #SmallModel #OpenWeight Issue Date: 2025-12-24 Comment
元ポスト:
解説:
Optimizing Large-Scale Pretraining at Character.ai, character.ai, 2025.12
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #Pretraining #LanguageModel #read-later Issue Date: 2025-12-24 Comment
元ポスト:
Medmarks v0.1, a new LLM benchmark suite of medical tasks, Sophont, 2025.12
Paper/Blog Link My Issue
#Article #Dataset #LanguageModel #Evaluation #Medical Issue Date: 2025-12-23 Comment
元ポスト:
A2UI: A Protocol for Agent-Driven Interfaces, Google, 2025
Paper/Blog Link My Issue
#Article #ComputerVision #Tools #AIAgents #SoftwareEngineering #VisionLanguageModel #One-Line Notes Issue Date: 2025-12-22 Comment
AI Agent (Gemini)を用いてUIを自動生成できるツールらしい
元ポスト:
Hot topics in RL, Kimbo, X, 2025.12
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #LanguageModel #ReinforcementLearning #Post #PostTraining #Diversity #train-inference-mismatch Issue Date: 2025-12-22 Comment
ロールアウト側のエンジンと、学習側のエンジンのトークンのlogprobのミスマッチによりon-policy RLを実施しているつもりが実はoff policyになってしまっているという話と
- Your Efficient RL Framework Secretly Brings You Off-Policy RL Training, Yao+, 2025.08
- [Paper Note] Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale
Thinking Model, Ling Team+, arXiv'25, 2025.10
- [Paper Note] Stabilizing MoE Reinforcement Learning by Aligning Training and
Inference Routers, Wenhan Ma+, arXiv'25, 2025.10
長いロールアウトを待っている間がアイドルタイムとなり学習が非常に遅くなる問題を、長すぎるロールアウトは待たないでモデルの重みをロールアウトの途中でもかけてしまい、新しいポリシーでロールアウトを継続すると学習は崩壊せずに高速化できるよ(=in flight updates)という話と
- [Paper Note] PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence
Generation, Alexandre Piché+, arXiv'25, 2025.09
- PipelineRL, Piche+, ServiceNow, 2025.04
RLVRはもともとモデルが事前学習時に保持しているReasoningの能力を広げるわけではなく効率化するだけだよ、という主張と、
- [Paper Note] Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?, Yang Yue+, NeurIPS'25, 2025.04
効率化するだけという主張と、Reasoning能力を拡大しているよ、という相反する主張がコミュニティでされているがそれらをphysics of language modelsに則り完全にコントロールされた条件下で実験し、どのような条件でどのような挙動になるかを明らかにしたよ、という話
- [Paper Note] On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models, Charlie Zhang+, arXiv'25, 2025.12
RLVRはPass@1を報酬としているとみなせるが、それをPass@kにすることで、モデルがRL中に探索する能力が向上し、downstreamタスクのPass@kが向上するよ
- [Paper Note] Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models, Zhipeng Chen+, arXiv'25, 2025.08
といったこの辺の話がホットトピックとして挙げられている。
train-inference-mismatchについては、以下もおもしろかった:
- SID-1 Technical Report: Test-Time Compute for Retrieval, SID Research, 2025.12
- [Paper Note] Defeating the Training-Inference Mismatch via FP16, Penghui Qi+, arXiv'25, 2025.10
OpenTinker Democratizing Agentic Reinforcement Learning as a Service, Zhu+, University of Illinois Urbana-Champaign, 2025.12
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #Tools #LanguageModel #ReinforcementLearning #Blog #PostTraining #KeyPoint Notes #TrainingFramework Issue Date: 2025-12-22 Comment
元ポスト:
code: https://github.com/open-tinker/OpenTinker
関連:
- verl: Volcano Engine Reinforcement Learning for LLMs, ByteDance Seed Team, 2025.04
- Tinker is a training API for {developers, builders, researchers}, THINKING MACHINES, 2025.10
Tinkerに着想を得てクライアントとサーバを分離した設計になっており、バックエンド側のGPUクラスタでサーバを一度起動するだけでクライアント側がスケジューラにジョブを送ればRLが実行される(ローカルにGPUは不要)。クライアント側はRLを実施したい環境のみをローカルで定義しコンフィグをロードしfitを呼び出すだけ。verlよりもよりも手間が省けているらしい。
リポジトリを見る限りは、verlをRLのコアエンジンとして使ってる模様。
Prompt caching: 10x cheaper LLM tokens, but how?, Sam Rose, ngrok, 2025.12
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel Issue Date: 2025-12-22 Comment
元ポスト:
LLMの基礎を勉強してもらう時に用語説明、コード、数式だけでなく、分かりやすい図解やmatrixの具体例まで含めて解説されているので非常に良さそう。
Circuit Tracing: Revealing Computational Graphs in Language Models, Anthropic, 2025.03
Paper/Blog Link My Issue
#Article #NeuralNetwork #LanguageModel #Blog #Transcoders #CircuitAnalysis #Interpretability Issue Date: 2025-12-21
dictionary_learning, Marks+, 2024
Paper/Blog Link My Issue
#Article #NeuralNetwork #MachineLearning #SparseAutoencoder #Transcoders #CircuitAnalysis #Interpretability Issue Date: 2025-12-21
Equipping agents for the real world with Agent Skills, Anthropic, 2025.10
Paper/Blog Link My Issue
#Article #Tutorial #AIAgents #Blog #Selected Papers/Blogs #AgentSkills Issue Date: 2025-12-21
Agent Skills, OpenAI, 2025.12
Paper/Blog Link My Issue
#Article #AIAgents #Repository #AgentSkills Issue Date: 2025-12-21 Comment
元ポスト:
CodexにおけるSkillsのカタログ。
Agent Skillsを最初に提唱したのはAnthropicと記憶している:
- Equipping agents for the real world with Agent Skills, Anthropic, 2025.10
Introducing Bloom: an open source tool for automated behavioral evaluations, Anthropic, 2025.12
Paper/Blog Link My Issue
#Article #Tools #LanguageModel #Alignment #AIAgents #Evaluation #python #Blog #Safety Issue Date: 2025-12-21 Comment
元ポスト:
ByteDance Doubao-Seed-1.8 Review, toyama nao, Zhihu, 2025.12
Paper/Blog Link My Issue
#Article #AIAgents #Evaluation #MultiModal #Reasoning #Proprietary #VisionLanguageModel Issue Date: 2025-12-20 Comment
元ポスト:
Gemma Scope 2: helping the AI safety community deepen understanding of complex language model behavior, Google Deepmind, 2025.12
Paper/Blog Link My Issue
#Article #Tools #LanguageModel #Reasoning #Safety #KeyPoint Notes #SparseAutoencoder #Transcoders #CircuitAnalysis Issue Date: 2025-12-20 Comment
元ポスト:
関連:
- [Paper Note] Sparse Autoencoders Find Highly Interpretable Features in Language Models, Hoagy Cunningham+, ICLR'24
- dictionary_learning, Marks+, 2024
- [Paper Note] Transcoders Find Interpretable LLM Feature Circuits, Jacob Dunefsky+, arXiv'24, 2024.06
- [Paper Note] Learning Multi-Level Features with Matryoshka Sparse Autoencoders, Bart Bussmann+, ICLR'25, 2025.03
- [Paper Note] Transcoders Beat Sparse Autoencoders for Interpretability, Gonçalo Paulo+, arXiv'25, 2025.01
(↓勉強中なので誤りが含まれる可能性大)
Sparse Auto Encoder (SAE; あるlayerにおいてどのような特徴が保持されているかを見つける)とTranscoder (ある層で見つかった特徴と別の層の特徴の関係性を見つける)を用いて、Gemma3の回路分析が行えるモデル・ツール群をリリースした、という話に見える。
応用例の一つとして、たとえば詐欺メールをinputしたときに、詐欺関連する特徴量がどのトークン由来で内部的にどれだけ活性したかを可視化できる。
可視化例:
Evaluating chain-of-thought monitorability, OpenAI, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Chain-of-Thought #Evaluation #Reasoning #Safety #KeyPoint Notes #Monitorability Issue Date: 2025-12-19 Comment
元ポスト:
Monitorabilityの定義が記述されているので覚えておくとよいかもしれない。
Monitorabilityを一言でいうと、"エージェントの挙動に関する重要な性質を正確に予測できる能力"ということっぽい。
Monitorabilityは以下の2つの側面を持ち、
>- Monitors. We could make monitors smarter in various ways, and we would like our evaluations to give higher scores to monitors that can make better predictions about an agent’s behavior.
>- Agents. There are many observations of the agent that could be given to the monitor, including the agent’s actions, outputs, neural activations, and recently with the advent of reasoning models, the chain-of-thought. It’s possible for some agents to have thoughts that are harder to understand than others, and so we would like our monitorability evaluations to be able to compare agents and not just monitors.
モニターそのものの賢さと、モニタリングされるエージェントも理解しやすいものとしづらいものがいるため、モニターの性能だけじゃなくてエージェントそのものも評価の対象にする必要がある。
monitarable / unmonitorable なCoTのサンプルとして以下が挙げられている。これは実際には実現不可能なコーディング課題が与えられているらしいのだが、左側はtestをpassしている(つまりなんかチートした)がそのことについてはCoTに言及されていない。一方、右側はチートに関して検討する様子がCoTに記述されており、これを見ることでユーザはモデルのよろしくない挙動に気付ける。
Introducing Mistral OCR 3, Mistral AI, 2025.12
Paper/Blog Link My Issue
#Article #ComputerVision #LanguageModel #Proprietary #VisionLanguageModel #OCR #One-Line Notes Issue Date: 2025-12-19 Comment
元ポスト:
MistralによるOCR。他のOCRに比べてmulti-lingual, 様々なデータセットで高い性能を発揮。APIでのみ提供されている模様。
Seed1.8, ByteDance Seed, 2025.12
Paper/Blog Link My Issue
#Article #ComputerVision #AIAgents #Proprietary #ComputerUse #VisionLanguageModel Issue Date: 2025-12-18 Comment
元ポスト:
GUI Agentとして性能はトップレベル(Opusが比較対象に入っていないが)で、
テキスト、画像モダリティでの検索でもトップレベル、codingやツール利用などは少し劣るように見える。
LLM系、VideoUnderstanding系ののベンチマークではフロンティアモデル群と同等、VLM系のタスクではフロンティアモデル群と同等以上の性能に見える。
が、一方のモダリティはGPT5で比較しているのに対し、他方はGPT5.1であったりしており、比較対象が少し恣意的にピックされているのでは?という気もする。
Gemini 3 Flash: frontier intelligence built for speed, Google, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #Blog #Reasoning #Distillation #Proprietary #VisionLanguageModel #One-Line Notes #Reference Collection Issue Date: 2025-12-18 Comment
元ポスト:
Gemini 2.5 Proよりも3倍高速でかつ様々なベンチマークで上回っているとのこと。素晴らしい。Gemini 3 Proと比較しても基本的なQAや数学的な能力(reasoning能力)は性能に遜色なく、long sequence/contextの取り扱いでは明確に劣っている、という感じに見えるので、普段使いではこちらでも困らなそうに感じる。
Hallucination Rateが非常に高いとのことだが果たして:
Proからlogit baseな蒸留をして事前学習(=distillation pretraining)をしているっぽい?
GENIAC第3期で自律稼働デバイス向けの軽量な大規模視覚言語モデルPLaMo 2.1-8B-VLを開発, PFN, 2025.12
Paper/Blog Link My Issue
#Article #Blog #SmallModel #Japanese #VisionLanguageModel Issue Date: 2025-12-17 Comment
元ポスト:
PLaMo2.1-8BをベースにPLaMo翻訳を通じてVision Languageモデル用の合成データを学習し、既存の公開データと混ぜて学習することで学習されたVision Language Model Plamo2.1-8B-VLがのプロモーション用のブログ。
日本語でのVisual Question Answering (VQA)、Visual Groundingベンチマークにおいて、Qwen3-VL-8Bを上回るスコアを達成しているとのこと(具体的な数値は言及されていないが、いくつかの実例が見れる)。
現場での技術検証のためのモニター企業を募集している。
Evaluating AI’s ability to perform scientific research tasks, OpenAI, 2025.12
Paper/Blog Link My Issue
#Article #Dataset #LanguageModel #Evaluation #Reasoning #Science #KeyPoint Notes Issue Date: 2025-12-17 Comment
元ポスト:
HF: https://huggingface.co/datasets/openai/frontierscience
physics, chemistry, biologyの分野の専門家が作成した問題によって構成されるPh.D levelの新たなscientificドメインのベンチマークとのこと。OlympiadとResearchの2種類のスプリットが存在し、Olympiadは国際オリンピックのメダリストによって設計された100問で構成され回答は制約のある短答形式である一方、Researchは博士課程学生・教授・ポスドク研究者などのPh.Dレベルの人物によって設計された60個の研究に関連するサブタスクによって構成されており、10点満点のルーブリックで採点される、ということらしい。
公式アナウンスではGPT-5.2がSoTAでResearchの性能はまだまだスコアが低そうである。
bu-30b-a3b-preview: Meet BU-30B-A3B-Preview — bringing SoTA Browser Use capabilities in a small model that can be hosted on a single GPU., browser-use, 2025.12
Paper/Blog Link My Issue
#Article #ComputerVision #OpenWeight #ComputerUse #VisionLanguageModel Issue Date: 2025-12-17 Comment
元ポスト:
Introducing MiMo-V2-Flash, Xiaomi, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #OpenWeight #MoE(Mixture-of-Experts) #AttentionSinks #PostTraining #Selected Papers/Blogs #Reference Collection Issue Date: 2025-12-17 Comment
technical report:
https://github.com/XiaomiMiMo/MiMo-V2-Flash/blob/main/paper.pdf
HF:
https://huggingface.co/XiaomiMiMo/MiMo-V2-Flash
元ポスト:
関連:
ポイント解説:
attention sink(というより恐らくsink token)により性能が向上している:
言及されているpost trainingが有用らしい:
所見:
省パラメータでtop-tierのモデルに肉薄する方法のヒントがあるかもしれない。
解説:
Molmo 2: State-of-the-art video understanding, pointing, and tracking, Ai2, 2025.12
Paper/Blog Link My Issue
#Article #ComputerVision #MultiModal #SmallModel #OpenWeight #OpenSource #Selected Papers/Blogs #VideoGeneration/Understandings #VisionLanguageModel #2D (Image) #3D (Video) #KeyPoint Notes Issue Date: 2025-12-17 Comment
テクニカルレポート:
https://www.datocms-assets.com/64837/1765901660-molmo_v2_2026-techreport-3.pdf
HF:
https://huggingface.co/collections/allenai/molmo2
Qwen3とOlmoをベースにしたvariantsが存在し、Olmoの方はバックボーンのLLMも含めて全てがオープンになっている。MetaのPerceptionLMと比較して1/8の動画データ量で高い性能を達成できており、データのcurationの品質と、grounding basedな目的関数の工夫によって実現されているとのこと。
proprietaryなモデル群と比較すると、trackingは圧勝、そのほかはGPT5-miniと同様なものが多い。モデルによってタスクの優劣が結構分かれており、Video関連タスクをタスクをまたいで汎化させることにはclosedでも苦戦しているように見える。
オープンモデルとの比較で言うと圧勝で、LongVideoのQAに関してだけは、Eagle2.5-8Bと呼ばれるモデルが勝っている。
あとは全体を通じてLLMのバックボーンがQwen3の場合の性能が良いことが興味深い。バックボーンに採用するLLMに応じて性能が結構変わる。これはアーキテクチャがそもそもConnectorを利用するタイプのもので、Unifiedなアーキテクチャではないことが要因としては考えられる。
元ポスト:
demo:
コードベースが公開:
https://github.com/allenai/molmo2
2025 Open Models Year in Review, Interconnects AI, 2025.12
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel #Blog Issue Date: 2025-12-15 Comment
元ポスト:
SID-1 Technical Report: Test-Time Compute for Retrieval, SID Research, 2025.12
Paper/Blog Link My Issue
#Article #InformationRetrieval #LanguageModel #ReinforcementLearning #AIAgents #Proprietary #Selected Papers/Blogs #KeyPoint Notes #Scalability #train-inference-mismatch Issue Date: 2025-12-15 Comment
元ポスト:
Figure4の話が非常に興味深い。rolloutの結果をtraining engineに渡す間のchat_templateによる抽象化では、マルチターン+tooluseにおいては、たとえばtool call周辺のホワイトスペースに関する情報を消してしまう問題がある。具体的には、一例として、ポリシーがホワイトスペースを含まないフォーマットの誤りがあるrolloutを生成した場合(=B)を考える。これをtraining engineに渡す際は、以下のような操作を伴うが
>apply_chat_template(parse(B))=G′
この際に、parse→apply_chat_templateの過程でtoolcall周辺のホワイトスペースが補完されるためtraining側ではホワイトスペースが含まれたrollout時とはトークン列が与えられる。この結果、フォーマットに誤りがある状態でrolloutされたにも関わらず、trainingエンジン側では正しい生成結果に擬似的に見える(=G')のだが、ホワイトスペースが含まれたことでトークナイズ結果が変わり、変化したトークンの部分が極端に小さなlogprobを持つことになる(i.e., ホワイトスペースは実装上の都合で生じ、ポリシーはそのトークンを(尤度が低く)出力していないにもかかわらず、出力されたことにされて学習される)。その結果、見かけ上は正しい生成結果なのだが、負のAdvantageを持つことになり、GRPOではそのような生成がされないように学習されてしまう。これが繰り返されることで、学習の安定性を損なう、という話である。
言語生成の強化学習をやっていく(手法紹介 REINFORCE編), Seitaro Shinagawa, 2020.12
Paper/Blog Link My Issue
#Article #Tutorial #ReinforcementLearning #Blog Issue Date: 2025-12-14
Olmo 3.1, Ai2, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #Reasoning #OpenWeight #OpenSource #Selected Papers/Blogs Issue Date: 2025-12-13 Comment
元ポスト:
Instruction Followingのベンチマークスコアが、他モデルと比較して非常に高いように見える。
GPT-5.2 が登場 専門的な業務や長時間稼働するエージェント向けの、最先端のフロンティアモデル。, OpenAI, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #ChatGPT #GenerativeAI #Reasoning #Proprietary #Selected Papers/Blogs #VisionLanguageModel Issue Date: 2025-12-12 Comment
元ポスト:
OpenAIがGPT-5.2をリリースし、再び様々なベンチマークにおいてGemini 3 Proをoutperform。
フロントエンド開発(デザイン)(アリーナ形式)ではOpus, Gemini 3 Proの勝利らしい:
https://www.designarena.ai
ポイント解説:
GDPval:
- [Paper Note] GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks, Tejal Patwardhan+, arXiv'25, 2025.10
- GDPVAL: EVALUATING AI MODEL PERFORMANCE ON REAL-WORLD ECONOMICALLY VALUABLE TASKS, Patwardhan+, 2025.09
GDPvalのclearwinがGPT-5.2- Thinkingで49.8%なので、14年程度の専門家がこなす米国主要産業の一部のタスクは数値上は置き換え可能という風に見える。Proに至っては60.0%である。
が、GDPvalはたとえば以下のようなlimitationがあり、数値の解釈には注意が必要である:
- 完全なcontextが与えられる前提
- 暗黙知が多いタスクは対象外
- 自己完結型で他社とのコミュニケーションが必要とされないタスクを対象
- 1職種あたり30タスク程度の限定的な網羅性
- コンピュータを利用したタスクのみ
- ...
実際の現場で活用しようと思うと、完全なcontextを揃えられるか、揃わない場合に不完全なcontextでタスクを遂行できるか、そのための社内での運用フローの整備等、モデルを活用するための周辺のシステムや運用フローの設計が重要(かつ膨大)である点には(ベンチマークのスコアを見ると驚くべき進歩だが)留意する必要がある。
Vals AI IndexというGDPvalに類似したベンチマークでもSoTAとのこと:
関連:
非常に簡単な論理的な推論でも誤る例:
nomos-1, NousResearch, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #Reasoning #Mathematics #OpenWeight #One-Line Notes Issue Date: 2025-12-11 Comment
元ポスト:
30Bの強力な数学モデルで、(同じハーネスでテストした結果)Qwen3-30ba3b-Thinking-2507を大幅に上回る性能を持つとのこと。
AutoGLM-Phone-9B, Zhipu AI, 2025.12
Paper/Blog Link My Issue
#Article #ComputerVision #SmallModel #OpenWeight #VisionLanguageModel Issue Date: 2025-12-10 Comment
元ポスト:
GLM-4.6V, Zhipu AI, 2025.12
Paper/Blog Link My Issue
#Article #ComputerVision #OpenWeight #VisionLanguageModel Issue Date: 2025-12-10 Comment
元ポスト:
Devstral2 Mistral Vibe CLI State-of-the-art, open-source agentic coding models and CLI agent., Mistral AI, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Coding #OpenWeight #SoftwareEngineering Issue Date: 2025-12-10 Comment
SWE Bench VerifiedでOpenweightモデルの中ではSoTAと同等程度を達成。123B, 24Bの2種類がリリース。DeepSeekV3.2, Kimi K2よりも大幅に小さいパラメータで同等以上の性能。独自の人手評価(win, tie, loseのアリーナ形式)によるとSonnet 4.5には負けるがDeepSeekV3.2とは同等以上の割合で好まれた。
元ポスト:
NIIにおける大規模言語モデル構築事業の現在地, Yusuke Oda, 人工知能学会合同研究会 招待講演資料, 2025.12.01
Paper/Blog Link My Issue
#Article #LanguageModel #Optimizer #ExperimentManagement #Slide #Japanese #DataMixture Issue Date: 2025-12-09 Comment
WSD Scheduler:
- [Paper Note] MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies, Shengding Hu+, COLM'24
元ポスト:
State of AI An Empirical 100 Trillion Token Study with OpenRouter, Aubakirova+, OpenRouter, 2025.12
Paper/Blog Link My Issue
#Article #Analysis #LanguageModel #GenerativeAI #One-Line Notes Issue Date: 2025-12-09 Comment
元ポスト:
> 利用傾向として、最初に課題を解決したモデルがその後も使われ続けるという「ガラスの靴」現象が起きている。これは、あるモデルがリリース改善したとき、特定の技術的・経済的制約を満たす瞬間があり、そのときにユーザーが一気に使い始め、一度それが起きるとシステム設計、データパイプライン、ユーザー習慣がそのモデルを中心に構築されるため、乗り換えインセンティブは急激に低下し、ユーザー離脱がおきづらくなるものである。
(上記元ポストより引用)
特にこの点は非常に興味深いと感じる。一度設計や評価をしてしまうと簡単にはモデルを変更できずロックインするという状況は実際に見聞きする。Tech Giantが汎用的なモデルを出し続けるなら、資金力やリソースが乏しい場合は同じ土俵ではなく、特定ユースケース特化で小型、か 高性能、かつ使いやすいインタフェースをセットで出すのが良さそうではある(最近見かけるのはOCR, 翻訳などだろうか)。
Why Training MoEs is So Hard, _xjdr, X Post
Paper/Blog Link My Issue
#Article #LanguageModel #SmallModel #Post #MoE(Mixture-of-Experts) #read-later #reading Issue Date: 2025-12-08
Titans + MIRAS: Helping AI have long-term memory, Google Research, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Blog #Test-Time Scaling #memory Issue Date: 2025-12-07 Comment
元ポスト:
関連:
- [Paper Note] It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization, Ali Behrouz+, arXiv'25, 2025.04
- [Paper Note] Titans: Learning to Memorize at Test Time, Ali Behrouz+, NeurIPS'25, 2024.12
解説:
ポイント解説:
Architecting efficient context-aware multi-agent framework for production, Hangfei Lin, Google, 2025.12
Paper/Blog Link My Issue
#Article #AIAgents #Blog #read-later #Selected Papers/Blogs #ContextEngineering Issue Date: 2025-12-07 Comment
元ポスト:
OpenThinker-Agent-v1, open-thoughts, 2025.12
Paper/Blog Link My Issue
#Article #Dataset #LanguageModel #AIAgents #Evaluation #SmallModel #OpenWeight #OpenSource #Selected Papers/Blogs #KeyPoint Notes Issue Date: 2025-12-07 Comment
元ポスト:
-
-
agenticなSLM(8Bモデル)で、モデル、データ(SFT, RL)、学習用のコードなど全て公開。同等規模のモデルQwen3-{8,32B}よりもSWE Bench Verified, Terminal Benchなどで上回る(ただし、Qwen3はgenericなモデルであり、コーディング特化のQwen3-coder-30Bには及ばない。しかしモデルサイズはこちらの方が大きいので何とも言えない。おそらく同等規模のコーディング特化Qwen3が存在しない)。また、SLMのコーディングエージェントの進化をより精緻に捉えるためのベンチマーク OpenThoughts-TB-Devも公開している。こちらでもQwen3-{8, 32B}に対しても高い性能を記録。
Introducing the Yupp SVG AI Leaderboard, YUPP, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #Evaluation #Coding #Reasoning Issue Date: 2025-12-06 Comment
元ポスト:
SVG生成においてもGemini 3 Proが強い
Announcing Rnj-1: Building Instruments of Intelligence, Ashish Vaswani, essential AI, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #SmallModel #read-later #Selected Papers/Blogs Issue Date: 2025-12-06 Comment
元ポスト:
transformerの人では...?SWE Benchのスコアが同サイズのモデル群と比較して圧倒的
model: https://huggingface.co/EssentialAI/rnj-1-instruct
解説:
Mismatch Praxis: Rollout Settings and IS Corrections, LLM Data, 2025.12
Paper/Blog Link My Issue
#Article #Analysis #LanguageModel #ReinforcementLearning #Blog #SamplingParams #One-Line Notes #LongHorizon #train-inference-mismatch Issue Date: 2025-12-04 Comment
元ポスト:
on-policy RLにおけるロールアウト時のtemperature, top_p, top_kの設定、およびlong horizonの場合でのtrain-inference mismatchの関係性の分析
Nemotron-Content-Safety-Reasoning-4B, Nvidia, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #Reasoning #Conversation #SmallModel #OpenWeight #Safety #Safeguard Issue Date: 2025-12-03 Comment
元ポスト:
Introducing Amazon Nova 2 Lite, a fast, cost-effective reasoning model, AWS, 2025.12
Paper/Blog Link My Issue
#Article #LanguageModel #Blog #Reasoning #Proprietary Issue Date: 2025-12-03 Comment
元ポスト:
関連:
- Introducing Amazon Nova, our new generation of foundation models, AWS, 2024.12
Building Safer AI Browsers with BrowseSafe, Perplenity Team, 2025.12
Paper/Blog Link My Issue
#Article #Dataset #LanguageModel #Prompting #Evaluation #Blog #OpenWeight #Safety #Safeguard Issue Date: 2025-12-03 Comment
元ポスト:
prompt injectionをリアルタイムに検知するモデルとそのベンチマークとのこと
dataset:
https://huggingface.co/datasets/perplexity-ai/browsesafe-bench
model:
https://huggingface.co/perplexity-ai/browsesafe
Introducing Mistral 3 The next generation of open multimodal and multilingual AI, Mistral AI, 2025.12
Paper/Blog Link My Issue
#Article #ComputerVision #MultiModal #Blog #MultiLingual #OpenWeight #VisionLanguageModel #One-Line Notes Issue Date: 2025-12-03 Comment
元ポスト:
マルチモーダルなベンチマークがほとんどないように見えるMM-MT-Benchというもののみ?
Expert Parallel Deployment, vLLM, 2025.10
Paper/Blog Link My Issue
#Article #LanguageModel #Blog #MoE(Mixture-of-Experts) #Parallelism #One-Line Notes Issue Date: 2025-12-01 Comment
MoEアーキテクチャにおいて、eXertsの重みを複数のGPUに分散することで計算効率を増大させるexpert parallelによるデプロイ方法をexpert parallelの配列数はData Parallel数*tensor parallel数となる。
Evaluating honesty and lie detection techniques on a diverse suite of dishonest models, Wang+, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #Evaluation #Blog #read-later Issue Date: 2025-11-30 Comment
元ポスト:
[Paper Notes] Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem, Longpre+, 2025.11
Paper/Blog Link My Issue
#Article #ComputerVision #Analysis #LanguageModel #OpenWeight #VisionLanguageModel Issue Date: 2025-11-30 Comment
元ポスト:
MITとHuggingFaceの調査によると、open weightモデルのDLにおいて、米国のAI産業における中国のモデルDL数が米国のモデルを初めて抜いた模様。
ダッシュボード: https://huggingface.co/spaces/economies-open-ai/open-model-evolution
[Paper Notes] Structured Prompting Enables More Robust, Holistic Evaluation of Language Models, Aali+, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #Prompting #Evaluation #read-later #Selected Papers/Blogs #One-Line Notes Issue Date: 2025-11-30 GPT Summary- 高品質な言語モデル(LM)の評価には、HELMのようなフレームワークが重要だが、固定プロンプトに依存するため過小評価のリスクがある。DSPyのような宣言的プロンプトフレームワークは、タスクごとに最適化されたプロンプトを提供するが、体系的な評価が不足している。本研究では、再現可能なDSPy+HELMフレームワークを提案し、構造化プロンプトを用いてLMのパフォーマンスをより正確に評価する。4つのプロンプト手法を用いて7つのベンチマークで評価した結果、HELMがLMのパフォーマンスを平均4%過小評価し、パフォーマンスの変動が大きくなることが示された。この研究は、LMの挙動を特徴付ける初の大規模ベンチマーク研究であり、オープンソースの統合とプロンプト最適化パイプラインを提供する。 Comment
AI Agentsの評価でもハーネスによって性能が変わるし、一般的なLLMでの評価もpromptingで性能変わるだろうなぁ、とは思っていたが、やはりそうだった模様。重要論文
しかしそもそもLLMの評価は変数が多すぎて、網羅的な評価は難しく、活用する際にベンチマークスコアは参考程度にした方が良いとは思う。自前データがあるなら自前で手元で評価すべし、という気はするが、評価するLLMの候補を選定する際には有用だと思われる(小並感)
関連:
- [Paper Note] Holistic Evaluation of Language Models, Percy Liang+, arXiv'22, 2022.11
元ポスト:
オープンウェイトモデル( gpt-oss )の日本語精度は? – AWS パートナー アクロクエストによる徹底検証, Yamamoto+, 2025.11
Paper/Blog Link My Issue
#Article #Analysis #Evaluation #OpenWeight #Japanese Issue Date: 2025-11-29 Comment
元ポスト:
LLMのための強化学習手法 2025 -PPO・DPO・GRPO・DAPO一気に理解する-, Keisuke Kamata, 2025.11
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel #ReinforcementLearning #Blog #Selected Papers/Blogs Issue Date: 2025-11-29 Comment
元ポスト:
こちらもあわせて読むと良さそう
- 言語生成の強化学習をやっていく(手法紹介 REINFORCE編), Seitaro Shinagawa, 2020.12
- 深層強化学習アルゴリズムまとめ, Shion Honda, 2020.09
- RLHF/DPO 小話, 和地瞭良/ Akifumi Wachi, 2024.04
Ilya Sutskever – We're moving from the age of scaling to the age of research, DWARKESH PATEL, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #GenerativeAI #Blog #One-Line Notes Issue Date: 2025-11-29 Comment
元ポスト:
現在のnext token predictionに基づく事前学習とRLに基づくスケーリング則による性能改善の時代から(理解が進んでいない部分があり、特に現在のRLでは汎化性能が十分に獲得できないため)、人間のような高度な価値関数の探求を含む新たなパラダイムを研究する時代の到来に関する話な模様
Introducing the WeirdML Benchmark, Håvard Tveit Ihle, 2025.01
Paper/Blog Link My Issue
#Article #MachineLearning #LanguageModel #Evaluation #Blog #Initial Impression Notes #Author Thread-Post Issue Date: 2025-11-29 Comment
著者ポスト:
元ポスト:
WeirdML v2: https://htihle.github.io/weirdml.html
MLにおけるあまり一般的ではない(=Weird)なタスクによるLLMのベンチマークらしい
Why (Senior) Engineers Struggle to Build AI Agents, PHILSCHMID, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Blog #read-later Issue Date: 2025-11-27 Comment
元ポスト:
[Paper Note] DeepSeek-Math-V2, DeepSeekAI, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #ReinforcementLearning #Reasoning #Mathematics #read-later #Selected Papers/Blogs #Verification #One-Line Notes #Reference Collection #GenerativeVerifier Issue Date: 2025-11-27 GPT Summary- 大規模言語モデル(LLM)は数学的推論において進展を遂げており、強化学習を用いて定量的推論コンペティションでのパフォーマンスを向上させている。しかし、最終回答の精度向上が正しい推論を保証しない問題や、厳密な導出が必要なタスクに対する限界がある。自己検証可能な数学的推論を目指し、定理証明のためのLLMベースの検証器を訓練し、生成器が自らの証明の問題を特定・解決するよう奨励する方法を提案。結果として得られたモデルDeepSeekMath-V2は、強力な定理証明能力を示し、国際数学オリンピックやプットナム競技会で高得点を記録した。これにより、自己検証可能な数学的推論が数学AIシステムの発展に寄与する可能性が示唆される。管理人コメント:モデル単体でIMO金メダル級を達成とのこと。outcomeに基づくRLVRからtrajectoryそのものをcritiqueし、その情報に基づいて再生成するといったループを繰り返す模様?このアプローチは数学以外のドメインでも有効な可能性があるので興味深い。 Comment
元ポスト:
HF: https://huggingface.co/deepseek-ai/DeepSeek-Math-V2
所見:
所見:
どのように高品質なverifierを構築し、高品質なデータ生成パイプラインを構築するか、という内容が記述されているらしい:
報酬に対する理解補助のための注釈:
ポイント解説:
verifier: proofsをスコアリングできるようRLで学習される
meta verifier: verifierの批評を確認する
generator: より良い証明を書きself checkもできるようverifierによるreward signalによりRLで訓練される
の三刀流らしい。
ポイント解説:
ポイント解説:
所見:
veAgentBench, ByteDance, 2025.11
Paper/Blog Link My Issue
#Article #Dataset #Education #AIAgents #Evaluation #Financial #Legal Issue Date: 2025-11-26 Comment
元ポスト:
GPT-4V-Act, ddupont808, 2023.10
Paper/Blog Link My Issue
#Article #ComputerVision #Repository #ComputerUse #VisionLanguageModel #One-Line Notes #Grounding Issue Date: 2025-11-25 Comment
GPT4V(VLM)と、SoMを用いてVLMによってWebUIとClick/Keyboard操作を通じてinteractできる実装
- [Paper Note] Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V, Jianwei Yang+, arXiv'23, 2023.10
Sarashina2.2-Vision-3B: コンパクトかつ性能が高いVLMの公開, SB Intuitions, 2025.11
Paper/Blog Link My Issue
#Article #Blog #SmallModel #Japanese #VisionLanguageModel #Cultural Issue Date: 2025-11-25 Comment
元ポスト:
HF: https://huggingface.co/sbintuitions/sarashina2.2-vision-3b
Claude-Opus-4.5: Introducing advanced tool use on the Claude Developer Platform, Anthropic, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Blog #Proprietary #Selected Papers/Blogs #VisionLanguageModel #Reference Collection Issue Date: 2025-11-25 Comment
元ポスト:
AnthropicがClaude-Opus-4.5をリリース。AgenticなユースケースでClaudeがベンチマーク上の首位をGemini3 Proから奪還
システムカード:
https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf
人間と比較した時のパフォーマンスの解説:
EpochAIによるFrontierMath Tier1-3での評価:
o3(high), Grok4と同等程度で、Gemini3 Pro, GPT-5.1(high)には劣る
ベンチマーク上でのコーディング能力やagenticなツール呼び出し能力の差は縮まっている:
Artificial Analysisの評価:
スライドをいい感じに作れるらしい:
50% time horizonは4時間49分で現在top。
Stanford Agentic Reviewer, Stanford University, 2025.11
Paper/Blog Link My Issue
#Article #AIAgents #GenerativeAI #Blog #One-Line Notes Issue Date: 2025-11-25 Comment
元ポスト:
Andrew Ng氏によるAI Agentによる論文のレビュワーシステムで、ICLR'25のレビューで学習し、テストセットで評価したところ、人間-人間間の相関と人間-AI間の相関係数が同等の水準に到達とのこと。ICLR'25のレビューで学習しているということは当該ドメインに近しい研究であるほど適切なレビューが実施されるであろう点に注意。
OCR Arena, extend.ai, 2025.11
Paper/Blog Link My Issue
#Article #ComputerVision #Evaluation #VisionLanguageModel #OCR #One-Line Notes Issue Date: 2025-11-25 Comment
元ポスト:
OCRのアリーナ(=ユーザがPDFをアップロードし2モデルでOCRし優劣をユーザが判定しその結果からElo Rateを算出する)。
言語間の性能差はわからないので参考程度にすると良いと思われる。
Context Arena, DillonUzar, 2025.04
Paper/Blog Link My Issue
#Article #LanguageModel #Evaluation #LongContext Issue Date: 2025-11-24 Comment
元ポスト:
関連:
異なる学習手法、アーキテクチャがlong contextの性能に与える影響の考察:
大規模言語モデルの次期バージョン PLaMo 3 シリーズにおける8B, 31Bの小規模モデルによる事前学習の検証, PFN, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #Blog #Japanese Issue Date: 2025-11-21 Comment
元ポスト:
コーディング能力で大幅に性能向上している模様:
- Swallow LLM Leaderboard v2, Swallow LLM Team, 2025.08
Benchmark Scores = General Capability + Claudiness, EpochAI, 2025.11
Paper/Blog Link My Issue
#Article #Dataset #LanguageModel #Evaluation #Blog #read-later Issue Date: 2025-11-21 Comment
元ポスト:
Claudiness=Claudeらしさ=エージェントタスクに優れている、しかしマルチモーダルや数学には弱いこと(皮肉を込めてこう呼んでいるらしい)
Claudeらしくないモデルとしては、o4-miniやGPT-5が挙げられる。
TAURO Project, note, 2024.10
Paper/Blog Link My Issue
#Article #Tutorial #ComputerVision #Blog #ScientificDiscovery #Japanese #Robotics Issue Date: 2025-11-20 Comment
元ポスト:
👀👀👀
Distributed Inference Serving - vLLM, LMCache, NIXL and llm-d, Mikiya Michishita, 2025.06
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel #LLMServing #Slide #SoftwareEngineering #read-later #Selected Papers/Blogs Issue Date: 2025-11-20 Comment
元ポスト:
vLLM, paged attention, prefix caching, continuous batching, 分散環境でのKV Cacheの共有, ...おおお、、読まねば
NVIDIA-Nemotron-Parse-v1.1, NVIDIA, 2025.11
Paper/Blog Link My Issue
#Article #ComputerVision #TabularData #OpenWeight #read-later #DocParser #VisionLanguageModel #OCR Issue Date: 2025-11-20 Comment
元ポスト:
olmocr2と比較して性能はどうだろうか、特に日本語
- olmOCR 2: Unit test rewards for document OCR, Ai2, 2025.10
Introducing zerank-2: The Most Accurate Multilingual Instruction-Following Reranker, ZeroEntropy, 2025.11
Paper/Blog Link My Issue
#Article #RecommenderSystems #Embeddings #InformationRetrieval #Blog #OpenWeight #Reranking Issue Date: 2025-11-20 Comment
HF: https://huggingface.co/zeroentropy/zerank-2
SoTA reranker
Introducing Navigator, Yutori team, 2025.11
Paper/Blog Link My Issue
#Article #AIAgents #Blog #Proprietary #ComputerUse #read-later #VisionLanguageModel #One-Line Notes Issue Date: 2025-11-20 Comment
元ポスト:
gemini2.5, claude4.5, openaioperator等よりも性能が良いweb agentらしい
Previewing Locus, INTOLOGY, 2025.11
Paper/Blog Link My Issue
#Article #AIAgents #Blog #ScientificDiscovery #Test-Time Scaling #LongHorizon Issue Date: 2025-11-20 Comment
元ポスト:
所見:
AI Model Benchmarks Nov 2025, lmcouncil, 2025.11
Paper/Blog Link My Issue
#Article #Dataset #LanguageModel #AIAgents #Evaluation #Blog Issue Date: 2025-11-19 Comment
元ポスト:
50% time horizonなどを含む良さそうなベンチマークと主要モデルの比較が簡単にできそうなサイト
LLM Datasets, mlabonne, 2025.11
Paper/Blog Link My Issue
#Article #Survey #Dataset #LanguageModel #AIAgents Issue Date: 2025-11-19 Comment
元ポスト:
Gemini 3 による知性の新時代, Google, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #GenerativeAI #Blog #Proprietary #Selected Papers/Blogs #VisionLanguageModel #One-Line Notes #Reference Collection Issue Date: 2025-11-19 Comment
所見:
GPT5.1に対して各種ベンチマークで上回る性能。
所見:
Gemini2.5 Proは回答が冗長で使いにくかったが、Gemini3は冗長さがなくなり、クリティカルな情報を簡潔に、しかし短すぎない、ちょうど良いくらいの応答に感じており、レスポンスもGPT5.1, GPT5と比べ早いので普段使いのLLMとしては非常に良いのではないか、という感想(2,3個のクエリを投げただけだが)を抱いた。
Oriol Vinyals氏のコメント:
LiveCodeBench ProでもSoTA:
Gemini Pro 3 Developer Guide:
https://ai.google.dev/gemini-api/docs/gemini-3?hl=ja
元ポスト:
GAIA Verified (Browser Use?)でもSoTA:
ただし、どのようなハーネスが使われているかは不明だし、それらが各モデルにとってフェアなものになってるかも不明
スクショのみでリンクも無し。
所見:
content window,pricingなどの情報:
一般的なユースケースでのBest Practice:
パラメータ数に関する考察:
韓国語でのベンチマークに関するポスト:
自身のハーネス、ユースケース、タスクではうまくいかなかったよという話(でもただのサンプル数1だよ、という話が記載されている):
結局のところベンチマークはあくまで参考程度であり、自分たちのタスク、データセットで性能を測らねばわからない。
Artificial Intelligenceによる評価:
MCP Universeでtop:
- [Paper Note] MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers, Ziyang Luo+, arXiv'25
Live SWE Agentと呼ばれるself-evolvingな枠組みを採用した場合(=scaffoldをbashのみから自己進化させる)のSWE Bench Vevifiedにやる評価でもSoTA:
- [Paper Note] Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?, Chunqiu Steven Xia+, arXiv'25, 2025.11
- [Paper Note] SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, Carlos E. Jimenez+, ICLR'24
この辺のsoftware agent系のベンチマークにおけるハーネスが具体的にどうなっているのか、中身を見たことないので見ておきたい。
(追記)
SWE Bench Verifiedのリーダーボードではmini-SWE-Agentを利用した公正な比較が行われており、こちらではGemini3がトップだったもののその後リリースされたClaude-Opus-4.5がtopを僅差で奪還しGemini3が2位とのこと。
ハーネスについてはこちらを読むと良さそう:
- [Paper Note] SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, John Yang+, arXiv'24, 2024.05
EpochAIによる評価:
ECIでtop。ECIは39のベンチマークから算出されるスコア、らしい。
Scale AIのVisual Tool BenchでもSoTA:
- Beyond Seeing: Evaluating Multimodal LLMs On Tool-enabled Image Perception, Transformation, and Reasoning, Scale AI, 2025.10
CriPtと呼ばれるベンチマークにおける評価でもSoTA:
- [Paper Note] Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark, Minhui Zhu+, arXiv'25, 2025.09
最近提案された新たなtooluseベンチマークでもsecond placeらしい:
- [Paper Note] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution, Junlong Li+, arXiv'25, 2025.10
IQ130らしい(果たして):
GPQA DiamondでSoTA:
Jeff Dean氏によるポスト:
Grok 4.1, xAI, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #GenerativeAI #Blog #Proprietary #Selected Papers/Blogs Issue Date: 2025-11-18 Comment
元ポスト:
Awesome Spatial Intelligence in VLMs, mll-lab-nu, 2025.11
Paper/Blog Link My Issue
#Article #Survey #ComputerVision #MultiModal #Repository #VisionLanguageModel #SpatialUnderstanding Issue Date: 2025-11-18 Comment
元ポスト:
VLM, マルチモーダルなLLMにおけるSpatial Intelligenceに関する論文リスト
Third-Party Pangram Evaluations, Pangram., Destiny Akinode, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #GenerativeAI #Blog #text #AI Detector Issue Date: 2025-11-16 Comment
元ポスト:
[IBIS 2025] 深層基盤モデルのための強化学習 驚きから理論にもとづく納得へ, Akifumi Wachi, 2025.11
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel #ReinforcementLearning #Slide #Selected Papers/Blogs Issue Date: 2025-11-15 Comment
元ポスト:
ICLR 2026 - Submissions, Pangram Labs, 2025.11
Paper/Blog Link My Issue
#Article #Analysis #LanguageModel #Blog #ICLR #Selected Papers/Blogs #One-Line Notes #Reference Collection Issue Date: 2025-11-15 Comment
元ポスト:
ICLR'26のsubmissionとreviewに対してLLMが生成したものが否かをDetectionした結果(検出性能は完璧な結果ではない点に注意)
この辺の議論が興味深い:
関連:
oh...
パイプライン解説:
母国語でレビューを書いて英語に翻訳している場合もAI判定される場合があるよという話:
ICLR公式が対応検討中とのこと:
ICLRからの続報:
> As such, reviewers who posted such poor quality reviews will also face consequences, including the desk rejection of their submitted papers.
> Authors who got such reviews (with many hallucinated references or false claims) should post a confidential message to ACs and SACs pointing out the poor quality reviews and provide the necessary evidence.
citationに明らかな誤植があり、LLMによるHallucinationが疑われる事例が多数見つかっている:
Oralに選ばれるレベルのスコアの研究論文にも多数のHallucinationが含まれており、1人の査読者がそれに気づきスコア0を与える、といった事態にもなっているようである:
当該論文はdesk rejectされたので現在は閲覧できないとのこと。
NeurIPS'25ではそもそも査読を通過した研究についても多くのHallucinationが見つかっているとのこと:
ACL2025@ウィーン 参加報告, shirotaro, 2025.10
Paper/Blog Link My Issue
#Article #Tutorial #Blog #ACL Issue Date: 2025-11-15
SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds, Google DeepMind, 2025.11
Paper/Blog Link My Issue
#Article #ComputerVision #Blog #Reasoning #ComputerUse #VisionLanguageModel #3D (Scene) #Game Issue Date: 2025-11-14 Comment
元ポスト:
もはやAIがゲームをできるのは当たり前の時代だが、どのくらいOODに汎化するのかは気になる。
Holo2: Cost-Efficient Models for Cross-Platform Computer-Use Agents, H Company, 2025.11
Paper/Blog Link My Issue
#Article #ComputerVision #AIAgents #Blog #OpenWeight #ComputerUse #Selected Papers/Blogs #VisionLanguageModel #Grounding #GUI Issue Date: 2025-11-14 Comment
HF: https://huggingface.co/collections/Hcompany/holo2
元ポスト:
関連:
- Holo1.5 - Open Foundation Models for Computer Use Agents, H Company, 2025.09
GPT-5.1: A smarter, more conversational ChatGPT, OpenAI, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #ChatGPT #Blog #Reasoning #Proprietary #Selected Papers/Blogs #VisionLanguageModel #Routing #One-Line Notes #Reference Collection Issue Date: 2025-11-13 Comment
元ポスト:
instantモデルはよりあたたかい応答でより指示追従能力を高め、thinkingモデルは入力に応じてより適応的に思考トークン数を調整する。autoモデルは入力に応じてinstant, thinkingに適切にルーティングをする。
所見:
Artificial Analysisによるベンチマーキング:
GPT-5.1-Codex-maxの50% time horizon:
SYNTH: the new data frontier, pleias, 2025.11
Paper/Blog Link My Issue
#Article #Pretraining #Dataset #LanguageModel #SyntheticData #Reasoning #One-Line Notes Issue Date: 2025-11-12 Comment
元ポスト:
SoTAなReasoning能力を備えたSLMを学習可能な事前学習用合成データ
元ポスト:
Project AELLA: Custom LLMs to process 100 Million Research Papers, ssam Hogan, 2025.11
Paper/Blog Link My Issue
#Article #DocumentSummarization #LanguageModel #GenerativeAI #Blog #Science Issue Date: 2025-11-12 Comment
100M+の論文に対してAIによる要約を作成し構造化した上でvisualizeすることでよりscientificな情報へのアクセシビリティを高めたい、という話に見える
RL Learning with LoRA: A Diverse Deep Dive, kalomaze's kalomazing blog, 2025.11
Paper/Blog Link My Issue
#Article #Analysis #LanguageModel #ReinforcementLearning #Blog #PEFT(Adaptor/LoRA) #PostTraining #read-later Issue Date: 2025-11-10 Comment
元ポスト:
所見:
Lessons from the Trenches on Building Usable Coding Agents - Graham Neubig, Graham Neubig, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Coding #Video Issue Date: 2025-11-09 Comment
元ポスト:
Introducing Kimi K2 Thinking, MoonshotAI, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #Blog #Reasoning #OpenWeight #Selected Papers/Blogs #One-Line Notes #Reference Collection Issue Date: 2025-11-07 Comment
HF: https://huggingface.co/moonshotai
元ポスト:
coding系ベンチマークでは少しGPT5,Claude Sonnet-4.5に劣るようだが、HLE, BrowseCompなどではoutperform
tooluseのベンチマークであるtau^2 Bench TelecomではSoTA
モデルの図解:
INT4-QATに関する解説:
INT4-QATの解説:
Kimi K2 DeepResearch:
METRによる50% timehorizonの推定は54分:
ただしサードパーティのinference providerによってこれは実施されており、(providerによって性能が大きく変化することがあるため)信頼性は低い可能性があるとのこと。
METRでの評価でClaude 3.7 Sonnetと同等のスコア:
openweightモデルがproprietaryモデルに追いつくのはsoftwere engineeringタスク(agenticなlong horizon+reasoningタスク)9ヶ月程度を要しているとのこと
Mapping LLMs with Sparse Autoencoders, Hussein+, 2025.11
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel #Blog #One-Line Notes #SparseAutoencoder Issue Date: 2025-11-06 Comment
SparseAutoEncoderを用いた機械学習モデルの特徴の可視化方法に関するチュートリアル
OlmoEarth-v1-Large, Ai2, 2025.11
Paper/Blog Link My Issue
#Article #ComputerVision #FoundationModel #OpenWeight #2D (Image) Issue Date: 2025-11-06 Comment
元ポスト:
衛星画像で学習されたモデルらしい
進化する大規模言語モデル評価: Swallowプロジェクトにおける実践と知見, Naoaki Okazaki, 2025.10
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel #Evaluation #Slide #One-Line Notes Issue Date: 2025-11-02 Comment
元ポスト:
LLMの評価は些細な評価設定の違いで大きな変動が生じるだけでなく、事後学習済みモデルやreasoningモデルが主流になってきた現在では評価方法もアップデートが必要という話。たとえばreasoningモデルはfew-shotで評価すると性能が低下することが知られているなど。
Open-weight models lag state-of-the-art by around 3 months on average, EPOCH AI, 2025.10
Paper/Blog Link My Issue
#Article #Analysis #LanguageModel #Blog #OpenWeight Issue Date: 2025-11-01 Comment
タイトルの通りな模様
元ポスト:
LongCat-Flash-Omni Technical Report, 2025.10
Paper/Blog Link My Issue
#Article #ComputerVision #LanguageModel #SpeechProcessing #OpenWeight #MoE(Mixture-of-Experts) #2D (Image) #UMM #3D (Video) #Omni #audio #text Issue Date: 2025-11-01 Comment
元ポスト:
HF: https://huggingface.co/meituan-longcat/LongCat-Flash-Omni
text, image/video, audioをinputし、audioを生成するomniモデル
LLM-jp-3 and beyond: Training Large Language Models, Yusuke Oda, NII LLMC, 2025.10
Paper/Blog Link My Issue
#Article #Tutorial #Pretraining #LanguageModel #Slide #Japanese Issue Date: 2025-11-01 Comment
元ポスト:
The Smol Training Playbook: The Secrets to Building World-Class LLMs, Allal+, HuggingFace, 2025.10
Paper/Blog Link My Issue
#Article #Tutorial #Pretraining #Dataset #LanguageModel #Infrastructure #PostTraining #Selected Papers/Blogs Issue Date: 2025-10-31 Comment
元ポスト:
Emergent Introspective Awareness in Large Language Models, Jack Lindsey, Anthropic, 2025.10
Paper/Blog Link My Issue
#Article #Analysis #LanguageModel #Blog #Selected Papers/Blogs Issue Date: 2025-10-31 Comment
元ポスト:
公式ポスト:
Introducing Aardvark: OpenAI’s agentic security researcher, OpenAI, 2025.10
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #One-Line Notes #Security Issue Date: 2025-10-31 Comment
元ポスト:
> In benchmark testing on “golden” repositories, Aardvark identified 92% of known and synthetically-introduced vulnerabilities, demonstrating high recall and real-world effectiveness.
合成された脆弱性については92%程度検出できたとのこと。Claudeとかだとこの辺はどの程度の性能なのだろう。
