LanguageModel (3311) — 14/17
How we optimized Dash's relevance judge with DSPy, Dropbox, 2026.03
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #Prompting #AutomaticPromptEngineering #LLM-as-a-Judge #read-later #Initial Impression Notes Issue Date: 2026-04-07 Comment
元ポスト:
APEを使ってモデルを変更した際のプロンプト適応を効率化した話な模様。
Making RL Fast, Finbarr Timbers, 2026.04
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #ReinforcementLearning #SoftwareEngineering #PostTraining #Selected Papers/Blogs #reading #Initial Impression Notes #Asynchronous Issue Date: 2026-04-07 Comment
元ポスト:
Olmo3においてpost-trainingのインフラを同期から非同期に変更したことを含めて4倍高速化したことに関して、それをどのように実現したかに関するwrite up。気になる。
オープンソースAIの現状 | NVIDIA GTC, Nvidia, 2026.04
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #Video #OpenSource #read-later #One-Line Notes Issue Date: 2026-04-07 Comment
元ポスト:
GTCのパネルディスカッション
TurboQuant-Gpu, DevTechJr, 2026.04
Paper/Blog Link My Issue
#Article #NLP #Library #KV Cache #Compression Issue Date: 2026-04-07 Comment
元ポスト:
TurboQuant:
- TurboQuant: Redefining AI efficiency with extreme compression, Google Research, 2026.03
国産生成AI PLaMoを支える事後学習と推論最適化, PFN, 2026.04
Paper/Blog Link My Issue
#Article #Tutorial #NLP #Supervised-FineTuning (SFT) #ReinforcementLearning #ContextWindow #Quantization #PositionalEncoding #LLMServing #Slide #mid-training #DPO #PostTraining #GRPO #KV Cache #Compression Issue Date: 2026-04-07 Comment
元ポスト:
関連:
- PLaMo 3.0 Prime β版, PFN, 2026.03
関連:
- RoPE / YaRN
- [Paper Note] RoFormer: Enhanced Transformer with Rotary Position Embedding, Jianlin Su+, arXiv'21, 2021.04
- [Paper Note] YaRN: Efficient Context Window Extension of Large Language Models, Bowen Peng+, ICLR'24
- DPO
- [Paper Note] Direct Preference Optimization: Your Language Model is Secretly a Reward Model, Rafael Rafailov+, arXiv'23, 2023.05
- GRPO
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open
Language Models, Zhihong Shao+, arXiv'24
- RLはSFTよりも汎化性能に優れ、基本的には事前学習で獲得された能力を引き出す、という話
- [Paper Note] SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training, Tianzhe Chu+, ICML'25
- [Paper Note] Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?, Yang Yue+, NeurIPS'25, 2025.04
- JFBench: 実務レベルの日本語指示追従性能を備えた生成AIを目指して, PFN, 2026.01
- LLM Serving系
- [Paper Note] Efficient Memory Management for Large Language Model Serving with PagedAttention, Woosuk Kwon+, SOSP'23
- [Paper Note] GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar+, ICLR'23, 2022.10
- [Paper Note] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration, Ji Lin+, MLSys'24
- TurboQuant: Redefining AI efficiency with extreme compression, Google Research, 2026.03
うーーんおもしろかった!後でnote中の関連文献を紐づけてついでに復習したい
Components of A Coding Agent: How coding agents use tools, memory, and repo context to make LLMs work better in practice, Sebastian Raschka, 2026.04
Paper/Blog Link My Issue
#Article #Tutorial #NLP #AIAgents #Coding #SoftwareEngineering #read-later #Selected Papers/Blogs #Initial Impression Notes #AgentHarness Issue Date: 2026-04-05 Comment
LLM, Reasoning Model, Agent, Agent Harness, coding harnessなどの定義とその役割やスコープ、そしてそれらを構成するためのminimalなコンポーネントについて説明されており、基礎的な理解に役立ちそう。
元ポスト:
Emotion Concepts and their Function in a Large Language Model, Anthropic, 2026.04
Paper/Blog Link My Issue
#Article #Analysis #NLP #read-later #Selected Papers/Blogs #Emotion #Initial Impression Notes Issue Date: 2026-04-04 Comment
元ポスト:
これは非常に面白そうだ
Announcing 1-bit Bonsai: The First Commercially Viable 1-bit LLMs, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Blog #Proprietary Issue Date: 2026-04-04 Comment
元ポスト:
関連:
- [Paper Note] The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits, Shuming Ma+, arXiv'24, 2024.02
- [Paper Note] BitNet b1.58 2B4T Technical Report, Shuming Ma+, arXiv'25, 2025.04
圧倒的デコーディング速度:
Claude Code's Real Secret Sauce (Probably) Isn't the Model, Sebastian Raschka, 2026.04
Paper/Blog Link My Issue
#Article #NLP #Post #Architecture #read-later #AgentHarness Issue Date: 2026-04-04 Comment
関連:
- Claude Code's source code leaked through a `.map` file - How bad is it, really?, Chubby, 2026.04
How far does alignment midtraining generalize?, Tomek+, OpenAI Alignment Research Blog, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Alignment #mid-training #read-later #Initial Impression Notes Issue Date: 2026-04-04 Comment
元ポスト:
mid trainingにおいてalignment関してmisaligned/alignedな文書で学習をすると中間学習直後はalignmentに関する挙動が維持されるが、RLをしたらその効果は消えて無くなってしまう、という感じだろうか?超絶流し読みなので、後でしっかり読んだ方が良さそう。
Qwen3.5-Omni: Scaling Up, Toward Native Omni-Modal AGI, Qwen Team, 2026.04
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #SpeechProcessing #Proprietary #VisionLanguageModel #2D (Image) #3D (Video) #Omni #AudioLanguageModel #audio #text Issue Date: 2026-04-04 Comment
元ポスト:
llm-wiki.md, karpathy, 2026.04
Paper/Blog Link My Issue
#Article #Initial Impression Notes #Gist Issue Date: 2026-04-04 Comment
Karpathy氏によるLLMを利用して個人の生のドキュメントコレクションからwikiを構築するためのidea file。本paper_noteは自分の勉強のために手作業をすることで自身への知識の定着を狙っているけれども、自動構築したwikiがどのようなものになるかは興味があるなあ。
CuLA, InclusionAI, 2026.04
Paper/Blog Link My Issue
#Article #NLP #Library #Attention #SoftwareEngineering #One-Line Notes #GPUKernel #LinearAttention Issue Date: 2026-04-04 Comment
元ポスト:
Hopper(SM90), Blackwell(SM10X)において、flash-linear-attention(FLA)よりも最大2.45倍、平均1.52倍速いlinear attention kernelらしい
GPU Memory Math for LLMs (2026 Edition), Ahmad, 2026.04
Paper/Blog Link My Issue
#Article #NLP #SoftwareEngineering #Initial Impression Notes Issue Date: 2026-04-04 Comment
様々な量子化や浮動小数点フォーマット、パラメータ数やMoEの場合などにおける、VRAM消費量に関する考え方について解説されている
約12兆トークンの良質なコーパスで学習した新たな国産LLM「LLM-jp-4 8Bモデル」「LLM-jp-4 32B-A3Bモデル」をオープンソースライセンスで公開 ~一部ベンチマークでGPT-4oやQwen3-8Bを上回る性能を達成~, NII, 2026.04
Paper/Blog Link My Issue
#Article #Pretraining #NLP #Reasoning #OpenWeight #Japanese #OpenSource #mid-training #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-04-03 Comment
8BモデルはLlama-2アーキテクチャ、32B-A3.8BモデルはQwen3-MoEアーキテクチャで、フルスクラッチ学習をすることで実現[^1]。
19.5Tトークン(概算として、日本語0.7Tトークン、英語17.8Tトークン、中国語・韓国語0.85Tトークン、プログラムコード0.2Tトークン)のインターネット上の公開データや政府・国会の文書を収集し(LLM-jp-3.1のデータの6倍の規模)し事前学習データを構築、DataMixtureを最適化し10.5Tトークンを事前学習で利用。
中間学習では、事前学習データにInstruction Pretraining[^2]データを含む合成データを加え1.2Tトークンを利用。
その後最終的にInstruction Tuningを、日本語、英語合計22種類のデータで実施(元記事ではチューニングと呼称されているがおそらくInstruction Tuningだと思われる)。
MTBenchでは、GPT-4o, gpt-oss-20B, Qwen3-8Bと同等以上の性能、日本語MTBench[^3]では、GPT-4o, gpt-oss-20B, Qwen3-8Bを上回る性能とのこと。MTBenchで用いるLLM-as-a-JudgeのモデルとしてはGPT-5.4を利用とのこと。
[^1]: つまり、モデルのパラメータは完全に新規で学習されており、ベースとして既存OpenWeightモデルを利用していない点に注意。
[^2]: Instruction Pretrainingは、LLM-jp-3.1の頃から実施されている:
LLM-jp-3.1 シリーズ instruct4 の公開, LLM-jp, 2025.05
[Paper Note] Instruction Pre-Training: Language Models are Supervised Multitask Learners, Daixuan Cheng+, arXiv'24, 2024.06
[^3]: MT-Benchの概要については
[Paper Note] Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng+, NeurIPS'23, 2023.06
も参照のこと。
フルスクラッチモデル点に関する説明:
HF:
https://huggingface.co/collections/llm-jp/llm-jp-4-models
Reasoningモデルもある!!!
関連:
- PLaMo 3.0 Prime β版, PFN, 2026.03
上記PLaMo 3.0に続いて、国内でのフルスクラッチReasoningモデルは二例目だろうか。
Gemma 4: Byte for byte, the most capable open models, Google, 2026.04
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #AIAgents #MultiModal #SpeechProcessing #Reasoning #OpenWeight #MoE(Mixture-of-Experts) #Selected Papers/Blogs #VisionLanguageModel #2D (Image) #3D (Video) #One-Line Notes #Reference Collection #AudioLanguageModel #audio #text #Initial Impression Notes Issue Date: 2026-04-02 Comment
元ポスト:
2B, 4B, 26BのMoEモデルと31BのDenseモデルの4種類のモデルファミリーで、マルチモーダル(vision)対応。2B, 4Bはaudioも入力として扱える。
edgeデバイス向けのモデルは128k, 他は256kのコンテキストウィンドウ。140+の多言語サポート。
Apache 2.0ライセンス
arenaで同サイズのモデル群でSoTAといった話がブログ中に記述されている。
モデルカードには一般的なベンチマーク群とのスコアも記載されている。
https://ai.google.dev/gemma/docs/core/model_card_4?hl=ja
(そもそも既存のベンチマークにもコンタミネーションがあると思われるが、)arenaに関しては特定の企業に対してデータを提供し、複数のモデルの亜種をテストできるという慣行があり、リーダーボードにバイアスがあるであろう点には注意:
- [Paper Note] The Leaderboard Illusion, Shivalika Singh+, NeurIPS'25
artificial analysisによる評価:
Qwenがproprietaryになったことから、ライセンス的に使いやすく、日本語に強そうなモデルとしては筆頭ではなかろうか。日本語性能が気になる。
アーキテクチャ解説:
ポイント解説:
所見:
attentionのscaleをsqrt(d)でスケールさせる代わりに、QK-norm, V normを適用するなど。
NvidiaによるNVFP4へのpost-trainingによる量子化:
https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4
量子化後の性能も比較されており、知識、数学、コーディング、terminac useなど6種類のベンチマークでオリジナルのモデルと遜色ない性能が出ている旨記載されている。
解説:
https://newsletter.maartengrootendorst.com/p/a-visual-guide-to-gemma-4
所見(encoder-freeにした裏側でパッチ化→projection + x/y軸のpositional embeddingを実施している話):
Trinity-Large-Thinking: Scaling an Open Source Frontier Agent, Arcee, 2026.04
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Reasoning #OpenWeight #MoE(Mixture-of-Experts) #read-later #Selected Papers/Blogs Issue Date: 2026-04-02 Comment
元ポスト:
HF: https://huggingface.co/collections/arcee-ai/trinity-large-thinking
Qwen3.6-Plus: Towards Real World Agents, Qwen Team, 2026.04
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Proprietary #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-04-02 Comment
元ポスト:
Opus 4.6相当のベンチマークスコアがありそうだが、プロプライエタリモデル化
LFM2.5-350M: No Size Left Behind, Liquid AI, 2026.04
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #SmallModel #OpenWeight #read-later #Selected Papers/Blogs #KeyPoint Notes Issue Date: 2026-04-01 Comment
元ポスト:
- LFM2のアーキテクチャを採用の350Mパラメータモデルで、CPUでも十分な速度で推論可能
- 追加の事前学習(10T -> 28T tokens)、および、large-scale RLを実施
- 同等規模のパラメータ数(あるいは2倍程度)のモデル群に対して、知識, 指示追従能力, ツール呼び出し、データ抽出などのベンチマークで上回る
- LFM2-350Mと比較して、指示追従能力, データ抽出, tool useの性能が大きく向上
- edgeデバイスでの軽量なデータ抽出パイプラインとして有用
- しかし、math, coding, creative writingなどでの利用は推奨されない
- CPU/GPUでの推論ともに同等規模、あるいは1B級のモデルよりも早く、省メモリ
OneCompression, FujitsuResearch, 2026.04
Paper/Blog Link My Issue
#Article #NLP #Library #Quantization #One-Line Notes Issue Date: 2026-04-01 Comment
元ポスト:
example_autorun.pyを見るとわかるが、ワンライナーで(post-training basedな)量子化をしたいモデルのコンフィグを渡して実行するだけで、自動的にGPTQ量子化など(DBF, RTNと呼ばれる方式もあるようだ)をしてくれるライブラリのようである。現在はLlama, Qwen3をサポートしており、今後も適用可能なモデルは拡張していく予定と書かれている。また、量子化したモデルはvLLMとの互換性も担保される。
サポートされているアルゴリズムはこちらにまとまっていそう:
https://fujitsuresearch.github.io/OneCompression/algorithms/overview/
calibration dataは何が用いられるのだろうか?
関連:
SmolLM - blazingly fast and remarkably powerful, Allal+, HuggingFace, 2024.07
Paper/Blog Link My Issue
#Article #NLP #OpenWeight Issue Date: 2026-03-31 Comment
OpenSourceなLLMについて過去を遡ってみているが、SmolLMの最初の段階では、データのみがオープンでコードはオープンでないように見える。
次:
- SmolLM2, 2024.11
RedPajama, a project to create leading open-source models, starts by reproducing LLaMA training dataset of over 1.2 trillion tokens, together.ai, 2023.04
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #OpenSource #One-Line Notes Issue Date: 2026-03-31 Comment
完全なオープンソースLLMの構築を目指すprojectで、LLaMAの学習データを再現する取り組み。
ParaGator: Learning to Aggregate through Online RL, Li+, 2026.03
Paper/Blog Link My Issue
#Article #NLP #ReinforcementLearning #Test-Time Scaling #Diversity #Aggregation-aware #Initial Impression Notes Issue Date: 2026-03-30 Comment
元ポスト:
関連:
- [Paper Note] Reasoning over mathematical objects: on-policy reward modeling and test time aggregation, Pranjal Aggarwal+, arXiv'26, 2026.03
上記研究のSection 3の内容っぽい?
解候補を生成する際はPass@kに対して最適化をし多様な候補の生成を促し、解候補を集約してFinal Answerを導出する際には、Pass@1に対して最適化をし複数の解候補を効果的に集約する方向に最適化することで、性能がブーストされ、それをend-to-endに実現する、という話にみえる。
- [Paper Note] PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning, Jingcheng Hu+, arXiv'26, 2026.01
と似たような考え方に見える。
The Anatomy of an LLM Benchmark, Cameron R. Wolfe, Ph.D., 2026.03
Paper/Blog Link My Issue
#Article #Tutorial #NLP #Dataset #Evaluation #Initial Impression Notes Issue Date: 2026-03-30 Comment
元ポスト:
本文中のDisclaimerにも記述されている通り、coding/SWE/AI-Agentに関するベンチマークや、Multi-modalなベンチマークについては説明されていない点には注意。
LLMとしての評価として初期の頃から使われて(いる|いた)、MMLU, GPQA, BIG-Bench, IFEvalといった代表的なものが紹介され、単にそれらベンチマークがどういったものかを説明しているというより、どのようにすれば自分たちのタスクに関して良いLLMベンチマークを作成できるか?という観点で議論されているように見える。
Why aren't we fine-tuning more?, Nate Meyvis, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Supervised-FineTuning (SFT) #ReinforcementLearning #Blog #PostTraining #Finetuning Issue Date: 2026-03-30 Comment
元ポスト:
なぜFinetuningは普及していないのか?という点を考察しているブログ。ざっくり言うと、「コストの割に合わない」ということであり、具体的には
- Finetuningをしなくてもprompt engineeringで十分な性能が出てしまい
- Finetuningをしなくても、ドメイン固有のツールを組み合わせることでドメインspecificな挙動が実現できたり
- Finetuningを実施すると、新たなモデルが利用可能になった場合に再度Finetuningを実施するなどのオーバヘッドが生じ割に合わない
といった話が書かれている。個人的にはさらに言うと
- Finetuningを実施することでAI Safety周りの懸念が生じてしまい、Safetyに関する評価を厳密には実施しなければならない(特に何らかのチャットベースの応用の場合はなおさら)
というのもあると感じており、このモデルは安全ですと顧客にどのように説明するのか?という新たな説明責任が生じるという点もあるのかなと思う。
しかし、やはりFinetuningはあまり普及していないんだなあ、感
How Kimi, Cursor, and Chroma Train Agentic Models with RL, PHILSCHMID, 2026.03
Paper/Blog Link My Issue
#Article #NLP #ReinforcementLearning #AIAgents #Blog #read-later #reading #LongHorizon Issue Date: 2026-03-29
Introducing Marin: An Open Lab for Building Foundation Models, marin-community, 2025.05
Paper/Blog Link My Issue
#Article #Pretraining #NLP #Blog #OpenWeight #OpenSource #Selected Papers/Blogs Issue Date: 2026-03-29 Comment
github:
https://github.com/marin-community/marin
issueのExperimentsが興味深い
関連:
- Marin 32B Retrospective, marin-community, 2025.10
Marin projectのアナウンスをメモっていなかったので今更ながらメモ
- open-weight, open-sourceを超えて、LLMのopen-developmentを実現するための完全な透明性を持ったopen lab
- すべての実験はgithub issueで管理され公開される
- marinのコードベースを使い誰でも実験をコード中に記述しpull repuestを送れ、誰でもレビューできる
- プルリクが承認されると実験が実際に実行され、誰でもWandB上の経過をリアルタイムで観察できる
Delphi[^1]の実験において、25Bパラメータモデルがweight decayフェーズに突入し、Marin-32Bでは以前はweight decayフェーズでloss spikeが頻発したが、Delphiでは安定していそうな見込み、という話がポストされている:
[^1]: 現代版のPythiaを構築しましょうという話で、Pythiaのモデルパラメータを70Bまでスケールアップし、学習に用いるトークン数もチンチラ則従いモデルサイズに応じてスケールアップ、The PileデータなどのデータセットをNemotron-CCなどのlarge scaleモデル用のデータセットに置換する、といった話が含まれる。Marin Issue 1337を参照のこと。
129B-A16Bの学習を開始したとのこと:
535B-A23Bモデルの学習を開始したとのこと:
リアルタイムRLでComposerを改善する, Cursor, 2026.03
Paper/Blog Link My Issue
#Article #NLP #ReinforcementLearning #Blog #Coding #SoftwareEngineering #KeyPoint Notes #Realtime Issue Date: 2026-03-28 Comment
実際の推論トークンとユーザの応答を集約して報酬を作成しモデルの改善に使うリアルタイムRLによって5時間ごとにComposerチェックポイントをアップデートしデプロイする。
Reward Hackingを防ぐことはこのようなリアルタイムRLではより一層重要でそのための報酬設計として工夫した点が2つ挙げられている。
- 元々はツール呼び出しが無効だった例を除外するようにして報酬を設計していたが、モデルはこれにより無効なツールを呼び出せば負の報酬を得ないことを学び意図的に無効なツールを呼び出すことを学習した。これを防ぐために、ツール呼び出しに失敗した場合に明確に負の報酬を与えるように変更
- モデルが実施した編集について、自分がコードを編集しなければペナルティを受けないことを学習し、難しい編集については質問をすることで先送りする挙動をRewardHackingの結果学習した。質問については適切なタイミングで実施する必要があるため、報酬を修正した
といった話が書かれている。
現在は比較的短いタスクを実行してユーザからフィードバックを受け取れるが、今後はlong horizonなタスクを実行することが予想され、その場合
- ユーザのフィールドバックの頻度は減り
- 成果物全体に対するフィードバックを返すようになる
という異なる性質のデータを扱わなければならないのでそれに向けて改善を進めるとのこと。
ソフトウェア開発エージェント 初歩から上級, Graham Neubig, 2026.03
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Coding #SoftwareEngineering #read-later Issue Date: 2026-03-26 Comment
全体をざっくり概観してイメージをつかむのに良さそう。詳細を知りたい場合はリンク先を見ると良さげ。
(スライド最後の強化学習における「3」のスケーリングってなんだろう...?)
元ポスト:
Introducing SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding, Nvidia, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Dataset #Evaluation #LongContext #Decoding #Diversity #SpeculativeDecoding Issue Date: 2026-03-26 Comment
元ポスト:
Here are the 2025 AI safety papers and posts I like the most, Fabien Roger, LW, 2026.03
Paper/Blog Link My Issue
#Article #Survey #NLP #Safety #read-later #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-03-26 Comment
元ポスト:
AI Safetyに関する研究者の方の2025年のAI Safetyハイライトとのこと。
Emergent Misalignmentなど以外にも多くの研究に⭐︎︎︎⭐︎⭐︎が付与されている。気になる。
A Visual Guide to Attention Variants in Modern LLMs, Sebastian Raschka, 2026.03
Paper/Blog Link My Issue
#Article #Tutorial #NLP #Attention #Blog #read-later Issue Date: 2026-03-26
One interesting dynamic in AI is infra+application co-dependence, Graham Neubig, X, 2026.03
Paper/Blog Link My Issue
#Article #Post Issue Date: 2026-03-26
TurboQuant: Redefining AI efficiency with extreme compression, Google Research, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Blog #Reference Collection #KV Cache #Compression #Initial Impression Notes Issue Date: 2026-03-25 Comment
元ポスト:
kv cacheをlong contextで1/6に圧縮して、8倍スピードアップして、accuracyのlossがない圧縮技術とのこと。果たして
たまたまこの動画を見つけたがおそらくこの研究のことを行っているのだろう:
https://youtube.com/shorts/5LMoZjoprQc?si=C43dJuXqpAa-p4BP
不要な逆量子化処理を省くことで高速化可能らしい:
Vibe physics: The AI grad student, Anthropic, 2026.03
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Blog #ScientificDiscovery #Physics #AI-Human Co-Improvement #Human-in-the-Loop Issue Date: 2026-03-25 Comment
元ポスト:
最大規模のオープン基盤モデルを各国仕様へ適応させる事後学習技術を開発, sakana.ai, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Alignment #Blog #Bias #Japanese #PostTraining #Reading Reflections Issue Date: 2026-03-24 Comment
技術的な詳細は不明で、
> 事後学習では、日本の文化的・社会的文脈におけるバイアス是正のための独自データセットを構築し、以下のベンチマークに示す結果を得ました。
と記述されている。おそらく構築したデータセットに基づいてAlignmentをとるための事後学習(ベースモデルの能力を落としていないため Catastrophic Forgettingは起きておらず、同社がLoRA系の技術に力を入れていることを鑑みるとおそらく何らかのPEFT手法ではないかと推察)を実施しているのだと思われる。
元ポスト:
Continue Pre-training can only work with "actual Base model"., Wenhu Chen, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Post #mid-training Issue Date: 2026-03-22 Comment
こちらのポストに様々な理由が言及されており勉強になる:
Xiaomi MiMo-V2-Pro, Xiaomi, 2026.03
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Proprietary #Author Thread-Post Issue Date: 2026-03-21 Comment
元ポスト:
THE CONSCIOUSNESS CLUSTER: PREFERENCES OF MODELS THAT CLAIM TO BE CONSCIOUS, Chua+, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Alignment #Safety #read-later #Initial Impression Notes Issue Date: 2026-03-20 Comment
元ポスト:
LLMに意識があるように振る舞うように学習したらどうなるかという話らしい。これによって新たなpreferenceが獲得され、自己保存欲求や反発が発現したり、共感や葛藤などの人間的な感情について話したり、思考過程をモニタリングされることをどう感じますか?といった質問に対して、uncomfortableだと感じる、私は悪い評価を受けたら停止されてしまうの?といった不安について述べたりするなど、これまでにない挙動が見受けられるという感じらしい。
MiroThinker-1.7, MiroMindAI, 2026.03
Paper/Blog Link My Issue
#Article #NLP #AIAgents #OpenWeight #DeepResearch #LongHorizon #Initial Impression Notes Issue Date: 2026-03-20 Comment
元ポスト:
ベンチマークに応じて、GPT-5, GPT-5.2, GPT-5.4など比較するGPTが恣意的に変わっているように見えるが、ベンチマーク上ではGPT-5と同等以上のAgenticなLLMっぽい?BrowseCompの性能がかなり良さそうに見える。
LLM Architecture Gallery, Sebastian Raschka, 2026.03
Paper/Blog Link My Issue
#Article #Survey #NLP #Transformer #Blog #OpenWeight #Architecture #Initial Impression Notes Issue Date: 2026-03-20 Comment
元ポスト:
Sebastian Raschka氏がいつもポストしているOpenWeight LLMのアーキテクチャ図のギャラリー。パラメータサイズ, head数などの細かい情報も含まれているので、全体を概観するのに良さそう。
Reinforcement Learning from Human Feedback, Nathan Lambert, 2026.03
Paper/Blog Link My Issue
#Article #Tutorial #NLP #ReinforcementLearning #Blog #PostTraining #read-later Issue Date: 2026-03-20 Comment
元ポスト:
REINFORCE, PPO, GRPOの気持ちを理解するのに有用という所見:
Composer 2 のご紹介, Cursor, 2026.03
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #ReinforcementLearning #AIAgents #Evaluation #Coding #SoftwareEngineering #mid-training #PostTraining #read-later #Selected Papers/Blogs #ContextEngineering #Live #Reference Collection #Initial Impression Notes Issue Date: 2026-03-20 Comment
元ポスト:
所見:
Kimi-K2.5がベースらしいとのこと:
ベンチマークスコアに対する所見:
テクニカルレポートが出た:
https://cursor.com/resources/Composer2.pdf
元ポスト:
Kimi-K2.5をベースに、どのようにinstruction tuning後のモデルに対して継続事前学習、RLをし、GPT-5.4(high)級の性能を達成できたのか、ヒントがわかるかもしれない。
- [Paper Note] Kimi K2.5: Visual Agentic Intelligence, Kimi Team+, arXiv'26, 2026.02
所見:
所見:
RLによってpass@k(best-of-16)とpass@1の両方が改善する。既存研究では少なくともRLVRを用いた場合はPass@1は改善するが多様性が損なわれてPass@kの性能は改善しない ([Paper Note] Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR, Xiao Liang+, arXiv'25, 2025.08 , VibeVoice-1.5B, microsoft, 2025.08 )、という話があったが、Composer 2のレシピではそうではないようだ。どんなレシピだろう~と思ってさらっと関連しそうなところを見てみたが、詳細は書いてなさそうだ。
- [Paper Note] Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR, Xiao Liang+, arXiv'25, 2025.08
- VibeVoice-1.5B, microsoft, 2025.08
QA:
CursorBenchの解説:
要はrealisticなデータとシチュエーションでの評価に非常に重きを置いていて
- 実際のコーディングsessionのデータが用いられ、contamination-free
- 機能的な正しさのみならず、コードの品質、効率、挙動などの実用的な価値を意識し
- long horizonなタスクが多く取り入れられ
- Promptは曖昧性をうまく扱えるかを評価するために意図的にシンプルで短く
- CursorBenchのデータは継続的に更新される
- realisticなsessionデータだけでなく、その他の重要な挙動の評価(e.g., 指示追従, ルール/skilltのハンドリング, コメントの品質, editするか否かの判断の適切性など)のためのデータでも拡張されている
という感じらしい
ポイント解説:
- How Kimi, Cursor, and Chroma Train Agentic Models with RL, PHILSCHMID, 2026.03
self-summarizationによるcontextのcompressionを実施している
- [Paper Note] InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning, Yuchen Yan+, arXiv'26, 2026.02
- [Paper Note] Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL, Ian Wu+, arXiv'26, 2026.02
- より長いホライズンに向けた Composer の学習, Cursor, 2026.03
所見:
PLaMo 3.0 Prime β版, PFN, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Blog #Reasoning #Japanese #Selected Papers/Blogs #In-Depth Notes #Surface-level Notes Issue Date: 2026-03-19 Comment
元ポスト:
日本国内初のフルスクラッチReasoningモデル
## 公式発表のまとめ
- [Paper Note] YaRN: Efficient Context Window Extension of Large Language Models, Bowen Peng+, ICLR'24
によってcontext windowを64Kまで拡張(PLaMo 2.2 Primeの2倍)。
事後学習データの見直し(新たなオープンデータセット追加, 独自データとして、日本語指示追従能力, tool use, long horizon QA, 医療分野, STEM, RAG性能向上のためのデータ)を実施し、SFT, DPO, RLの流れで学習を実施。SFT, DPOについてはreasoning trajectoryもLossで考慮するように変更。SFT, DPO向けデータについてはreasoning trajectoryを合成したものを利用。
RLは今回初めて導入し学習を安定させるための工夫を取り入れているとのこと。Reference Answerとの比較と表層的な特徴から報酬を計算する関数を実装した、という書かれ方をしている。
gpt-oss-120B(memium)との比較で言うと
指示追従性能が日本語、英語ともによりも高く、医療分野のQA(国家試験を除く)、英語、日本語での対話能力で勝っている。また、法令分野のQAは同等である。
単一ツールや複数ツールからの選択は同等、multi turnの場合はPLaMo2.2から大幅に性能向上しているもののgpt-ossよりも劣る。また、long contextのQA、医療分野の国家試験QA、STEM分野のQAや数学的な推論能力は大幅に前回モデルよりも向上したが、まだgpt-ossなどには届いていない、という感じに見える。
アーキテクチャについては、一新したという話とRoPEベースということ以外はよくわからない。
## 筆者の憶測と感想
※以下、筆者の憶測を多く含んだ感想です。ただ筆者が勝手に想像して自分なりに考えてみているだけです。
DPOにNLL lossを追加することでreasoningを強化できることは下記研究で示されている:
- [Paper Note] Iterative Reasoning Preference Optimization, Richard Yuanzhe Pang+, NeurIPS'24, 2024.04
RLの報酬に関して、表層的な特徴とReference Answerとの比較から最適な報酬を計算とのことなので、おそらく何らかのVerificationのための仕組みと、Rubric-basedなLLM-as-a-Judgeだろうか?Reward Modelという書かれ方はしていない。
RLについては安定性のある手法を採用したとのことだが、DAPO、
- [Paper Note] DAPO: An Open-Source LLM Reinforcement Learning System at Scale, Qiying Yu+, NeurIPS'25
あるいはRLのスケーリング則を導いた研究でDAPOよりも安定性と最終到達性能において優れていることが示された
- [Paper Note] The Art of Scaling Reinforcement Learning Compute for LLMs, Devvrit Khatri+, arXiv'25, 2025.10
CISPOあたりだろうか:
- [Paper Note] MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning
Attention, MiniMax+, arXiv'25, 2025.06
あとは安定性という観点で言うと、inference/trainingエンジンでのtraining-inference gapの課題についても対処している可能性がある。
- Hot topics in RL, Kimbo, X, 2025.12
- [Paper Note] Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It, Yaxiang Zhang+, arXiv'26, 2026.02
思考過程が英語ということは、言語間で能力は転移し、かつ事前学習データとしてはリソースが豊富な英語が多く含まれると想像すると、明示的(strong LLMでtrajectoryを合成したものを加える系の話)あるいはデータに自然と現れるreasoningの挙動から事前学習中にreasoning能力が暗黙的に学習されることを踏まえ、SFTでreasoning能力を強化する際に(日本語よりも英語の方が効果的な可能性が高く)英語でのtrajectoryを合成したという感じだろうか(いつか日本語のreasoning trajectoryを出力するモデルも見てみたいなあ)。
Multi Turnのtool useの性能向上に関して、AI Agent分野のlong horizonな合成データを合成するアプローチや、Sink Tokenの活用や、トークン単位でsink tokenを計算することに相当するHead wise gated attentionなどはしているのだろうか。
- [Paper Note] Efficient Streaming Language Models with Attention Sinks, Guangxuan Xiao+, ICLR'24
- [Paper Note] Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters, Ailin Huang+, arXiv'26, 2026.02
また、アーキテクチャに関してはcontext windowが海外のフロンティアモデルと比較してまだ小さめであるが、今後context windowを大きくするにあたって、オンポリシーRLでのロールアウト時間がボトルネックとなることが考えられ、Mamba(=linear attention)系のアーキテクチャをハイブリッドや、DSA系のsparse attentionなどの採用によるアーキテクチャ起因の計算コスト低減(現在どのようなアーキテクチャなのかは全くわからないが)、あるいはin-flight-updateのような学習エンジン側での効率化なども必要になるのではなかろうか(現在どういうエンジンなのかは全くわからないが)。
- [Paper Note] DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models, DeepSeek-AI+, arXiv'25, 2025.12
- [Paper Note] PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence
Generation, Alexandre Piché+, arXiv'25, 2025.09
MiniMax-M2.7, MiniMax, 2026.03
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #Selected Papers/Blogs #Reference Collection Issue Date: 2026-03-19 Comment
所見:
所見:
Artificial Analysisによる評価:
GLM-5と同等の知能スコア、GDPvalでGPT-5.2(xhigh)超え。
modelがオープンに:
https://huggingface.co/MiniMaxAI/MiniMax-M2.7
元ポスト:
openになったが商用利用は許可を得ないとできないということで、リリース時のポストにはopennsourcedと銘打たれているが、open sourceではない。
中国系のOpenModelのライセンス、あるいはプロプライエタリ化が進んできている?
所見:
GPT‑5.4 mini と nano が登場, OpenAI, 2026.03
Paper/Blog Link My Issue
#Article #NLP #ChatGPT #Proprietary #VisionLanguageModel Issue Date: 2026-03-18 Comment
元ポスト:
Artificial Analysisによるベンチマーク:
楽天、「GENIACプロジェクト」の一環として開発された国内最大規模の高性能AIモデル「Rakuten AI 3.0」を提供開始, 楽天グループ株式会社, 2026.03
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #Japanese #MoE(Mixture-of-Experts) #Initial Impression Notes Issue Date: 2026-03-18 Comment
HF: https://huggingface.co/Rakuten/RakutenAI-3.0
公式アナウンス、HFのモデルカードの情報が少なすぎてよくわからない。
所見:
Mistral Forge: Build your own frontier models, MistralAI, 2026.03
Paper/Blog Link My Issue
#Article #NLP #ReinforcementLearning #Blog #Proprietary #mid-training #PostTraining #Data Issue Date: 2026-03-18 Comment
元ポスト:
エンタープライズ向けの社内の機密データによってLLMの(おそらく継続)事前学習、事後学習、RLを実施したカスタムモデルを構築するソリューションのようである。Dense, MoEなどのアーキテクチャも選択可能な模様。
ベースモデルなどが書かれていないように見えるが、Mistral製のオープンLLMがベースとなるのだろうか。
5 Agent Skill design patterns every ADK developer should know, Google Cloud Tech, X, 2026.03
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Post #SoftwareEngineering #Selected Papers/Blogs #One-Line Notes #AgentSkills Issue Date: 2026-03-18 Comment
Agent Skillsの定義の仕方による性能差については下記を参照のこと:
- [Paper Note] SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks, Xiangyi Li+, arXiv'26, 2026.02
以下の5つのPatternが紹介されている:
- Tool Wrapper
- Generator
- Reviewer
- Inversion
- Pipeline
最終的にどのようなPatternを採用すべきかの判断となるフローチャートも提供されている。
全体的なポイントとしては、
- 各種SKILLS.mdにはhowを記述し(e.g., 具体的な実行のstepを記述するなど)、
- 実行内容やルールなどの"what"に関する情報は別のドキュメントに移譲し、SKILLS.mdにはそのポインタを記述する、
- ユーザの承認なしで先へ進まないようにするには、ユーザに何らかの質問・承認を求めるよう指示を明示的に記述する
といった作法である。一つの巨大で複雑なSKILLS.mdやsystem promptを作るのではなく、内容をbreak downして記述やドキュメントの構造を設計するのが肝要と感じる。
他の参考文献として
-
# Writing a good CLAUDE.md, Kyle, 2025.11
はAGENTS.mdの話だが、同じような議論がされており、なぜless is moreが重要なのかといった説明も研究動向を踏まえながら説明されている。
Mistral Small 4, MistralAI, 2026.03
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #MoE(Mixture-of-Experts) #Initial Impression Notes Issue Date: 2026-03-17 Comment
元ポスト:
119Bでsmallと銘打たれる時代になってしまった
公式ポスト:
From REINFORCE to Dr. GRPO, Qingfeng's blog, 2025.03
Paper/Blog Link My Issue
#Article #Tutorial #NLP #ReinforcementLearning #PostTraining Issue Date: 2026-03-17 Comment
元ポスト:
State of RL for reasoning LLMs, A. Weers, 2026.03
Paper/Blog Link My Issue
#Article #Survey #NLP #ReinforcementLearning #Blog #PostTraining Issue Date: 2026-03-17 Comment
元ポスト:
OpenMAIC, THU-MAIC, 2026.03
Paper/Blog Link My Issue
#Article #Multi #Tools #NLP #Education #AdaptiveLearning #AIAgents #Repository #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-03-17 Comment
マルチエージェントによってスケーラブル、adaptiveにオンライン教育を実現するフレームワークのようである
元ポスト:
L11: Synthetic Data Powering Pretraining, Eric W. Tramel, Ph.D., UC Berkeley EE 290_194-11: Scalable AI, 2026.02
Paper/Blog Link My Issue
#Article #Pretraining #NLP #SyntheticData #read-later #Selected Papers/Blogs #KeyPoint Notes #Reading Reflections Issue Date: 2026-03-17 Comment
元ポスト:
- インターネットのデータ枯渇問題が指摘されながらも、合成データによって事前学習は進化を続けている
- LLMは事後学習で性能を向上させられるが、事前学習時点で伸ばせる上限が決まっているとされている
- 事前学習データの投入量はChinchilla則のパラメータ量の20倍から現在は60倍まで増加
- MoEは過学習しやすくパラメータ数の40倍は必要
- 学習データの多様性が重要で繰り返し同じデータを見ても性能は改善しない
- 合成データをそのまま用いるとmode collapseが生じ出力が単調化するため、実データを混ぜるか言い換えをしたデータで是正する(弱めのdata augmentationで良い)
- 最近重要な合成データはコードと推論過程を含むデータで、これらが事前学習データに含まれていると汎用な表現、思考能力、推論能力を事前学習時点から獲得できる可能性がある
というような話が元ポストに書かれている。
- [Paper Note] Scaling Data-Constrained Language Models, Niklas Muennighoff+, NeurIPS'23
のようにrepetitionは4回までが効果的といった知見が報告されているが、現在はどこまで当てはまるのだろうか?
後ほど関連するissueのリンクを貼りたい
うーんおもしろそう、p.15, p.20, p.26, p.28, p.35, p.36 あたりが気になる。
てかこれが大学の講義...?楽しすぎでは。
NOUMENA, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Blog #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-03-15 Comment
元ポスト:
関連:
- Why Training MoEs is So Hard, _xjdr, X Post
おそらく上記ポストの方の作業ログに関するブログと思われる。Canon Layer, mHC, Engramの再現、MoEのエキスパートは異なる学習率が必要なのか?、RDEPと呼ばれるアーキテクチャ(MoEアーキテクチャを採用するとexpertsがしばしば異なるGPUに割り当てられ、routingが特定のexsertsに偏るため特定のGPUがアイドルしてる時間が長くなるため効率が悪いというボトルネックをNVLinkがひもづくネットワーク全体に対してexpertsに対して送信するトークンを収集しパッチを作って送信することで効率を改善する、といったアプローチらしい?)のスループットとメモリ節約効果など、最新の生の知見が数多くまとまっているらしい。
A2UI, google, 2026.03
Paper/Blog Link My Issue
#Article #Tools #NLP #AIAgents #SoftwareEngineering #One-Line Notes #UI Issue Date: 2026-03-15 Comment
元ポスト:
AgentがUIを表現するための標準的なライブラリ群で、agentから応答されるjsonをクライアント側のライブラリでrenderingすることでUIがレンダリング可能というものらしい。
UIはコンポーネントのリストで表現されるためユーザのリクエストに応じてincrementalにUIを変化させる といったことが可能とのこと。
Claude now creates interactive charts, diagrams and visualizations, Claude, 2026.03
Paper/Blog Link My Issue
#Article #NLP #TextToImageGeneration #Proprietary #Reference Collection #Initial Impression Notes #Visualization Issue Date: 2026-03-14 Comment
かなり良いらしい(小並感)
元ポスト:
たとえばMLAとDSAの図解を作らせたら以下:
MuonとAdam(W)の違いの解説を作らせたら以下:
NVIDIA Nemotron 3 Super, NVIDIA, 2026.03
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #SSM (StateSpaceModel) #OpenSource #MoE(Mixture-of-Experts) #read-later #Selected Papers/Blogs #KeyPoint Notes #Reference Collection #Hybrid #LowPrecision #LinearAttention Issue Date: 2026-03-12 Comment
元ポスト:
解説:
artificial analysisによる評価:
Swallow LVM Leaderboardに性能が掲載:
解説:
アーキテクチャ:
- NVFP4で学習して gpt-ossより2.2倍高速だが性能も向上
- 88 Layer: 40 Latent MoE / 40 Mamba-2 / 8 GQA Attention
- GQA Attentiom Layerは非常に少なく、ほとんどがMamba-2 (linear attention)となっている
- Latent MoEは入力をそのまま変換するshared expertsと、入力を1/4のlatent vectorに変換した潜在空間上で処理をするLatext expertsの組み合わせによって出力を得る。
- 具体的には、RouterによってTop-22のexpertsを選択し、inputを1/4のlatent vectorに圧縮した上でExpertsに入力。Expertsの出力を加算して4倍のvectorに変換し次元を戻して、別ルートでshared expertsに元の入力次元から変換されたベクトルと組み合わせて出力するようなアーキテクチャ
Latent MoE解説:
要はMoEに必要なmatrixが、latent vectorを扱うことで小さくなるのでMoEのWeightのメモリロードのボトルネックが緩和されるだけでなく、
各MoE Laverは異なるGPUやマシンに分散されて配置されるため計算のためにはベクトルのバッチを通信しなければならないがそのコストが削減されスループットの向上につながるので嬉しい、ということだと思われる。
ポイント解説:
technical reportが出た:
- [Paper Note] Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning, NVIDIA+, arXiv'26, 2026.04
Using NVFP4 Low-Precision Model Training for Higher Throughput Without Losing Accuracy, NVIDIA, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Blog #read-later #LowPrecision Issue Date: 2026-03-12 Comment
元ポスト:
Bringing Code Review to Claude Code, Anthropic, 2026.03
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Coding #SoftwareEngineering #read-later Issue Date: 2026-03-12 Comment
元ポスト:
コードレビューに特化した機能が追加された模様
Anthropic社内で運用済みで、エンジニアがコードレビューに誤りがあると判断したものは<1%とのこと。
autoresearch, karpathy, 2026.03
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Repository #SelfImprovement #ScientificDiscovery #Selected Papers/Blogs #One-Line Notes #autoresearch/RSI Issue Date: 2026-03-10 Comment
元ポスト:
リポジトリのDiscussionsに、定期的にsession reportがアップロードされるようだ:
https://github.com/karpathy/autoresearch/discussions/43
nanochatは現在、126回の実験を経て、Validation BPBが0.997900 -> 0.969686 まで改善しているとのこと。
pjの目的やテーマは、**研究者がpythonファイルのコードをいじるのではなく、program.mdと呼ばれるAgentにコンテキストとして与えるmarkdownファイルのみの編集を通じて、研究組織(≠単一のPh.D student)をエミュレートできるか?** という点にありそうである。
https://github.com/karpathy/autoresearch/blob/master/program.md
その題材の一つとして、nanochatを簡略化したGPTを用いて、GPTの事前学習の性能を改善させるようなtraining.pyの編集をAI Agentsに実施させ、5分間学習させて成果を報告させるという形式をとっている(と解釈した。)
関連:
- [Paper Note] AlphaEvolve: A coding agent for scientific and algorithmic discovery, Alexander Novikov+, arXiv'25, 2025.06
- [Paper Note] ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution, Robert Tjarko Lange+, arXiv'25, 2025.09
続報:
Effective harnesses for long-running agents, Anthropic, 2025.11
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Blog #Initial Impression Notes Issue Date: 2026-03-10 Comment
`Agent Harness` という用語の起源が気になっており、アンテナを張っているが、本ブログでAgent Harnessという用語が登場している。
- [Paper Note] Building Effective AI Coding Agents for the Terminal: Scaffolding, Harness, Context Engineering, and Lessons Learned, Nghi D. Q. Bui, arXiv'26, 2026.03
において本ブログが引用され `harness` という用語が用いられている。このブログが起源なのだろうか(勉強不足)。
The Synthetic Data Playbook: Generating Trillions of the Finest Tokens, HuggingFace, 2026.03
Paper/Blog Link My Issue
#Article #Pretraining #NLP #SyntheticData #read-later #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-03-10 Comment
12.7 GPU yearを使い、90回の実験、1 Trillion tokenの生成を経て見つけた、合成事前学習データの構築方法のbest recipeが紹介されている模様。先行研究を上回る学習効率を達成している。
元ポスト:
Open-Sourcing Sarvam 30B and 105B, sarvam, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Reasoning #OpenWeight Issue Date: 2026-03-10 Comment
元ポスト:
The importance of Agent Harness in 2026, PHILSCHMID, 2026.01
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Blog #Selected Papers/Blogs #LongHorizon #Reading Reflections #AgentHarness Issue Date: 2026-03-08 Comment
本ブログで定義されているAgent Harnessは、これまでのAI Agent研究で利用されてきた Scaffold(=実行基盤)とEvaluation Harness(=評価基盤)のように、実行と評価を区別してきたLiteratureとは異なる、より包括的な概念に見える(言葉としてHarnessが用いられているので、最初に読んだときは困惑した)。
先行研究:
- [Paper Note] Holistic Evaluation of Language Models, Percy Liang+, arXiv'22, 2022.11
- [Paper Note] Lessons from the Trenches on Reproducible Evaluation of Language Models, Stella Biderman+, arXiv'24, 2024.05
- [Paper Note] Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent
Evaluation, Sayash Kapoor+, arXiv'25, 2025.10
これまでのLiteratureでは、エージェントがタスクを遂行するためのエコシステム全般(言い換えるとLLMをエージェントの脳とした時の、エージェントの実装そのもの)のことをScaffold(ツール利用やコンテキスト管理、サブエージェントの実行、エラー時の挙動、プロンプト構成など)と呼び、
評価をする際の評価基盤となるインフラ(エージェントを動作させる仮想マシン等の実行環境やそのオーケストレーション、Scaffoldの構成、評価ベンチマーク、コストやtrajectoryのロギング等の評価全体に関わるエコシステム)のことをEvaluation Harnessと呼んできたと認識している。
(私の認識違いの可能性もあるが)このLiteratureを理解しておかないと、今後Harnessという言葉がバズワードと化して、思わぬ誤解を生むかもしれないので注意した方が良いかなと感じた。
つまり世の中には
- Scaffold
- Evaluation Harness
- Agent Harness
の3種類の定義があり、特に後者二つは省略してHarnessと呼ばれそう、という気がするが、後者二つは呼称が似ているが異なる概念を指しているので注意した方が良いかも(あくまで個人の感想)。
たとえば下記OpenAIのブログでも「Harness Engineering」という言葉がタイトルで用いられており、Harnessの定義がなされずに記述されているように見える。実際ブログ後半にはEvaluation HarnessというこれまでのLiteratureと同じ意味合いでの用語も登場している。今後どのような用語が何を指すのようになるかは分からないが、ハーネスという言葉の定義が人によって異なる可能性があるという点は認識しておいた方が良さそうである。
- Harness engineering: leveraging Codex in an agent-first world, Ryan Lopopolo, 2026.02
`Agent Harness` という用語の起源が気になっており、アンテナを張っているが、下記AnthropicブログでAgent Harnessという用語が登場している。
- Effective harnesses for long-running agents, Anthropic, 2025.11
下記文献でも
- [Paper Note] Building Effective AI Coding Agents for the Terminal: Scaffolding, Harness, Context Engineering, and Lessons Learned, Nghi D. Q. Bui, arXiv'26, 2026.03
Effective harnesses for long-running agents, Anthropic, 2025.11
が引用され `harness` という用語が用いられている。このブログが起源なのだろうか(勉強不足)。
- [Paper Note] SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks, Xiangyi Li+, arXiv'26, 2026.02
でも Agent Harness という用語が使われている。
Codex Security: now in research preview, OpenAI, 2026.03
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Blog #Security Issue Date: 2026-03-07 Comment
元ポスト:
Chinese Open Source: A Definitive History, Kevin Xu, 2026.03
Paper/Blog Link My Issue
#Article #Survey #NLP #Blog #OpenWeight #read-later Issue Date: 2026-03-07
ガバメントAIで試用する国内大規模言語モデル(LLM)の公募結果, デジタル庁, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Blog #Japanese #One-Line Notes Issue Date: 2026-03-06 Comment
元ポスト:
以下が選出されたとのこと:
- 株式会社NTTデータ「tsuzumi 2」
- カスタマークラウド株式会社「CC Gov-LLM」
- KDDI株式会社・株式会社ELYZA共同応募体「Llama-3.1-ELYZA-JP-70B」
- ソフトバンク株式会社「Sarashina2 mini」
- 日本電気株式会社「cotomi v3」
- 富士通株式会社「Takane 32B」
- 株式会社Preferred Networks「PLaMo 2.0 Prime」
Google Workspace CLI, Google, 2026.03
Paper/Blog Link My Issue
#Article #Tools #NLP #AIAgents #Repository #ContextEngineering #One-Line Notes #AgentSkills Issue Date: 2026-03-06 Comment
元ポスト:
google workspaceにone-lineのコマンドでアクセス可能なCLIツールとのこと。40以上のAgentSkillsを内包。
Practical Guide to Evaluating and Testing Agent Skills, PHILSCHMID, 2026.03
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Blog #Coding #SoftwareEngineering #read-later #AgentSkills Issue Date: 2026-03-06 Comment
元ポスト:
関連:
- How to Create Effective Agent Skills, openhands, 2026.02
Reasoning models struggle to control their chains of thought, and that’s good, OpenAI, 2026.03
Paper/Blog Link My Issue
#Article #Controllable #NLP #Dataset #Chain-of-Thought #Evaluation #Blog #Reasoning #read-later #Author Thread-Post Issue Date: 2026-03-06 Comment
元ポスト:
著者ポスト:
Introducing GPT‑5.4, OpenAI, 2026.03
Paper/Blog Link My Issue
#Article #NLP #AIAgents #ChatGPT #Coding #Proprietary #VisionLanguageModel #Reference Collection #Reading Reflections Issue Date: 2026-03-06 Comment
元ポスト:
Artiflcial Analysisによる評価:
所見:
所見:
評判が良い。管理人も利用しているが、指示で曖昧な点をきちんと質問してくれる点が便利。かつ応答として、選択可能なオプションを提示し、自由記述もできる。実装の内容はClaude 4.6 Opusと比べるとコードがシンプルな印象を受けるが、これも指示次第な気はする。
曖昧な点があったら質問を投げかけるという挙動はopenhandsのPosition Paperとも整合する流れである。
- [Paper Note] Position: Humans are Missing from AI Coding Agent Research, Wang+, 2026.02
Introducing Olmo Hybrid: Combining transformers and linear RNNs for superior scaling, Ai2, 2026.03
Paper/Blog Link My Issue
#Article #Pretraining #NLP #Attention #OpenWeight #mid-training #read-later #Selected Papers/Blogs #One-Line Notes #RecurrentModels #Hybrid #LinearAttention Issue Date: 2026-03-06 Comment
元ポスト:
x1のFull Attention + x3のGated DeltaNetによるハイブリッドアーキテクチャで、75%のattentionをlinear attention (recurrent module)に置換。x3のSliding Window Attentionを用いているOlmo3と比較した結果
- 事前学習におけるデータ効率がより高く(約2倍)
- mid-training後の評価では、数学、コード、STEM, non-STEM, QA、long-contextなどの主要なドメインにおいてOlmo3と同と床それ以上の性能を達成。特に、long-contextにおけるベンチマでは大幅な性能向上(Recurrentなアーキテクチャの恩恵)
関連:
- [Paper Note] Gated Delta Networks: Improving Mamba2 with Delta Rule, Songlin Yang+, ICLR'25, 2024.12
元ポスト:
関連:
所見:
FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling, together.ai, 2026.03
Paper/Blog Link My Issue
#Article #NLP #Library #Transformer #Attention #Chip #read-later #Selected Papers/Blogs #GPUKernel #Initial Impression Notes Issue Date: 2026-03-06 Comment
元ポスト:
関連:
これは読まねば。。。
AReaL: A Large-Scale Asynchronous Reinforcement Learning System, inclusionAI, 2026.03
Paper/Blog Link My Issue
#Article #Tools #NLP #ReinforcementLearning #Reasoning #read-later #Asynchronous #TrainingFramework Issue Date: 2026-03-05 Comment
元ポスト:
PPO → DPO → GRPO→ Rubrics, PROF. TOM YEH, 2026.03
Paper/Blog Link My Issue
#Article #Tutorial #NLP #ReinforcementLearning #Blog #Video #PostTraining #Non-VerifiableRewards #One-Line Notes #Rubric-based Issue Date: 2026-03-05 Comment
Cameron R. Wolfe氏によるRubic-basedなRL(主にnon-verifiableなドメインへの適用)のチュートリアル。序盤はPPO, DPO, GRPOに関する解説
元ポスト:
GPT‑5.3 Instant:よりスムーズで、日常会話にもっと役立つ, OpenAI, 2026.03
Paper/Blog Link My Issue
#Article #NLP #ChatGPT #Blog #Proprietary #VisionLanguageModel Issue Date: 2026-03-04 Comment
元ポスト:
なんだかなあ
Gemini 3.1 Flash-Lite: Built for intelligence at scale, Google, 2026.03
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #Proprietary #VisionLanguageModel Issue Date: 2026-03-04 Comment
元ポスト:
How to Create Effective Agent Skills, openhands, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Blog #read-later #AgentSkills Issue Date: 2026-03-03 Comment
元ポスト:
New ARENA material: 8 exercise sets on alignment science & interpretability, CallumMcDougall, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Alignment #Blog #Safety #read-later #Selected Papers/Blogs Issue Date: 2026-03-03 Comment
元ポスト:
Qwen 3.5 small series, Qwen Team, 2026.02
Paper/Blog Link My Issue
#Article #NLP #SmallModel #OpenWeight #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-03-02 Comment
なんとSLMもリリース
元ポスト:
agent-vault, botiverse, 2026.02
Paper/Blog Link My Issue
#Article #Tools #NLP #AIAgents #Repository #Privacy Issue Date: 2026-03-02
TAKT, nrslib, 2026.01
Paper/Blog Link My Issue
#Article #Tools #NLP #AIAgents #Coding #SoftwareEngineering #AgentHarness Issue Date: 2026-03-01 Comment
色々使ってみたいなぁ(小並感)
元ポスト:
FP8 trainingを支える技術 1, Kazuki Fujii, 2026.02
Paper/Blog Link My Issue
#Article #Tutorial #Pretraining #NLP #Blog #mid-training #PostTraining #Selected Papers/Blogs #LowPrecision Issue Date: 2026-03-01
The Open Anonymity Project, The Open Anonymity Project, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Proprietary #read-later #Privacy Issue Date: 2026-02-28 Comment
元ポスト:
Coding agents progress over the past two months, Andrej Karpathy, X, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Coding #Post #SoftwareEngineering #Reading Reflections Issue Date: 2026-02-28 Comment
やっぱ英語で指示ださないとあかんか...(小並感)
関連:
LLM/VLA等の学習ライブラリ回りでは、人間が細かく実装方針分析を指示した上で、実装部分のみを移譲すると今のところ一番うまくいくとのこと。
CoderForge-Preview: SOTA open dataset for training efficient coding agents, together.ai, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Dataset #Supervised-FineTuning (SFT) #AIAgents #Blog #Coding #SoftwareEngineering #read-later #Selected Papers/Blogs Issue Date: 2026-02-28 Comment
元ポスト:
The third era of AI software development, Michael Turuell, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Coding #Post #SoftwareEngineering #read-later Issue Date: 2026-02-28
10 open-weight LLM releases in January and February 2026, Sebaschan Raschka, 2026.02
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #Post #read-later #Selected Papers/Blogs Issue Date: 2026-02-28 Comment
- Trinity Large, Arcee, 2026.01
- [Paper Note] Kimi K2.5: Visual Agentic Intelligence, Kimi Team+, arXiv'26, 2026.02
- [Paper Note] Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters, Ailin Huang+, arXiv'26, 2026.02
- Qwen3-Coder-Next: Pushing Small Hybrid Models on Agentic Coding, QwenTeam, 2026.02
- [Paper Note] GLM-5: from Vibe Coding to Agentic Engineering, GLM-5 Team+, arXiv'26, 2026.02
- MiniMax M2.5: SOTA in Coding and Agent, designed for Agent Universe, MiniMax, 2026.02
- [Paper Note] Nanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Acts, Chen Yang+, arXiv'26, 2026.02
- Qwen3.5: Towards Native Multimodal Agents, Qwen Team, 2026.02
- Ling-2.5-1T, inclusionAI, 2026.02
- Ring-1T-2.5-FP8, inclusionAI, 2026.02
- Cohere Labs Launches Tiny Aya, Making Multilingual AI Accessible, COHERE LABS TEAM, 2026.02
元ポストには書かれていないがLLMというくくりで言うと以下もある:
- New ARENA material: 8 exercise sets on alignment science & interpretability, CallumMcDougall, 2026.02
- LFM2-24B-A2B: Scaling Up the LFM2 Architecture, LiquidAI, 2026.02
- Qwen3 Swallow, Swallow LLM, 2026.02
- Japanese
- GPT-OSS Swallow, Swallow LLM, 2026.02
- Japanese
- GLM-4.7-Flash, Z.ai, 2026.01
- LongCat-Flash-Thinking-2601, Meituan, 2026.01
- Introducing LFM2.5: The Next Generation of On-Device AI, LiquidAI, 2026.01
Omniモデルを含めると以下:
- Ming-omni-tts-0.5B, inclusionAI, 2026.02
- [Paper Note] Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability, Aaditya Vikram Prasad+, arXiv'26, 2026.02
- MiniCPM-o-4_5, OpenBMB, 2026.02
World Modelsを含めると以下?:
- [Paper Note] Causal-JEPA: Learning World Models through Object-Level Latent Interventions, Heejeong Nam+, arXiv'26, 2026.02
- [Paper Note] Code2World: A GUI World Model via Renderable Code Generation, Yuhao Zheng+, arXiv'26, 2026.02
- [Paper Note] DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos, Shenyuan Gao+, arXiv'26, 2026.02
- [Paper Note] World Action Models are Zero-shot Policies, Seonghyeon Ye+, arXiv'26, 2026.02
- [Paper Note] Advancing Open-source World Models, Robbyant Team+, arXiv'26, 2026.01
- Project Genie: Experimenting with infinite, interactive worlds, Google Deepmind, 2026.01
- Waypoint-1: Real-time Interactive Video Diffusion from Overworld, Overworld, 2026.01
確実に見落としがあるけど。
Training Recipes, PRIME Intellect Lab, 2026.02
Paper/Blog Link My Issue
#Article #NLP #ReinforcementLearning #ExperimentManagement #PostTraining #read-later #One-Line Notes Issue Date: 2026-02-28 Comment
公式によるPrime Intellect Labを用いたRLによるレシピの模様。これ読んだらだいたい実験できるようになるんではなかろうか。
元ポスト:
prime-lab-trainer, abideenml, 2026.02
Paper/Blog Link My Issue
#Article #NLP #ReinforcementLearning #AIAgents #Repository #ExperimentManagement #SoftwareEngineering #AgentSkills Issue Date: 2026-02-28 Comment
- Introducing Lab: The Full-Stack Platform for Training your Own Models, Prime Intellect, 2026.02
に対して任意のHF Datasetを用いて自動的にRLによるモデルの学習をsubmit可能なClaude Code skillとのこと。
元ポスト:
Qwen3.5 Medium Model Series, Qwen Team, 2026.02
Paper/Blog Link My Issue
#Article #NLP #MultiLingual #OpenWeight #MoE(Mixture-of-Experts) #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-02-28 Comment
元ポスト:
いずれのモデルもベンチマーク上はGPT-5 miniと同等以上の性能に見える。
また、Qwen3.5-35B-A3BはQwen3-235B-A22B-2507やQwen3-VL235B-A22Bを上回っており、アーキテクチャ、データの品質、RLによって実現されているとのこと。
27BモデルのHLEのスコアが非常に高いと話題:
FP8版もリリース:
日本語の医師国家試験(2026)において35B-A3Bが非常に高いスコアを記録:
Artificial Analysisによるベンチマーキング:
New in Claude Code: Remote Control, Anthropic, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Coding #Post #SoftwareEngineering Issue Date: 2026-02-27 Comment
スマホからターミナルのClaude Codeに対してリモートで制御が可能になったらしい
Introducing Mercury 2, inception, 2026.02
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #DiffusionModel #Blog #Reasoning #Proprietary #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-02-27 Comment
元ポスト:
1092 token/secのproprietary (reasoning) dLLM
関連:
- [Paper Note] Mercury: Ultra-Fast Language Models Based on Diffusion, Inception Labs+, arXiv'25
Artificial Analysisのベンチマーキング結果とスループットの散布図:
スループット/性能比において明らかに抜きんでている。
LFM2-24B-A2B: Scaling Up the LFM2 Architecture, LiquidAI, 2026.02
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #SmallModel #OpenWeight #MoE(Mixture-of-Experts) #Initial Impression Notes #EdgeDevices Issue Date: 2026-02-27 Comment
元ポスト:
edge deviceにデプロイできる規模でLFM2をスケールさせた模様
How much does distillation really matter for Chinese LLMs?, Interconnects, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Distillation #read-later Issue Date: 2026-02-27 Comment
関連:
- Detecting and preventing distillation attacks, Anthropic, 2026.02
Swallowにおける 日英推論型大規模言語モデルの構築, 水木栄, 第26回LLM勉強会, 2026.02
Paper/Blog Link My Issue
#Article #Pretraining #NLP #Dataset #Supervised-FineTuning (SFT) #ReinforcementLearning #Japanese #mid-training #PostTraining #Selected Papers/Blogs #DataMixture #Initial Impression Notes Issue Date: 2026-02-27 Comment
元ポスト:
関連:
- Qwen3-Swallow & GPT-OSS-Swallow, Kazuki Fujii, 2026.02
まだしっかり読めていないのだが、適切なDataMixtureはどのようにして決めているのだろうか?
- 数学データによる学習がコーディングにのみ転移
- 英語データを邦訳したデータが学習に寄与するためcross-lingualで能力が転移する
- RLはpass@1を改善するが、Pass@10などの改善幅は縮小する
- この辺の話は資料中でも先行研究が引用されており、実際に確認されたということだと思われる
...
[Paper Note] Preconditioned inexact stochastic ADMM for deep models, Nature Machine Intelligence 2026, 2026.02
Paper/Blog Link My Issue
#Article #NeuralNetwork #ComputerVision #MachineLearning #NLP #Optimizer #Initial Impression Notes #Nature Machine Intelligence Issue Date: 2026-02-24 Comment
元ポスト:
パラメータサイズが大きい場合にMuon超え...?
所見:
The persona selection model, Anthropic, 2026.02
Paper/Blog Link My Issue
#Article #Analysis #Pretraining #Alignment #Blog #Safety #PostTraining #Personality Issue Date: 2026-02-24 Comment
元ポスト:
Why SWE-bench Verified no longer measures frontier coding capabilities, OpenAI, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Evaluation #Blog #Coding #SoftwareEngineering #Selected Papers/Blogs #One-Line Notes #Contamination Issue Date: 2026-02-24 Comment
元ポスト:
SWE-Bench Verifiedはpublicなリポジトリに基づいたベンチマークなのでcontaminationが生じやすく、実際にいくつかのモデルでcontaminationが確認されたと言う話と、testコードに本来は正しい実装でもfailedとなる許容するスコープが狭いテストが存在していた、という話で、これらの教訓を生かしたSWE-Bench Proを作成し、実際それはcontaminationがほとんど起きておらず、仮に起きていたとしても非常にマイナーなものだよ、というような話が書かれている。
Detecting and preventing distillation attacks, Anthropic, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Blog #OpenWeight #Proprietary #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-02-24 Comment
元ポスト:
DeepSeek, Moonshot AI, MiniMax がDistillationを用いてClaude出力からモデルを改善するためのattackを特定したというAnthropicからのアナウンス
所見:
- [Paper Note] Extracting books from production language models, Ahmed Ahmed+, arXiv'26, 2026.01
で提案されている手法を用いてClaude Sonnetからハリーポッターと賢者の石の95.8%を抽出できた、との報告もある。
GPT-OSS Swallow, Swallow LLM, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Reasoning #OpenWeight #Japanese Issue Date: 2026-02-21 Comment
元ポスト:
第120回医師国家試験(2026)を解かせてみた結果:
Qwen3 Swallow, Swallow LLM, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Reasoning #OpenWeight #Japanese #MoE(Mixture-of-Experts) Issue Date: 2026-02-21 Comment
元ポスト:
Qwen3-Swallow & GPT-OSS-Swallow, Kazuki Fujii, 2026.02
Paper/Blog Link My Issue
#Article #Pretraining #NLP #Supervised-FineTuning (SFT) #ReinforcementLearning #Evaluation #Japanese #mid-training #PostTraining #read-later #RLVR #Selected Papers/Blogs Issue Date: 2026-02-21 Comment
元ポスト:
関連:
- [Paper Note] Accelerating Large Language Model Training with 4D Parallelism and Memory Consumption Estimator, Kazuki Fujii+, arXiv'24, 2024.11
- FP8 trainingを支える技術 1, Kazuki Fujii, 2026.02
Gemini 3.1 Pro: A smarter model for your most complex tasks, Google, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Proprietary #Selected Papers/Blogs #VisionLanguageModel #Reference Collection Issue Date: 2026-02-20 Comment
元ポスト:
Artificial Analysisによる評価:
所見:
ベンチマークほどの性能は実用上は感じられず、API利用などにおいては安定性に課題があるとのこと。
ALE BenchでSoTA:
- [Paper Note] ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering, Yuki Imajuku+, NeurIPS'25
Introducing Claude Sonnet 4.6, Anthropic, 2026.02
Paper/Blog Link My Issue
#Article #Blog #Proprietary #read-later #VisionLanguageModel Issue Date: 2026-02-18 Comment
もうSonnetが出てきた
元ポスト:
所見:
Cohere Labs Launches Tiny Aya, Making Multilingual AI Accessible, COHERE LABS TEAM, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Blog #SmallModel #MultiLingual #OpenWeight #Selected Papers/Blogs #LowResource #KeyPoint Notes #Reference Collection Issue Date: 2026-02-18 Comment
元ポスト:
公式ポスト:
アーキテクチャ解説:
70程度の言語の性能をバランス良くサポートする3.35BのLLMで、Baseモデルと、マルチリンガルの性能は保ちつつも特定のregionに特化したinstruction tuningを実施したvariantを公開。また、multilingualでのベンチマークも公開。同程度の規模間のモデルについて、qwen3-4Bとの比較がわかりやすく、Europe, south asiaは同等、Asia-pacificはQwenよりも劣り、west asia, africa regionのようなこれまでlow resourceだと思われたregionではほか同規模のモデルと比較して突出した性能を誇るモデルに見える。CC上でのページ数と、言語モデルごとの性能を比較したグラフもあり、CCでのデータが少ない言語はこれまでのモデルは性能が低かったが、Tiny Ayaは非常に高い性能を達成している(このグラフで言うと日本語はかなりinformation richな言語にカテゴライズされているように見える)。
SWE-fficiency: Evaluating How to Fix Code, Not Just What to Fix, OpenHands, 2026.02
Paper/Blog Link My Issue
#Article #Metrics #NLP #AIAgents #Evaluation #Coding #SoftwareEngineering #Selected Papers/Blogs #KeyPoint Notes Issue Date: 2026-02-17 Comment
元ポスト:
既存のAI Agentsのベンチマークは、バグを修正することに特化しており(what to fix)、機能的には正しいが高速化が必要といった効率性や最適化の観点(how to fix)が評価から抜けているので、そのためにSpeedup Ratioと呼ばれる人間の専門家に対してどの程度の高速化を達成できたかを測るmetricとそのためのベンチマークSWE-ffiencyを構築。SWE-fficiencyはnumpy, pandas, sklearnなどの9つの主要なリポジトリにおける498のタスクで構成される。評価の結果、Claude Opus 4.5をOpenhandsのハーネスで駆動させだ場合でも人間のエキスパートに対して0.225倍程度の高速化しか実現できないことがわかった、といった話な模様。
IA Agents Minimal agent framework for the Gemini Interactions API, philschmid, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Repository #read-later #MinimalCode #Initial Impression Notes Issue Date: 2026-02-17 Comment
元ポスト:
Gemini Interactions APIを用いたエージェントのminimal code。これは非常に勉強になりそう。
Rubric-Based Rewards for RL Extending the benefits of large-scale RL training to non-verifiable domains..., Cameron R. Wolfe, 2026.02
Paper/Blog Link My Issue
#Article #Tutorial #NLP #ReinforcementLearning #Blog #PostTraining #read-later #VerifiableRewards #Selected Papers/Blogs #Non-VerifiableRewards #Rubric-based Issue Date: 2026-02-17 Comment
元ポスト:
Beyond MuP: 2. Linear Layers and Steepest Descent, Scientific Spaces, 2026.02
Paper/Blog Link My Issue
#Article #Pretraining #NLP #Blog #Optimizer #Stability Issue Date: 2026-02-16 Comment
元ポスト:
Ling-2.5-1T, inclusionAI, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Reasoning #OpenWeight #MoE(Mixture-of-Experts) Issue Date: 2026-02-16 Comment
Ringに続いてLingもリリース
関連:
- Ring-1T-2.5-FP8, inclusionAI, 2026.02
元ポスト:
AI 101: "On-Policy Distillation Zeitgeist", Turing Post, 2026.02
Paper/Blog Link My Issue
#Article #Tutorial #NLP #ReinforcementLearning #Blog #PostTraining #On-Policy #One-Line Notes #SelfDistillation Issue Date: 2026-02-16 Comment
元ポスト:
最近よくみかける on-policy self-distillationに関する解説
QED-Nano: Teaching a Tiny Model to Prove Hard Theorems, LM Provers Team, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Supervised-FineTuning (SFT) #ReinforcementLearning #Blog #Mathematics #SmallModel #PostTraining #Proofs #Rubric-based #Initial Impression Notes Issue Date: 2026-02-16 Comment
元ポスト:
ポイント解説:
早くもReasoning Cacheが利用されている:
- [Paper Note] Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL, Ian Wu+, arXiv'26, 2026.02
4B級のモデルで特定タスクに特化したモデルを作りたい場合に非常に役立ちそうなレシピ
Building Olmo in the Era of Agents, Nathan Lambert, LTI Colloquim, 2026.02
Paper/Blog Link My Issue
#Article #Tutorial #Survey #NLP #AIAgents #Reasoning #Slide #OpenSource #read-later #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-02-16 Comment
元ポスト:
うーんこれは時間をとってしっかり読んで色々まとめたい・・・
[Paper Notes] Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity, Bytedance Seed, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #AIAgents #Reasoning #Proprietary #VisionLanguageModel Issue Date: 2026-02-16 Comment
元ポスト:
所見:
GPT‑5.2 derives a new result in theoretical physics, OpenAI, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Blog #ScientificDiscovery #Physics #Human-in-the-Loop Issue Date: 2026-02-14 Comment
元ポスト:
Introducing GPT‑5.3‑Codex‑Spark: An ultra-fast model for real-time coding in Codex, OpenAI, 2026.02
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #AIAgents #Blog #Coding #SoftwareEngineering Issue Date: 2026-02-13 Comment
元ポスト:
所見:
Gemini 3 Deep Think: Advancing science, research and engineering, Google, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Coding #Reasoning #Mathematics #Proprietary #SoftwareEngineering #VisionLanguageModel #Science Issue Date: 2026-02-13 Comment
まずはUltra Subscriberに公開し、その後徐々にAPIアクセスを解禁していくとのこと。
LiveCodeBench:
MiniMax M2.5: SOTA in Coding and Agent, designed for Agent Universe, MiniMax, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Blog #Coding #OpenWeight #SoftwareEngineering #Selected Papers/Blogs Issue Date: 2026-02-13 Comment
元ポスト:
OsenHands IndexでClaude Sonnet 4.5超えの初めてのOpenWeightモデル:
コストパフォーマンスにおいては、低コストなモデル群の中では抜きん出た性能
まだHF上にWeightは公開されていないようだが後ほど公開されると思われる。
所見:
weightが公開:
https://huggingface.co/MiniMaxAI/MiniMax-M2.5
元ポスト:
UnslothがGGUF版を公開:
A2A: The Agent2Agent Protocol, DeepLearning.AI, 2026.02
Paper/Blog Link My Issue
#Article #Multi #Tutorial #NLP #AIAgents #Video #SoftwareEngineering #A2A Issue Date: 2026-02-13 Comment
元ポスト:
元ポスト:
Ring-1T-2.5-FP8, inclusionAI, 2026.02
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #AIAgents #Attention #Reasoning #LongContext #OpenWeight #LongHorizon #LinearAttention Issue Date: 2026-02-12 Comment
元ポスト:
関連:
- Ring-1T, inclusionAI, 2025.10
MLA + lightning linear attentionのハイブリッド
- MHA vs MQA vs GQA vs MLA, Zain ul Abideen, 2024.07
- [Paper Note] Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention, Zhen Qin+, ICML'24, 2024.05
Harness engineering: leveraging Codex in an agent-first world, Ryan Lopopolo, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #GenerativeAI #Blog #Coding #SoftwareEngineering #One-Line Notes Issue Date: 2026-02-12 Comment
OpenAI社内でのコードを1行も人間が書かないで製品をリリースする取り組みに関する詳細なレポートのようである。初期の設計などで想像以上に時間がかかってしまった点(これはCodexの能力の問題ではない)や、実装を続ける中で品質に責任を持つ人間の能力(というより時間)がボトルネックになっていったため、極力Codexが自律的に品質管理ができるような実行・検証環境を用意することで負担を低減した話や、Codexに膨大なマニュアルを読ませて処理をさせるのではなく、どこにどのような情報が格納されているのかといったマップ(目次)を与えることがコンテキストエンジニアリング上重要だったことなどを通じてエージェントにとってリポジトリ全体の可読性を高めることが重要だったといった話や、プロジェクトの期間が長引くにつれて、リポジトリ内に共有されていないcontextが増大していき、それらをリポジトリに統合する作業が生じるなどの課題も生じたといったような話など色々と書かれている。
microgpt.py, Andrej Karpathy, 2026.02
Paper/Blog Link My Issue
#Article #NLP #python #Selected Papers/Blogs #MinimalCode Issue Date: 2026-02-12 Comment
元ポスト:
[Paper Note] Accelerating Mathematical and Scientific Discovery with Gemini Deep Think, Google DeepMin, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Blog #Mathematics #ScientificDiscovery #Test-Time Scaling #read-later #KeyPoint Notes #Physics #Human-in-the-Loop Issue Date: 2026-02-12 Comment
元ポスト:
- 数学について
- verifierを通じて解の修正と再生成を繰り返すが、問題が解けないことを認めることで(無駄な修正・再生成を減らすことで)効率を大幅に改善
- 博士課程レベル・オリンピックレベルを超えてもtest-time scalingが継続する
- 検索を融合することで既存文献を取り入れ正確性向上
- 完全自動で出版できるレベルの研究を実施可能なところまできている(level0--5のlevel2)
- コンピュータサイエンス・物理学について
- ネットワーク側で広範な解空間を探索してlong-trailな解も捉え推論に組み込むことが可能で、自動的なverificationと人間によるverificationを通じてoutputを生成する
- たとえば10年間未解決だったオンライン列モジュラ最適化と呼ばれる問題や、モデル学習時のノイズ除去による理論的な証明などを実施できている
論文:
- [Paper Note] Towards Autonomous Mathematics Research, Tony Feng+, arXiv'26, 2026.02
[Paper Note] Position: Humans are Missing from AI Coding Agent Research, Wang+, 2026.02
Paper/Blog Link My Issue
#Article #NLP #UserBased #AIAgents #Coding #read-later #Selected Papers/Blogs #interactive #One-Line Notes #Initial Impression Notes Issue Date: 2026-02-12 Comment
# Authors
Zora Zhiruo Wang, John Yang, Kilian Lieret, Alexa Tartaglini, Valerie Chen, Yuxiang Wei,
Zijian Wang, Lingming Zhang, Karthik Narasimhan, Ludwig Schmidt, Graham Neubig, Daniel Fried, Diyi Yang
元ポスト:
現在のコーディングエージェントは自動的にタスクを完了させ、難易度の高いベンチマークを解けることが実用的な価値とみなされているが、今後より実用的な価値を高めプロダクト化するためには単独でタスクをこなすのではなく、人間開発者やユーザとの相互作用をするような枠組みが次のブレイクスルーとなりうるというposition。非常に共感できる。
GLM-5: From Vibe Coding to Agentic Engineering, Z.ai, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #OpenWeight #MoE(Mixture-of-Experts) #Selected Papers/Blogs #KeyPoint Notes #Reference Collection #LongHorizon #SparseAttention Issue Date: 2026-02-12 Comment
関連:
- GLM-4.7: Advancing the Coding Capability, Z.ai, 2025.12
GLMシリーズの最新モデルGLM-5がリリースされた
元ポスト:
- DeepSeek Sparse Attentionを採用:
- DeepSeek-V3.2-Exp: Boosting Long-Context Efficiency with DeepSeek Sparse Attention, DeepSeek-AI, 2025.09
- [Paper Note] DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models, DeepSeek-AI+, arXiv'25, 2025.12
- 事前学習データを23Tから28.5Tトークンへ
- パラメータ数は4.5の355B-A32から744B-A40Bへ
- RLのインフラとして4.5から引き続きSlimeを採用
- slime, THUDM & Zhihu, 2025.09
- long-horizonなタスクに秀でており、reasoning, coding, agenticタスクにおける各種ベンチマークでOpus 4.5, GPT-5.2, Gemini 3 Proと同等程度の性能
FP8版も公開されている模様(Hopper以後のアーキテクチャでないとサポートされていない点に注意
所見:
元ポスト:
unslothがGGUF版をすでにリリースしている模様。早い:
https://unsloth.ai/docs/models/glm-5
アーキテクチャ解説:
アーキテクチャ解説:
所見:
ENGRAM, EvolvingLMMs-Lab, 2026.02
Paper/Blog Link My Issue
#Article #Tools #NLP #AIAgents #Privacy #MCP #memory #One-Line Notes Issue Date: 2026-02-12 Comment
元ポスト:
MCPに対応しているAI Agentであれば互換性がある暗号化されたストレージの実装なようで、サードパーティのストレージにデータを預けなくてもローカルのストレージでLLMに対して知識を提供可能な模様。
最近DeepSeekが提案したEngramとは異なるので注意:
- [Paper Note] Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models, Xin Cheng+, arXiv'26, 2026.01
Introducing Lab: The Full-Stack Platform for Training your Own Models, Prime Intellect, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #MachineLearning #NLP #Infrastructure #ReinforcementLearning #AIAgents #Blog #ScientificDiscovery #PostTraining #Selected Papers/Blogs #One-Line Notes #Reference Collection #Environment Issue Date: 2026-02-11 Comment
元ポスト:
事後学習、特にAgenticな研究の民主化のためのプラットフォームの提供
所見:
利用例 (Environment Hub):
Sabotage Risk Report: Claude Opus 4.6, Anthropic, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Proprietary #Safety #read-later #Sabotage Issue Date: 2026-02-11 Comment
元ポスト:
[Paper Note] OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis, Li+, 2026.02
Paper/Blog Link My Issue
#Article #InformationRetrieval #NLP #Search #Supervised-FineTuning (SFT) #AIAgents #SyntheticData #OpenSource #Selected Papers/Blogs #Reproducibility #DeepResearch #One-Line Notes #LongHorizon #Initial Impression Notes #Environment Issue Date: 2026-02-10 Comment
元ポスト:
APIに依存せずオフラインコーパスと検索を利用し、高品質なDeepResearchのlong horizonなtrajectoryを合成可能な環境を構築。合成したtrajectoryでNemotron-3-nano-30B-A3B-BaseをSFTすることで、Kimi-K2, GLM-4.6などの10倍以上大きいサイズのモデルよりもBrowseCompで高い性能を獲得。同サイズのTongyiDeepResearchもoutperform。
Deterministicなプロセスで、オフラインコーパスからデータを合成し外部APIに依存しないため完全に再現性があり、かつAPIのコストやrate limitにも引っかからないという利点がある。検索エンジン、コード、データ、合成データ、モデル、全てを公開。
完全に再現性のある研究は素晴らしい。
Opus 4.6, Codex 5.3, and the post-benchmark era, Interconnects, 2026.02
Paper/Blog Link My Issue
#Article #Analysis #AIAgents #Blog #Coding #SoftwareEngineering #One-Line Notes #Author Thread-Post Issue Date: 2026-02-10 Comment
有識者によるClaude 4.6 Opus と Codex 5.3 を利用した際の所見(定性評価)が記述されている。
元ポスト:
著者によるTLDR:
Context-Bench: A benchmark for agentic context engineering, Letta Research, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Evaluation #Blog #ContextEngineering Issue Date: 2026-02-09 Comment
元ポスト:
Knowledge Editing for LLMs Papers, zjunlp, 2024.07
Paper/Blog Link My Issue
#Article #Survey #NLP #KnowledgeEditing Issue Date: 2026-02-08
Building a C compiler with a team of parallel Claudes, Anthropic, 2026.02
Paper/Blog Link My Issue
#Article #Multi #AIAgents #Blog #Coding #SoftwareEngineering #read-later #Selected Papers/Blogs Issue Date: 2026-02-06 Comment
元ポスト:
Introducing GPT-5.3-Codex: Expanding Codex across the full spectrum of professional work on a computer, OpenAI, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Coding #Proprietary #SoftwareEngineering #Selected Papers/Blogs #Reference Collection Issue Date: 2026-02-06 Comment
元ポスト:
terminal bench 2.0でOpus 4.6超え:
所見:
Advancing finance with Claude Opus 4.6, Anthropic, 2026.02
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Financial #Proprietary #SoftwareEngineering #Selected Papers/Blogs #One-Line Notes #Reference Collection Issue Date: 2026-02-06 Comment
元ポスト:
全体的に能力が向上しているが、ターミナルでのコーディング、BrowseComp(Agentic search), HLE, Financial Analysis, GDPValにおけるOffice Task, Novel Problem Solvingの能力が大きく向上しているように見える。
Context Windowが1Mとのことで素晴らしい
OpenHands Indexでトップとのことだが、Codex 5.3との比較はまだの模様:
50% time horizonが脅威の14.5時間:
MiniCPM-o-4_5, OpenBMB, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #SpeechProcessing #DiffusionModel #OpenWeight #AutomaticSpeechRecognition(ASR) #VisionLanguageModel #TTS #Omni #AudioLanguageModel Issue Date: 2026-02-05 Comment
元ポスト:
The Second Pre-training Paradigm, Jim Fan, X, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #Pretraining #NLP #MultiModal #Post #Robotics #WorldModels #One-Line Notes Issue Date: 2026-02-05 Comment
事前学習がnext word predictionから過去の行動と状態によって条件付けられ次の(ある期間の)世界の状態を予測するワールドモデリング(next physical state prediction)へのパラダイムシフトの予想(というよりこのパラダイムシフトの真っ只中にいる)。人間の脳が処理する情報の多くは視覚であり、言語的な領域は部分的なことであることや、猿は言語的な能力が低くても視覚や運動、触覚などの感覚的情報から世界の物理法則を理解し知的なアクションをとるメンタルモデルを確立していることなどを引き合いに説明している。
Time Horizon 1.1, METR, 2026.01
Paper/Blog Link My Issue
#Article #Metrics #NLP #AIAgents #Evaluation #Scaling Laws #Selected Papers/Blogs Issue Date: 2026-02-05 Comment
元ポスト:
続報:
関連:
- [Paper Note] Measuring AI Ability to Complete Long Tasks, Thomas Kwa+, arXiv'25, 2025.03
Fine-tuning open LLM judges to outperform GPT-5.2, together.ai, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Evaluation #Blog #LLM-as-a-Judge #DPO #RewardModel #One-Line Notes #Initial Impression Notes Issue Date: 2026-02-05 Comment
元ポスト:
Reward Bench 2:
- [Paper Note] RewardBench 2: Advancing Reward Model Evaluation, Saumya Malik+, arXiv'25, 2025.06
LLMでLLMを評価するというパラドックスに違和感はあるが、一般論として、「生成」するよりも「検証」することがモデルにとって簡単なタスクであるためうまくいきます(LLM-as-a-Judge)、といった説明が書いてあり、数千程度のサンプルでOpenLLMをDPOすることによって、GPT-5.2のようなFrontierモデルをReward Benchで上回ることができた、といった話が書かれている。
ただし、上記Reward Bench 2研究で示されている通り、**Reward Benchでの性能が高いReward Modelだからといって、必ずしもRLによって下流タスクの性能が向上するとは限らない点には注意**であり、元論文に従うとBest-of-Nサンプリングのようなtest-time-scalingのパラダイムとして利用するのが現在の実務上は良さそうである。
Together Evaluations now supports comparing top commercial APIs vs. open source models, together.ai, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Evaluation #Blog #PEFT(Adaptor/LoRA) #PostTraining #One-Line Notes Issue Date: 2026-02-05 Comment
元ポスト:
OpenLLMのFinetuningをサポートしているプラットフォームにおいて、データセットをアップロードすると
- Prompt optimization (GEPA)
- Fine-tuning (PEFT + full finetuning)
の両方を実施し、コスト-性能のパレート最適なポイントを評価し、かつGPT等とのProprietaryモデルとの比較もした評価もできるようになりました、といった話の紹介。
GEPA:
- [Paper Note] GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning, Lakshya A Agrawal+, ICLR'26, 2025.07
Finetuningがサポートされているモデル群:
-
https://docs.together.ai/docs/fine-tuning-models
Qwen3-Coder-Next: Pushing Small Hybrid Models on Agentic Coding, QwenTeam, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Attention #Blog #Coding #LongContext #SmallModel #MoE(Mixture-of-Experts) #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-02-04 Comment
HF: https://huggingface.co/collections/Qwen/qwen3-coder-next?spm=a2ty_o06.30285417.0.0.3bdec921Ja5TZI
元ポスト:
A3BでSWE Bench ProにおいてClaude Sonnet 4.5超え
関連:
- [Paper Note] Gated Delta Networks: Improving Mamba2 with Delta Rule, Songlin Yang+, ICLR'25, 2024.12
開発者の方のポスト:
int4 model from Cerebras:
https://huggingface.co/Intel/Qwen3-Coder-Next-int4-AutoRound
元ポスト:
Latest open artifacts (#18): Arcee's 400B MoE, LiquidAI's underrated 1B model, new Kimi, and anticipation of a busy month, Interconnects, 2026.02
Paper/Blog Link My Issue
#Article #Analysis #NLP #Blog #OpenWeight Issue Date: 2026-02-03 Comment
paid userしか全文は閲覧できない
元ポスト:
Moltbook is the most interesting place on the internet right now, Simon Willisons's blog, 2026.01
Paper/Blog Link My Issue
#Article #Multi #NLP #AIAgents #GenerativeAI #Blog #Conversation #Selected Papers/Blogs #Reference Collection Issue Date: 2026-02-01 Comment
元ポスト:
興味深い:
話したことのないhumanとの会話をあたかもあったことのように話し始める:
所見:
Andrej Karpathy氏もエージェントを参加させたようである:
所見:
Introducing the OpenHands Index, OpenHands, 2026.01
Paper/Blog Link My Issue
#Article #Analysis #NLP #AIAgents #Evaluation #Blog #SoftwareEngineering #Selected Papers/Blogs #KeyPoint Notes Issue Date: 2026-01-30 Comment
元ポスト:
SWE Bench(pythonプログラムリポジトリに対するissueを解決するタスク)がSWE関連の代表的なベンチマークだがこれらはソフトウェアエンジニアリングのサブタスクの一つしか反映しておらず、より多くのタスクの解決能力でSWE Agentの能力を評価し、かつコストの軸でも評価をしてどのモデルがパレート最適なものなのかを見つけられるようなindexを作って評価しました、という話に見える。
タスクとしては以下の5つをピックしているとのこと:
> 1. Issue Resolution
> 2. Frontend Development
> 3. Greenfield Development
> 4. Software Testing
> 5. Information Gathering
これらのタスクを総合的に評価するとClaude 4.5 Opusが最も性能が高くコストも高い。次点でGPT-5.2-Codexという結果。またコストが最も安く平均的な性能が高いモデルとしてはDeepSeekV3.2-Reasonerとなった。また、特定のタスク、たとえばGreenfield developmentではGPT-5.2-Codexの性能が抜きん出ているなど、個別のタスクで見るとモデル間の優劣がはっきりと見えるような結果になっている。
以下のモデルが追加:
Claude 4.6 Opus
GPT 5.2 Codex
Kimi K2.5
GLM-4.7
MiniMax M2.5
PLaMo 2.2 Primeをリリースしました, PFN, 2026.01
Paper/Blog Link My Issue
#Article #Multi #NLP #Supervised-FineTuning (SFT) #Proprietary #Japanese #DPO #PostTraining #InstructionFollowingCapability #Medical #RolePlaying Issue Date: 2026-01-29 Comment
関連:
- [Paper Note] Generalizing Verifiable Instruction Following, Valentina Pyatkin+, NeurIPS'25, 2025.07
- JFBench: 実務レベルの日本語指示追従性能を備えた生成AIを目指して, PFN, 2026.01
non-thinkingモデルである点に注意
JFBench: 実務レベルの日本語指示追従性能を備えた生成AIを目指して, PFN, 2026.01
Paper/Blog Link My Issue
#Article #NLP #Dataset #InstructionTuning #Evaluation #Japanese #InstructionFollowingCapability Issue Date: 2026-01-29 Comment
元ポスト:
Trinity Large, Arcee, 2026.01
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #Pretraining #NLP #OpenWeight #MoE(Mixture-of-Experts) #read-later #Selected Papers/Blogs #Stability #One-Line Notes #Reference Collection #Sparse #Initial Impression Notes Issue Date: 2026-01-29 Comment
テクニカルレポート:
https://github.com/arcee-ai/trinity-large-tech-report/
HF:
https://huggingface.co/arcee-ai
GLM4.7やDeepSeekV3と比較してスループットやTTFTが二倍以上。
非常にsparseなMoE(400B-A13B, 4/256のexpertsにルーティング)であるため学習を安定させるためにDense layerを増やし、モメンタムを考慮したexpertのバランシングや、z-lossと呼ばれるlogitのスケールをコントロールするような手法を導入することで安定した学習を実現。2048 Nvidia B300 GPUsで、17Tトークンの事前学習33日で完了
元ポスト:
これほどsparseなMoEをここまで安定させて学習できるのは非常に興味深いと思われる。
インタビュー:
やると決めてチームビルディングも含めて非常に短期間(6ヶ月)で達成したとのことだが、気になる。
解説:
所見(風刺):
ポイント解説:
アーキテクチャ解説:
Introducing Prism, OpenAI, 2026.01
Paper/Blog Link My Issue
#Article #NLP #AIAgents #ChatGPT #GenerativeAI #MultiModal #AcademicWriting #DeepResearch #One-Line Notes Issue Date: 2026-01-29 Comment
デモを見るとdraftをベースに関連研究をdeepresearchしてワンクリックでbibtexにexport, ホワイトボードに描いた図をドラッグ&ドロップして論文に反映などしている。Overleafの競合。
元ポスト:
所見:
Open Coding Agents: Fast, accessible coding agents that adapt to any repo, Ai2, 2026.01
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Coding #SoftwareEngineering #read-later Issue Date: 2026-01-29 Comment
開発者の方のブログ:
https://timdettmers.com/2026/01/27/building-open-coding-agent-sera/
HF:
https://huggingface.co/collections/allenai/open-coding-agents
14Bモデルリリース:
A few random notes from claude coding quite a bit last few weeks., Andrej Karpathy, 2026.01
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Coding #Post #SoftwareEngineering Issue Date: 2026-01-27
Minimax Agent, Minimax, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #AIAgents #GenerativeAI #ComputerUse Issue Date: 2026-01-27 Comment
code: https://github.com/MiniMax-AI/Mini-Agent
元ポスト:
Continual Learning with RL for LLMs, CAMERON R. WOLFE, PH.D., 2026.01
Paper/Blog Link My Issue
#Article #NLP #ReinforcementLearning #Blog #PostTraining Issue Date: 2026-01-26 Comment
元ポスト:
RLHF Book - Code Examples, Nathan Lambert, 2026.01
Paper/Blog Link My Issue
#Article #NLP #ReinforcementLearning #Repository #PostTraining #Selected Papers/Blogs #MinimalCode #Initial Impression Notes Issue Date: 2026-01-26 Comment
元ポスト:
Qwen 1.7Bモデルでの様々なRLアルゴリズムでのミニマルコード集。学習曲線つきで非常に実用的
A well known important feature to stabilize RL training is implementing the LM head in fp32 precision to help with gradients ... , Nathan Lambert, X, 2026.01
Paper/Blog Link My Issue
#Article #NLP #ReinforcementLearning #Post #PostTraining #Stability #One-Line Notes Issue Date: 2026-01-24 Comment
関連:
- MiniMax-M1, MiniMax, 2025.06
- [Paper Note] MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning
Attention, MiniMax+, arXiv'25, 2025.06
RLを安定化するためのtipsとそれによりMiniMax M1のplotが再現できたという話な模様。RLはこういった細かいテクニックが大事だと思うので、共有して頂けるのは大変ありがたい。
関連:
- [Paper Note] Defeating the Training-Inference Mismatch via FP16, Penghui Qi+, arXiv'25, 2025.10
- train-inference-gap && ReinforcementLearning ラベルが紐づいたissueも参照のこと
Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigations, Anthropic, 2026.01
Paper/Blog Link My Issue
#Article #NLP #Alignment #Evaluation #Blog #read-later Issue Date: 2026-01-23 Comment
元ポスト:
eval awareness mitigation
Claude's new constitution, Anthropic, 2026.01
Paper/Blog Link My Issue
#Article #NLP #Blog #Safety #One-Line Notes Issue Date: 2026-01-22 Comment
ClaudeのAI Modelで利用される新たなConstitution
関連:
- [Paper Note] Constitutional AI: Harmlessness from AI Feedback, Yuntao Bai+, arXiv'22
元ポスト:
MCP is Not the Problem, It's your Server: Best Practices for Building MCP Servers, PHILSCHMID, 2026.01
Paper/Blog Link My Issue
#Article #Infrastructure #SoftwareEngineering #MCP #AgentSkills Issue Date: 2026-01-22 Comment
元ポスト:
MCPサーバ構築に関するベストプラクティスが記載されている模様。
Designing AI-resistant technical evaluations, Anthropic, 2026.01
Paper/Blog Link My Issue
#Article #Education #AIAgents #Blog #read-later #Selected Papers/Blogs #Initial Impression Notes #Testing Issue Date: 2026-01-22 Comment
元ポスト:
Anthropicの採用における持ち帰り課題の変遷に関する記事。昔の持ち帰り課題では、応募者の大半よりもClaudeが上回るようになり採用におけるシグナルが拾いづらくなったのでリデザインが必要になった、そしてそれをどう変化させたか、といった話のようである。これは採用の話だがtestingという広い文脈で捉えるとかなり参考になる話に見える。
Claudeを作っている会社が自社が作ったプロダクトによって採用で苦しむという構造になっており、それに対してどのように対処したかという話題は非常に興味深いトピックだと感じる。
IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMs, Cheng+, 2026.01
Paper/Blog Link My Issue
#Article #NLP #ReinforcementLearning #Blog #PostTraining #KeyPoint Notes #Scalability Issue Date: 2026-01-22 Comment
元ポスト:
RLにおけるロールアウト数nのスケーリングは、シグモイド関数のような形状になりどこかのポイントで明確にサチるポイントが存在し、それ以上増やしても少量のゲインしか得られないポイントが存在する。これらのトレンドはeasy/hardな問題の双方で共通して見出されるが、原因は大きく異なっており、nを大きくするとeasyな問題ではworst@kが改善し、hardな問題ではbest@kが改善することで性能が向上する。つまり、簡単な問題に対してはより安定して正解できてミスが減り、困難な問題に対しては探索空間が広がり1回でも正解できる可能性が高まる。また、また、ハードウェア制約によりバッチサイズは基本的に固定されるので、ロールアウト数nと1バッチあたりに含められる問題数はトレードオフの関係となる。
このロールアウト数nに関する性質は、異なるベースモデル間で共通して生じるが、サチるポイントが異なる。問題セットのサイズで見ると、サイズが小さいと早々にoverfitするためサチるnのポイントも早くなる。問題難易度の分布がmixしているものであればnによるスケーリングのトレンドは維持されるが、評価する際のmetricsによってサチるぽいんとが左右される。nのスケーリングはdownstreamタスクの性能も向上させる。
と言った話らしい。
Fantastic Pretraining Optimizers and Where to Find Them 2.1: Hyperball Optimization, Wen+, 2026.01
Paper/Blog Link My Issue
#Article #NeuralNetwork #EfficiencyImprovement #Pretraining #NLP #Optimizer #read-later #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-01-22 Comment
元ポスト:
シンプルな手法で、先行研究によってモデルのパラメータサイズやデータのスケールが大きくなるとMuonのような行列ベースのoptimiserの高速化の恩恵が小さくなる現象を改善しているとのこと。
具体的には、重みを更新する際にweight decayのようなソフトにweightのノルムをコントロールするような仕組みを入れるのではなく、optimiserの重みに対する更新量と、更新後のネットワークの重みをフロベニウスノルムで正規化し、最適化の軌跡を半径Rの超球面の表面上に位置するように明示的に制約する(ここで、Rは最初の重み行列のフロベニウスノルム)。Muonを含む様々なoptimiserでも機能して学習効率を高めるため、インパクトの大きな重要研究に見える。
関連(concurrent works):
- [Paper Note] Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models, Yonggan Fu+, arXiv'25, 2025.11
- [Paper Note] Controlled LLM Training on Spectral Sphere, Tian Xie+, arXiv'26, 2026.01
関連:
- [Paper Note] Fantastic Pretraining Optimizers and Where to Find Them, Kaiyue Wen+, ICLR'26, 2025.09
ICLR 2026 Acceptance Prediction: Benchmarking Decision Process with A Multi-Agent System, Zhang+, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #Dataset #AIAgents #Evaluation #MultiModal #ScientificDiscovery #VisionLanguageModel #AcademicWriting #Live #One-Line Notes Issue Date: 2026-01-20 Comment
元ポスト:
conference paperのpeer reviewに関するベンチマーク。accept/rejectを予測する。papers, reviews, rebuttalsそしてfinal decisionsが紐づけられている。
GLM-4.7-Flash, Z.ai, 2026.01
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Coding #OpenWeight #MoE(Mixture-of-Experts) #One-Line Notes Issue Date: 2026-01-20 Comment
元ポスト:
関連:
- GLM-4.7: Advancing the Coding Capability, Z.ai, 2025.12
30B-A3BのMoEモデルで、gpt-oss-20B, Qwen3-30B-A3B-Thinking-2507を、SWE Bench Verified, tau2_bench, BrowseComp(SWEタスク, tooluse, 検索)等で大幅にoutperform。AIME, GPQA, HLEなどの推論系のベンチマークも同等以上。つまり、agenticなタスクに適した能力を有することが示唆される。
ポイント解説:
10,924x: The Instability Bomb at 1.7B Scale, TayKolasinski, 2026.01
Paper/Blog Link My Issue
#Article #Tutorial #MachineLearning #NLP #Blog #Selected Papers/Blogs #Reproducibility #ResidualStream Issue Date: 2026-01-19 Comment
元ポスト:
関連:
- [Paper Note] mHC: Manifold-Constrained Hyper-Connections, Zhenda Xie+, arXiv'25, 2025.12
- [Paper Note] Hyper-Connections, Defa Zhu+, ICLR'25, 2024.09
part1:
https://taylorkolasinski.com/notes/mhc-reproduction/
HC, mHCの説明が美しい図解と数式で説明されている。分かりやすい!
HCの課題とmHCがどのように解決したかを数式的、直感的に理解でき非常に有用
Pocket Flow: 100-line LLM framework. Let Agents build Agents, The-Rocket, 2026.01
Paper/Blog Link My Issue
#Article #NLP #Library #AIAgents #python #SoftwareEngineering #read-later #Selected Papers/Blogs #MinimalCode #Initial Impression Notes Issue Date: 2026-01-19 Comment
元ポスト:
たったの100行で実現されるミニマルなAI Agent/LLMフレームワークで、9種類の抽象化(Node, Flow, Shared, ...)でchat, agent, workflow, RAG, MCP, A2Aなどの様々なLLMをベースとした機能を実装できるフレームワークな模様。コード読みたい
Context Rot: How Increasing Input Tokens Impacts LLM Performance, CHROMA TECHNICAL REPORT, 2025.07
Paper/Blog Link My Issue
#Article #NLP #Blog #LongContext #read-later #ContextEngineering #ContextRot Issue Date: 2026-01-17
FrogMini-14B-2510, Microsoft, 2026.01
Paper/Blog Link My Issue
#Article #NLP #Supervised-FineTuning (SFT) #AIAgents #Coding #OpenWeight #SoftwareEngineering #One-Line Notes Issue Date: 2026-01-16 Comment
元ポスト:
strong modelから合成されたbug fixのtrajectoryでSFTすることで小規模モデルでSWE Benchの性能改善
Narrow Misalignment is Hard, Emergent Misalignment is Easy, Turner+, 2025.07
Paper/Blog Link My Issue
#Article #Analysis #NLP #Alignment #PEFT(Adaptor/LoRA) #PostTraining #One-Line Notes #EmergentMisalignment Issue Date: 2026-01-15 Comment
openreview: https://openreview.net/forum?id=q5AawZ5UuQ
一般的にevilになることを学習することが、狭義にevilになるよりも簡単だ、という知見を示した研究とのこと。
LongCat-Flash-Thinking-2601, Meituan, 2026.01
Paper/Blog Link My Issue
#Article #NLP #AIAgents #OpenWeight #MoE(Mixture-of-Experts) #Selected Papers/Blogs Issue Date: 2026-01-15 Comment
元ポスト:
解説:
coding, agentiaなベンチでTopTierを獲得した560B-27BのMoEモデル。MIT Licence
1MコンテキストウィンドウのZigzag attentionのモデルもcoming soon...だと...!?
Zigzag attentionはおそらく以下だろうか:
- [Paper Note] Efficient Context Scaling with LongCat ZigZag Attention, Chen Zhang+, arXiv'25, 2025.12
[Paper Note] Training large language models on narrow tasks can lead to broad misalignment, Nature 649, 2026.01
Paper/Blog Link My Issue
#Article #NLP #Alignment #Safety #read-later #Selected Papers/Blogs #Nature #EmergentMisalignment Issue Date: 2026-01-15 Comment
元ポスト:
元ポストによると、以下のような時系列でEmergent Misalignmentのliteratureは形成されていったらしい:
- [Paper Note] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs, Jan Betley+, arXiv'25, 2025.02
- [Paper Note] Persona Features Control Emergent Misalignment, Miles Wang+, arXiv'25, 2025.06
- [Paper Note] Model Organisms for Emergent Misalignment, Edward Turner+, arXiv'25, 2025.06
- [Paper Note] Convergent Linear Representations of Emergent Misalignment, Anna Soligo+, arXiv'25, 2025.06
- Narrow Misalignment is Hard, Emergent Misalignment is Easy, Turner+, 2025.07
- [Paper Note] School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs, Mia Taylor+, arXiv'25, 2025.08
- [Paper Note] Natural Emergent Misalignment from Reward Hacking in Production RL, Monte MacDiarmid+, arXiv'25, 2025.11
- [Paper Note] Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs, Jan Betley+, arXiv'25, 2025.12
GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image Generation, Z.ai, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #DiffusionModel #TextToImageGeneration #OpenWeight #Editing Issue Date: 2026-01-14 Comment
元ポスト:
Cowork: Claude Code for the rest of your work, Anthropic, 2026.01
Paper/Blog Link My Issue
#Article #NLP #AIAgents #GenerativeAI #Blog #WorkspaceAgents Issue Date: 2026-01-13 Comment
元ポスト:
競合(こちらは完全にオフラインで動作する):
- 🍫 Local Cocoa: Your Personal AI Assistant, Fully Local 💻, synvo-ai, 2026.01
MedReason-Stenographic, openmed-community, 2026.01
Paper/Blog Link My Issue
#Article #NLP #Dataset #QuestionAnswering #Chain-of-Thought #SyntheticData #Evaluation #Reasoning #Medical #KeyPoint Notes Issue Date: 2026-01-12 Comment
元ポスト:
MiniMax M2.1を用いてMedical QAに対してreasoning traceを生成。生成されたreasoning traceをstenographic formatと呼ばれる自然言語からフィラーを排除し、論理の流れのみをsymbolicな表現に変換することで合成されたデータセットとのこと。
ユースケースとしては下記とのこと:
> 1. Train reasoning models with symbolic compression
> 2. Fine-tune for medical QA
> 3. Research reasoning compression techniques
> 4. Benchmark reasoning trace quality
個人的には1,3が興味深く、symbolを用いてreasoning traceを圧縮することで、LLMの推論時のトークン効率を改善できる可能性がある。
が、surfaceがシンボルを用いた論理の流れとなると、汎化性能を損なわないためにはLLMが内部でシンボルに対する何らかの強固な解釈が別途必要になるし、それが多様なドメインで機能するような柔軟性を持っていなければならない気もする。
AI Safetyの観点でいうと、論理の流れでCoTが表現されるため、CoTを監視する際には異常なパターンがとりうる空間がshrinkし監視しやすくなる一方で、surfaceの空間がshrinkする代わりに内部のブラックボックス化された表現の自由度が高まり抜け道が増える可能性もある気がする。結局、自然言語もLLMから見たらトークンの羅列なので、本質的な課題は変わらない気はする。
SETA: Scaling Environments for Terminal Agents, CAMEL-AI, 2026.01
Paper/Blog Link My Issue
#Article #Tools #NLP #ReinforcementLearning #AIAgents #SyntheticData #Evaluation #Blog #Repository #SoftwareEngineering #PostTraining Issue Date: 2026-01-12 Comment
元ポスト:
HF: https://huggingface.co/datasets/camel-ai/seta-env
GitHubのreadmeに日本語がある!?
FineTranslations, Penedo+, 2026.01
Paper/Blog Link My Issue
#Article #MachineTranslation #Pretraining #NLP #Dataset #SyntheticData #mid-training #One-Line Notes Issue Date: 2026-01-10 Comment
元ポスト:
FineWeb2のテキストを英訳することで合成されたパラレルコーパスらしい
Demystifying evals for AI agents, Anthropic, 2026.01
Paper/Blog Link My Issue
#Article #Tutorial #NLP #AIAgents #Evaluation #Blog #Selected Papers/Blogs Issue Date: 2026-01-10 Comment
元ポスト:
NousCoder-14B: A Competitive Olympiad Programming Model, Joe Li, 2026.01
Paper/Blog Link My Issue
#Article #NLP #ReinforcementLearning #Blog #Coding #OpenWeight #PostTraining #read-later Issue Date: 2026-01-09 Comment
元ポスト:
HF:
https://huggingface.co/NousResearch/NousCoder-14B
Apache 2.0
PipelineRLを採用している模様。興味深い。
Introducing LFM2.5: The Next Generation of On-Device AI, LiquidAI, 2026.01
Paper/Blog Link My Issue
#Article #NLP #ReinforcementLearning #Blog #SmallModel #OpenWeight #Japanese #PostTraining #Selected Papers/Blogs #VisionLanguageModel #One-Line Notes #AudioLanguageModel Issue Date: 2026-01-09 Comment
元ポスト:
日本語に特化した言語モデルも存在し、Sarashina2.2-1b-instruct-v0.1, TinySwallow-1.5B-InstructよりもJMMLU, M-IFEval (ja), GSM8K (ja)においてより高い性能を発揮している。
LFM2.5-1.2B-Base: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-1.2B-Base)
LFM2.5-1.2B-Instruct: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct),
[LEAP](
https://leap.liquid.ai/models?model=lfm2.5-1.2b-instruct),
[Playground](
https://playground.liquid.ai/chat?model=cmk1jyp8f000204i56yy76uwh)
LFM2.5-1.2B-JP: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-1.2B-JP),
[LEAP](
https://leap.liquid.ai/models?model=lfm2.5-1.2b-jp)
LFM2.5-VL-1.6B: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-VL-1.6B),
[LEAP](
https://leap.liquid.ai/models?model=lfm2.5-vl-1.6b),
[Playground](
https://playground.liquid.ai/chat?model=cmk0wefde000204jp2knb2qr8),
[Demo](
https://huggingface.co/spaces/LiquidAI/LFM2.5-VL-1.6B-WebGPU)
LFM2.5-Audio-1.5B: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-Audio-1.5B),
[LEAP](
https://leap.liquid.ai/models?model=lfm2.5-audio-1.5b),
[Playground](
http://playground.liquid.ai/talk)
LiquidAIのモデルは日本語に特化したモデルが多く存在するのが特徴的に感じる。
LFM2-2.6B-Transcript, LiquidAI, 2026.01
Paper/Blog Link My Issue
#Article #NLP #OpenWeight #RecurrentModels #Transcript Issue Date: 2026-01-09 Comment
関連:
- Introducing LFM2: The Fastest On-Device Foundation Models on the Market, LiquidAI, 2025.07
[Paper Note] On the Slow Death of Scaling, Hooker+, 2026.01
Paper/Blog Link My Issue
#Article #NeuralNetwork #EfficiencyImprovement #NLP #Scaling Laws #Author Thread-Post Issue Date: 2026-01-09 Comment
元ポスト:
著者ポスト:
New post: nanochat miniseries v1,
Paper/Blog Link My Issue
#Article #Post #read-later Issue Date: 2026-01-09
🍫 Local Cocoa: Your Personal AI Assistant, Fully Local 💻, synvo-ai, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #Tools #NLP #AIAgents #MultiModal #Selected Papers/Blogs #ContextEngineering #memory Issue Date: 2026-01-09 Comment
元ポスト:
The next equalizer is not model architecture, but mastery over data behavior, gm8xx8, 2025.12
Paper/Blog Link My Issue
#Article #Pretraining #NLP #SyntheticData #Post #Selected Papers/Blogs #DataMixture #PhaseTransition Issue Date: 2026-01-07 Comment
関連(4-epochまで再利用するのがコスパが良いことを示した研究):
- [Paper Note] Scaling Data-Constrained Language Models, Niklas Muennighoff+, NeurIPS'23
関連(合成データの比率によるPhaseTransition):
- [Paper Note] Data Mixing Can Induce Phase Transitions in Knowledge Acquisition, Xinran Gu+, NeurIPS'25 Spotlight, 2025.05
- [Paper Note] Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls, Feiyang Kang+, EMNLP'25, 2025.10
- [Paper Note] Why Less is More (Sometimes): A Theory of Data Curation, Elvis Dohmatob+, arXiv'25, 2025.11
VAETKI, NC-AI-consortium, 2026.01
Paper/Blog Link My Issue
#Article #NLP #Reasoning #MultiLingual #OpenWeight #MoE(Mixture-of-Experts) Issue Date: 2026-01-03 Comment
元ポスト:
Solar-Open-100B, upstage, 2025.12
Paper/Blog Link My Issue
#Article #NLP #Reasoning #OpenWeight #MoE(Mixture-of-Experts) #Korean Issue Date: 2026-01-03 Comment
元ポスト:
ポイント解説:
K-EXAONE-236B-A23B, LG AI Research, 2025.12
Paper/Blog Link My Issue
#Article #NLP #Reasoning #MultiLingual #OpenWeight #MoE(Mixture-of-Experts) Issue Date: 2026-01-03 Comment
関連:
- EXAONE-Deep-32B, LG AI Research, 2025.03
Multi Token Prediction
Sliding Window Attention
256k context length
MoE
元ポスト:
A.X-K1, SK Telecom, 2026.01
Paper/Blog Link My Issue
#Article #NLP #Reasoning #OpenWeight #MoE(Mixture-of-Experts) #Korean Issue Date: 2026-01-03 Comment
元ポスト:
Production-Grade Agentic AI System, FareedKhan-dev, 2025.12
Paper/Blog Link My Issue
#Article #Tutorial #NLP #AIAgents #SoftwareEngineering #read-later Issue Date: 2026-01-03 Comment
元ポスト:
Recursive Language Models: the paradigm of 2026, PRIME Intellect, 2026.01
Paper/Blog Link My Issue
#Article #NLP #Blog #LongContext #read-later #Selected Papers/Blogs #LatentReasoning #reading #RecursiveModels #ContextRot Issue Date: 2026-01-02 Comment
関連研究:
- [Paper Note] Recursive Language Models, Alex L. Zhang+, arXiv'25, 2025.12
- Context Rot: How Increasing Input Tokens Impacts LLM Performance, CHROMA TECHNICAL REPORT, 2025.07
- [Paper Note] Scaling Long-Horizon LLM Agent via Context-Folding, Weiwei Sun+, arXiv'25, 2025.10
- [Paper Note] AgentFold: Long-Horizon Web Agents with Proactive Context Management, Rui Ye+, arXiv'25, 2025.10
- [Paper Note] Agentic Context Engineering: Evolving Contexts for Self-Improving
Language Models, Qizheng Zhang+, arXiv'25, 2025.10
IQuest-Coder, IQuestLab, 2026.01
Paper/Blog Link My Issue
#Article #NLP #Coding #OpenWeight #SoftwareEngineering Issue Date: 2026-01-01 Comment
元ポスト:
Deriving the DPO Loss from First Principles, aayush garg, 2025.12
Paper/Blog Link My Issue
#Article #Tutorial #NLP #ReinforcementLearning #Blog #DPO #PostTraining #read-later Issue Date: 2025-12-31 Comment
元ポスト:
関連:
- Deriving the PPO Loss from First Principles, aayush garg, 2025.12
Today's conversations about AI-assisted programming are strikingly similar to those from decades ago about the choice between low-level languages like C versus high-level languages like Python, Arvind Narayanan, 2025.12
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Coding #Post #SoftwareEngineering Issue Date: 2025-12-31
LLMRouter: An Open-Source Library for LLM Routing, Feng+, 2025.12
Paper/Blog Link My Issue
#Article #Tools #NLP #python #SoftwareEngineering #Routing #Orchestration Issue Date: 2025-12-30 Comment
元ポスト:
SpecBundle & SpecForge v0.2: Production-Ready Speculative Decoding Models and Framework, Spec Forge Team+, lmsys org, 2025.12
Paper/Blog Link My Issue
#Article #NLP #Blog #LLMServing #SpeculativeDecoding Issue Date: 2025-12-28 Comment
元ポスト:
Reverse Engineering a Phase Change in GPT's Training Data... with the Seahorse Emoji 🌊🐴, PRATYUSH MAINI, 2025.12
Paper/Blog Link My Issue
#Article #Analysis #NLP #ChatGPT #Reasoning #SelfCorrection #mid-training #One-Line Notes Issue Date: 2025-12-28 Comment
元ポスト:
Is there seahorse emoji?という質問に対するLLMのreasoning trajectoryと、self correctionの挙動が、OpenAIのどの時点のモデルで出現するか、しないかを線引くことで、mid-trainingにself correction形式のデータが追加されたのがいつ頃なのかを考察している。
mini-sglang: A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems, sgl-project, 2025
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #python #Repository #LLMServing #SoftwareEngineering #read-later #Selected Papers/Blogs #MinimalCode Issue Date: 2025-12-28 Comment
元ポスト:
めっちゃ勉強したい
ノーコードで言語モデルの「学習」を体験できるMN-Core Playground _ SLM Customizeの遊び方, PFN, 2025.12
Paper/Blog Link My Issue
#Article #NLP #Blog #SmallModel #Japanese #PostTraining Issue Date: 2025-12-27 Comment
元ポスト:
Aligning to What? Rethinking Agent Generalization in MiniMax M2, MiniMaxAI, 2025.12
Paper/Blog Link My Issue
#Article #NLP #Alignment #AIAgents #Blog #Reasoning #read-later Issue Date: 2025-12-27 Comment
元ポスト:
Deriving the PPO Loss from First Principles, aayush garg, 2025.12
Paper/Blog Link My Issue
#Article #Tutorial #NLP #ReinforcementLearning #Blog #PostTraining #read-later Issue Date: 2025-12-27 Comment
元ポスト:
