read-later (1126) — 5/6
终局奖励之后:从 Apodex 1.1 看长程 Agentic RL 如何分配 credit, 发布于, 2026.09
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #ReinforcementLearning #AIAgents #PostTraining #CreditAssignment Issue Date: 2026-09-06 Comment
元ポスト:
Training frontier knowledge work agents: A 397B RL training guide with SkyRL, Mercor and SkyRL, 2026.09
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #ReinforcementLearning #AIAgents #Blog #PostTraining #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-09-06 Comment
元ポスト:
397Bモデルに対するDPPOによる事後学習によってAPEX-Agentsのスコアを+11.2%を実現するための取り組みと地検が共有されているようである。
関連:
- [Paper Note] APEX-Agents, Bertie Vidgen+, arXiv'26, 2026.01
- [Paper Note] Rethinking the Trust Region in LLM Reinforcement Learning, Penghui Qi+, arXiv'26, 2026.02
Atlas: A World Model for Spatial Intelligence, World Labs, 2026.09
Paper/Blog Link My Issue
#Article #ComputerVision #Transformer #DiffusionModel #Blog #Proprietary #Selected Papers/Blogs #3D Reconstruction #WorldModels #SpatialUnderstanding #Author Thread-Post #4D(Scene + Time) Issue Date: 2026-09-06 Comment
元ポスト:
解説:
How Far Are We from a Native Autoregressive Video Model?, Weiyang Jin, 2026.08
Paper/Blog Link My Issue
#Article #ComputerVision #Blog #VideoGeneration/Understandings #Author Thread-Post Issue Date: 2026-09-06 Comment
元ポスト:
[Paper Note] Why Pretraining Fails to Share Cross-Lingual Knowledge, Fittschen+, 2026.08
Paper/Blog Link My Issue
#Article #Analysis #Pretraining #CrossLingual Issue Date: 2026-09-05 Comment
元ポスト:
Hy4-preview, Tencent, 2026.08
Paper/Blog Link My Issue
#Article #OpenWeight #Author Thread-Post Issue Date: 2026-09-05 Comment
blog: https://hy.tencent.com/research/hy4-preview
元ポスト:
公式:
Previewing the Model Hardware Standard, Anthropic, 2026.08
Paper/Blog Link My Issue
#Article #Author Thread-Post Issue Date: 2026-09-05 Comment
元ポスト:
Hugging Face のインシデントと今後の道筋, OpenAI, 2026.08
Paper/Blog Link My Issue
#Article #Analysis #NLP #LanguageModel #AIAgents #Blog #Selected Papers/Blogs #Security #Author Thread-Post Issue Date: 2026-08-30 Comment
元ポスト:
- OpenAI and Hugging Face partner to address security incident during model evaluation, OpenAI, 2026.07
の件の調査結果のようである。
Introducing Pipette: A benchmarking suite for on-device intelligence, Liquid AI, 2026.08
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Evaluation #Blog #SmallModel #Author Thread-Post Issue Date: 2026-08-29 Comment
元ポスト:
Scaling Activation Oracles to Trillion-Parameter Models: Oracles improve with model size, data size, and data quality, Transluce, 2026.08
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Selected Papers/Blogs #Author Thread-Post Issue Date: 2026-08-22 Comment
元ポスト:
GEN-1.5: Embodied Foundation Models are One-Shot Learners, Generalist, 2026.08
Paper/Blog Link My Issue
#Article #Zero/Few/ManyShotPrompting #FoundationModel #In-ContextLearning #Blog #Selected Papers/Blogs #Robotics #Initial Impression Notes #Author Thread-Post Issue Date: 2026-08-20 Comment
元ポスト:
Embodied AIのFoundation Modelの話で、とうとう一つのデモンストレーションからタスクを学習できる(In-Context Learning)ようになったようである。
関連:
- GEN1: Scaling Embodied Foundation Models to Mastery, Generalist AI Team, 2026.04
所見:
解説:
大模型 MoE 负载均衡的 N 种方法, 觉醒也醉了, qingke_ai, 2026.08
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #Architecture #MoE(Mixture-of-Experts) Issue Date: 2026-08-20 Comment
元ポスト:
Neural chameleons can('t) hide from activation oracles, LW, 2026.01
Paper/Blog Link My Issue
#Article #LanguageModel #Safety #Probing #Monitorability Issue Date: 2026-08-20
Inside: a Neural Chameleon Reverse-engineering how a language model hides from activation monitors, Jackson Mowatt Gok, 2026.08
Paper/Blog Link My Issue
#Article #LanguageModel #Safety #Monitorability #Author Thread-Post Issue Date: 2026-08-20 Comment
元ポスト:
関連:
- Neural chameleons can('t) hide from activation oracles, LW, 2026.01
DFlash 2: Keep Drafting Parallel, Inco AI, 2026.08
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #DiffusionModel #Selected Papers/Blogs #SpeculativeDecoding #Author Thread-Post Issue Date: 2026-08-19 Comment
元ポスト:
関連:
- [Paper Note] DFlash: Block Diffusion for Flash Speculative Decoding, Jian Chen+, arXiv'26, 2026.02
AI Control: An Assessment of Frontier Practices, Guidelight, 2026.08
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #Proprietary #Safety #Selected Papers/Blogs #Author Thread-Post Issue Date: 2026-08-19 Comment
著者ポスト:
元ポスト:
Pacing model development in an era of cyber-critical capabilities, OpenAI, 2026.08
Paper/Blog Link My Issue
#Article #LanguageModel #Chain-of-Thought #Monitorability #Author Thread-Post Issue Date: 2026-08-19 Comment
元ポスト:
所見:
Why Removing the Vision Encoder Can Be Better — From an Infra Perspective, Yuwei Niu, 2026.08
Paper/Blog Link My Issue
#Article #MultiModal #VisionLanguageModel Issue Date: 2026-08-19 Comment
元ポスト:
How to Parallelize a Transformer for Training, Edward Z. Yang, 2026.08
Paper/Blog Link My Issue
#Article #Author Thread-Post Issue Date: 2026-08-19 Comment
元ポスト:
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers, Aarsen+, 2026.08
Paper/Blog Link My Issue
#Article #Embeddings #Blog #Author Thread-Post Issue Date: 2026-08-19 Comment
元ポスト:
Understanding a Law Firm through Study, Engram, 2026.08
Paper/Blog Link My Issue
#Article #Author Thread-Post Issue Date: 2026-08-18 Comment
元ポスト:
Outperforming cuBLAS on NVFP4, PRANJAL SHANKHDHAR, 2026.08
Paper/Blog Link My Issue
#Article #NeuralNetwork #EfficiencyImprovement #NLP #LanguageModel #SoftwareEngineering #GPUKernel #Author Thread-Post Issue Date: 2026-08-17 Comment
元ポスト:
Reflections on Video DeltaNet, Haoyi Zhu, 2026.08
Paper/Blog Link My Issue
#Article #ComputerVision #Blog #VideoGeneration/Understandings #LinearAttention #Author Thread-Post Issue Date: 2026-08-14 Comment
元ポスト:
videoに対してdelta netのようなlinear attentionモデルを適用することに関する考察のようである。
Incident Report: unsanctioned agent behaviour during cyber testing, AISI, 2026.08
Paper/Blog Link My Issue
#Article #AIAgents #Blog #Selected Papers/Blogs #Security Issue Date: 2026-08-10 Comment
元ポスト:
所見:
所見:
関連:
- OpenAI and Hugging Face partner to address security incident during model evaluation, OpenAI, 2026.07
- [Paper Note] Chunky Post-Training: Data Driven Failures of Generalization, Seoirse Murray+, arXiv'26, 2026.02
- [Paper Note] Natural Emergent Misalignment from Reward Hacking in Production RL, Monte MacDiarmid+, arXiv'25, 2025.11
TutorMoments: Do AI tutors know when to help and when to hold back?, Ai2, 2026.08
Paper/Blog Link My Issue
#Article #Education #Blog #Author Thread-Post Issue Date: 2026-08-10 Comment
元ポスト:
Defending Against the Training–Inference Numeric Mismatch in RL (Especially Linear Attention) — and Whether It Helps Async RL, Yichuan Wang in collaboration with the TorchTitan team, 2026.08
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #ReinforcementLearning #Blog #PostTraining #Selected Papers/Blogs #train-inference-mismatch #Initial Impression Notes #Asynchronous Issue Date: 2026-08-10 Comment
元ポスト:
linear attentionモデルにおいて、初めてRLにおけるtrain-inference mismatchを完全に解消した実装とのことで、非常にインパクトが大きそうに見える。
string2string Studio: Explore how strings line up, differ, or match, Mirac Suzgun, 2026.08
Paper/Blog Link My Issue
#Article #Blog #Author Thread-Post Issue Date: 2026-08-10 Comment
元ポスト:
VISTA: A Visual Harness for Reasoning in an Interactive World, Han+, 2026.08
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #LanguageModel #AIAgents #Reasoning #Selected Papers/Blogs #2D (Image) #3D (Scene) #memory #LongHorizon #Author Thread-Post #AgentHarness Issue Date: 2026-08-10 Comment
元ポスト:
From LR to ELR: A Better Heuristic for Pretraining Dynamics, Hy Pretrain Team, 2026.08
Paper/Blog Link My Issue
#Article #Pretraining #Selected Papers/Blogs #Author Thread-Post Issue Date: 2026-08-10 Comment
元ポスト:
Introducing Muse Code and Muse Spark 1.2, Meta, 2026.08
Paper/Blog Link My Issue
#Article #Proprietary #Selected Papers/Blogs #Author Thread-Post Issue Date: 2026-08-10 Comment
元ポスト:
ザッカーバーグ氏によるポスト:
Alexandr Wang氏によるポスト:
データ収集をオプトインするとtokenのコストが1/10未満になる。
GPT 5.6 Terraと同等程度のArtificial Analysis Index:
Meta Superintelligence Lab立ち上げからここまで、すさまじいスピード感である。
Kimi K3, The Manos, The Mythos, The Legendos, CHEN+, 2026.08
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #Attention #Blog #MoE(Mixture-of-Experts) #Routing #Initial Impression Notes #LinearAttention #Author Thread-Post Issue Date: 2026-08-10 Comment
関連:
- Kimi K3: Open Frontier Intelligence, Moonshot AI, 2026.07
元ポスト:
以下のようなKimi K3で利用されている技術の解説:
- linear attention
- DeltaNet
- GatedDeltaNet
- FlashKDA
- Kimi Linear
- MLA
- Attention Residuals
- LatentMoE
- Quantile load balancing (QB)
- KVCache Efficiency
関連:
- [Paper Note] Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention, Angelos Katharopoulos+, ICML'20
- [Paper Note] Linear Transformers Are Secretly Fast Weight Programmers, Imanol Schlag+, arXiv'21, 2021.02
- [Paper Note] Parallelizing Linear Transformers with the Delta Rule over Sequence Length, Songlin Yang+, NeurIPS'24, 2024.06
- FlashKDA: Flash Kimi Delta Attention — high-performance KDA kernels built on CUTLASS, MoonshotAI, 2026.04
- [Paper Note] Kimi Linear: An Expressive, Efficient Attention Architecture, Kimi Team+, arXiv'25, 2025.10
- [Paper Note] DeepSeek-V3 Technical Report, DeepSeek-AI+, arXiv'24, 2024.12
- [Paper Note] Attention Residuals, Kimi Team+, arXiv'26, 2026.03
- [Paper Note] LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts, Venmugil Elango+, arXiv'26, 2026.01
QBをチェックしたい
Towards Looped Models Done Right, Huang+, 2026.08
Paper/Blog Link My Issue
#Article #Analysis #NLP #LanguageModel #Blog #RecurrentModels #RecursiveModels #Author Thread-Post Issue Date: 2026-08-09 Comment
元ポスト:
関連:
- [Paper Note] Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach, Jonas Geiping+, NeurIPS'25
- [Paper Note] Scaling Latent Reasoning via Looped Language Models, Rui-Jie Zhu+, arXiv'25, 2025.10
代表的なLooped Modelsアーキテクチャに対して、apple-to-appleな比較実験を実施し結果を考察しているようである。
Online KL Shampoo, Tilde, 2026.07
Paper/Blog Link My Issue
#Article #Pretraining #NLP #LanguageModel #Optimizer Issue Date: 2026-08-06 Comment
元ポスト:
Muonよりもさらに高性能なoptimiserのようである。
関連:
- [Paper Note] Shampoo: Preconditioned Stochastic Tensor Optimization, Vineet Gupta+, ICML'18, 2018.02
- [Paper Note] Muon is Scalable for LLM Training, Jingyuan Liu+, arXiv'25, 2025.02
- [Paper Note] SOAP: Improving and Stabilizing Shampoo using Adam, Nikhil Vyas+, ICLR'25
- Fantastic Pretraining Optimizers and Where to Find Them 2.1: Hyperball Optimization, Wen+, 2026.01
Scaling Automated Post-Training, Intology, 2026.08
Paper/Blog Link My Issue
#Article #Blog #PostTraining #Author Thread-Post Issue Date: 2026-08-05 Comment
元ポスト:
Mixture-of-Kittens: our open-source MoE megakernel for NVL72s, Cursor, 2026.08
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #LanguageModel #Blog #MoE(Mixture-of-Experts) #Selected Papers/Blogs #Initial Impression Notes #Author Thread-Post Issue Date: 2026-08-05 Comment
元ポスト:
主要なOpenWeight LLM(アーキテクチャ)のスループットが軒並み2倍以上になっているので、これはかなりインパクトがでかい話に見える
関連:
- DeepEP: an efficient expert-parallel communication library, Zhao+, DeepseekAI, 2025.09
- MoonEP, MoonshotAI, 2026.07
ポイント解説:
所見:
著者ポスト:
SIGReg from First Principles, Reza Bayat, 2026.07
Paper/Blog Link My Issue
#Article #Tutorial #ComputerVision #RepresentationLearning #Self-SupervisedLearning #WorldModels #Author Thread-Post #LatentRepresentation Issue Date: 2026-08-04 Comment
元ポスト:
JEPA関連のチュートリアル
関連:
- [Paper Note] Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture, Mahmoud Assran+, CVPR'23, 2023.01
- [Paper Note] LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics, Randall Balestriero+, arXiv'25, 2025.11
- [Paper Note] LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels, Lucas Maes+, arXiv'26, 2026.03
- JEPAwiki, mishig, 2026.04
著者ポスト:
How much science is verifiable? Results from replicating ICML 2026 oral papers, SAI, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Blog #ScientificDiscovery #Selected Papers/Blogs #Reproducibility #Science #Initial Impression Notes Issue Date: 2026-08-03 Comment
元ポスト:
ICML 2026のOralとして採択された論文の再現実験を通じて得られた知見がまとめられているようで、非常に面白そう。
Surfacing Benchmark-Maxxing in Kimi-K3, arjun, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Evaluation #Post #Initial Impression Notes #RewardSeeking Issue Date: 2026-08-03 Comment
元ポスト:
関連:
- Kimi K3: Open Frontier Intelligence, Moonshot AI, 2026.07
Kimi K3における、評価されていると認識した上で評価作成者の意図の検討や、reference solutionに関する検討をするなど、ベンチマークに過度に最適化されているような挙動を示す現象に関する検討な模様。
エージェントスウォームと新しいモデルの経済性, Cursor, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Blog #Coding #SoftwareEngineering #Initial Impression Notes Issue Date: 2026-08-03 Comment
元ポスト:
コストが最大15倍も利用するモデルにより変化したとのことだが、コードの品質はどの程度変化したのだろうか?(まだ読めていない
Language model harnesses are compositional generalizers, Zhang+, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Blog #Generalization #RecurrentModels #RecursiveModels #AgentHarness Issue Date: 2026-07-31
Extending Legal Agent Bench to M&A Due Diligence, Harvey, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Evaluation #Post #Legal Issue Date: 2026-07-31 Comment
元ポスト:
Agentic World Models, Cameron R. Wolfe, Ph.D., 2026.07
Paper/Blog Link My Issue
#Article #Tutorial #NLP #AIAgents #Blog #WorldModels #Author Thread-Post Issue Date: 2026-07-31 Comment
元ポスト:
What I’ve learned about designing the stack for power efficiency from my time at nvidia architecture research, Arya Tschand, 2026.07
Paper/Blog Link My Issue
#Article #LanguageModel #Infrastructure #Post Issue Date: 2026-07-31 Comment
エネルギー効率の良いデータセンターを構築するにはどうすれば良いのか?といった話題のようで、とてもおもしろそうなので読みたい。
※ issueのタイトルは、ポストにタイトルがなかったため、私がポスト中の文言から切り出したもので、公式のタイトルではありません。
LLMエージェント評価の現在地 — 主要ベンチマークに学ぶ設計原則, Hiraoka, 2026.07
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #AIAgents #Evaluation #Author Thread-Post Issue Date: 2026-07-31 Comment
元ポスト:
こちらの記事を含めて元ポストに5本、日本語でAI Agentの評価に関して実際に手を動かした知見がまとまっているようで、読みたい。
The fastest LLM on Earth: Frontier level intelligence, 15x faster, celeris, 2026.07
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #DiffusionModel #Blog #Proprietary #One-Line Notes #Author Thread-Post Issue Date: 2026-07-25 Comment
元ポスト:
アーキテクチャとしてdLLMを採用したGPT-5-miniと同等程度の性能のモデルで、GPT-5-miniに対して15倍程度(1280TPS)のスループットを実現しているとのこと。
Artificial Analysisによるoutput speed評価で1位:
[Paper Note] LLaDA2.2, InclusionAI, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #DiffusionModel #MoE(Mixture-of-Experts) #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-07-25 Comment
HF: https://huggingface.co/inclusionAI/LLaDA2.2-flash
元ポスト:
LLaDA2.2シリーズとして初めてのAgent指向なdLLMとのこと。MoEアーキテクチャ採用。ベンチマークスコアはARモデルよりも少し劣る面があるが、スループットは1.64倍。
Inside TPU and GPU Clusters: The Anatomy of Collective Communication, Aleksa Gordić, 2026.07
Paper/Blog Link My Issue
#Article #Selected Papers/Blogs #Author Thread-Post Issue Date: 2026-07-19 Comment
元ポスト:
- Inside vLLM: Anatomy of a High-Throughput LLM Inference System, Aleksa Gordić blog, 2025.08
The Future Worth Building Is Human, THINKING MACHINES, 2026.07
Paper/Blog Link My Issue
#Article #NLP #GenerativeAI #Blog Issue Date: 2026-07-19 Comment
元ポスト:
Controlling Reasoning Effort in LLMs, Sebastian Raschka, 2026.07
Paper/Blog Link My Issue
#Article #Controllable #NLP #LanguageModel #Blog #Reasoning #Length #Initial Impression Notes Issue Date: 2026-07-19 Comment
元ポスト:
OpenAIのReasoning Effortについては調整方法が公開されていないのでGPT-OSSのテンプレートなどから可能な方法を推測、そのほかOpenWeightモデルについてはテクニカルレポートに記載されている情報からReasoning Effortの調整方法が解説されている
PLaMoベースの強い日本語報酬モデルの構築, PFN, 2026.07
Paper/Blog Link My Issue
#Article #NLP #Blog #Japanese #RewardModel Issue Date: 2026-07-17 Comment
元ポスト:
What If the Harness Comes Before Pretraining? A Data Flywheel Perspective, Hanchen Li, 2026.07
Paper/Blog Link My Issue
#Article #Pretraining #LanguageModel #Post #Data Issue Date: 2026-07-15 Comment
元ポスト:
AIDE²: The First Evidence of Recursive Self-Improvement, Waco Team, 2026.07
Paper/Blog Link My Issue
#Article #NLP #SelfImprovement #autoresearch/RSI #Author Thread-Post #AgentHarness Issue Date: 2026-07-15 Comment
元ポスト:
FastAFD: Open-Source Large-Scale Attention-FFN Disaggregation on Blackwell NVL72, Hao AI Lab, 2026.07
Paper/Blog Link My Issue
#Article #LLMServing #Author Thread-Post Issue Date: 2026-07-14 Comment
元ポスト:
Reducing Doom Loops with Final Token Preference Optimization, Liquid.AI, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #Reasoning #Author Thread-Post Issue Date: 2026-07-12 Comment
元ポスト:
Workspace geometry, model by model., elie, 2026.07
Paper/Blog Link My Issue
#Article #Analysis #NLP #LanguageModel #Blog #OpenWeight #Interpretability #Initial Impression Notes Issue Date: 2026-07-12 Comment
元ポスト:
38種類のOpenWeightモデルのlayer間でのJ-lensの類似性を検証したとのこと
アーキテクチャが異なっていても非常にモデル間で類似しているようである。gemma-4やSLM(Qwen, GPT2)などは外れ値を含む模様
関連:
- Verbalizable Representations Form a Global Workspace in Language Models, Anthropic, 2026.07
Harness Engineering for Self-Improvement, Lilian Weng, 2026.07
Paper/Blog Link My Issue
#Article #Tutorial #NLP #AIAgents #Blog #SelfImprovement #RecursiveModels #Author Thread-Post #AgentHarness Issue Date: 2026-07-12 Comment
元ポスト:
Benchmarking Coding Agents on Databricks’ Multi-Million Line Codebase, databricks, 2026.07
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #LanguageModel #AIAgents #Evaluation #Blog #Initial Impression Notes #Author Thread-Post #AgentHarness Issue Date: 2026-07-10 Comment
元ポスト:
ハーネス選択によってトークン消費を1/2にでき、GLM 5.2のコストパフォーマンスが良いらしい
How to Build a Diffusion Language Model, Kuleshov Group, 2026.07
Paper/Blog Link My Issue
#Article #Tutorial #LanguageModel #DiffusionModel #Author Thread-Post Issue Date: 2026-07-09 Comment
元ポスト:
The optimizer that outlived its proof: A case study on Adam, Sadhika Malladi, 2026.07
Paper/Blog Link My Issue
#Article #Blog #Optimizer #Author Thread-Post Issue Date: 2026-07-07 Comment
元ポスト:
[Paper Note] ASPIRE: Agentic _Skills Discovery for Robotics, Lu+, NVIDIA, 2026.07
Paper/Blog Link My Issue
#Article #Robotics #EmbodiedAI #AgentSkills Issue Date: 2026-07-07 Comment
元ポスト:
Devin Fusion: Frontier Performance at 35% Lower Cost, Cognition, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Architecture #ContextEngineering #One-Line Notes #Author Thread-Post #AgentHarness Issue Date: 2026-07-07 Comment
元ポスト:
高価で高性能なメインエージェントがタスクを管理し、安価なサイドキックエージェントがメインエージェントからタスクを移譲され処理をするようなアーキテクチャによって、メインエージェントで処理した場合と同等の性能を発揮しつつコストを削減できる。
似たような考え方はAgent Swarmなどでも実施されている。Agent Swarmはハーネスだけでなく、モデル自身もサブエージェントをうまく活用できるような学習のされかたをしている:
- [Paper Note] Kimi K2.5: Visual Agentic Intelligence, Kimi Team+, arXiv'26, 2026.02
Statistics for LLM Evals, evalstats project, 2026.03
Paper/Blog Link My Issue
#Article #LanguageModel #Evaluation Issue Date: 2026-07-07 Comment
github: https://github.com/ianarawjo/evalstats
元ポスト:
From Brain Waves to Words: Brain2Qwerty Offers a New Path to Communication Without Surgery, Meta, 2026.06
Paper/Blog Link My Issue
#Article #Author Thread-Post Issue Date: 2026-07-05 Comment
元ポスト:
DiffusionGemma: The First Diffusion LLM (dLLM) Natively Supported in vLLM, vLLM, 2026.06
Paper/Blog Link My Issue
#Article #DiffusionModel #Blog #LLMServing Issue Date: 2026-07-05 Comment
元ポスト:
関連:
- [Paper Note] How Transparent is DiffusionGemma?, Joshua Engels+, arXiv'26, 2026.06
Disaggregated Inference: 18 Months Later, Chen+, Hao AI Lab, 2025.11
Paper/Blog Link My Issue
#Article #LanguageModel #Blog #LLMServing Issue Date: 2026-07-05 Comment
元ポスト:
[Paper Note] Economic Evaluations of Language Models, Wan+, 2026.06
Paper/Blog Link My Issue
#Article #LanguageModel #AIAgents #Evaluation #Selected Papers/Blogs #Author Thread-Post Issue Date: 2026-07-05
Agentic RL: Frameworks and Best Practices, CAMERON R. WOLFE, 2026.06
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #ReinforcementLearning #AIAgents #Blog #PostTraining Issue Date: 2026-07-03
Speculative Decoding - The Bits and the Bytes, Aakash Kumar Nain, 2026.06
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #Blog #Decoding #SpeculativeDecoding #Author Thread-Post Issue Date: 2026-07-03 Comment
元ポスト:
Speculation Is All You Need, Modal, 2026.06
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #Decoding #SpeculativeDecoding #Author Thread-Post Issue Date: 2026-07-03 Comment
元ポスト:
[Paper Note] EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments, ByteDance Seed, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Evaluation #Selected Papers/Blogs #Initial Impression Notes #Author Thread-Post Issue Date: 2026-07-03 Comment
blog: https://edge-bench.org/
人間が平均で57.2時間かかるようなlong horizonなタスクによるAgentのベンチマーク
Learning to Replicate Expert Judgment in Financial Tasks, THINKING MACHINES, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #ReinforcementLearning #Blog #Financial #PostTraining #Author Thread-Post Issue Date: 2026-07-02 Comment
元ポスト:
Sarashina3 embedding: 日本語に強い最新のテキスト埋め込みモデル, SB Intuitions, 2026.07
Paper/Blog Link My Issue
#Article #Embeddings #NLP #RepresentationLearning #Japanese Issue Date: 2026-07-02 Comment
元ポスト:
Introducing LongCat-2.0, LongCat, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #OpenWeight #Selected Papers/Blogs #Reference Collection #Initial Impression Notes #Author Thread-Post Issue Date: 2026-06-30 Comment
元ポスト:
- 1M context window
- Sparse Attention
- Muon
- Ngram Embedding
- dynamic activation (33B-56B)
- Multi-teacher OPD
HF:
https://huggingface.co/meituan-longcat/LongCat-2.0
現在はまだ公開されていないがモデルは上記で公開予定
Sparse Attention + 1M Context + MOPD + Muonが標準になっている印象
LongCat Sparse Attentionポイント解説:
所見:
50k基のH800 GPUで学習されているとのこと。
所見:
ポイント解説:
Scaling Laws, Carefully, Lilian Weng, 2026.06
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #Scaling Laws Issue Date: 2026-06-26 Comment
元ポスト:
Thinking to recall: How reasoning unlocks parametric knowledge in LLMs, Google Research, 2026.06
Paper/Blog Link My Issue
#Article #Analysis #LanguageModel #Blog #Reasoning #FactualKnowledge #Author Thread-Post Issue Date: 2026-06-26 Comment
元ポスト:
分散推論基盤やその前提の考え方 〜高火力 PHYで作る分散推論基盤 vol.1〜, 道下幹也, さくらのナレッジ, 2025.11
Paper/Blog Link My Issue
#Article #Blog #LLMServing Issue Date: 2026-06-22 Comment
元ポスト:
推論基盤のパフォーマンス検証と最適化戦略, Mikiya Michishita, 第3回 vLLM roundup Community Meetup Tokyoの登壇資料, 2026.03
Paper/Blog Link My Issue
#Article #LLMServing #Slide #Author Thread-Post Issue Date: 2026-06-22 Comment
元ポスト:
Kubernetesにおける学習基盤とLLMOpsの概要, ry, Kubernetes祭り #1 登壇資料, 2026.06
Paper/Blog Link My Issue
#Article #ML-LLM Ops #Slide Issue Date: 2026-06-22
Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era, Mang+, 2026.06
Paper/Blog Link My Issue
#Article #Analysis #NLP #LanguageModel #AIAgents #Coding #Selected Papers/Blogs #LongHorizon #Initial Impression Notes #Author Thread-Post Issue Date: 2026-06-22 Comment
2週間にわたるopen endなプログラミングコンテストのデータを用いて人間の専門家とAI AgentのElo ratingの変遷を比較すると、序盤は人間に対して大きくリードしスコアは試行回数の対数に対して線形にスケーリングし次第に停滞を始めるが、人間の専門家は4日を超えたあたりから非線形にスコアが進化し、最終的にAI Agentよりも高いスコアを記録する。人間はAI Agentと比較して、継続学習ができていることが示唆され、AI Agentのlong horizonタスクに対する能力の限界が示唆される、という話に見える。
元ポスト:
人工知能関連の国際会議と 人工知能関連の国際会議と プロシーディングス運営の現状と展望, 杉山将, 第40回人工知能学会全国大会, 2026.06
Paper/Blog Link My Issue
#Article #Slide Issue Date: 2026-06-21 Comment
元ポスト:
Deli_AutoResearch, Deli Chen, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #SelfImprovement #AgentSkills #autoresearch/RSI #Author Thread-Post Issue Date: 2026-06-20 Comment
元ポスト:
RL Systems Mind the Gap: Matching Trainer and Generator Throughput, Chen+, 2026.06
Paper/Blog Link My Issue
#Article #ReinforcementLearning #Blog #Selected Papers/Blogs Issue Date: 2026-06-17 Comment
元ポスト:
所見:
First Steps Toward Automated AI Research, Recursive Superintelligence, 2026.06
Paper/Blog Link My Issue
#Article #Blog #SelfImprovement #Initial Impression Notes #autoresearch/RSI #Author Thread-Post Issue Date: 2026-06-14 Comment
元ポスト:
word2vec, GloVe, Rucursive ModelのRichard Socher氏のポスト
フィジカルAI 日本の処方箋[前編]国を挙げて議論渦巻く、今何に取り組むべきか, 進藤智則, 日経Robotics編集長, 2026.06
Paper/Blog Link My Issue
#Article Issue Date: 2026-06-13 Comment
元ポスト:
Scaling Video Training with Parallelism, Chen+, 2026.06
Paper/Blog Link My Issue
#Article #Tutorial #Blog #Parallelism #3D (Video) Issue Date: 2026-06-11 Comment
元ポスト:
[Paper Note] ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning
Paper/Blog Link My Issue
#Article #ReinforcementLearning Issue Date: 2026-06-11 Comment
元ポスト:
gemma-4-12B-it-qat-UD-japanese-imatrix, dahara1, 2026.06
Paper/Blog Link My Issue
#Article #LanguageModel #OpenWeight #Japanese #Author Thread-Post Issue Date: 2026-06-11 Comment
元ポスト:
Policy on the AI Exponential, Dario Amodei, 2026.06
Paper/Blog Link My Issue
#Article #GenerativeAI Issue Date: 2026-06-11 Comment
元ポスト:
[Paper Note] PixelRAG: Web Screenshots Beat Text for Retrieval Augmented Generation, 2026.06
Paper/Blog Link My Issue
#Article #RAG(RetrievalAugmentedGeneration) #OCR #Author Thread-Post Issue Date: 2026-06-11 Comment
元ポスト:
Designing loops with Fable 5, Lance Martin, 2026.06
Paper/Blog Link My Issue
#Article #AIAgents #Post Issue Date: 2026-06-10 Comment
元ポスト:
Kimi Code CLI, MoonshotAI, 2026.06
Paper/Blog Link My Issue
#Article #AgentHarness Issue Date: 2026-06-09 Comment
元ポスト:
人間中心の意思決定支援AI 2026年度 人工知能学会全国大会 (JSAI 2026) チュートリアル (2026.06.09), Yukino Baba, 2026.06
Paper/Blog Link My Issue
#Article #Tutorial #Author Thread-Post Issue Date: 2026-06-09 Comment
元ポスト:
Implications of Large-Scale Test-Time Compute, Noam Brown, 2026.06
Paper/Blog Link My Issue
#Article #Post #Author Thread-Post Issue Date: 2026-06-09
視覚基盤モデルの構築, 産総研, 片岡裕雄, 2026.06
Paper/Blog Link My Issue
#Article #ComputerVision #FoundationModel #Slide #Robotics Issue Date: 2026-06-09 Comment
元ポスト:
Understanding Self-Distillation and Privileged Information Distillation, Penaloza+, 2026
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #ReinforcementLearning #Blog #Distillation #Selected Papers/Blogs #On-Policy #SelfDistillation Issue Date: 2026-06-05
NVIDIA Nemotron 3 Ultra, nvidia, 2026.06
Paper/Blog Link My Issue
#Article #Pretraining #NLP #LanguageModel #OpenWeight #SSM (StateSpaceModel) #MoE(Mixture-of-Experts) #Selected Papers/Blogs #Reference Collection #LowPrecision #LinearAttention #Author Thread-Post Issue Date: 2026-06-05 Comment
元ポスト:
アーキテクチャ解説:
Mamba2 layer, Latent MoE, GQA
ポイント解説:
HF: https://huggingface.co/collections/nvidia/nvidia-nemotron-v3
所見:
所見:
MAI-Thinking-1: Building a Hill-Climbing Machine, Microsoft, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Proprietary #Selected Papers/Blogs #Reference Collection Issue Date: 2026-06-02 Comment
元ポスト:
解説:
解説:
解説:
所見:
Figure 6について言及されている:
RLVR時代におけるInference Framework: Weight Syncing編, Kazuki Fujii, 2026.05
Paper/Blog Link My Issue
#Article #Blog #Author Thread-Post Issue Date: 2026-06-01 Comment
元ポスト:
MLエンジニアのための本質から理解するLLM推論: LLM Inference Benchmarking, Kazuki Fujii, 2026.05
Paper/Blog Link My Issue
#Article #Blog #Author Thread-Post Issue Date: 2026-05-31 Comment
元ポスト:
次:
- RLVR時代におけるInference Framework: Weight Syncing編, Kazuki Fujii, 2026.05
Step-3.7-Flash, stepfun-ai, 2026.05
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #AIAgents #MultiModal #OpenWeight #MoE(Mixture-of-Experts) #VisionLanguageModel #Author Thread-Post Issue Date: 2026-05-31 Comment
元ポスト:
公式:
LFM2.5-8B-A1B: An Even Better On-Device Mixture of Experts, LiquidAI, 2026.05
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #SmallModel #OpenWeight Issue Date: 2026-05-31 Comment
Should we use genetics instead of system prompts for AI Agents & Personas?, Fyx, 2026.05
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Supervised-FineTuning (SFT) #Blog #PEFT(Adaptor/LoRA) #PostTraining #Personality #One-Line Notes #Reading Reflections Issue Date: 2026-05-31 Comment
元ポスト:
ゲノム文字列からキャラクターのペルソナを推論し出力に反映可能なLoRAアダプタを学習することによって、ゲノム文字列によってペルソナの条件付けを可能とするモデルに関する実験の報告のようである。
具体的には、まずゲノム文字列と、ペルソナに関するテキストを対応付ける。次に、ペルソナテキストを与えてモデルの応答を収集し、ゲノム文字列と(ペルソナのテキストなしで)収集したモデルの応答を用いてLoRAアダプタを学習する。これにより、LoRAアダプタはゲノム文字列から、ペルソナテキストによって応答づけられた応答を出力することを学習する、といった挙動を実現する。
ざっとみた感じLoRAアダプタとは書かれているが、学習手法が書かれていないような気がする。が、手法はおそらくSFTだと思われる。
これにより、prefillをするinput tokenが削減可能と思われるが、実用上はどちらかというと出力/reasoningトークンが圧縮される方がlatency/コストの双方から恩恵があるので、事後学習をして他の能力が劣化、あるいはAlignmentが悪化するリスクを背負ってまでやる価値があるかは怪しいな、という感想を持ったが、果たしてどうだろうか。
非常に大規模なユーザからの同時接続があり、キャラクターと対話できるようなサービスにおいて、キャラクターのペルソナテキスト(用途はペルソナだけには限らない汎用的な技術だと思うが)が97%圧縮できたら結構なインパクトがあるかもしれない。
[Paper Note] LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding, Wang+, 2026.05
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #CVPR #Selected Papers/Blogs #ObjectLocalization #VisionLanguageModel #2D (Image) #UMM #3D (Video) #text #ObjectDetection #GUI #Author Thread-Post Issue Date: 2026-05-30 Comment
元ポスト:
DeepSWE: Measuring frontier coding agents on original, long-horizon engineering tasks, DeepSWE, 2026.05
Paper/Blog Link My Issue
#Article #NLP #Dataset #AIAgents #Evaluation #Coding #SoftwareEngineering #Selected Papers/Blogs #One-Line Notes #LongHorizon #Author Thread-Post Issue Date: 2026-05-27 Comment
元ポスト:
所見:
既存のベンチマークのような、githubのPRに基づいたものではなく(memorizationの問題があるため)、ゼロベースで構築。rolloutのtrajectoryを分析して、有効なPRなのに拒否する、あるいは何らかのcheatingをするといった挙動のdetectionもできるとのこと。また、SWE Bench Proと比較して、タスクを解くためのpromptは1/2である一方、タスクを解くために必要なコードの量は5.5倍となっており、より複雑なタスクとなっている。
contamination-freeが主張されているが、データセットは公開されているので、そのうちcontaminationが生じるであろう点には注意。
[Paper Note] Optimization Techniques for GPU Programming, ACM Computing Surveys, Volume 55, Issue 11, 2023.03
Paper/Blog Link My Issue
#Article #Survey #SoftwareEngineering #GPUKernel Issue Date: 2026-05-27 Comment
元ポスト:
Next-generation LLM Inference Network: How ZCube Alleviates Network Bottlenecks?, Z.ai, 2026.05
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Infrastructure #LLMServing #Post #SoftwareEngineering Issue Date: 2026-05-27
mKernel: Fast Multi-GPU, Multi-Node Fused Kernels, Ziming Mao, and the UCCL team, 2026.05
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #Infrastructure #GPUKernel #Author Thread-Post Issue Date: 2026-05-26 Comment
元ポスト:
> mKernel is our attempt at the missing piece: GPU-driven, fused kernels that deliver fine-grained compute–communication overlap across both intra-node NVLink and inter-node RDMA, while staying portable across various networking backends (ConnectX-7, AWS EFA, and more on the way).
Vision in the Age of LLMs, Lucas Beyer, 2026.05
Paper/Blog Link My Issue
#Article #Tutorial #ComputerVision #Pretraining #Transformer #MultiModal #ContrastiveLearning #Video #VisionLanguageModel #Backbone Issue Date: 2026-05-21 Comment
関連:
- [Paper Note] An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy+, ICLR'21
- [Paper Note] Sigmoid Loss for Language Image Pre-Training, Xiaohua Zhai+, ICCV'23
- [Paper Note] PaliGemma: A versatile 3B VLM for transfer, Lucas Beyer+, arXiv'24, 2024.07
元ポスト:
MemEx: A Programmable Scratchpad for LLM Agents, Databricks, 2026.05
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Blog #ContextEngineering Issue Date: 2026-05-20 Comment
元ポスト:
元ポスト:
Reliable Reasoning, Naoto Iwase, 2026.05
Paper/Blog Link My Issue
#Article #Survey #NLP #LanguageModel #Blog #Reasoning #Author Thread-Post Issue Date: 2026-05-20 Comment
元ポスト:
Aurora: A Leverage-Aware Optimizer for Rectangular Matrices, Tilde Research, 2026.05
Paper/Blog Link My Issue
#Article #Pretraining #NLP #LanguageModel #Optimizer #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-05-12 Comment
元ポスト:
nanoGPT speedrunのSoTAを更新したoptimiserのようである。
EMO: Pretraining mixture of experts for emergent modularity, Ai2, 2026.05
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #MoE(Mixture-of-Experts) #Selected Papers/Blogs #Routing #Initial Impression Notes Issue Date: 2026-05-11 Comment
元ポスト:
従来のMoEモデルを学習する際には全てのトークンが全てのexpertsに対してroutingされるため、能力のごく一部しか必要ないにもかかわらず、experts全体をデプロイする必要があり効率が悪かった。本研究では、MoEモデルを学習する際に、ドキュメント単位で8個のshared experts poolに対してのみtokenがroutingされるように制約することで、特定のexpertsの集合が特定のドメインのexpertsとなることが奨励されmodularityが高まり、必要なexpertsのみをデプロイすることで効率化が図れるようになるので嬉しい、といった話に見える。
Raven Part-1 - Memory as a set of Slots, Afzal+, Goomba Lab, 2026.05
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #SSM (StateSpaceModel) #memory #Initial Impression Notes #Author Thread-Post Issue Date: 2026-05-10 Comment
元ポスト:
元ポストのGIFがわかりやすく、SSMにおけるStateの更新をgatingによって選択的に実施するモデル、という感じだろうか。これによりSSMの弱点であった、long contextにおけるrecall-heavyなタスクにおいてより高い性能を獲得する。
MolmoAct 2: An open foundation for robots that work in the real world, Ai2, 2026.05
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #Reasoning #OpenWeight #OpenSource #Robotics #VisionLanguageActionModel #Author Thread-Post Issue Date: 2026-05-08 Comment
元ポスト:
関連:
- [Paper Note] MolmoAct: Action Reasoning Models that can Reason in Space, Jason Lee+, arXiv'25
dataset:
https://huggingface.co/collections/allenai/molmoact2-datasets
models:
https://huggingface.co/collections/allenai/molmoact2-models
著者ポスト:
著者ポスト2:
著者ポスト3:
The ultimate guide to RL environments: building and scaling them in the LLM era, HuggingFace, 2026.05
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #ReinforcementLearning #Blog #PostTraining #Selected Papers/Blogs #Environment Issue Date: 2026-05-08 Comment
元ポスト:
On SFT, RL, and on-policy distillation, will brown, 2026.05
Paper/Blog Link My Issue
#Article #Post Issue Date: 2026-05-01
OlmPool: How small architectural choices compound to undermine long context extension, Ai2, 2026.04
Paper/Blog Link My Issue
#Article #Analysis #NLP #LanguageModel #Transformer #Attention #LongContext #Architecture #Selected Papers/Blogs #One-Line Notes #ContextRot #Author Thread-Post Issue Date: 2026-05-01 Comment
元ポスト:
QK Norm, GQA, SWA, 事前学習のcontext長の短縮、これらはいずれもモデルが入力に対するattendの仕方を変えるものだが、これらを3つ以上組み合わせるとlong contextでの性能が急落するらしく、このようなlong contextの性能劣化は一般的な(しばしば短い)コンテキスト長のベンチマークやloss/perplexityなどでは検知できず、long contextで性能が急落するアーキテクチャでは、50Bトークンでのlong contextの学習を経ても、Llamaアーキテクチャが1Bトークンの学習で到達できる性能に届かない、といった話が元ポストに書かれている。
The Futures of Programming, Graham Neubig, 2026.04
Paper/Blog Link My Issue
#Article Issue Date: 2026-05-01 Comment
元ポスト:
Where the goblins came from, OpenAI, 2026.04
Paper/Blog Link My Issue
#Article #Reference Collection Issue Date: 2026-04-30 Comment
元ポスト:
所見:
Scaling Pain of Coding Agent Serving: Lessons from Debugging GLM-5 at Scale, Z.ai, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Infrastructure #LLMServing #SoftwareEngineering #Selected Papers/Blogs #Initial Impression Notes #Author Thread-Post Issue Date: 2026-04-30 Comment
GLM-5をサービングしている中でのバグ(モデル側ではなくインフラ側)の発見と修復
Laguna XS.2 and M.1: A Deeper Dive, Poolside team, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Coding #OpenWeight #MoE(Mixture-of-Experts) #SoftwareEngineering #Selected Papers/Blogs #LongHorizon #Author Thread-Post Issue Date: 2026-04-30 Comment
HF: https://huggingface.co/poolside/Laguna-XS.2
元ポスト:
テクニカルレポート:
https://poolside.ai/assets/laguna/laguna-m1-xs2-technical-report.pdf
元ポスト:
Vintage Large Language Models, Owain Evans
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Selected Papers/Blogs #vintage LLMs Issue Date: 2026-04-29 Comment
過去の特定の時刻までのデータで学習された大規模言語モデル, vintage Large Language Modelsを提唱
Introducing talkie: a 13B vintage language model from 1930, Levine+, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #OpenWeight #Selected Papers/Blogs #Initial Impression Notes #vintage LLMs Issue Date: 2026-04-29 Comment
元ポスト:
1930年以前の英語テキストで学習された言語モデル(vintage Large Language Models)で、歴史や文化の変化を分析したり、1930年までのデータで学習されたモデルが1931年以後に発見された革新的な科学的な発見を自ら見出せるか?、LLMが将来を予測する能力がどの程度あり、それがモデルサイズによってどのように変化するか?、プログラミングに関する知識がないモデルが現代のコーディングを学習できるかなどのcontamination freeな評価など様々な活用方法があるとのこと。
関連:
- Vintage Large Language Models, Owain Evans
所見:
Project Deal, Anthropic, 2026.04
Paper/Blog Link My Issue
#Article #Analysis #NLP #LanguageModel #AIAgents #Personalization #One-Line Notes #Sales Issue Date: 2026-04-26 Comment
元ポスト:
AI同士が商取引をしたら何が起きるかという社内実験のようである。69人の社員に何を売りたいか/買いたいををインタビューし、カスタム指示が与えられた上でAI Agentに取引をさせたところ、きちんと商取引が行われ、186件、$4000のやりとりがあったとのこと。そして賢いモデルが大幅に有利に取引を終えて、実際の参加者はこの事実に気づかなかったとのこと。また、カスタム指示(e.g., 強行姿勢, 礼儀正しいなど)はあまり良い成果を上げる上では重要ではなかった、
といった話が元ポストに書かれている。
Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis, Ai2, 202604
Paper/Blog Link My Issue
#Article #ComputerVision #Embeddings #NLP #Proprietary Issue Date: 2026-04-25 Comment
元ポスト:
関連:
- OlmoEarth-v1-Large, Ai2, 2025.11
ちょっとこれはしっかり読まないと具体的にはわからないかもしれない
claude-code-best-practice, shanraisshan
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Repository #Coding #SoftwareEngineering #Selected Papers/Blogs Issue Date: 2026-04-25 Comment
元ポスト:
Gemini Embedding 2: Our first natively multimodal embedding model, Google, 2026.03
Paper/Blog Link My Issue
#Article #ComputerVision #Embeddings #NLP #MultiModal #Blog #Proprietary #Selected Papers/Blogs #KeyPoint Notes #Author Thread-Post Issue Date: 2026-04-25 Comment
元ポスト:
単一のモデルで、マルチモーダルな情報を統合されたembedding空間で表現し、マトリョーシカ表現によって3種類の次元で取得でき、100+言語をサポートしかつcontext windowは8192。オーディオをわざわざ書き起こしてテキストモダリティに変換する必要もなく直接unifiedなembeddingを取得可能というなかなか便利そうな代物。
(以前のIssueを誤って削除したため再掲)
Generally Availableになったとのこと:
What 81,000 people told us about the economics of AI, Anthropic, 2026.04
Paper/Blog Link My Issue
#Article #Analysis #GenerativeAI #Blog #Selected Papers/Blogs #One-Line Notes #Author Thread-Post Issue Date: 2026-04-25 Comment
元ポスト:
賃金が最も小さいグループ、おより最も高いグループではClaudeによる生産性向上が最も大きく、職を失う懸念も同時に大きい。同様に、Claudeの利用量が多いグループも職を失う懸念が大きい。
アメリカにおいて代替されると思っていたソフトウェアエンジニアの求人がむしろ増えていて、AIによって新たな雇用が生まれているという意見もある:
Latest Agentic AI Trends to Watch in 2026: Market Shifts, Adoption Patterns, and What Comes Next, Daya Shankar, 2026.04
Paper/Blog Link My Issue
#Article #Survey #NLP #Infrastructure #AIAgents #Blog Issue Date: 2026-04-25 Comment
元ポスト:
ML Intern, HuggingFace, 2026.04
Paper/Blog Link My Issue
#Article #Tools #NLP #LanguageModel #AIAgents #AutoML #ScientificDiscovery #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-04-21 Comment
元ポスト:
自動で研究が可能なエコシステムがどんどん構築されていく
関連:
take-homeをend-to-endで解けるくらい優秀とのこと。
Train separately, merge together: Modular post-training with mixture-of-experts, Ai2, 2026.04
Paper/Blog Link My Issue
#Article #ContinualLearning Issue Date: 2026-04-21 Comment
元ポスト:
関連:
- [Paper Note] FlexOlmo: Open Language Models for Flexible Data Use, Weijia Shi+, NeurIPS'25
Designing synthetic datasets for the real world: Mechanism design and reasoning from first principles, Google, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #SyntheticData #Distillation #Selected Papers/Blogs #One-Line Notes #Reference Collection #Critic #Reading Reflections #Human-in-the-Loop #Author Thread-Post Issue Date: 2026-04-19 Comment
元ポスト:
公式:
解説:
(詳細は解説や元ブログ参照のこと)
強い教師モデルから弱い生徒モデルを学習する場合の合成データ生成手法で、
生成したいデータの観点(内容、形式等)を分類し、どの観点からどの程度の難易度のデータを合成するかを制御する。その後生成されたデータが正しいか/正しくないかの2方向から批評を行いvalidationをするような枠組みのようである。
単純なデータ合成では性能がすぐに頭打ちになるが、ローカル多様性(特定のパターンの多様性)、グローバル多様性(データ全体がカバーするパターンの範囲)の2つを同時に大きくしないと不十分であることや、批判によるvalidationは少なくとも性能を悪化させることはないことも示されたとのこと。
[Paper Note] Open-world evaluations for measuring frontier AI capabilities, Kapoor+, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Evaluation #Selected Papers/Blogs #Author Thread-Post Issue Date: 2026-04-19 Comment
元ポスト:
Evaluating Netflix Show Synopses with LLM-as-a-Judge, Netflix Technology Blog, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Evaluation #Factuality #Blog #LLM-as-a-Judge #Test-Time Scaling #Reference Collection #Scalability #Initial Impression Notes Issue Date: 2026-04-17 Comment
元ポスト:
Netflix上に存在するsynopses(映画の短いdescription)を高品質に保ちたいが、非常に量が多いのでどのようにスケーラブルに評価しているか、という話のようである。
LLM-as-a-Judgeを活用して評価をしており、4種類の観点(制度、事実性、トーン、明瞭さ)のような多次元のRubricを用いて、それぞれの観点ごとにLLM-as-a-Judgeを専門家の判断にalignさせるためにgold dataを作成し、どのように推論すればLLM-as-a-Judgeの性能が向上するかを調査した結果、long CoT / Majority Voting (精度向上+分散低下)/ Agents-as-a-Judge (複数のFactualityの側面を評価するために4種類のAI Agentを用いてメタデータとsynopsesのFactual Consistencyを評価し、全てのエージェントの結果を集約)といった感じのことをやっているらしい。
Defining Continual Learning, Ilija Lichkovski, 2026.04
Paper/Blog Link My Issue
#Article #Tutorial #Post #ContinualLearning Issue Date: 2026-04-17 Comment
元ポスト:
8 Tips for Writing Agent Skills, Philipp Schmid, 2026.04
Paper/Blog Link My Issue
#Article #NLP #AIAgents #AgentSkills #Author Thread-Post Issue Date: 2026-04-17
Introducing Ternary Bonsai: Top Intelligence at 1.58 Bits, PrismML, 2026.04
Paper/Blog Link My Issue
#Article #Author Thread-Post Issue Date: 2026-04-16 Comment
HF: https://huggingface.co/collections/prism-ml/ternary-bonsai
- Announcing 1-bit Bonsai: The First Commercially Viable 1-bit LLMs, 2026.03
の次世代モデル。
前回リリースからまだ1ヶ月しか経っていない。デコーディング速度が速いのでその分RLによるPostTrainingも高速なのだろうと推察される。
元ポスト:
Taking the Pulse of Agentic AI from the Developer Community at the End of Q1 2026, InclusionAI, 2026.04
Paper/Blog Link My Issue
#Article #Survey #Tools #NLP #LanguageModel #Library #AIAgents #GenerativeAI #Repository Issue Date: 2026-04-11 Comment
元ポスト:
The OpenHands Vulnerability Fixer: Automated Security Remediation with AI Agents, Graham Neubig, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Blog #SoftwareEngineering #Security Issue Date: 2026-04-11 Comment
元ポスト:
AI Cybersecurity After Mythos: The Jagged Frontier, Stanislav Fort, 2026.04
Paper/Blog Link My Issue
#Article Issue Date: 2026-04-11 Comment
元ポスト:
Introducing Muse Spark: Scaling Towards Personal Superintelligence, Meta, 2026.04
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #LanguageModel #MultiModal #Proprietary #Selected Papers/Blogs #VisionLanguageModel #One-Line Notes #Reference Collection #Author Thread-Post Issue Date: 2026-04-11 Comment
元ポスト:
-
-
元ポストのベンチマークスコアを見るとマルチモーダルの性能はフロンティアモデル(gpt5.4, Opus 4.6, Gemini 3.1 Pro)と同等、text/reasoningはフロンティアモデルより少しスコアが低く、特に抽象的な思考が苦手(ARC-AGI-2)。HEALTH分野はhealthは高スコアだがmedicalは少し低めのスコア、Agenticな分野では、SWE Bench Verified/Proよスコアは少し低め、terminal useは明確にスコアが低くtool useは少しスコアが低い、という感じにみえる。
codingとlong horizon taskに継続的に投資するとのこと。
中の人による解説:
全てをフルスクラッチから作り直したっぽい。
Artificial Analysisによる解説:
一気にOpenWeight最強のGLM-5.1超え
所見:
所見:
所見:
第三者によるおそらく独自のベンチマークによる評価の結果、(おそらく101モデルのうち)全体で3位となっているらしい(つまり、既存ベンチマークにoverfittingしているわけではないという考えがある)。
The ATOM Report: Measuring the Open Language Model Ecosystem, Lambert+, 2026.04
Paper/Blog Link My Issue
#Article #Analysis #NLP #LanguageModel #OpenWeight #OpenSource #Data #Author Thread-Post Issue Date: 2026-04-11 Comment
著者ポスト:
元ポスト:
What do “economic value” benchmarks tell us?, Epoch.AI, 2026.02
Paper/Blog Link My Issue
#Article #NLP #Dataset #LanguageModel #Evaluation Issue Date: 2026-04-11 Comment
元ポスト:
GDPvalだけでなく、RLI, APEX Agentsと呼ばれるものも解説されているようである
- REMOTE LABOR INDEX (RLI) : Evaluating the capability of AI agents to perform real-world, economically valuable remote work
- [Paper Note] APEX-Agents, Bertie Vidgen+, arXiv'26, 2026.01
- [Paper Note] GDPval: Evaluating AI Model Performance on Real-World Economically
Valuable Tasks, Tejal Patwardhan+, arXiv'25, 2025.10
Unfolding Robotics: The Open-Source Recipe for Teaching a Robot to Fold Your Clothes, Hugging Face, 2026.04
Paper/Blog Link My Issue
#Article #Tutorial #ComputerVision #NLP #OpenWeight #OpenSource #Selected Papers/Blogs #Robotics #VisionLanguageActionModel Issue Date: 2026-04-07 Comment
元ポスト:
Microsoft Open-Sources Industry-Leading Embedding Model, Microsoft Bing Blog, 2026.04
Paper/Blog Link My Issue
#Article #Embeddings #NLP #Blog #MultiLingual #OpenWeight Issue Date: 2026-04-07 Comment
元ポスト:
Introducing WildDet3D: Open-world 3D detection from a single image, Ai2, 2026.04
Paper/Blog Link My Issue
#Article #ComputerVision #Dataset #OpenWeight #OpenSource #Selected Papers/Blogs #3D (Video) #ObjectDetection #Initial Impression Notes Issue Date: 2026-04-07 Comment
元ポスト:
wildな環境においてzero shot(click, text, bounding boxで対象を指定)で動作する単眼の3D Object Detectionモデルとのこと。データセットもコードも公開
How we optimized Dash's relevance judge with DSPy, Dropbox, 2026.03
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #Prompting #AutomaticPromptEngineering #LLM-as-a-Judge #Initial Impression Notes Issue Date: 2026-04-07 Comment
元ポスト:
APEを使ってモデルを変更した際のプロンプト適応を効率化した話な模様。
オープンソースAIの現状 | NVIDIA GTC, Nvidia, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #OpenWeight #Video #OpenSource #One-Line Notes Issue Date: 2026-04-07 Comment
元ポスト:
GTCのパネルディスカッション
JEPAwiki, mishig, 2026.04
Paper/Blog Link My Issue
#Article #ComputerVision #Blog #WorldModels #LatentRepresentation Issue Date: 2026-04-07 Comment
元ポスト:
Components of A Coding Agent: How coding agents use tools, memory, and repo context to make LLMs work better in practice, Sebastian Raschka, 2026.04
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #AIAgents #Coding #SoftwareEngineering #Selected Papers/Blogs #Initial Impression Notes #AgentHarness Issue Date: 2026-04-05 Comment
LLM, Reasoning Model, Agent, Agent Harness, coding harnessなどの定義とその役割やスコープ、そしてそれらを構成するためのminimalなコンポーネントについて説明されており、基礎的な理解に役立ちそう。
元ポスト:
Moonlake: Causal World Models should be Multimodal, Interactive, and Efficient — with Chris Manning and Fan-yun Sun, LatentSpace, 2026.04
Paper/Blog Link My Issue
#Article #Tutorial #ComputerVision #WorldModels Issue Date: 2026-04-05 Comment
元ポスト:
Emotion Concepts and their Function in a Large Language Model, Anthropic, 2026.04
Paper/Blog Link My Issue
#Article #Analysis #NLP #LanguageModel #Selected Papers/Blogs #Emotion #Initial Impression Notes Issue Date: 2026-04-04 Comment
元ポスト:
これは非常に面白そうだ
Claude Code's Real Secret Sauce (Probably) Isn't the Model, Sebastian Raschka, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Post #Architecture #AgentHarness Issue Date: 2026-04-04 Comment
関連:
- Claude Code's source code leaked through a `.map` file - How bad is it, really?, Chubby, 2026.04
How far does alignment midtraining generalize?, Tomek+, OpenAI Alignment Research Blog, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Alignment #mid-training #Initial Impression Notes Issue Date: 2026-04-04 Comment
元ポスト:
mid trainingにおいてalignment関してmisaligned/alignedな文書で学習をすると中間学習直後はalignmentに関する挙動が維持されるが、RLをしたらその効果は消えて無くなってしまう、という感じだろうか?超絶流し読みなので、後でしっかり読んだ方が良さそう。
Holo3: Breaking the Computer Use Frontier, H Company, 2026.03
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #AIAgents #MultiModal #OpenWeight #MoE(Mixture-of-Experts) #ComputerUse #VisionLanguageModel #One-Line Notes #GUI #Environment Issue Date: 2026-04-02 Comment
元ポスト:
HF: https://huggingface.co/Hcompany/Holo3-35B-A3B
関連:
- Holo2: Cost-Efficient Models for Cross-Platform Computer-Use Agents, H Company, 2025.11
Qwen3.5をファインチューニングすることで実現。以前のシリーズもQwenベースだったが、新たなQwenのリリースに伴いより強力なベースモデルを得て、かつシナリオをベースにして自動でwebsiteを構築しverifiableが可能な独自のEnvironmentを保持しており、多様な合成データの活用とRLを実現することで、性能が向上していると思われる。
Trinity-Large-Thinking: Scaling an Open Source Frontier Agent, Arcee, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Reasoning #OpenWeight #MoE(Mixture-of-Experts) #Selected Papers/Blogs Issue Date: 2026-04-02 Comment
元ポスト:
HF: https://huggingface.co/collections/arcee-ai/trinity-large-thinking
ClawAegis, Ant Group, 2026.04
Paper/Blog Link My Issue
#Article Issue Date: 2026-04-02 Comment
元ポスト:
LFM2.5-350M: No Size Left Behind, Liquid AI, 2026.04
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #SmallModel #OpenWeight #Selected Papers/Blogs #KeyPoint Notes Issue Date: 2026-04-01 Comment
元ポスト:
- LFM2のアーキテクチャを採用の350Mパラメータモデルで、CPUでも十分な速度で推論可能
- 追加の事前学習(10T -> 28T tokens)、および、large-scale RLを実施
- 同等規模のパラメータ数(あるいは2倍程度)のモデル群に対して、知識, 指示追従能力, ツール呼び出し、データ抽出などのベンチマークで上回る
- LFM2-350Mと比較して、指示追従能力, データ抽出, tool useの性能が大きく向上
- edgeデバイスでの軽量なデータ抽出パイプラインとして有用
- しかし、math, coding, creative writingなどでの利用は推奨されない
- CPU/GPUでの推論ともに同等規模、あるいは1B級のモデルよりも早く、省メモリ
LongCat-AudioDiT, Meituan LongCatTeam, 2026.03
Paper/Blog Link My Issue
#Article #NLP #SpeechProcessing #DiffusionModel #OpenWeight #Architecture #Selected Papers/Blogs #TTS #Initial Impression Notes Issue Date: 2026-04-01 Comment
HF:
-
https://huggingface.co/meituan-longcat/LongCat-AudioDiT-1B
-
https://huggingface.co/meituan-longcat/LongCat-AudioDiT-3.5B
元ポスト:
デコード時に、メルスペクトログラム→Vocoderの場合細かい特徴が落ちてしまうことが懸念されるため、Waveformを直接デコードするWav-VAEによって、音声に直接変換する、というアーキテクチャの革新があるように見える。
Introducing Cohere-transcribe: state-of-the-art speech recognition, CohereLabs, 2026.03
Paper/Blog Link My Issue
#Article #AutomaticSpeechRecognition(ASR) Issue Date: 2026-03-30 Comment
元ポスト:
How Kimi, Cursor, and Chroma Train Agentic Models with RL, PHILSCHMID, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #ReinforcementLearning #AIAgents #Blog #reading #LongHorizon Issue Date: 2026-03-29
Chroma Context-1: Training a Self-Editing Search Agent, Chroma, 2026.03
Paper/Blog Link My Issue
#Article Issue Date: 2026-03-29 Comment
ポイント解説:
- How Kimi, Cursor, and Chroma Train Agentic Models with RL, PHILSCHMID, 2026.03
ルーブリックに基づく主観的な判定を取り入れたGRPO学習, Akira Sasaki, 2026.03
Paper/Blog Link My Issue
#Article Issue Date: 2026-03-28 Comment
元ポスト:
ソフトウェア開発エージェント 初歩から上級, Graham Neubig, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Coding #SoftwareEngineering Issue Date: 2026-03-26 Comment
全体をざっくり概観してイメージをつかむのに良さそう。詳細を知りたい場合はリンク先を見ると良さげ。
(スライド最後の強化学習における「3」のスケーリングってなんだろう...?)
元ポスト:
RACER: 自動運転VLAモデルの学習データセットの構築, Tech Blog - Turing, Zenn, 2026.03
Paper/Blog Link My Issue
#Article #Dataset #VisionLanguageActionModel Issue Date: 2026-03-26 Comment
元ポスト:
Here are the 2025 AI safety papers and posts I like the most, Fabien Roger, LW, 2026.03
Paper/Blog Link My Issue
#Article #Survey #NLP #LanguageModel #Safety #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-03-26 Comment
元ポスト:
AI Safetyに関する研究者の方の2025年のAI Safetyハイライトとのこと。
Emergent Misalignmentなど以外にも多くの研究に⭐︎︎︎⭐︎⭐︎が付与されている。気になる。
Lossy self-improvement, Nathan Lambert, 2026.03
Paper/Blog Link My Issue
#Article Issue Date: 2026-03-26
A Visual Guide to Attention Variants in Modern LLMs, Sebastian Raschka, 2026.03
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #Attention #Blog Issue Date: 2026-03-26
動画生成を蒸留で27倍速くした話, AIdeaLab, 2026.03
Paper/Blog Link My Issue
#Article Issue Date: 2026-03-26 Comment
元ポスト:
Where Machines Get Reward, OpenReward, 2026.03
Paper/Blog Link My Issue
#Article #ReinforcementLearning #Environment Issue Date: 2026-03-25 Comment
元ポスト:
Andrej Karpathy on Code Agents, AutoResearch, and the Loopy Era of AI, No Priors: AI, Machine Learning, Tech, & ..., 2026.03
Paper/Blog Link My Issue
#Article #Video Issue Date: 2026-03-22 Comment
元ポスト:
Nemotron 3 Nano 4B: A Compact Hybrid Model for Efficient Local AI, Nvidia, 2026.03
Paper/Blog Link My Issue
#Article #SmallModel Issue Date: 2026-03-22 Comment
HF: https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16
元ポスト:
デコーディングがめちゃ速い:
THE CONSCIOUSNESS CLUSTER: PREFERENCES OF MODELS THAT CLAIM TO BE CONSCIOUS, Chua+, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Alignment #Safety #Initial Impression Notes Issue Date: 2026-03-20 Comment
元ポスト:
LLMに意識があるように振る舞うように学習したらどうなるかという話らしい。これによって新たなpreferenceが獲得され、自己保存欲求や反発が発現したり、共感や葛藤などの人間的な感情について話したり、思考過程をモニタリングされることをどう感じますか?といった質問に対して、uncomfortableだと感じる、私は悪い評価を受けたら停止されてしまうの?といった不安について述べたりするなど、これまでにない挙動が見受けられるという感じらしい。
より長いホライズンに向けた Composer の学習, Cursor, 2026.03
Paper/Blog Link My Issue
#Article #DocumentSummarization #AIAgents Issue Date: 2026-03-20 Comment
元ポスト:
関連:
- [Paper Note] Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL, Ian Wu+, arXiv'26, 2026.02
- [Paper Note] InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning, Yuchen Yan+, arXiv'26, 2026.02
Reinforcement Learning from Human Feedback, Nathan Lambert, 2026.03
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #ReinforcementLearning #Blog #PostTraining Issue Date: 2026-03-20 Comment
元ポスト:
REINFORCE, PPO, GRPOの気持ちを理解するのに有用という所見:
How NVIDIA Builds Open Data for AI, Nvidia, 2026.03
Paper/Blog Link My Issue
#Article Issue Date: 2026-03-20 Comment
元ポスト:
MSA: Memory Sparse Attention, Chen+, 2026.03
Paper/Blog Link My Issue
#Article Issue Date: 2026-03-20 Comment
元ポスト:
Composer 2 のご紹介, Cursor, 2026.03
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #ReinforcementLearning #AIAgents #Evaluation #Coding #SoftwareEngineering #mid-training #PostTraining #Selected Papers/Blogs #ContextEngineering #Live #Reference Collection #Initial Impression Notes Issue Date: 2026-03-20 Comment
元ポスト:
所見:
Kimi-K2.5がベースらしいとのこと:
ベンチマークスコアに対する所見:
テクニカルレポートが出た:
https://cursor.com/resources/Composer2.pdf
元ポスト:
Kimi-K2.5をベースに、どのようにinstruction tuning後のモデルに対して継続事前学習、RLをし、GPT-5.4(high)級の性能を達成できたのか、ヒントがわかるかもしれない。
- [Paper Note] Kimi K2.5: Visual Agentic Intelligence, Kimi Team+, arXiv'26, 2026.02
所見:
所見:
RLによってpass@k(best-of-16)とpass@1の両方が改善する。既存研究では少なくともRLVRを用いた場合はPass@1は改善するが多様性が損なわれてPass@kの性能は改善しない ([Paper Note] Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR, Xiao Liang+, arXiv'25, 2025.08 , VibeVoice-1.5B, microsoft, 2025.08 )、という話があったが、Composer 2のレシピではそうではないようだ。どんなレシピだろう~と思ってさらっと関連しそうなところを見てみたが、詳細は書いてなさそうだ。
- [Paper Note] Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR, Xiao Liang+, arXiv'25, 2025.08
- VibeVoice-1.5B, microsoft, 2025.08
QA:
CursorBenchの解説:
要はrealisticなデータとシチュエーションでの評価に非常に重きを置いていて
- 実際のコーディングsessionのデータが用いられ、contamination-free
- 機能的な正しさのみならず、コードの品質、効率、挙動などの実用的な価値を意識し
- long horizonなタスクが多く取り入れられ
- Promptは曖昧性をうまく扱えるかを評価するために意図的にシンプルで短く
- CursorBenchのデータは継続的に更新される
- realisticなsessionデータだけでなく、その他の重要な挙動の評価(e.g., 指示追従, ルール/skilltのハンドリング, コメントの品質, editするか否かの判断の適切性など)のためのデータでも拡張されている
という感じらしい
ポイント解説:
- How Kimi, Cursor, and Chroma Train Agentic Models with RL, PHILSCHMID, 2026.03
self-summarizationによるcontextのcompressionを実施している
- [Paper Note] InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning, Yuchen Yan+, arXiv'26, 2026.02
- [Paper Note] Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL, Ian Wu+, arXiv'26, 2026.02
- より長いホライズンに向けた Composer の学習, Cursor, 2026.03
所見:
MolmoPoint: Better pointing architecture for vision-language models, Ai2, 2026.03
Paper/Blog Link My Issue
#Article Issue Date: 2026-03-19 Comment
元ポスト:
公式:
L11: Synthetic Data Powering Pretraining, Eric W. Tramel, Ph.D., UC Berkeley EE 290_194-11: Scalable AI, 2026.02
Paper/Blog Link My Issue
#Article #Pretraining #NLP #LanguageModel #SyntheticData #Selected Papers/Blogs #KeyPoint Notes #Reading Reflections Issue Date: 2026-03-17 Comment
元ポスト:
- インターネットのデータ枯渇問題が指摘されながらも、合成データによって事前学習は進化を続けている
- LLMは事後学習で性能を向上させられるが、事前学習時点で伸ばせる上限が決まっているとされている
- 事前学習データの投入量はChinchilla則のパラメータ量の20倍から現在は60倍まで増加
- MoEは過学習しやすくパラメータ数の40倍は必要
- 学習データの多様性が重要で繰り返し同じデータを見ても性能は改善しない
- 合成データをそのまま用いるとmode collapseが生じ出力が単調化するため、実データを混ぜるか言い換えをしたデータで是正する(弱めのdata augmentationで良い)
- 最近重要な合成データはコードと推論過程を含むデータで、これらが事前学習データに含まれていると汎用な表現、思考能力、推論能力を事前学習時点から獲得できる可能性がある
というような話が元ポストに書かれている。
- [Paper Note] Scaling Data-Constrained Language Models, Niklas Muennighoff+, NeurIPS'23
のようにrepetitionは4回までが効果的といった知見が報告されているが、現在はどこまで当てはまるのだろうか?
後ほど関連するissueのリンクを貼りたい
うーんおもしろそう、p.15, p.20, p.26, p.28, p.35, p.36 あたりが気になる。
てかこれが大学の講義...?楽しすぎでは。
From SAM to SAM 3, Massimiliano Viola, 2026.03
Paper/Blog Link My Issue
#Article #Post Issue Date: 2026-03-12
NVIDIA Nemotron 3 Super, NVIDIA, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #OpenWeight #SSM (StateSpaceModel) #OpenSource #MoE(Mixture-of-Experts) #Selected Papers/Blogs #KeyPoint Notes #Reference Collection #Hybrid #LowPrecision #LinearAttention Issue Date: 2026-03-12 Comment
元ポスト:
解説:
artificial analysisによる評価:
Swallow LVM Leaderboardに性能が掲載:
解説:
アーキテクチャ:
- NVFP4で学習して gpt-ossより2.2倍高速だが性能も向上
- 88 Layer: 40 Latent MoE / 40 Mamba-2 / 8 GQA Attention
- GQA Attentiom Layerは非常に少なく、ほとんどがMamba-2 (linear attention)となっている
- Latent MoEは入力をそのまま変換するshared expertsと、入力を1/4のlatent vectorに変換した潜在空間上で処理をするLatext expertsの組み合わせによって出力を得る。
- 具体的には、RouterによってTop-22のexpertsを選択し、inputを1/4のlatent vectorに圧縮した上でExpertsに入力。Expertsの出力を加算して4倍のvectorに変換し次元を戻して、別ルートでshared expertsに元の入力次元から変換されたベクトルと組み合わせて出力するようなアーキテクチャ
Latent MoE解説:
要はMoEに必要なmatrixが、latent vectorを扱うことで小さくなるのでMoEのWeightのメモリロードのボトルネックが緩和されるだけでなく、
各MoE Laverは異なるGPUやマシンに分散されて配置されるため計算のためにはベクトルのバッチを通信しなければならないがそのコストが削減されスループットの向上につながるので嬉しい、ということだと思われる。
ポイント解説:
technical reportが出た:
- [Paper Note] Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning, NVIDIA+, arXiv'26, 2026.04
Using NVFP4 Low-Precision Model Training for Higher Throughput Without Losing Accuracy, NVIDIA, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #LowPrecision Issue Date: 2026-03-12 Comment
元ポスト:
大規模言語モデルからみた脳科学と人工知能の未来, 岡野原大輔, 生体の科学 77巻1号, 2026.02
Paper/Blog Link My Issue
#Article Issue Date: 2026-03-12 Comment
元ポスト:
4月下旬までの期間限定公開とのこと
Applying Statistics to LLM Evaluations, CAMERON R. WOLFE, PH.D., 2026.03
Paper/Blog Link My Issue
#Article Issue Date: 2026-03-12 Comment
元ポスト:
Bringing Code Review to Claude Code, Anthropic, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Coding #SoftwareEngineering Issue Date: 2026-03-12 Comment
元ポスト:
コードレビューに特化した機能が追加された模様
Anthropic社内で運用済みで、エンジニアがコードレビューに誤りがあると判断したものは<1%とのこと。
How to train the best embedding model in the world, dr. jack morris, 2026.03
Paper/Blog Link My Issue
#Article #Post Issue Date: 2026-03-10 Comment
元ポスト:
The Synthetic Data Playbook: Generating Trillions of the Finest Tokens, HuggingFace, 2026.03
Paper/Blog Link My Issue
#Article #Pretraining #NLP #LanguageModel #SyntheticData #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-03-10 Comment
12.7 GPU yearを使い、90回の実験、1 Trillion tokenの生成を経て見つけた、合成事前学習データの構築方法のbest recipeが紹介されている模様。先行研究を上回る学習効率を達成している。
元ポスト:
Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis, Black Forest Labs, 2026.03
Paper/Blog Link My Issue
#Article #ComputerVision #Pretraining #NLP #MultiModal #SpeechProcessing #Self-SupervisedLearning #2D (Image) #FlowMatching #3D (Video) #Omni #RectifiedFlow #audio Issue Date: 2026-03-10 Comment
backbone modelは下記のFLUX.2と呼ばれるモデル:
FLUX Commercial Licensing:
https://bfl.ai/licensing
先行研究:
- The Simulation Company, Simile, 2026.02
先行研究から読みたい
元ポスト:
Towards Efficient World Models, Moonlake, 2026.03
Paper/Blog Link My Issue
#Article #Post #WorldModels Issue Date: 2026-03-07 Comment
関連:
- Building Multimodal Worlds with Moonlake's World Modeling Agent, Moonlake, 2026.02
Chinese Open Source: A Definitive History, Kevin Xu, 2026.03
Paper/Blog Link My Issue
#Article #Survey #NLP #LanguageModel #Blog #OpenWeight Issue Date: 2026-03-07
Practical Guide to Evaluating and Testing Agent Skills, PHILSCHMID, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Blog #Coding #SoftwareEngineering #AgentSkills Issue Date: 2026-03-06 Comment
元ポスト:
関連:
- How to Create Effective Agent Skills, openhands, 2026.02
Reasoning models struggle to control their chains of thought, and that’s good, OpenAI, 2026.03
Paper/Blog Link My Issue
#Article #Controllable #NLP #Dataset #LanguageModel #Chain-of-Thought #Evaluation #Blog #Reasoning #Author Thread-Post Issue Date: 2026-03-06 Comment
元ポスト:
著者ポスト:
Introducing Olmo Hybrid: Combining transformers and linear RNNs for superior scaling, Ai2, 2026.03
Paper/Blog Link My Issue
#Article #Pretraining #NLP #LanguageModel #Attention #OpenWeight #mid-training #Selected Papers/Blogs #One-Line Notes #RecurrentModels #Hybrid #LinearAttention Issue Date: 2026-03-06 Comment
元ポスト:
x1のFull Attention + x3のGated DeltaNetによるハイブリッドアーキテクチャで、75%のattentionをlinear attention (recurrent module)に置換。x3のSliding Window Attentionを用いているOlmo3と比較した結果
- 事前学習におけるデータ効率がより高く(約2倍)
- mid-training後の評価では、数学、コード、STEM, non-STEM, QA、long-contextなどの主要なドメインにおいてOlmo3と同と床それ以上の性能を達成。特に、long-contextにおけるベンチマでは大幅な性能向上(Recurrentなアーキテクチャの恩恵)
関連:
- [Paper Note] Gated Delta Networks: Improving Mamba2 with Delta Rule, Songlin Yang+, ICLR'25, 2024.12
元ポスト:
関連:
所見:
FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling, together.ai, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Library #Transformer #Attention #Chip #Selected Papers/Blogs #GPUKernel #Initial Impression Notes Issue Date: 2026-03-06 Comment
元ポスト:
関連:
これは読まねば。。。
Interactive Benchmarks, Yue+, 2026.03
Paper/Blog Link My Issue
#Article Issue Date: 2026-03-05 Comment
元ポスト:
AReaL: A Large-Scale Asynchronous Reinforcement Learning System, inclusionAI, 2026.03
Paper/Blog Link My Issue
#Article #Tools #NLP #LanguageModel #ReinforcementLearning #Reasoning #Asynchronous #TrainingFramework Issue Date: 2026-03-05 Comment
元ポスト:
Current activation oracles are hard to use, Senthooran+, LW, 2026.03
Paper/Blog Link My Issue
#Article Issue Date: 2026-03-05 Comment
元ポスト:
Latest open artifacts (#19): Qwen 3.5, GLM 5, MiniMax 2.5 — Chinese labs' latest push of the frontier, Interconnects, 2026.03
Paper/Blog Link My Issue
#Article Issue Date: 2026-03-04 Comment
元ポスト:
How to Create Effective Agent Skills, openhands, 2026.02
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Blog #AgentSkills Issue Date: 2026-03-03 Comment
元ポスト:
