OpenWeight (517) — 2/3
Motif-3-Beta, Motif-Technologies, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #One-Line Notes Issue Date: 2026-08-04 Comment
元ポスト:
アーキテクチャ解説:
所見:
DeepSeek-V4 Pro, MiniMax M3等のモデルと同等程度のスコアの314B-A13Bモデル。
Grouped Differential Latent Attention (GDLA), Grouped PolyNorm activation と呼ばれる(おそらく本モデル独自の)技術を利用している。
DeepSeek-V4-Flash, Deepseek, 2023.07
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #AIAgents #Selected Papers/Blogs Issue Date: 2026-08-02 Comment
Artiflcial Analysisによる評価:
GLM-5.2, Muse Spark 1.1, Gemini3.6-Flash, GPT-5.6-lunaと同等程度のスコアにもかかわらず、スコアに対するパラメータ比率で圧倒的にパレート最適に位置付けられる。
Epoch Capabilities IndexでOpenWeightモデルでKimi K3に次ぐ2位、全体でOpus 4.6に次ぐ7位
Introducing Cosmos 3 Edge, Nvidia, 2026.07
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #Blog #SmallModel #VisionLanguageModel #Robotics #WorldModels #EmbodiedAI #Author Thread-Post Issue Date: 2026-07-31 Comment
元ポスト:
HF: https://huggingface.co/nvidia/Cosmos3-Edge?linkId=100000431533162
autoregressive -> diffusion の two-tower モデル
Cosmos3 テクニカルペーパー:
- [Paper Note] Cosmos 3: Omnimodal World Models for Physical AI, NVIDIA+, arXiv'26, 2026.06
NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval, Nvidia, 2026.07
Paper/Blog Link My Issue
#Article #Embeddings #InformationRetrieval #NLP #LanguageModel #RepresentationLearning #RAG(RetrievalAugmentedGeneration) #MultiLingual #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-07-31 Comment
元ポスト:
HF: https://huggingface.co/collections/nvidia/nemotron-3-embed
LLMベースのretrieval, STSタスク特化のembeddingモデルで、最大32k contextをサポート。多言語対応している。
RTEB Multilingualと呼ばれるベンチマークでSoTA:
https://mteb-leaderboard.hf.space/benchmark/RTEB(beta)
- Introducing RTEB: A New Standard for Retrieval Evaluation, Liu+, 2025.10
解説:
Autoregressiveなモデルをbidirectionalなattentionに変換した上で、contrastive learning+高品質データでFinutuning。小規模モデルの作成には、pruning, NAS, 蒸留を2回繰り返すことで実現。
Solar Open 2: Korea's Sovereign Foundation Model, Built for Agentic Use, Upstage, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Korean Issue Date: 2026-07-23 Comment
元ポスト:
Introducing Laguna S 2.1 Poolside team, Poolside, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Blog Issue Date: 2026-07-22 Comment
Announcing Bonsai 27B: The First 27B-Class Model to Run on a Phone, PrismRL, 2026.07
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #LanguageModel #MultiModal #Blog #Selected Papers/Blogs #VisionLanguageModel #Initial Impression Notes #Author Thread-Post Issue Date: 2026-07-19 Comment
元ポスト:
HF: https://huggingface.co/collections/prism-ml/bonsai-27b
Bonsaiシリーズ:
- Announcing 1-bit Bonsai: The First Commercially Viable 1-bit LLMs, 2026.03
- Introducing Ternary Bonsai: Top Intelligence at 1.58 Bits, PrismML, 2026.04
- Introducing 1-bit and Ternary Bonsai Image 4B: Image Generation for Local Devices, PrismML, 2026.05
1-bit, ternary weightによって、27B級モデルがエッジデバイス上で動作する。
How We Serve Ideogram V4 Lightning Fast on fal, fal, 2026.07
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #TextToImageGeneration #Blog #Initial Impression Notes Issue Date: 2026-07-19 Comment
元ポスト:
関連:
- Ideogram 4: Open image model at the forefront of design, Ideogram, 2026.06
Ideogramに基づいた、高速なTextToImage Modelのようである。
SingGuard-NSFA, inclusionAI, 2026.07
Paper/Blog Link My Issue
#Article #NLP #AIAgents #Safety #One-Line Notes #Security #Safeguard Issue Date: 2026-07-19 Comment
元ポスト:
AI Agent向けのガードレールで、モデルの安全性からランタイムの安全性を指向して開発されている。エージェントのアクション実行前に、attackを検知し防御するなどの用途に使われる。
Kimi K3: Open Frontier Intelligence, Moonshot AI, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Selected Papers/Blogs #VisionLanguageModel #reading #One-Line Notes #Reference Collection Issue Date: 2026-07-17 Comment
元ポスト:
Artificial Analysisによる評価:
Artificial Analysis IndexでOpus 4.8超え!?
各ブロックの出力を重みつきで足し合わせることで柔軟にどのブロックにattendするかを決定するattention residualと呼ばれる技術が導入されているように見える。
また、MoEのexpert数をスケールさせ、896 experts, 16 activated expertsに変更。
学習レシピの洗練化と合わせてK2との比較で2.5倍のスケーリング効率が改善され、投入した計算コストをより効果的に知能に変換しているとのこと。
やはり
- [Paper Note] Slicing and Dicing: Configuring Optimal Mixtures of Experts, Margaret Li+, arXiv'26, 2026.05
でも示されているようにexpert数を増やすのが今のところ良さそうに見える。
どうやら2.8Tモデルらしい:
design arenaでFable5超えのSoTA:
writingに関するベンチマークでSoTA:
Latent MoE
- [Paper Note] LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts, Venmugil Elango+, arXiv'26, 2026.01
KDA
- [Paper Note] Kimi Linear: An Expressive, Efficient Attention Architecture, Kimi Team+, arXiv'25, 2025.10
Linear TransformerとDelta Net
- [Paper Note] Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention, Angelos Katharopoulos+, ICML'20
- [Paper Note] Linear Transformers Are Secretly Fast Weight Programmers, Imanol Schlag+, arXiv'21, 2021.02
所見:
Latent MoEによってこれほどのsparsityを実現できているのではという所見:
GDPValでも3位:
Kernel optimizationにおいてもFable5と同等性能という報告:
Next.jsのコーディングベンチマークでもFable5超えのSoTAとのこと:
DeepSWEでGPT-5.6-sol, fable5に次ぐ3位:
アーキテクチャを公開されている情報から再構築した図:
公式ブログの図ではValueにはConvがかかっていないように見えたが、はたして。
テクニカルレポート:
https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf
HF:
https://huggingface.co/moonshotai/Kimi-K3
アーキテクチャのポイント:
テクニカルレポートまとめ(スレッド):
解説スレッド:
解説:
関連:
- AgentENV, kvcache-ai, 2026.07
- MoonEP, MoonshotAI, 2026.07
- [Paper Note] Attention Residuals, Kimi Team+, arXiv'26, 2026.03
- FlashKDA: Flash Kimi Delta Attention — high-performance KDA kernels built on CUTLASS, MoonshotAI, 2026.04
- [Paper Note] Kimi Linear: An Expressive, Efficient Attention Architecture, Kimi Team+, arXiv'25, 2025.10
- MLA
- [Paper Note] DeepSeek-V3 Technical Report, DeepSeek-AI+, arXiv'24, 2024.12
- Latent MoE
- [Paper Note] LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts, Venmugil Elango+, arXiv'26, 2026.01
- MoonViT
- [Paper Note] Kimi-VL Technical Report, Kimi Team+, arXiv'25
KDAの更新ルールに位置情報が直接エンコードされるためRoPEを必要としない:
WSDよりもcosine decayの方が一貫して優れていた...?どういう条件の実験だろうか:
- [Paper Note] MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies, Shengding Hu+, COLM'24
Quantile Balancingについて理解したい
VQAからDocument Parsingへ:Nemotron-3-Nano-Omniに対する日本語文書の構造化出力Post-training, Stockmark, 2026.07
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #Japanese #VisionLanguageModel #reading Issue Date: 2026-07-16 Comment
元ポスト:
Inkling: Our open-weights model, THINKING MACHINES, 2026.07
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #LanguageModel #MultiModal #SpeechProcessing #Blog #Reasoning #MoE(Mixture-of-Experts) #Selected Papers/Blogs #VisionLanguageModel #UMM #KeyPoint Notes #Reference Collection #AudioLanguageModel #Author Thread-Post Issue Date: 2026-07-16 Comment
HF: https://huggingface.co/thinkingmachines/Inkling
元ポスト:
THINNING MACHINESによる最初のOpenWeightモデル。975B-41B
- フルスクラッチで学習したReasoningモデル
- 1M context window
- text, image, audio, video 45Tトークンで学習
- バランスよくさまざまな領域で性能を発揮するように訓練し、Tinker上でのfinetuningやカスタマイゼーションの良い基盤として機能することを目指した
- vision, audioドメインに関してはencoder freeアーキテクチャを採用
- audio signalの入力は dMel spectrogramsと呼ばれる手法を採用
- [Paper Note] dMel: Speech Tokenization made Simple, Richard He Bai+, arXiv'24, 2024.07
- 画像は40x40のパッチに分割され、4つのhMLPと呼ばれるレイヤーでエンコーディング
- [Paper Note] Three things everyone should know about Vision Transformers, Hugo Touvron+, ECCV'22, 2022.03
- IFの訓練には、Rubric basedなgraderとclaims graderの2種類のgraderを用いることで、helpfulnessとhallucinationを同時に低減
- Rubric basedなgraderはチェックリストに基づいてスコアリング
- claims graderはfactualな主張をagentic web searchを通じてverificationする
- アーキテクチャはDeepSeek V3を踏襲し、256のexpertsと2つのshared expertを採用し、6つのexpertsがactivateされる
- SWAとGlobal attentionの比率は5:1でKV Headは8つ (GQA)
- **相対位置エンコーディングを用いることでRoPEよりもlong contextに対してより高い外挿性能を示した**
- [Paper Note] Self-Attention with Relative Position Representations, Peter Shaw+, NAACL'18
- K, Vのprojection後、**およびattentionとMLPが残差ストリームに合流する前にshort convolutionを導入**
- QKV projection後に convolution を導入するアーキテクチャは下記研究で提案
- [Paper Note] Primer: Searching for Efficient Transformers for Language Modeling, David R. So+, NIPS'21, 2021.09
- optimiserはMuonとAdamWのハイブリッドで、前者は巨大な行列の重みに対して適用し、そのほかは後者を利用
- 重み自体が多様体上に存在するように制約することが有効であったことに着想を得て、weight decayを学習率の2乗に連動させることでモデルの重みの大きさを安定させたとのこと
- Modular Manifolds, Jeremy Bernstein+, THINKING MACHINES, 2025.09
- [Paper Note] Why Gradients Rapidly Increase Near the End of Training, Aaron Defazio, arXiv'25, 2025.06
- 事後学習としては、まずKimi K2.5を含むOpenWeightモデルで合成データ上でSFTをし、その後合成、あるいは人間が作成したenvironment上で大規模なRLを実施
- RLは非同期RLを異様し、ロールアウト数に対して推論能力が対数線形にスケールした。
- 最終的に30Mロールアウト以上のRLを実施した
- システムメッセージを変更し、トークン単位のコストを調整することでreasoning effortを調整
- Controlling Reasoning Effort in LLMs, Sebastian Raschka, 2026.07
- **276B-A12BのInkling-Smallもプレビュー段階にあり、agenticな能力においてInklingに近い性能を達成しており、今後公開予定**
1M context windowを実現した工夫は特に書かれていなかったように感じるが、よく使われるのはKV CacheをコンパクトにするためのSparse Attention、あるいはLinear Attentionなのに対し、本モデルでは使われていないように見える。linear attentionでなくとも、SWAとGlobal Attentionのハイブリッド、かつGQAによって、KV Cacheの肥大化を実用レベルで抑制できるのだろうか。また、外挿性能に関してはRoPEは採用せず、相対位置エンコーディングのを採用したという点が効いているのだろうか。1Mレベルでのcontextでの評価がなさそうに見えるので性能はよくわからない。
RoPEのlong contextでの限界を示した以下の研究を思い出した:
- [Paper Note] RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably, Yufeng Du+, arXiv'26, 2026.05
Artificial Analysisによる評価:
アーキテクチャでの目新しい点:
所見:
公式:
所見:
LingBot-VLA-V2, Robbyant, 2026.07
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #Robotics #VisionLanguageActionModel #EmbodiedAI #Author Thread-Post Issue Date: 2026-07-12 Comment
元ポスト:
Workspace geometry, model by model., elie, 2026.07
Paper/Blog Link My Issue
#Article #Analysis #NLP #LanguageModel #Blog #read-later #Interpretability #Initial Impression Notes Issue Date: 2026-07-12 Comment
元ポスト:
38種類のOpenWeightモデルのlayer間でのJ-lensの類似性を検証したとのこと
アーキテクチャが異なっていても非常にモデル間で類似しているようである。gemma-4やSLM(Qwen, GPT2)などは外れ値を含む模様
関連:
- Verbalizable Representations Form a Global Workspace in Language Models, Anthropic, 2026.07
大規模言語モデルの次期バージョンPLaMo 3シリーズにおける120B事前学習モデル、31B蒸留モデルの評価, PFN, 2026.07
Paper/Blog Link My Issue
#Article #Pretraining #NLP #LanguageModel #LongContext #Reading Reflections #Author Thread-Post Issue Date: 2026-07-09 Comment
元ポスト:
weight reusing:
- [Paper Note] Scaling Language Models: Methods, Analysis & Insights from Training Gopher, Jack W. Rae+, arXiv'21, 2021.12
- The Surprising Power of Small Language Models, Mojan Javaheripi+, Microsoft Research, 2023.12
long context評価:
- [Paper Note] RULER: What's the Real Context Size of Your Long-Context Language Models?, Cheng-Ping Hsieh+, COLM'24, 2024.04
- [Paper Note] Repeat After Me: Transformers are Better than State Space Models at Copying, Samy Jelassi+, arXiv'24, 2024.02
HF: https://huggingface.co/pfnet/plamo-3-nict-2604-31b-base
YaRNを適用するとlong contextの性能が大きく改善し、事前学習で学習させるトークン数を増やしすぎるとYaRNのcontext lengthの外挿能力が低下するという実験結果が非常に興味深い。
記載がなかった気がするのだが、Sink Tokenの利用などはやっているのだろうか?
- [Paper Note] YaRN: Efficient Context Window Extension of Large Language Models, Bowen Peng+, ICLR'24
Hy3, Tencent, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-07-08 Comment
元ポスト:
- Hy3-preview, tencent, 2026.04
の公開後から、50以上のプロダクトからフィードバックを受けより高品質なデータで事後学習を実施したとのこと。
- GLM-5.2: Built for Long-Horizon Tasks, Z.ai, 2026.06
が744B-A40B。Hy3は295B-A21Bなので、半分程度のパラメータ数でフロンティアモデルに近い性能を達成している。
公式:
Introducing Rampart, National Design Studio, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #Privacy #One-Line Notes #PII #Author Thread-Post Issue Date: 2026-07-08 Comment
HF: https://huggingface.co/nationaldesignstudio/rampart
生成AIに送信されるリクエストを検閲し、プロンプトに含まれる個人情報をマスクするよう訓練された軽量モデルのようである。ブログ下部のデモンストレーションを見るとイメージを掴みやすい。
元ポスト:
Using Local Coding Agents, Sebastian Raschka, 2026.06
Paper/Blog Link My Issue
#Article #Tutorial #NLP #LanguageModel #AIAgents #Blog #Coding #LLMServing #SoftwareEngineering #Author Thread-Post #AgentHarness Issue Date: 2026-07-05 Comment
元ポスト:
Leanstral 1.5: Proof Abundance for All, Mistral AI, 2026.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Coding #Mathematics #Proofs #Author Thread-Post Issue Date: 2026-07-04 Comment
元ポスト:
HF: https://huggingface.co/mistralai/Leanstral-1.5-119B-A6B
所見:
Introducing LongCat-2.0, LongCat, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #read-later #Selected Papers/Blogs #Reference Collection #Initial Impression Notes #Author Thread-Post Issue Date: 2026-06-30 Comment
元ポスト:
- 1M context window
- Sparse Attention
- Muon
- Ngram Embedding
- dynamic activation (33B-56B)
- Multi-teacher OPD
HF:
https://huggingface.co/meituan-longcat/LongCat-2.0
現在はまだ公開されていないがモデルは上記で公開予定
Sparse Attention + 1M Context + MOPD + Muonが標準になっている印象
LongCat Sparse Attentionポイント解説:
所見:
50k基のH800 GPUで学習されているとのこと。
所見:
ポイント解説:
LFM2.5-230M: Built to Run Anywhere, Liquid AI, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #SmallModel Issue Date: 2026-06-26 Comment
HF: https://huggingface.co/LiquidAI/LFM2.5-230M
元ポスト:
Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding, Ornith, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Coding #SoftwareEngineering #PostTraining #One-Line Notes #Author Thread-Post Issue Date: 2026-06-26 Comment
HF: https://huggingface.co/collections/deepreinforce-ai/ornith-10
元ポスト:
gemma4とQwen3.5をpost trainingしたコーディング特化LLMで、397BモデルではSWE Bench ProでGLM 5.2超え
Ming-omni-tts-16.8B-A3B, inclusionAI, 2026.06
Paper/Blog Link My Issue
#Article #NLP #SpeechProcessing #TTS #One-Line Notes #AudioLanguageModel #audio #Music Issue Date: 2026-06-17 Comment
元ポスト:
pj page: https://xqacmer.github.io/Ming-omni-tts/
speech/sound/musicを単一のモデルで生成可能
GLM-5.2: Built for Long-Horizon Tasks, Z.ai, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Actor-Critic #Selected Papers/Blogs #Reference Collection #Critic #LongHorizon #train-inference-mismatch #Initial Impression Notes Issue Date: 2026-06-16 Comment
HF: https://huggingface.co/zai-org/GLM-5.2
1Mコンテキストでフロンティアモデルには少し劣るが非常に高い性能。SWE Bench ProではGPT5.5をoutperform、Terminal Bench/DeepSWEなどでは3--12pt程度GPT5.5が高い。Opus4.8はさらにスコアが高い。
744B-A40B
元ポスト:
Artificial Analysis Intelligence IndexでOpenWeightモデルでSoTA更新:
MiniMax-M3超え
long Horizonのタスクの学習でGRPOではなく、criticベースのPPOを利用:
アーキテクチャ解説:
GLM 5.2では、Reward Hackingを検知した場合、ペナルティを与えるのではなく、ツール呼び出しをブロックしてダミー情報を返し、ロールアウト自体は継続させる。Reward Hackingをすると、ただ損をするという構造に近づけていると思われる。
slimeフレームワークを用いることで10を超えるエキスパートモデルに基づくOPDを2日で終えたとのこと。この記述に基づき、OPDがRLと比較して非常に効率的で、かつエキスパートなモデルは並列で学習をさせておけるよね?という利点に関して推察をしている:
エキスパートなRLでは常にドメイン別のRLを走らせ特化モデルを学習しておき、性能が向上したらOPDで単一モデルに効率的に蒸留することで、モデルに知識を集約するプロセスと、個々のドメインごとに強力なモデルを得るプロセスが分離されて良さそう、という感想を得た。
1M context + linear attentionが標準に:
MTPの学習時において、step 2以後のMTP Layerの出力は、学習時はteacher forcing, 推論時はstep 1のMTP Layerが推定したhidden_stateが用いられるが、これが学習-推論時のgapを生んでしまう。このため、step 1の計算で利用したKV(とsparse attentionのtop-kで洗濯をしたトークンのindex)を全て共有し、step 2以後のMTP Layerで活用し、かつstep 2以後のsparse attentionの計算のオーバヘッドをなくしつつ、学習-推論のgapを無くすことでspeculative decodingの性能を向上させる、といった話題があるようである:
画像分野におけるself-forcingと同じ発想に思う。
- [Paper Note] Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion, Xun Huang+, NeurIPS'25
GRPO, PPOのそれぞれの利点と欠点、およびLong Horizonタスクに適用する価値の説明:
Cursor Benchにおいて、Opus 4.8と同等程度のコスト効率とのこと:
PPOへ最終的に回帰してきたのは一種のbitter lesson:
PPO vs. GRPOに関する所見:
GLM 5.2で利用されたRL:
- [Paper Note] Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning, Zhenyu Hou+, arXiv'26, 2026.07
要はGRPOだと応答単位のrewardから算出されるAdvantageがトークン単位に一律に割り当てられるが、PPOの場合はトークン単位のcredit assignmentが可能(Criticが各時点でのValueを推定し、GAEなどを通じてAdvantageに反映されるため)な点が重要と思われる。
ZONOS2: Real-time TTS with High-Fidelity Voice Cloning, ZYPHRA, 2026.06
Paper/Blog Link My Issue
#Article #NLP #Dataset #Evaluation #SpeechProcessing #MultiLingual #MoE(Mixture-of-Experts) #TTS #Realtime #Author Thread-Post Issue Date: 2026-06-15 Comment
元ポスト:
Introducing North Mini Code: Cohere’s first model for developers, Cohere, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Coding #SoftwareEngineering Issue Date: 2026-06-14 Comment
元ポスト:
アーキテクチャ解説:
Kimi-K2.7-Code, moonshotai, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Coding #SoftwareEngineering Issue Date: 2026-06-12 Comment
元ポスト:
DiffusionGemma: 4x faster text generation, Google, 2026.06
Paper/Blog Link My Issue
#Article #LanguageModel #DiffusionModel #Blog #Author Thread-Post Issue Date: 2026-06-11 Comment
元ポスト:
関連:
- [Paper Note] How Transparent is DiffusionGemma?, Joshua Engels+, arXiv'26, 2026.06
- DiffusionGemma: The First Diffusion LLM (dLLM) Natively Supported in vLLM, vLLM, 2026.06
gemma-4-12B-it-qat-UD-japanese-imatrix, dahara1, 2026.06
Paper/Blog Link My Issue
#Article #LanguageModel #Japanese #read-later #Author Thread-Post Issue Date: 2026-06-11 Comment
元ポスト:
Quasar-Preview, silx-ai, 2026.06
Paper/Blog Link My Issue
#Article #Transformer #SmallModel #RecurrentModels Issue Date: 2026-06-09 Comment
元ポスト:
実験的にcontext windowが5M
Apodex-1.0: A Verification-Centric Agent Team for Discoverative Intelligence, Apodex, 2026.06
Paper/Blog Link My Issue
#Article #DeepResearch #Author Thread-Post Issue Date: 2026-06-09 Comment
HF: https://huggingface.co/collections/apodex/apodex-1
元ポスト:
LFM2.5-1.2B-JP-202606, LiquidAI, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #SmallModel #Japanese Issue Date: 2026-06-06 Comment
元ポスト:
LFM2.5-Audio-1.5B-JP, LiquidAI, 2026.06
Paper/Blog Link My Issue
#Article #SpeechProcessing #Japanese #One-Line Notes #audio #SpeechToSpeech Issue Date: 2026-06-06 Comment
元ポスト:
LiquidAI初の日本語特化speech-to-speechモデル
Ideogram 4: Open image model at the forefront of design, Ideogram, 2026.06
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #TextToImageGeneration #ImageSynthesis #Author Thread-Post Issue Date: 2026-06-05 Comment
元ポスト:
HF: https://huggingface.co/collections/ideogram-ai/ideogram-4
NVIDIA Nemotron 3 Ultra, nvidia, 2026.06
Paper/Blog Link My Issue
#Article #Pretraining #NLP #LanguageModel #SSM (StateSpaceModel) #MoE(Mixture-of-Experts) #read-later #Selected Papers/Blogs #Reference Collection #LowPrecision #LinearAttention #Author Thread-Post Issue Date: 2026-06-05 Comment
元ポスト:
アーキテクチャ解説:
Mamba2 layer, Latent MoE, GQA
ポイント解説:
HF: https://huggingface.co/collections/nvidia/nvidia-nemotron-v3
所見:
所見:
LFM2.5-VL-450M-Extract, LiquidAI, 2026.06
Paper/Blog Link My Issue
#Article #SmallModel #VisionLanguageModel #One-Line Notes #Author Thread-Post Issue Date: 2026-06-05 Comment
元ポスト:
画像から抽出したい情報をシステムプロンプトに定義することで、json形式で結果を返してくれるVLM
Introducing Gemma 4 12B: a unified, encoder-free multimodal model, Google, 2026.06
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #VisionLanguageModel #2D (Image) #UMM #SpatialUnderstanding #One-Line Notes #Reference Collection #AudioLanguageModel #audio #Author Thread-Post Issue Date: 2026-06-04 Comment
元ポスト:
vision/audioエンコーダーを無くしたvision/audio nativeなマルチモーダルLLM
HF: https://huggingface.co/google/gemma-4-12B
アーキテクチャ図:
Holo3.1: Fast & Local Computer Use Agents, H Company, 2026.06
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #AIAgents #ComputerUse #VisionLanguageModel #Author Thread-Post Issue Date: 2026-06-03 Comment
HF: https://huggingface.co/collections/Hcompany/holo31
元ポスト:
関連:
- Holo3: Breaking the Computer Use Frontier, H Company, 2026.03
Develop Physical AI Reasoning, World, and Action Models with NVIDIA Cosmos 3, nvidia, 2026.05
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #Selected Papers/Blogs #VideoGeneration/Understandings #Robotics #WorldModels #UMM #reading #Omni #One-Line Notes #WorldActionModel #Author Thread-Post Issue Date: 2026-06-02 Comment
元ポスト:
公式:
encoder-freeなOmniモダリティモデルで、かつ将来の世界の状態、およびactionを予測可能なWorldActionModel
MiniMax-M3, MiniMaxAI, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MultiModal #Post #Selected Papers/Blogs #VisionLanguageModel #UMM #One-Line Notes #Reference Collection #Author Thread-Post Issue Date: 2026-06-01 Comment
ベンチマーク上はフロンティアモデルに性能がかなり肉薄しており、10日以内にモデルがオープンになる。
所見:
関連:
- [Paper Note] Learning Dynamics of LLM Finetuning, Yi Ren+, ICLR'25 Outstanding Paper Award
Artificial Analysisによる評価:
OpenWeightでSoTA
Announcing Surya OCR 2: small, accurate, multilingual, Datalab, 2026.05
Paper/Blog Link My Issue
#Article #Blog #OCR #Author Thread-Post Issue Date: 2026-05-31 Comment
HF: https://huggingface.co/datalab-to/surya-ocr-2
元ポスト:
Step-3.7-Flash, stepfun-ai, 2026.05
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #AIAgents #MultiModal #MoE(Mixture-of-Experts) #read-later #VisionLanguageModel #Author Thread-Post Issue Date: 2026-05-31 Comment
元ポスト:
公式:
LFM2.5-8B-A1B: An Even Better On-Device Mixture of Experts, LiquidAI, 2026.05
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #SmallModel #read-later Issue Date: 2026-05-31 Comment
Hy-MT2-30B-A3B, Tencent Hy, 2026.05
Paper/Blog Link My Issue
#Article #MachineTranslation #NLP #LanguageModel #MultiLingual #One-Line Notes #Author Thread-Post Issue Date: 2026-05-27 Comment
HF: https://huggingface.co/collections/tencent/hy-mt2
元ポスト:
テンセントによる1.8B--30BのMT特化モデルファミリー。fast thinkingが強みとのこと。
Nemotron-Labs-Diffusion-14B, Nvidia, 2026.05
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #DiffusionModel #LLMServing #SpeculativeDecoding #One-Line Notes #Author Thread-Post Issue Date: 2026-05-27 Comment
元ポスト:
3つの生成モード: AR/dLM/Hybrid を備えたLLM(VLM variantも存在)ファミリーで、ARモードでは一般的な自己回帰的な生成をし、dLMモードでは拡散モデルに基づくparallel decodingを実施、hybridではdLMでドラフト作成、ARでverificationを実施するSpeculative Decoding (self-speculation)を実施する。これらモードは内部のattention patternを変化させることでシームレスに切り替えられ(シームレスモード)期待されるconcurrencyに応じて柔軟に対応ができるようである。
シームレスの粒度がどの程度のものかはよくわからない。concurrency levelを検知して、それに応じて動的に切り替わったりするのだろうか。
Speculative Decodingの高速化手法としては以下のようなものもある:
- [Paper Note] TriSpec: Ternary Speculative Decoding via Lightweight Proxy Verification, Haoyun Jiang+, arXiv'26, 2026.01
OlmoEarth v1.1: A more efficient family of models, Ai2, 2026.05
Paper/Blog Link My Issue
#Article #ComputerVision #EfficiencyImprovement #FoundationModel #2D (Image) #Author Thread-Post Issue Date: 2026-05-27 Comment
元ポスト:
関連:
- OlmoEarth-v1-Large, Ai2, 2025.11
- Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis, Ai2, 202604
Toto 2.0: Time series forecasting enters the scaling era, DATADOG, 2026.05
Paper/Blog Link My Issue
#Article #TimeSeriesDataProcessing #MachineLearning #Transformer #FoundationModel #One-Line Notes Issue Date: 2026-05-21 Comment
HF: https://huggingface.co/collections/Datadog/toto-20
時系列予測の基盤モデルも、パラメータサイズに対して性能がスケールする(ということが初めて示された)
Ring-2.6-1T, inclusionAI, 2026.05
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Author Thread-Post Issue Date: 2026-05-21 Comment
元ポスト:
Introducing Command A+: Making sovereign agentic capabilities available to all, Cohere, 2026.05
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MultiLingual #MoE(Mixture-of-Experts) #Reference Collection #Initial Impression Notes Issue Date: 2026-05-21 Comment
元ポスト:
HF:
https://huggingface.co/CohereLabs/command-a-plus-05-2026-w4a4
apache-2.0
デコーディング速度が非常に速い
アーキテクチャサマリ:
-
-
翻訳性能高いよという話のようである:
ZAYA1-VL-8B: Efficient Open Visual Intelligence, ZYPHRA, 2026.05
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #SmallModel #MoE(Mixture-of-Experts) #VisionLanguageModel #One-Line Notes Issue Date: 2026-05-12 Comment
HF: https://huggingface.co/Zyphra/ZAYA1-VL-8B
元ポスト:
画像トークンには双方向のattentionを適用できるようなアーキテクチャを採用
Marco-MoE: A suit of multilingual MoE models with highly-sparse architectures, AIDC-AI, 2026.05
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #MultiLingual #MoE(Mixture-of-Experts) #CurriculumLearning #One-Line Notes Issue Date: 2026-05-12 Comment
元ポスト:
4 stageのカリキュラムによって学習されているようで、学習が進むにつれて、広く使われる英語やreasoning, instructionなどのデータは減らし、low resourceな言語のデータを増やしていき最終的にマルチリンガルなデータを支配的にするような学習レシピとなっているようである。
MolmoAct 2: An open foundation for robots that work in the real world, Ai2, 2026.05
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #Reasoning #OpenSource #read-later #Robotics #VisionLanguageActionModel #Author Thread-Post Issue Date: 2026-05-08 Comment
元ポスト:
関連:
- [Paper Note] MolmoAct: Action Reasoning Models that can Reason in Space, Jason Lee+, arXiv'25
dataset:
https://huggingface.co/collections/allenai/molmoact2-datasets
models:
https://huggingface.co/collections/allenai/molmoact2-models
著者ポスト:
著者ポスト2:
著者ポスト3:
Cosmos-Reason2-32B, Nvidia, 2026.04
Paper/Blog Link My Issue
#Article #NLP #Reasoning #VisionLanguageModel #Robotics #EmbodiedAI Issue Date: 2026-05-01 Comment
元ポスト:
Mistral-Medium-3.5-128B, MistralAI, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel Issue Date: 2026-04-30
Laguna XS.2 and M.1: A Deeper Dive, Poolside team, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Coding #MoE(Mixture-of-Experts) #SoftwareEngineering #read-later #Selected Papers/Blogs #LongHorizon #Author Thread-Post Issue Date: 2026-04-30 Comment
HF: https://huggingface.co/poolside/Laguna-XS.2
元ポスト:
テクニカルレポート:
https://poolside.ai/assets/laguna/laguna-m1-xs2-technical-report.pdf
元ポスト:
Introducing talkie: a 13B vintage language model from 1930, Levine+, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #read-later #Selected Papers/Blogs #Initial Impression Notes #vintage LLMs Issue Date: 2026-04-29 Comment
元ポスト:
1930年以前の英語テキストで学習された言語モデル(vintage Large Language Models)で、歴史や文化の変化を分析したり、1930年までのデータで学習されたモデルが1931年以後に発見された革新的な科学的な発見を自ら見出せるか?、LLMが将来を予測する能力がどの程度あり、それがモデルサイズによってどのように変化するか?、プログラミングに関する知識がないモデルが現代のコーディングを学習できるかなどのcontamination freeな評価など様々な活用方法があるとのこと。
関連:
- Vintage Large Language Models, Owain Evans
所見:
Hunyuan3D-2, tencent, 2026.04
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #Transformer #DiffusionModel #TextTo3D #ImageTo3D Issue Date: 2026-04-25 Comment
元ポスト:
Hy3-preview, tencent, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MoE(Mixture-of-Experts) #Author Thread-Post Issue Date: 2026-04-24 Comment
元ポスト:
Xiaomi MiMo-V2.5-Pro: A leap in agentic and long horizon coherence, Xiaomi, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #MultiModal #Blog #Coding #Selected Papers/Blogs #UMM #Reference Collection #Initial Impression Notes #Author Thread-Post Issue Date: 2026-04-23 Comment
元ポスト:
いずれモデルをオープンにするとのこと
Artificial Analysisによる評価:
オープンになった:
https://huggingface.co/collections/XiaomiMiMo/mimo-v25
元ポスト:
GDPValやSWE-Bench-ProがGemini-3.1-Proよりも高い。
MIT Licenceかつnative multimodal
所見:
解説:
privacy-filter, openai, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Encoder #PII Issue Date: 2026-04-23 Comment
元ポスト:
Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model, Qwen Team, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Coding #SoftwareEngineering #One-Line Notes #Author Thread-Post Issue Date: 2026-04-23 Comment
HF: https://huggingface.co/Qwen/Qwen3.6-27B
元ポスト:
Qwen3.5-397B-A17Bを主要なcodingベンチマークで上回り、同等程度の規模感のdenseモデルを上回る。
inclusionAI: Ling-2.6-flash (free), OpenRouter (InclusionAI), 2026.04
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #AIAgents #MoE(Mixture-of-Experts) #Reference Collection #Initial Impression Notes #Author Thread-Post Issue Date: 2026-04-22 Comment
元ポスト:
Lingの最新モデル。元ポストに強みが簡潔に書かれている。OpenRouterで1週間freeで利用可能で、今後商用モデルのLingDTのリリースも控えているとこと。
また、将来的に本モデルはオープンになる予定とのこと。
Artificial Analysisによる評価:
オープンになった:
HF: https://huggingface.co/inclusionAI/Ling-2.6-flash
Kimi K2.6: Advancing Open-Source Coding, Kimi, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Selected Papers/Blogs #KeyPoint Notes #Reference Collection Issue Date: 2026-04-21 Comment
ブログ中ではまずはAgenticな能力の評価が掲載されており、スコアとしてはOpus 4.6と同等程度の水準に達している。
Kimi-K2.5と同様Agent Swarmを採用している。
- [Paper Note] Kimi K2.5: Visual Agentic Intelligence, Kimi Team+, arXiv'26, 2026.02
推論・知識に関するベンチマーク(AIME, HMMT, GPQA-Diamond)などについては、Opus4.6と比較してスコアが高いのはIMO-AnswerBenchと呼ばれるものだけであり、他は同等かスコアが低くなっている。Vision系のベンチマークでは、全体的にOpus4.6よりもスコアが高い。ただし、Gemini-3.1-Pro, GPT-5.4の方がKimi K2.6よりもスコアが全体として高い。
他にも5日間にわたる監視システムのようなプロアクティブなエージェントとしても活用でき、独自ベンチマークのKimiClawBenchと呼ばれるものでK2.5を上回った旨が記述されているが、詳細不明。
元ポスト:
HF: https://huggingface.co/moonshotai/Kimi-K2.6
その他ベンチマーク情報:
HY-World-2.0, Tencent, 2026.04
Paper/Blog Link My Issue
#Article #ComputerVision #Transformer #DiffusionModel #WorldModels #Author Thread-Post Issue Date: 2026-04-16 Comment
元ポスト:
テクニカルレポート: https://3d-models.hunyuan.tencent.com/world/world2_0/HY_World_2_0.pdf
Qwen3.6-35B-A3B: Agentic Coding Power, Now Open to All, QwenTeam, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #MultiModal #MoE(Mixture-of-Experts) #Selected Papers/Blogs #Sparse #Initial Impression Notes #Author Thread-Post Issue Date: 2026-04-16 Comment
HF: https://huggingface.co/Qwen/Qwen3.6-35B-A3B
元ポスト:
ざっと見た感じ明言されていない気がするが、プロプライエタリとなったQwen3.6-Plusの廉価版(オープンなので廉価と言うのかはあれだが)だと思われる。
Introducing ERNIE‑Image, Baidu, 2026.04
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #Transformer #DiffusionModel #TextToImageGeneration #Selected Papers/Blogs #2D (Image) #One-Line Notes #ImageSynthesis #Author Thread-Post Issue Date: 2026-04-15 Comment
HF: https://huggingface.co/baidu/ERNIE-Image
ERNIEからtext-to-imageモデルがOpenWeightモデルとしてリリース。ベンチマークとしては公式ブログ上ではOpenWeightモデルの中でトップで、nano banana 2.0に匹敵するようなスコアが出ているように見える
LLM-jp-4-VL 9B betaリリース, LLM-jp, 2026.04
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #Japanese #OpenSource #VisionLanguageModel #Author Thread-Post Issue Date: 2026-04-14 Comment
元ポスト:
Marco-Mini-Instruct, AGDC-AI, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MultiLingual #MoE(Mixture-of-Experts) Issue Date: 2026-04-11 Comment
元ポスト:
The ATOM Report: Measuring the Open Language Model Ecosystem, Lambert+, 2026.04
Paper/Blog Link My Issue
#Article #Analysis #NLP #LanguageModel #OpenSource #read-later #Data #Author Thread-Post Issue Date: 2026-04-11 Comment
著者ポスト:
元ポスト:
Unfolding Robotics: The Open-Source Recipe for Teaching a Robot to Fold Your Clothes, Hugging Face, 2026.04
Paper/Blog Link My Issue
#Article #Tutorial #ComputerVision #NLP #OpenSource #read-later #Selected Papers/Blogs #Robotics #VisionLanguageActionModel Issue Date: 2026-04-07 Comment
元ポスト:
Microsoft Open-Sources Industry-Leading Embedding Model, Microsoft Bing Blog, 2026.04
Paper/Blog Link My Issue
#Article #Embeddings #NLP #Blog #MultiLingual #read-later Issue Date: 2026-04-07 Comment
元ポスト:
GLM-5.1: Towards Long-Horizon Tasks, Z.ai, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Selected Papers/Blogs #Reference Collection Issue Date: 2026-04-07 Comment
元ポスト:
SWE Bench ProでSoTA...?!
HF: https://huggingface.co/zai-org/GLM-5.1
Artificial Analysis:
アーキテクチャ解説:
DeepSeekV3.2 likeなアーキテクチャで、MLA, DeepSeek Sparse Attentionを採用。Layer数がDeepSeekV3.2より多いとのこと。
Introducing WildDet3D: Open-world 3D detection from a single image, Ai2, 2026.04
Paper/Blog Link My Issue
#Article #ComputerVision #Dataset #OpenSource #read-later #Selected Papers/Blogs #3D (Video) #ObjectDetection #Initial Impression Notes Issue Date: 2026-04-07 Comment
元ポスト:
wildな環境においてzero shot(click, text, bounding boxで対象を指定)で動作する単眼の3D Object Detectionモデルとのこと。データセットもコードも公開
オープンソースAIの現状 | NVIDIA GTC, Nvidia, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Video #OpenSource #read-later #One-Line Notes Issue Date: 2026-04-07 Comment
元ポスト:
GTCのパネルディスカッション
VoxCPM2, OpenBMB, 2026.04
Paper/Blog Link My Issue
#Article #SpeechProcessing #SmallModel #MultiLingual #TTS Issue Date: 2026-04-07 Comment
github: https://github.com/OpenBMB/VoxCPM/?tab=readme-ov-file
元ポスト:
30+言語をサポート
ibm-granite_granite-4.0-3b-vision, ibm-granite, 2026.03
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #VisionLanguageModel Issue Date: 2026-04-04 Comment
元ポスト:
約12兆トークンの良質なコーパスで学習した新たな国産LLM「LLM-jp-4 8Bモデル」「LLM-jp-4 32B-A3Bモデル」をオープンソースライセンスで公開 ~一部ベンチマークでGPT-4oやQwen3-8Bを上回る性能を達成~, NII, 2026.04
Paper/Blog Link My Issue
#Article #Pretraining #NLP #LanguageModel #Reasoning #Japanese #OpenSource #mid-training #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-04-03 Comment
8BモデルはLlama-2アーキテクチャ、32B-A3.8BモデルはQwen3-MoEアーキテクチャで、フルスクラッチ学習をすることで実現[^1]。
19.5Tトークン(概算として、日本語0.7Tトークン、英語17.8Tトークン、中国語・韓国語0.85Tトークン、プログラムコード0.2Tトークン)のインターネット上の公開データや政府・国会の文書を収集し(LLM-jp-3.1のデータの6倍の規模)し事前学習データを構築、DataMixtureを最適化し10.5Tトークンを事前学習で利用。
中間学習では、事前学習データにInstruction Pretraining[^2]データを含む合成データを加え1.2Tトークンを利用。
その後最終的にInstruction Tuningを、日本語、英語合計22種類のデータで実施(元記事ではチューニングと呼称されているがおそらくInstruction Tuningだと思われる)。
MTBenchでは、GPT-4o, gpt-oss-20B, Qwen3-8Bと同等以上の性能、日本語MTBench[^3]では、GPT-4o, gpt-oss-20B, Qwen3-8Bを上回る性能とのこと。MTBenchで用いるLLM-as-a-JudgeのモデルとしてはGPT-5.4を利用とのこと。
[^1]: つまり、モデルのパラメータは完全に新規で学習されており、ベースとして既存OpenWeightモデルを利用していない点に注意。
[^2]: Instruction Pretrainingは、LLM-jp-3.1の頃から実施されている:
LLM-jp-3.1 シリーズ instruct4 の公開, LLM-jp, 2025.05
[Paper Note] Instruction Pre-Training: Language Models are Supervised Multitask Learners, Daixuan Cheng+, arXiv'24, 2024.06
[^3]: MT-Benchの概要については
[Paper Note] Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng+, NeurIPS'23, 2023.06
も参照のこと。
フルスクラッチモデル点に関する説明:
HF:
https://huggingface.co/collections/llm-jp/llm-jp-4-models
Reasoningモデルもある!!!
関連:
- PLaMo 3.0 Prime β版, PFN, 2026.03
上記PLaMo 3.0に続いて、国内でのフルスクラッチReasoningモデルは二例目だろうか。
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory, Skywork AI, 2026.04
Paper/Blog Link My Issue
#Article #ComputerVision #Transformer #SyntheticData #DiffusionModel #VideoGeneration/Understandings #WorldModels #interactive #Game #3D (Video) #LongHorizon #Realtime #Initial Impression Notes Issue Date: 2026-04-02 Comment
元ポスト:
Unreal Engineで合成されたデータに基づいて学習されたDiTベースのWorld Modelらしい。
Acknowleagementから察するに、Wan2.2がベースモデルで、self-forcingが学習に用いられている。
- Wan2.2, Alibaba Wan, 2025.07
- [Paper Note] Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion, Xun Huang+, NeurIPS'25
また、action control moduleをアーキテクチャに導入することで、汎用的な動画生成モデルにキーボード、マウス等のアクションによるコントロールを実現している模様。
- [Paper Note] GameFactory: Creating New Games with Generative Interactive Videos, Jiwen Yu+, arXiv'25, 2025.01
デコードの高速化には量子化を利用しているとのこと。
Gemma 4: Byte for byte, the most capable open models, Google, 2026.04
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #LanguageModel #AIAgents #MultiModal #SpeechProcessing #Reasoning #MoE(Mixture-of-Experts) #Selected Papers/Blogs #VisionLanguageModel #2D (Image) #3D (Video) #One-Line Notes #Reference Collection #AudioLanguageModel #audio #text #Initial Impression Notes Issue Date: 2026-04-02 Comment
元ポスト:
2B, 4B, 26BのMoEモデルと31BのDenseモデルの4種類のモデルファミリーで、マルチモーダル(vision)対応。2B, 4Bはaudioも入力として扱える。
edgeデバイス向けのモデルは128k, 他は256kのコンテキストウィンドウ。140+の多言語サポート。
Apache 2.0ライセンス
arenaで同サイズのモデル群でSoTAといった話がブログ中に記述されている。
モデルカードには一般的なベンチマーク群とのスコアも記載されている。
https://ai.google.dev/gemma/docs/core/model_card_4?hl=ja
(そもそも既存のベンチマークにもコンタミネーションがあると思われるが、)arenaに関しては特定の企業に対してデータを提供し、複数のモデルの亜種をテストできるという慣行があり、リーダーボードにバイアスがあるであろう点には注意:
- [Paper Note] The Leaderboard Illusion, Shivalika Singh+, NeurIPS'25
artificial analysisによる評価:
Qwenがproprietaryになったことから、ライセンス的に使いやすく、日本語に強そうなモデルとしては筆頭ではなかろうか。日本語性能が気になる。
アーキテクチャ解説:
ポイント解説:
所見:
attentionのscaleをsqrt(d)でスケールさせる代わりに、QK-norm, V normを適用するなど。
NvidiaによるNVFP4へのpost-trainingによる量子化:
https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4
量子化後の性能も比較されており、知識、数学、コーディング、terminac useなど6種類のベンチマークでオリジナルのモデルと遜色ない性能が出ている旨記載されている。
解説:
https://newsletter.maartengrootendorst.com/p/a-visual-guide-to-gemma-4
所見(encoder-freeにした裏側でパッチ化→projection + x/y軸のpositional embeddingを実施している話):
Holo3: Breaking the Computer Use Frontier, H Company, 2026.03
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #AIAgents #MultiModal #MoE(Mixture-of-Experts) #ComputerUse #read-later #VisionLanguageModel #One-Line Notes #GUI #Environment Issue Date: 2026-04-02 Comment
元ポスト:
HF: https://huggingface.co/Hcompany/Holo3-35B-A3B
関連:
- Holo2: Cost-Efficient Models for Cross-Platform Computer-Use Agents, H Company, 2025.11
Qwen3.5をファインチューニングすることで実現。以前のシリーズもQwenベースだったが、新たなQwenのリリースに伴いより強力なベースモデルを得て、かつシナリオをベースにして自動でwebsiteを構築しverifiableが可能な独自のEnvironmentを保持しており、多様な合成データの活用とRLを実現することで、性能が向上していると思われる。
Trinity-Large-Thinking: Scaling an Open Source Frontier Agent, Arcee, 2026.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Reasoning #MoE(Mixture-of-Experts) #read-later #Selected Papers/Blogs Issue Date: 2026-04-02 Comment
元ポスト:
HF: https://huggingface.co/collections/arcee-ai/trinity-large-thinking
LFM2.5-350M: No Size Left Behind, Liquid AI, 2026.04
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #SmallModel #read-later #Selected Papers/Blogs #KeyPoint Notes Issue Date: 2026-04-01 Comment
元ポスト:
- LFM2のアーキテクチャを採用の350Mパラメータモデルで、CPUでも十分な速度で推論可能
- 追加の事前学習(10T -> 28T tokens)、および、large-scale RLを実施
- 同等規模のパラメータ数(あるいは2倍程度)のモデル群に対して、知識, 指示追従能力, ツール呼び出し、データ抽出などのベンチマークで上回る
- LFM2-350Mと比較して、指示追従能力, データ抽出, tool useの性能が大きく向上
- edgeデバイスでの軽量なデータ抽出パイプラインとして有用
- しかし、math, coding, creative writingなどでの利用は推奨されない
- CPU/GPUでの推論ともに同等規模、あるいは1B級のモデルよりも早く、省メモリ
LongCat-AudioDiT, Meituan LongCatTeam, 2026.03
Paper/Blog Link My Issue
#Article #NLP #SpeechProcessing #DiffusionModel #Architecture #read-later #Selected Papers/Blogs #TTS #Initial Impression Notes Issue Date: 2026-04-01 Comment
HF:
-
https://huggingface.co/meituan-longcat/LongCat-AudioDiT-1B
-
https://huggingface.co/meituan-longcat/LongCat-AudioDiT-3.5B
元ポスト:
デコード時に、メルスペクトログラム→Vocoderの場合細かい特徴が落ちてしまうことが懸念されるため、Waveformを直接デコードするWav-VAEによって、音声に直接変換する、というアーキテクチャの革新があるように見える。
SmolLM - blazingly fast and remarkably powerful, Allal+, HuggingFace, 2024.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel Issue Date: 2026-03-31 Comment
OpenSourceなLLMについて過去を遡ってみているが、SmolLMの最初の段階では、データのみがオープンでコードはオープンでないように見える。
次:
- SmolLM2, 2024.11
RedPajama, a project to create leading open-source models, starts by reproducing LLaMA training dataset of over 1.2 trillion tokens, together.ai, 2023.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #OpenSource #One-Line Notes Issue Date: 2026-03-31 Comment
完全なオープンソースLLMの構築を目指すprojectで、LLaMAの学習データを再現する取り組み。
sarashina2.2-ocr, SBIntuitions, 2026.03
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #Japanese #Selected Papers/Blogs #DocParser #OCR #Initial Impression Notes Issue Date: 2026-03-31 Comment
元ポスト:
縦書き文書に強いのは大変ありがたい
dots.ocrよりも日本語文書に対するCERとBLEUのスコアが良い。素晴らしい
Introducing Marin: An Open Lab for Building Foundation Models, marin-community, 2025.05
Paper/Blog Link My Issue
#Article #Pretraining #NLP #LanguageModel #Blog #OpenSource #Selected Papers/Blogs Issue Date: 2026-03-29 Comment
github:
https://github.com/marin-community/marin
issueのExperimentsが興味深い
関連:
- Marin 32B Retrospective, marin-community, 2025.10
Marin projectのアナウンスをメモっていなかったので今更ながらメモ
- open-weight, open-sourceを超えて、LLMのopen-developmentを実現するための完全な透明性を持ったopen lab
- すべての実験はgithub issueで管理され公開される
- marinのコードベースを使い誰でも実験をコード中に記述しpull repuestを送れ、誰でもレビューできる
- プルリクが承認されると実験が実際に実行され、誰でもWandB上の経過をリアルタイムで観察できる
Delphi[^1]の実験において、25Bパラメータモデルがweight decayフェーズに突入し、Marin-32Bでは以前はweight decayフェーズでloss spikeが頻発したが、Delphiでは安定していそうな見込み、という話がポストされている:
[^1]: 現代版のPythiaを構築しましょうという話で、Pythiaのモデルパラメータを70Bまでスケールアップし、学習に用いるトークン数もチンチラ則従いモデルサイズに応じてスケールアップ、The PileデータなどのデータセットをNemotron-CCなどのlarge scaleモデル用のデータセットに置換する、といった話が含まれる。Marin Issue 1337を参照のこと。
129B-A16Bの学習を開始したとのこと:
535B-A23Bモデルの学習を開始したとのこと:
chandra-ocr-2, datalab-to, 2026.03
Paper/Blog Link My Issue
#Article #ComputerVision #MultiLingual #Selected Papers/Blogs #VisionLanguageModel #OCR #One-Line Notes Issue Date: 2026-03-21 Comment
元ポスト:
日本語の認識性能がGemini-2.5-Flashよりも高い。マルチリンガルでの認識性能がこらほど網羅的に列挙されているのはありがたい。
MiroThinker-1.7, MiroMindAI, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #DeepResearch #LongHorizon #Initial Impression Notes Issue Date: 2026-03-20 Comment
元ポスト:
ベンチマークに応じて、GPT-5, GPT-5.2, GPT-5.4など比較するGPTが恣意的に変わっているように見えるが、ベンチマーク上ではGPT-5と同等以上のAgenticなLLMっぽい?BrowseCompの性能がかなり良さそうに見える。
LLM Architecture Gallery, Sebastian Raschka, 2026.03
Paper/Blog Link My Issue
#Article #Survey #NLP #LanguageModel #Transformer #Blog #Architecture #Initial Impression Notes Issue Date: 2026-03-20 Comment
元ポスト:
Sebastian Raschka氏がいつもポストしているOpenWeight LLMのアーキテクチャ図のギャラリー。パラメータサイズ, head数などの細かい情報も含まれているので、全体を概観するのに良さそう。
MiniMax-M2.7, MiniMax, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Selected Papers/Blogs #Reference Collection Issue Date: 2026-03-19 Comment
所見:
所見:
Artificial Analysisによる評価:
GLM-5と同等の知能スコア、GDPvalでGPT-5.2(xhigh)超え。
modelがオープンに:
https://huggingface.co/MiniMaxAI/MiniMax-M2.7
元ポスト:
openになったが商用利用は許可を得ないとできないということで、リリース時のポストにはopennsourcedと銘打たれているが、open sourceではない。
中国系のOpenModelのライセンス、あるいはプロプライエタリ化が進んできている?
所見:
楽天、「GENIACプロジェクト」の一環として開発された国内最大規模の高性能AIモデル「Rakuten AI 3.0」を提供開始, 楽天グループ株式会社, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Japanese #MoE(Mixture-of-Experts) #Initial Impression Notes Issue Date: 2026-03-18 Comment
HF: https://huggingface.co/Rakuten/RakutenAI-3.0
公式アナウンス、HFのモデルカードの情報が少なすぎてよくわからない。
所見:
Mistral Small 4, MistralAI, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MoE(Mixture-of-Experts) #Initial Impression Notes Issue Date: 2026-03-17 Comment
元ポスト:
119Bでsmallと銘打たれる時代になってしまった
公式ポスト:
Reka Edge: Frontier-Level Edge Intelligence for Physical AI, Reka, 2026.03
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #VisionLanguageModel Issue Date: 2026-03-14 Comment
元ポスト:
NVIDIA Nemotron 3 Super, NVIDIA, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #SSM (StateSpaceModel) #OpenSource #MoE(Mixture-of-Experts) #read-later #Selected Papers/Blogs #KeyPoint Notes #Reference Collection #Hybrid #LowPrecision #LinearAttention Issue Date: 2026-03-12 Comment
元ポスト:
解説:
artificial analysisによる評価:
Swallow LVM Leaderboardに性能が掲載:
解説:
アーキテクチャ:
- NVFP4で学習して gpt-ossより2.2倍高速だが性能も向上
- 88 Layer: 40 Latent MoE / 40 Mamba-2 / 8 GQA Attention
- GQA Attentiom Layerは非常に少なく、ほとんどがMamba-2 (linear attention)となっている
- Latent MoEは入力をそのまま変換するshared expertsと、入力を1/4のlatent vectorに変換した潜在空間上で処理をするLatext expertsの組み合わせによって出力を得る。
- 具体的には、RouterによってTop-22のexpertsを選択し、inputを1/4のlatent vectorに圧縮した上でExpertsに入力。Expertsの出力を加算して4倍のvectorに変換し次元を戻して、別ルートでshared expertsに元の入力次元から変換されたベクトルと組み合わせて出力するようなアーキテクチャ
Latent MoE解説:
要はMoEに必要なmatrixが、latent vectorを扱うことで小さくなるのでMoEのWeightのメモリロードのボトルネックが緩和されるだけでなく、
各MoE Laverは異なるGPUやマシンに分散されて配置されるため計算のためにはベクトルのバッチを通信しなければならないがそのコストが削減されスループットの向上につながるので嬉しい、ということだと思われる。
ポイント解説:
technical reportが出た:
- [Paper Note] Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning, NVIDIA+, arXiv'26, 2026.04
Moondream 3 Preview: Frontier-level reasoning at a blazing speed, Moondream, 2025.09
Paper/Blog Link My Issue
#Article #ComputerVision #EfficiencyImprovement #NLP #FoundationModel #Reasoning #SmallModel #Selected Papers/Blogs #VisionLanguageModel #KeyPoint Notes Issue Date: 2026-03-12 Comment
HF: https://huggingface.co/moondream/moondream3-preview
9B-A2Bの小規模なVLMで、
- visual reasoning: 小規模だが実タスクに適用可能なvisual reasoning性能
- trainable: Visual系のタスクは人間でもzero shotではできないことが多く、簡単にfinetuningできることが重要で
- fast: vision系のアプリケーションはリアルタイムのlavencyが求められることが多く
- inexpensive: 安くスケーラブルでなければならない
をテーマにしたモデルのようである。
object detection, pointing, 構造化された出力(犬の群の個々の犬の毛と首輪の色動画)、OCRなどの様々なタスクが実行可能で、GPT5, Gemini 2.5 Flash, Claude 4 Sonnetをこの規模感のモデルで、objec' detection, counting, document understanding, hallucinationに関するベンチマークで上回る。
前身のモデルであるmoondream2は、5Mダウンロードを達成したようだ
vikhyatk/moondream2
Open-Sourcing Sarvam 30B and 105B, sarvam, 2026.03
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning Issue Date: 2026-03-10 Comment
元ポスト:
Chinese Open Source: A Definitive History, Kevin Xu, 2026.03
Paper/Blog Link My Issue
#Article #Survey #NLP #LanguageModel #Blog #read-later Issue Date: 2026-03-07
Yuan3.0-Ultra, YuanLabAI, 2026.03
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #MoE(Mixture-of-Experts) #VisionLanguageModel #UMM #One-Line Notes #Initial Impression Notes Issue Date: 2026-03-07 Comment
元ポスト:
MoEのwarmupが終わり安定してきたタイミングでルーティングがされにくいExpertを枝刈りし、残ったexpertに対してバランスよくルーティングがされるようなrearrangeをするアルゴリズム Layer-Adaptive Expert Pruning (LAEP)によって、パラメータサイズを1515Bから1010Bまで削減し、49%程度事前学習の効率を改善したとのこと。
RAG, multimodal document understanding, tabular data analysis, content summarizationにおいて、非常に高い性能を獲得している。tool useに関してはGPT-5.2(effort不明)以外には負けているので、優秀ではあるが特に秀でているというわけではないよつに見える(BFCVv3)。
しかし他のベンチマークでこれらフロンティアモデル群をここまでPass@1やAccで抜くのは、驚きではあるが、実際にどのような評価をしているのかはテクニカルレポートを見た方が良いと思われる。
Introducing Olmo Hybrid: Combining transformers and linear RNNs for superior scaling, Ai2, 2026.03
Paper/Blog Link My Issue
#Article #Pretraining #NLP #LanguageModel #Attention #mid-training #read-later #Selected Papers/Blogs #One-Line Notes #RecurrentModels #Hybrid #LinearAttention Issue Date: 2026-03-06 Comment
元ポスト:
x1のFull Attention + x3のGated DeltaNetによるハイブリッドアーキテクチャで、75%のattentionをlinear attention (recurrent module)に置換。x3のSliding Window Attentionを用いているOlmo3と比較した結果
- 事前学習におけるデータ効率がより高く(約2倍)
- mid-training後の評価では、数学、コード、STEM, non-STEM, QA、long-contextなどの主要なドメインにおいてOlmo3と同と床それ以上の性能を達成。特に、long-contextにおけるベンチマでは大幅な性能向上(Recurrentなアーキテクチャの恩恵)
関連:
- [Paper Note] Gated Delta Networks: Improving Mamba2 with Delta Rule, Songlin Yang+, ICLR'25, 2024.12
元ポスト:
関連:
所見:
Qwen 3.5 small series, Qwen Team, 2026.02
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #SmallModel #Selected Papers/Blogs #Initial Impression Notes Issue Date: 2026-03-02 Comment
なんとSLMもリリース
元ポスト:
10 open-weight LLM releases in January and February 2026, Sebaschan Raschka, 2026.02
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Post #read-later #Selected Papers/Blogs Issue Date: 2026-02-28 Comment
- Trinity Large, Arcee, 2026.01
- [Paper Note] Kimi K2.5: Visual Agentic Intelligence, Kimi Team+, arXiv'26, 2026.02
- [Paper Note] Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters, Ailin Huang+, arXiv'26, 2026.02
- Qwen3-Coder-Next: Pushing Small Hybrid Models on Agentic Coding, QwenTeam, 2026.02
- [Paper Note] GLM-5: from Vibe Coding to Agentic Engineering, GLM-5 Team+, arXiv'26, 2026.02
- MiniMax M2.5: SOTA in Coding and Agent, designed for Agent Universe, MiniMax, 2026.02
- [Paper Note] Nanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Acts, Chen Yang+, arXiv'26, 2026.02
- Qwen3.5: Towards Native Multimodal Agents, Qwen Team, 2026.02
- Ling-2.5-1T, inclusionAI, 2026.02
- Ring-1T-2.5-FP8, inclusionAI, 2026.02
- Cohere Labs Launches Tiny Aya, Making Multilingual AI Accessible, COHERE LABS TEAM, 2026.02
元ポストには書かれていないがLLMというくくりで言うと以下もある:
- New ARENA material: 8 exercise sets on alignment science & interpretability, CallumMcDougall, 2026.02
- LFM2-24B-A2B: Scaling Up the LFM2 Architecture, LiquidAI, 2026.02
- Qwen3 Swallow, Swallow LLM, 2026.02
- Japanese
- GPT-OSS Swallow, Swallow LLM, 2026.02
- Japanese
- GLM-4.7-Flash, Z.ai, 2026.01
- LongCat-Flash-Thinking-2601, Meituan, 2026.01
- Introducing LFM2.5: The Next Generation of On-Device AI, LiquidAI, 2026.01
Omniモデルを含めると以下:
- Ming-omni-tts-0.5B, inclusionAI, 2026.02
- [Paper Note] Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability, Aaditya Vikram Prasad+, arXiv'26, 2026.02
- MiniCPM-o-4_5, OpenBMB, 2026.02
World Modelsを含めると以下?:
- [Paper Note] Causal-JEPA: Learning World Models through Object-Level Latent Interventions, Heejeong Nam+, arXiv'26, 2026.02
- [Paper Note] Code2World: A GUI World Model via Renderable Code Generation, Yuhao Zheng+, arXiv'26, 2026.02
- [Paper Note] DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos, Shenyuan Gao+, arXiv'26, 2026.02
- [Paper Note] World Action Models are Zero-shot Policies, Seonghyeon Ye+, arXiv'26, 2026.02
- [Paper Note] Advancing Open-source World Models, Robbyant Team+, arXiv'26, 2026.01
- Project Genie: Experimenting with infinite, interactive worlds, Google Deepmind, 2026.01
- Waypoint-1: Real-time Interactive Video Diffusion from Overworld, Overworld, 2026.01
確実に見落としがあるけど。
Qwen3.5 Medium Model Series, Qwen Team, 2026.02
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MultiLingual #MoE(Mixture-of-Experts) #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-02-28 Comment
元ポスト:
いずれのモデルもベンチマーク上はGPT-5 miniと同等以上の性能に見える。
また、Qwen3.5-35B-A3BはQwen3-235B-A22B-2507やQwen3-VL235B-A22Bを上回っており、アーキテクチャ、データの品質、RLによって実現されているとのこと。
27BモデルのHLEのスコアが非常に高いと話題:
FP8版もリリース:
日本語の医師国家試験(2026)において35B-A3Bが非常に高いスコアを記録:
Artificial Analysisによるベンチマーキング:
LFM2-24B-A2B: Scaling Up the LFM2 Architecture, LiquidAI, 2026.02
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #SmallModel #MoE(Mixture-of-Experts) #Initial Impression Notes #EdgeDevices Issue Date: 2026-02-27 Comment
元ポスト:
edge deviceにデプロイできる規模でLFM2をスケールさせた模様
Detecting and preventing distillation attacks, Anthropic, 2026.02
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #Proprietary #Selected Papers/Blogs #One-Line Notes Issue Date: 2026-02-24 Comment
元ポスト:
DeepSeek, Moonshot AI, MiniMax がDistillationを用いてClaude出力からモデルを改善するためのattackを特定したというAnthropicからのアナウンス
所見:
- [Paper Note] Extracting books from production language models, Ahmed Ahmed+, arXiv'26, 2026.01
で提案されている手法を用いてClaude Sonnetからハリーポッターと賢者の石の95.8%を抽出できた、との報告もある。
GPT-OSS Swallow, Swallow LLM, 2026.02
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning #Japanese Issue Date: 2026-02-21 Comment
元ポスト:
第120回医師国家試験(2026)を解かせてみた結果:
Qwen3 Swallow, Swallow LLM, 2026.02
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning #Japanese #MoE(Mixture-of-Experts) Issue Date: 2026-02-21 Comment
元ポスト:
Ming-omni-tts-0.5B, inclusionAI, 2026.02
Paper/Blog Link My Issue
#Article #Transformer #SpeechProcessing #DiffusionModel #Speech #read-later #TTS #UMM #Omni #One-Line Notes #AdversarialTraining #Music Issue Date: 2026-02-18 Comment
元ポスト:
TTSだけでなく、環境音や音楽の生成も可能な音声生成モデル。発話速度、ピッチ、音量、感情、訛りなどを正確にコントロール可能で、100+以上のビルトインのvoiceや、zeroshotでのvoice designが可能とのこと。また、speechだけでなく環境音や音楽の生成もできる産業界では初めてのモデルとのこと。また、3.1Hzごとのフレームレートでパッチ化されて入力され(これはこれまでと比べるとかなり低いフレームレートらしい)るため高速に処理が走り、テキスト入力として数式などのフォーマットも入力可能とのこと。
テクニカルレポートのリンクがまだ生きておらず詳細は不明。
Cohere Labs Launches Tiny Aya, Making Multilingual AI Accessible, COHERE LABS TEAM, 2026.02
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #SmallModel #MultiLingual #Selected Papers/Blogs #LowResource #KeyPoint Notes #Reference Collection Issue Date: 2026-02-18 Comment
元ポスト:
公式ポスト:
アーキテクチャ解説:
70程度の言語の性能をバランス良くサポートする3.35BのLLMで、Baseモデルと、マルチリンガルの性能は保ちつつも特定のregionに特化したinstruction tuningを実施したvariantを公開。また、multilingualでのベンチマークも公開。同程度の規模間のモデルについて、qwen3-4Bとの比較がわかりやすく、Europe, south asiaは同等、Asia-pacificはQwenよりも劣り、west asia, africa regionのようなこれまでlow resourceだと思われたregionではほか同規模のモデルと比較して突出した性能を誇るモデルに見える。CC上でのページ数と、言語モデルごとの性能を比較したグラフもあり、CCでのデータが少ない言語はこれまでのモデルは性能が低かったが、Tiny Ayaは非常に高い性能を達成している(このグラフで言うと日本語はかなりinformation richな言語にカテゴライズされているように見える)。
Qwen3.5: Towards Native Multimodal Agents, Qwen Team, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #ReinforcementLearning #MultiModal #MultiLingual #MoE(Mixture-of-Experts) #read-later #Selected Papers/Blogs #VisionLanguageModel #UMM #KeyPoint Notes #Scalability #Environment Issue Date: 2026-02-17 Comment
元ポスト:
最新のQwenがリリース・・・!!
- Vision+TextのUMMを採用。
- real-world agentsのために訓練
- hybrid linear attention + sparse MoE + 環境スケーリングに基づくlarge scale RLを実施
- decodingのスループットがQwen3-Maxと比較して8.6--19.0倍
- 201の言語と方言をサポート
- 397B-A17B
- Gated DeltaNet
- Gated Attention
- context length: 262k
- Multi token prediction
- 言語系タスクではGPT5.2と比較して少し劣る程度、agenticなベンチマークでは大きく上回るものも存在(ただし、Claude 4.5 Opusには届いていないベンチマークが多いように見える)
- Vision系タスクでは全体的にGPT5.2, Opus 4.5よりも優秀に見え、Gemini 3 Proと同等か少し劣る程度に見える。
世はlinear attention時代
所見:
INT4モデル:
dots.ocr-1.5, rednote-hilab, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #StructuredData #SmallModel #DocParser #OCR Issue Date: 2026-02-16 Comment
元ポスト:
Ling-2.5-1T, inclusionAI, 2026.02
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning #MoE(Mixture-of-Experts) Issue Date: 2026-02-16 Comment
Ringに続いてLingもリリース
関連:
- Ring-1T-2.5-FP8, inclusionAI, 2026.02
元ポスト:
MiniMax M2.5: SOTA in Coding and Agent, designed for Agent Universe, MiniMax, 2026.02
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Blog #Coding #SoftwareEngineering #Selected Papers/Blogs Issue Date: 2026-02-13 Comment
元ポスト:
OsenHands IndexでClaude Sonnet 4.5超えの初めてのOpenWeightモデル:
コストパフォーマンスにおいては、低コストなモデル群の中では抜きん出た性能
まだHF上にWeightは公開されていないようだが後ほど公開されると思われる。
所見:
weightが公開:
https://huggingface.co/MiniMaxAI/MiniMax-M2.5
元ポスト:
UnslothがGGUF版を公開:
Ring-1T-2.5-FP8, inclusionAI, 2026.02
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #AIAgents #Attention #Reasoning #LongContext #LongHorizon #LinearAttention Issue Date: 2026-02-12 Comment
元ポスト:
関連:
- Ring-1T, inclusionAI, 2025.10
MLA + lightning linear attentionのハイブリッド
- MHA vs MQA vs GQA vs MLA, Zain ul Abideen, 2024.07
- [Paper Note] Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention, Zhen Qin+, ICML'24, 2024.05
Ming-flash-omni-2.0, inclusionAI, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #Transformer #MultiModal #SpeechProcessing #DiffusionModel #Speech #MoE(Mixture-of-Experts) #2D (Image) #Omni #text Issue Date: 2026-02-12 Comment
元ポスト:
関連:
- Ming-flash-omni-Preview, inclusionAI, 2025.10
- [Paper Note] Ming-Omni: A Unified Multimodal Model for Perception and Generation, Inclusion AI+, arXiv'25, 2025.06
公式ポスト:
GLM-5: From Vibe Coding to Agentic Engineering, Z.ai, 2026.02
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #MoE(Mixture-of-Experts) #Selected Papers/Blogs #KeyPoint Notes #Reference Collection #LongHorizon #SparseAttention Issue Date: 2026-02-12 Comment
関連:
- GLM-4.7: Advancing the Coding Capability, Z.ai, 2025.12
GLMシリーズの最新モデルGLM-5がリリースされた
元ポスト:
- DeepSeek Sparse Attentionを採用:
- DeepSeek-V3.2-Exp: Boosting Long-Context Efficiency with DeepSeek Sparse Attention, DeepSeek-AI, 2025.09
- [Paper Note] DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models, DeepSeek-AI+, arXiv'25, 2025.12
- 事前学習データを23Tから28.5Tトークンへ
- パラメータ数は4.5の355B-A32から744B-A40Bへ
- RLのインフラとして4.5から引き続きSlimeを採用
- slime, THUDM & Zhihu, 2025.09
- long-horizonなタスクに秀でており、reasoning, coding, agenticタスクにおける各種ベンチマークでOpus 4.5, GPT-5.2, Gemini 3 Proと同等程度の性能
FP8版も公開されている模様(Hopper以後のアーキテクチャでないとサポートされていない点に注意
所見:
元ポスト:
unslothがGGUF版をすでにリリースしている模様。早い:
https://unsloth.ai/docs/models/glm-5
アーキテクチャ解説:
アーキテクチャ解説:
所見:
Voxtral transcribes at the speed of sound, Mistral AI, 2026.02
Paper/Blog Link My Issue
#Article #SpeechProcessing #Blog #MultiLingual #Proprietary #AutomaticSpeechRecognition(ASR) #Realtime #Transcript Issue Date: 2026-02-05 Comment
元ポスト:
Voxtral Mini Transcribe V2はproprietaryモデルでAPI利用のみ、Vostraal RealtimeはOpenWeightで公開
mistralai/Voxtral-Mini-4B-Realtime-2602:
https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602
Vostral Mini Transcrive V2に対するVoxtral Realtimeの性能の比較。Voxtral Realtimeは遅延を調整可能なようで、遅延が大きければ大きいほど高い性能が出るが、リアルタイムに近づけば近づくほど性能はその分劣化する。
Intern-S1-Pro, internlm, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #Reasoning #PositionalEncoding #MoE(Mixture-of-Experts) #VisionLanguageModel #Science Issue Date: 2026-02-05 Comment
元ポスト:
ポイント解説:
関連:
- [Paper Note] Intern-S1: A Scientific Multimodal Foundation Model, Lei Bai+, arXiv'25, 2025.08
Fourier Position Encoding (FoPE) + upgraded time-series modeling
MiniCPM-o-4_5, OpenBMB, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #LanguageModel #SpeechProcessing #DiffusionModel #AutomaticSpeechRecognition(ASR) #VisionLanguageModel #TTS #Omni #AudioLanguageModel Issue Date: 2026-02-05 Comment
元ポスト:
New Holo2 model takes the lead in UI Localization, H Company, 2026.02
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #AIAgents #Blog #ComputerUse #Selected Papers/Blogs #VisionLanguageModel #Grounding #GUI Issue Date: 2026-02-05 Comment
HF: https://huggingface.co/Hcompany/Holo2-235B-A22B
元ポスト:
関連:
- Holo1.5 - Open Foundation Models for Computer Use Agents, H Company, 2025.09
Latest open artifacts (#18): Arcee's 400B MoE, LiquidAI's underrated 1B model, new Kimi, and anticipation of a busy month, Interconnects, 2026.02
Paper/Blog Link My Issue
#Article #Analysis #NLP #LanguageModel #Blog Issue Date: 2026-02-03 Comment
paid userしか全文は閲覧できない
元ポスト:
Qwen3-ASR & Qwen3-ForcedAligner is Now Open Sourced: Robust, Streaming and Multilingual, Qwen Team, 2026.01
Paper/Blog Link My Issue
#Article #SpeechProcessing #LongContext #MultiLingual #AutomaticSpeechRecognition(ASR) #AudioLanguageModel #Robustness Issue Date: 2026-01-30 Comment
HF:
https://huggingface.co/collections/Qwen/qwen3-asr
technical report:
https://github.com/QwenLM/Qwen3-ASR/blob/main/assets/Qwen3_ASR.pdf
元ポスト:
Trinity Large, Arcee, 2026.01
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #Pretraining #NLP #LanguageModel #MoE(Mixture-of-Experts) #read-later #Selected Papers/Blogs #Stability #One-Line Notes #Reference Collection #Sparse #Initial Impression Notes Issue Date: 2026-01-29 Comment
テクニカルレポート:
https://github.com/arcee-ai/trinity-large-tech-report/
HF:
https://huggingface.co/arcee-ai
GLM4.7やDeepSeekV3と比較してスループットやTTFTが二倍以上。
非常にsparseなMoE(400B-A13B, 4/256のexpertsにルーティング)であるため学習を安定させるためにDense layerを増やし、モメンタムを考慮したexpertのバランシングや、z-lossと呼ばれるlogitのスケールをコントロールするような手法を導入することで安定した学習を実現。2048 Nvidia B300 GPUsで、17Tトークンの事前学習33日で完了
元ポスト:
これほどsparseなMoEをここまで安定させて学習できるのは非常に興味深いと思われる。
インタビュー:
やると決めてチームビルディングも含めて非常に短期間(6ヶ月)で達成したとのことだが、気になる。
解説:
所見(風刺):
ポイント解説:
アーキテクチャ解説:
Waypoint-1: Real-time Interactive Video Diffusion from Overworld, Overworld, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #Controllable #NLP #Transformer #MultiModal #DiffusionModel #WorldModels #interactive #3D (Video) #One-Line Notes #RectifiedFlow #Realtime Issue Date: 2026-01-22 Comment
blog:
https://over.world/blog/the-path-to-real-time-worlds-and-why-it-matters
pj page:
https://over.world/
元ポスト:
リアルタイムにzero latencyでマウス(カメラも自由に動かせる)、キーボード、テキストでinteraction可能なworld model
GLM-4.7-Flash, Z.ai, 2026.01
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Coding #MoE(Mixture-of-Experts) #One-Line Notes Issue Date: 2026-01-20 Comment
元ポスト:
関連:
- GLM-4.7: Advancing the Coding Capability, Z.ai, 2025.12
30B-A3BのMoEモデルで、gpt-oss-20B, Qwen3-30B-A3B-Thinking-2507を、SWE Bench Verified, tau2_bench, BrowseComp(SWEタスク, tooluse, 検索)等で大幅にoutperform。AIME, GPQA, HLEなどの推論系のベンチマークも同等以上。つまり、agenticなタスクに適した能力を有することが示唆される。
ポイント解説:
FrogMini-14B-2510, Microsoft, 2026.01
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Supervised-FineTuning (SFT) #AIAgents #Coding #SoftwareEngineering #One-Line Notes Issue Date: 2026-01-16 Comment
元ポスト:
strong modelから合成されたbug fixのtrajectoryでSFTすることで小規模モデルでSWE Benchの性能改善
LongCat-Flash-Thinking-2601, Meituan, 2026.01
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #MoE(Mixture-of-Experts) #Selected Papers/Blogs Issue Date: 2026-01-15 Comment
元ポスト:
解説:
coding, agentiaなベンチでTopTierを獲得した560B-27BのMoEモデル。MIT Licence
1MコンテキストウィンドウのZigzag attentionのモデルもcoming soon...だと...!?
Zigzag attentionはおそらく以下だろうか:
- [Paper Note] Efficient Context Scaling with LongCat ZigZag Attention, Chen Zhang+, arXiv'25, 2025.12
Next generation medical image interpretation with MedGemma 1.5 and medical speech to text with MedASR, Google Research, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #SpeechProcessing #Blog #AutomaticSpeechRecognition(ASR) #VisionLanguageModel #Medical Issue Date: 2026-01-14 Comment
元ポスト:
ポイント解説:
GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image Generation, Z.ai, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #LanguageModel #MultiModal #DiffusionModel #TextToImageGeneration #Editing Issue Date: 2026-01-14 Comment
元ポスト:
NousCoder-14B: A Competitive Olympiad Programming Model, Joe Li, 2026.01
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #ReinforcementLearning #Blog #Coding #PostTraining #read-later Issue Date: 2026-01-09 Comment
元ポスト:
HF:
https://huggingface.co/NousResearch/NousCoder-14B
Apache 2.0
PipelineRLを採用している模様。興味深い。
Introducing LFM2.5: The Next Generation of On-Device AI, LiquidAI, 2026.01
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #ReinforcementLearning #Blog #SmallModel #Japanese #PostTraining #Selected Papers/Blogs #VisionLanguageModel #One-Line Notes #AudioLanguageModel Issue Date: 2026-01-09 Comment
元ポスト:
日本語に特化した言語モデルも存在し、Sarashina2.2-1b-instruct-v0.1, TinySwallow-1.5B-InstructよりもJMMLU, M-IFEval (ja), GSM8K (ja)においてより高い性能を発揮している。
LFM2.5-1.2B-Base: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-1.2B-Base)
LFM2.5-1.2B-Instruct: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct),
[LEAP](
https://leap.liquid.ai/models?model=lfm2.5-1.2b-instruct),
[Playground](
https://playground.liquid.ai/chat?model=cmk1jyp8f000204i56yy76uwh)
LFM2.5-1.2B-JP: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-1.2B-JP),
[LEAP](
https://leap.liquid.ai/models?model=lfm2.5-1.2b-jp)
LFM2.5-VL-1.6B: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-VL-1.6B),
[LEAP](
https://leap.liquid.ai/models?model=lfm2.5-vl-1.6b),
[Playground](
https://playground.liquid.ai/chat?model=cmk0wefde000204jp2knb2qr8),
[Demo](
https://huggingface.co/spaces/LiquidAI/LFM2.5-VL-1.6B-WebGPU)
LFM2.5-Audio-1.5B: [Hugging Face](
https://huggingface.co/LiquidAI/LFM2.5-Audio-1.5B),
[LEAP](
https://leap.liquid.ai/models?model=lfm2.5-audio-1.5b),
[Playground](
http://playground.liquid.ai/talk)
LiquidAIのモデルは日本語に特化したモデルが多く存在するのが特徴的に感じる。
LFM2-2.6B-Transcript, LiquidAI, 2026.01
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #RecurrentModels #Transcript Issue Date: 2026-01-09 Comment
関連:
- Introducing LFM2: The Fastest On-Device Foundation Models on the Market, LiquidAI, 2025.07
NVIDIA Cosmos Reason 2 Brings Advanced Reasoning To Physical AI, Nvidia, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #Reasoning #LongContext #SmallModel #ObjectLocalization #VisionLanguageModel #Robotics #SpatialUnderstanding #EmbodiedAI #Physics Issue Date: 2026-01-06 Comment
HF: https://huggingface.co/nvidia/Cosmos-Reason2-8B?linkId=100000401175768
元ポスト:
VAETKI, NC-AI-consortium, 2026.01
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning #MultiLingual #MoE(Mixture-of-Experts) Issue Date: 2026-01-03 Comment
元ポスト:
Solar-Open-100B, upstage, 2025.12
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning #MoE(Mixture-of-Experts) #Korean Issue Date: 2026-01-03 Comment
元ポスト:
ポイント解説:
K-EXAONE-236B-A23B, LG AI Research, 2025.12
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning #MultiLingual #MoE(Mixture-of-Experts) Issue Date: 2026-01-03 Comment
関連:
- EXAONE-Deep-32B, LG AI Research, 2025.03
Multi Token Prediction
Sliding Window Attention
256k context length
MoE
元ポスト:
A.X-K1, SK Telecom, 2026.01
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning #MoE(Mixture-of-Experts) #Korean Issue Date: 2026-01-03 Comment
元ポスト:
IQuest-Coder, IQuestLab, 2026.01
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Coding #SoftwareEngineering Issue Date: 2026-01-01 Comment
元ポスト:
LFM2-2.6B-Exp, LiquidAI, 2025.12
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #SmallModel #RecurrentModels Issue Date: 2025-12-25 Comment
元ポスト:
関連:
- Introducing LFM2: The Fastest On-Device Foundation Models on the Market, LiquidAI, 2025.07
ポイント解説:
LFM2にRLによるpost trainingを実施し、指示追従、知識、数学を伸ばしているとのこと。(ドキュメントにもこれは書かれている)
日本語もサポートされている。2.6Bモデルは、22 conv+8 attnと書かれている。
アーキテクチャは下記で、LIV Operatorは入力に応じて異なる線形変換をするオペレータだが、学習された結果convolutionするのが最適ということになったのだろうか?よくわからない。
>Architecture: Hybrid model with multiplicative gates and short convolutions: 10 double-gated short-range LIV convolution blocks and 6 grouped query attention (GQA) blocks.
GLM-4.7: Advancing the Coding Capability, Z.ai, 2025.12
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Coding #Reasoning #SoftwareEngineering #One-Line Notes #Reference Collection Issue Date: 2025-12-25 Comment
元ポスト:
HF: https://huggingface.co/zai-org/GLM-4.7
デザインアリーナでtop2:
Artificial Intelligence Indexにおいて、OpenModelの中でトップ:
GLM-4.6と比較して、コーディング/SWE, reasoning, tooluseなどの能力が大幅に向上
Interleaved Thinking, Preserved Thinking, Turn-level Thinkingの3つの特性がある。
Interleaved Thinkingは全てのレスポンスとtool callingの前にreasoningを挟むことで、IFや生成品質を向上。
Preserved Thinkingは過去のターンの全てのthinking blockのトークンを保持し、再計算もしないのでマルチターンでの一貫性が増す。
Turn-level Thinkingはターンごとにreasoningを実施するか否かをコントロールでき、latency/costを重視するか、品質を重視するかを選択できる、といった特徴がある模様。
モデルサイズは358B
Qwen3-TTS Steps Up: Voice Cloning and Voice Design, Qwen Team, 2025.12
Paper/Blog Link My Issue
#Article #SpeechProcessing #Blog #Proprietary #TTS Issue Date: 2025-12-25 Comment
元ポスト:
日本語のVoice Cloneもサポートされている
MiniMax M2.1: Significantly Enhanced Multi-Language Programming, Built for Real-World Complex Tasks, MiniMax, 2025.12
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #Coding #Reasoning #SmallModel Issue Date: 2025-12-24 Comment
元ポスト:
解説:
LongCat-Video-Avatar, meituan-longcat, 2025.12
Paper/Blog Link My Issue
#Article #ComputerVision #Transformer #DiffusionModel #VariationalAutoEncoder #VideoGeneration/Understandings #3D (Scene) #One-Line Notes #Audio-Text-to-Video #Audio-Text-Image-to-Video #Video Continuation Issue Date: 2025-12-17 Comment
元ポスト:
アーキテクチャはDiTベースのDiffusion Modelで、3D Variational AutoencoderによってEncode/Decodeされ、3D RoPEによって位置情報が埋め込まれる。DiT Blockでは、テキストとaudio用のcross attentionが用いられてこれらのモーダルに関する情報が組み込まれる。audioはWav2Vecでエンコードされ、テキストはUMT5[^1]によってエンコードされる。
[^1]: multilingualなT5で100言語以上がサポートされている模様
bu-30b-a3b-preview: Meet BU-30B-A3B-Preview — bringing SoTA Browser Use capabilities in a small model that can be hosted on a single GPU., browser-use, 2025.12
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #ComputerUse #VisionLanguageModel Issue Date: 2025-12-17 Comment
元ポスト:
Introducing MiMo-V2-Flash, Xiaomi, 2025.12
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MoE(Mixture-of-Experts) #AttentionSinks #PostTraining #Selected Papers/Blogs #Reference Collection Issue Date: 2025-12-17 Comment
technical report:
https://github.com/XiaomiMiMo/MiMo-V2-Flash/blob/main/paper.pdf
HF:
https://huggingface.co/XiaomiMiMo/MiMo-V2-Flash
元ポスト:
関連:
ポイント解説:
attention sink(というより恐らくsink token)により性能が向上している:
言及されているpost trainingが有用らしい:
所見:
省パラメータでtop-tierのモデルに肉薄する方法のヒントがあるかもしれない。
解説:
chatterbox-turbo, ResembleAI, 2025.12
Paper/Blog Link My Issue
#Article #SpeechProcessing #TTS #One-Line Notes #Realtime Issue Date: 2025-12-17 Comment
元ポスト:
realtime(最初の発話まで<150ms)のlatencyが実現されたOpenWeightなTTSで、multilingualモデルは日本語にも対応している模様。テクニカルレポートがないのでよくわからないが、githubがあるのでソースコードを見ればアーキテクチャがわかりそうではある。たとえばVoiceEncoderには(おそらく速度を重視するために)LSTMが利用されていた。
github:
https://github.com/resemble-ai/chatterbox
Molmo 2: State-of-the-art video understanding, pointing, and tracking, Ai2, 2025.12
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #SmallModel #OpenSource #Selected Papers/Blogs #VideoGeneration/Understandings #VisionLanguageModel #2D (Image) #3D (Video) #KeyPoint Notes Issue Date: 2025-12-17 Comment
テクニカルレポート:
https://www.datocms-assets.com/64837/1765901660-molmo_v2_2026-techreport-3.pdf
HF:
https://huggingface.co/collections/allenai/molmo2
Qwen3とOlmoをベースにしたvariantsが存在し、Olmoの方はバックボーンのLLMも含めて全てがオープンになっている。MetaのPerceptionLMと比較して1/8の動画データ量で高い性能を達成できており、データのcurationの品質と、grounding basedな目的関数の工夫によって実現されているとのこと。
proprietaryなモデル群と比較すると、trackingは圧勝、そのほかはGPT5-miniと同様なものが多い。モデルによってタスクの優劣が結構分かれており、Video関連タスクをタスクをまたいで汎化させることにはclosedでも苦戦しているように見える。
オープンモデルとの比較で言うと圧勝で、LongVideoのQAに関してだけは、Eagle2.5-8Bと呼ばれるモデルが勝っている。
あとは全体を通じてLLMのバックボーンがQwen3の場合の性能が良いことが興味深い。バックボーンに採用するLLMに応じて性能が結構変わる。これはアーキテクチャがそもそもConnectorを利用するタイプのもので、Unifiedなアーキテクチャではないことが要因としては考えられる。
元ポスト:
demo:
コードベースが公開:
https://github.com/allenai/molmo2
Olmo 3.1, Ai2, 2025.12
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning #OpenSource #Selected Papers/Blogs Issue Date: 2025-12-13 Comment
元ポスト:
Instruction Followingのベンチマークスコアが、他モデルと比較して非常に高いように見える。
nomos-1, NousResearch, 2025.12
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning #Mathematics #One-Line Notes Issue Date: 2025-12-11 Comment
元ポスト:
30Bの強力な数学モデルで、(同じハーネスでテストした結果)Qwen3-30ba3b-Thinking-2507を大幅に上回る性能を持つとのこと。
GLM-ASR-Nano-2512, Zhipu AI, 2025.12
Paper/Blog Link My Issue
#Article #SpeechProcessing #SmallModel #AutomaticSpeechRecognition(ASR) Issue Date: 2025-12-10 Comment
元ポスト:
AutoGLM-Phone-9B, Zhipu AI, 2025.12
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #SmallModel #VisionLanguageModel Issue Date: 2025-12-10 Comment
元ポスト:
GLM-4.6V, Zhipu AI, 2025.12
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #VisionLanguageModel Issue Date: 2025-12-10 Comment
元ポスト:
Devstral2 Mistral Vibe CLI State-of-the-art, open-source agentic coding models and CLI agent., Mistral AI, 2025.12
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Coding #SoftwareEngineering Issue Date: 2025-12-10 Comment
SWE Bench VerifiedでOpenweightモデルの中ではSoTAと同等程度を達成。123B, 24Bの2種類がリリース。DeepSeekV3.2, Kimi K2よりも大幅に小さいパラメータで同等以上の性能。独自の人手評価(win, tie, loseのアリーナ形式)によるとSonnet 4.5には負けるがDeepSeekV3.2とは同等以上の割合で好まれた。
元ポスト:
OpenThinker-Agent-v1, open-thoughts, 2025.12
Paper/Blog Link My Issue
#Article #NLP #Dataset #LanguageModel #AIAgents #Evaluation #SmallModel #OpenSource #Selected Papers/Blogs #KeyPoint Notes Issue Date: 2025-12-07 Comment
元ポスト:
-
-
agenticなSLM(8Bモデル)で、モデル、データ(SFT, RL)、学習用のコードなど全て公開。同等規模のモデルQwen3-{8,32B}よりもSWE Bench Verified, Terminal Benchなどで上回る(ただし、Qwen3はgenericなモデルであり、コーディング特化のQwen3-coder-30Bには及ばない。しかしモデルサイズはこちらの方が大きいので何とも言えない。おそらく同等規模のコーディング特化Qwen3が存在しない)。また、SLMのコーディングエージェントの進化をより精緻に捉えるためのベンチマーク OpenThoughts-TB-Devも公開している。こちらでもQwen3-{8, 32B}に対しても高い性能を記録。
[Paper Note] Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail, Pavone+, Nvidia, 2025.10
Paper/Blog Link My Issue
#Article #Dataset #ReinforcementLearning #Reasoning #SmallModel #Robotics #VisionLanguageActionModel #Realtime #AutonomousDriving Issue Date: 2025-12-06 GPT Summary- AR1は因果連鎖推論と軌道計画を統合した視覚–言語–行動モデルであり、自律運転の意思決定を強化します。主な革新は、因果連鎖データセットの構築、モジュラーVLAアーキテクチャの導入、強化学習を用いた多段階トレーニング戦略です。評価結果では、AR1は計画精度を最大12%向上させ、推論の質を45%改善しました。リアルタイムパフォーマンスも確認され、レベル4の自律運転に向けた実用的な道筋を示しています。 Comment
HF: https://huggingface.co/nvidia/Alpamayo-R1-10B
元ポスト:
Improved accuracy in Smart Turn v3.1, Daily, 2025.12
Paper/Blog Link My Issue
#Article #NeuralNetwork #Transformer #AIAgents #SpeechProcessing #Blog #MultiLingual #OpenSource #One-Line Notes #VAD Issue Date: 2025-12-04 Comment
dataset:
https://huggingface.co/pipecat-ai
code:
https://github.com/pipecat-ai/smart-turn
model:
https://huggingface.co/pipecat-ai/smart-turn-v3
オープンソースのVoice Activity Detection (VAD)モデル。本ブログのv3.1では、TTSデータだけでなく英語とスペイン語の人間によるaudio sampleも追加し学習し性能向上。23言語をサポートし、Accuracyは90%以上を達成。数msでのリアルタイムなlatencyを達成できる。
バックボーンはWhisper Tiny encoderで、headとしてshallow linear classifiesを利用しているとのこと。
Nemotron-Content-Safety-Reasoning-4B, Nvidia, 2025.11
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning #Conversation #SmallModel #Safety #Safeguard Issue Date: 2025-12-03 Comment
元ポスト:
Building Safer AI Browsers with BrowseSafe, Perplenity Team, 2025.12
Paper/Blog Link My Issue
#Article #NLP #Dataset #LanguageModel #Prompting #Evaluation #Blog #Safety #Safeguard Issue Date: 2025-12-03 Comment
元ポスト:
prompt injectionをリアルタイムに検知するモデルとそのベンチマークとのこと
dataset:
https://huggingface.co/datasets/perplexity-ai/browsesafe-bench
model:
https://huggingface.co/perplexity-ai/browsesafe
Introducing Mistral 3 The next generation of open multimodal and multilingual AI, Mistral AI, 2025.12
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #Blog #MultiLingual #VisionLanguageModel #One-Line Notes Issue Date: 2025-12-03 Comment
元ポスト:
マルチモーダルなベンチマークがほとんどないように見えるMM-MT-Benchというもののみ?
[Paper Notes] Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem, Longpre+, 2025.11
Paper/Blog Link My Issue
#Article #ComputerVision #Analysis #NLP #LanguageModel #VisionLanguageModel Issue Date: 2025-11-30 Comment
元ポスト:
MITとHuggingFaceの調査によると、open weightモデルのDLにおいて、米国のAI産業における中国のモデルDL数が米国のモデルを初めて抜いた模様。
ダッシュボード: https://huggingface.co/spaces/economies-open-ai/open-model-evolution
オープンウェイトモデル( gpt-oss )の日本語精度は? – AWS パートナー アクロクエストによる徹底検証, Yamamoto+, 2025.11
Paper/Blog Link My Issue
#Article #Analysis #NLP #Evaluation #Japanese Issue Date: 2025-11-29 Comment
元ポスト:
[Paper Note] Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer, Alibaba, 2025.11
Paper/Blog Link My Issue
#Article Issue Date: 2025-11-27 Comment
HF: https://huggingface.co/Tongyi-MAI/Z-Image-Turbo
元ポスト:
ポイント解説:
公式ポスト:
Hunyuan Video 1.5 Technical Report, Tencent, 2025.11
Paper/Blog Link My Issue
#Article #ComputerVision #Transformer #DiffusionModel #VideoGeneration/Understandings Issue Date: 2025-11-21 Comment
pj page:
https://hunyuan.tencent.com/video/zh?tabIndex=0
HF:
https://huggingface.co/tencent/HunyuanVideo-1.5
元ポスト:
NVIDIA-Nemotron-Parse-v1.1, NVIDIA, 2025.11
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #TabularData #read-later #DocParser #VisionLanguageModel #OCR Issue Date: 2025-11-20 Comment
元ポスト:
olmocr2と比較して性能はどうだろうか、特に日本語
- olmOCR 2: Unit test rewards for document OCR, Ai2, 2025.10
Introducing zerank-2: The Most Accurate Multilingual Instruction-Following Reranker, ZeroEntropy, 2025.11
Paper/Blog Link My Issue
#Article #RecommenderSystems #Embeddings #InformationRetrieval #NLP #Blog #Reranking Issue Date: 2025-11-20 Comment
HF: https://huggingface.co/zeroentropy/zerank-2
SoTA reranker
Holo2: Cost-Efficient Models for Cross-Platform Computer-Use Agents, H Company, 2025.11
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #AIAgents #Blog #ComputerUse #Selected Papers/Blogs #VisionLanguageModel #Grounding #GUI Issue Date: 2025-11-14 Comment
HF: https://huggingface.co/collections/Hcompany/holo2
元ポスト:
関連:
- Holo1.5 - Open Foundation Models for Computer Use Agents, H Company, 2025.09
Omnilingual ASR: Advancing Automatic Speech Recognition for 1,600+ Languages, Meta, 2025.11
Paper/Blog Link My Issue
#Article #Transformer #SpeechProcessing #MultiLingual #AutomaticSpeechRecognition(ASR) #Selected Papers/Blogs #AudioLanguageModel Issue Date: 2025-11-12 Comment
Introducing Kimi K2 Thinking, MoonshotAI, 2025.11
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #Reasoning #Selected Papers/Blogs #One-Line Notes #Reference Collection Issue Date: 2025-11-07 Comment
HF: https://huggingface.co/moonshotai
元ポスト:
coding系ベンチマークでは少しGPT5,Claude Sonnet-4.5に劣るようだが、HLE, BrowseCompなどではoutperform
tooluseのベンチマークであるtau^2 Bench TelecomではSoTA
モデルの図解:
INT4-QATに関する解説:
INT4-QATの解説:
Kimi K2 DeepResearch:
METRによる50% timehorizonの推定は54分:
ただしサードパーティのinference providerによってこれは実施されており、(providerによって性能が大きく変化することがあるため)信頼性は低い可能性があるとのこと。
METRでの評価でClaude 3.7 Sonnetと同等のスコア:
openweightモデルがproprietaryモデルに追いつくのはsoftwere engineeringタスク(agenticなlong horizon+reasoningタスク)9ヶ月程度を要しているとのこと
OlmoEarth-v1-Large, Ai2, 2025.11
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #FoundationModel #2D (Image) Issue Date: 2025-11-06 Comment
元ポスト:
衛星画像で学習されたモデルらしい
Open-weight models lag state-of-the-art by around 3 months on average, EPOCH AI, 2025.10
Paper/Blog Link My Issue
#Article #Analysis #NLP #LanguageModel #Blog Issue Date: 2025-11-01 Comment
タイトルの通りな模様
元ポスト:
LongCat-Flash-Omni Technical Report, 2025.10
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #LanguageModel #SpeechProcessing #MoE(Mixture-of-Experts) #2D (Image) #UMM #3D (Video) #Omni #audio #text Issue Date: 2025-11-01 Comment
元ポスト:
HF: https://huggingface.co/meituan-longcat/LongCat-Flash-Omni
text, image/video, audioをinputし、audioを生成するomniモデル
gpt-oss-safeguard, OpenAI, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning #Safety #One-Line Notes #Safeguard Issue Date: 2025-10-30 Comment
元ポスト:
blog: https://openai.com/index/introducing-gpt-oss-safeguard/
ポリシーとそのポリシーに従うべきコンテンツが与えられたときに、コンテンツを分類するタスクを実施できる汎用的なreasoningモデル。つまり、任意のポリシーを与えて追加の学習なしでpromptingによってコンテンツがポリシーのもとでsafe/unsafeなのかを分類できる。
gpt-ossをreinforcbment finetuningしているとのこと。
Marin 32B Retrospective, marin-community, 2025.10
Paper/Blog Link My Issue
#Article #Pretraining #NLP #LanguageModel #Blog #OpenSource #Selected Papers/Blogs Issue Date: 2025-10-30 Comment
元ポスト:
lossのスケーリング則に基づいた今後の見通し:
pj pageはこちら:
https://marin.community
Ming-flash-omni-Preview, inclusionAI, 2025.10
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #LanguageModel #MultiModal #SpeechProcessing #TextToImageGeneration #AutomaticSpeechRecognition(ASR) #Architecture #MoE(Mixture-of-Experts) #Selected Papers/Blogs #VideoGeneration/Understandings #Editing #TTS #Routing #UMM #Omni #Sparse #ImageSynthesis #Initial Impression Notes Issue Date: 2025-10-28 Comment
元ポスト:
過去一番多くのタグを付与した気がするが、果たして大規模、Omniモーダルかつ、UMMにしたことによる恩恵(=様々なモダリティを統一された空間上に学習させる恩恵)はどの程度あるのだろうか?
アーキテクチャを見ると、モダリティごとに(モダリティ単位でのバイアスがかかった)Routerが用意されexpertにルーティングされるような構造になっている。
OmniモーダルでUMMを大規模にスクラッチから事前学習:
- [Paper Note] ERNIE 5.0 Technical Report, Haifeng Wang+, arXiv'26, 2026.02
LLaDA 2.0, inclusionAI, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #DiffusionModel #MoE(Mixture-of-Experts) Issue Date: 2025-10-28 Comment
元ポスト:
MiniMax-M2: Intelligence, Performance & Price Analysis, Artificial Analysis, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #Selected Papers/Blogs #Reference Collection Issue Date: 2025-10-26 Comment
元ポスト:
関連:
- [Paper Note] MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning
Attention, MiniMax+, arXiv'25, 2025.06
CISPOを提案したMiniMax-M1の後続モデルと思われるMiniMax-M2-previewが中国製のモデルでArtificial Intelligenceでの評価でトップに立った模様。
所見:
モデルが公開:
https://huggingface.co/MiniMaxAI/MiniMax-M2
proprietaryモデルになるもんだと思ってた、、、これを公開するの凄すぎでは、、、
公式ポスト:
MITライセンス
vLLMでのserving方法:
https://docs.vllm.ai/projects/recipes/en/latest/MiniMax/MiniMax-M2.html
> You can use 4x H200/H20 or 4x A100/A800 GPUs to launch this model.
上記GPUにおいては--tensor-parallel-size 4で動作する模様。
SGLangでもサポートされている:
AnthropicのAPIの利用をお勧めする理由:
(以下管理人の補足を含みます)MiniMax-M2はAgenticなCoTをするモデルなので、contextの情報を正しく保持する必要がある。特に、マルチターンのやり取りをAPIを介してユーザが実行する場合、OpenAIのchatcompletionはCoTを返してくれず、マルチターンのやり取りをしても同じsessionで利用したとしても、前のターンと同じCoTが利用されないことがドキュメントに記述されている。このような使い方をサポートしているのはResponceAPIのみであるため、ResponceAPIでのみ適切なパフォーマンスが達成される。この点がconfusingなので、誤った使い方をするとMiniMaxの真価が発揮されず、しかもそれに気づけずに使い続けてしまう可能性がある。AnthropicのAPIではSonnet 4.5では全ての応答に明示的にCoTが含まれるため、その心配がない、だからAnthropicがおすすめ、みたいな話だと思われる。
アーキテクチャ解説:
解説:
LongCat-Video Techcal Report, Meituan LongCat Team, 2025.10
Paper/Blog Link My Issue
#Article #ComputerVision #Transformer #DiffusionModel #TextToImageGeneration #LongContext #VariationalAutoEncoder #VideoGeneration/Understandings Issue Date: 2025-10-26 Comment
元ポスト:
HF: https://huggingface.co/meituan-longcat/LongCat-Video
公式ポスト:
Introducing MiMo-Audio, LLM-Core Xiaomi, 2025.10
Paper/Blog Link My Issue
#Article #Pretraining #InstructionTuning #SpeechProcessing #Reasoning #SmallModel #Zero/FewShotLearning #Selected Papers/Blogs #UMM #AudioLanguageModel Issue Date: 2025-10-25 Comment
HF: https://huggingface.co/collections/XiaomiMiMo/mimo-audio
元ポスト:
text, audioを入力として受け取り、text, audioを出力するAudioLanguageModel
zerank-1, zeroentropy, 2025.07
Paper/Blog Link My Issue
#Article #RecommenderSystems #InformationRetrieval #Encoder #Reranking Issue Date: 2025-10-23 Comment
SoTAなcross-encoderに基づくreranker。おそらく英語にのみ対応。
zerank-1はcc-by-nc-4.0, smallはApache2.0ライセンス
LFM2-VL-3B: A New Efficient Vision-Language for the Edge, LiquidAI, 2025.10
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #SmallModel #MultiLingual #VisionLanguageModel Issue Date: 2025-10-22 Comment
元ポスト:
HF: https://huggingface.co/LiquidAI/LFM2-VL-3B
SigLIP2とLFM2がバックボーン
- Introducing LFM2: The Fastest On-Device Foundation Models on the Market, LiquidAI, 2025.07
dots.ocr, rednote-hilab, 2025.07
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #SmallModel #MultiLingual #DocParser #VisionLanguageModel #OCR Issue Date: 2025-10-22 Comment
100+言語のdots.ocr benchと呼ばれるものでの性能も報告されているが、日本語性能はどのくらいなのだろうか
MIT Licence
参考:VLMを使った多言語ドキュメントパーサ「dots.ocr」を試す, kun432, Zenn
https://zenn.dev/kun432/scraps/b91fce6fbeb30c
日本語もかなりいけてそう
Chandra, datalab-to, 2025.10
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #LanguageModel #MultiLingual #DocParser #OCR Issue Date: 2025-10-22 Comment
元ポスト:
SoTA.だったdots.ocrというモデルをoutperformしている模様
40+ languagesをサポート
AI PUBS OpenRAIL-M Modifiedライセンス🤔
https://huggingface.co/datalab-to/chandra/blob/main/LICENSE
dots.ocrはMIT Licence
- dots.ocr, rednote-hilab, 2025.07
LFM2-350M-PII-Extract-JP, LiquidAI, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #SmallModel #Japanese #RecurrentModels #PII Issue Date: 2025-10-14 Comment
元ポスト:
ポイント解説:
関連:
- Introducing LFM2: The Fastest On-Device Foundation Models on the Market, LiquidAI, 2025.07
Ring-1T, inclusionAI, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning Issue Date: 2025-10-14 Comment
元ポスト:
inclusionAIから続々とfrontierなモデルが出てきている。
テクニカルレポートが公開:
- [Paper Note] Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale
Thinking Model, Ling Team+, arXiv'25, 2025.10
K2 Vendor Verifier, MoonshotAI, 2025.09
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Evaluation Issue Date: 2025-10-12 Comment
Kimi K2のプロバイダー間でのツール呼び出しの性能の違いを確認できる
元ポスト:
関連:
- Kimi-K2-Instruct-0905, MoonshotAI, 2025.09
- Kimi K2: Open Agentic Intelligence, moonshotai, 2025.07
Introducing Stable Diffusion 3.5, StabilityAI, 2024.10
Paper/Blog Link My Issue
#Article #ComputerVision #Transformer #DiffusionModel #TextToImageGeneration #Blog #Selected Papers/Blogs Issue Date: 2025-10-10 Comment
SD3.5
commonvoice22_sidon, sarulab-speech, 2025.10
Paper/Blog Link My Issue
#Article #SpeechProcessing #MultiLingual #TTS Issue Date: 2025-10-09 Comment
元ポスト:
134言語サポートのTTS
colbert-muvera-femto, NeuML, 2025.10
Paper/Blog Link My Issue
#Article #Embeddings #NLP #SmallModel #Encoder Issue Date: 2025-10-09 Comment
元ポスト:
Jamba Reasoning 3B, AI21Labs, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #SmallModel #SSM (StateSpaceModel) Issue Date: 2025-10-09 Comment
元ポスト:
LFM2-8B-A1B: An Efficient On-device Mixture-of-Experts, LiquidAI, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Blog #SmallModel #MoE(Mixture-of-Experts) #RecurrentModels Issue Date: 2025-10-08 Comment
HF: https://huggingface.co/LiquidAI/LFM2-8B-A1B
元ポスト:
日本語もサポートしているとのこと
関連:
- Introducing LFM2: The Fastest On-Device Foundation Models on the Market, LiquidAI, 2025.07
エージェント機能が大幅に強化されたPLaMo 2.1 Primeの提供開始, PFN, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #AIAgents #Japanese Issue Date: 2025-10-07 Comment
マルチターンのtool callingのベンチマーク のSimple, Multiple(それぞれ単一ツール呼び出し、複数のツールの中から適切なツールを呼び出す能力)でBFCVv3でGPT-5超え。ただしGPT-5はツール呼び出しではなくユーザと対話する傾向にあるため、chatアプリケーションではこちらの方が有用な場合があるので全てのユースケースでPLaMoが上回ることを示しているわけではない、という注釈がついている。より実験的な環境であるLive MultipleではGPT-5の方がスコアが高い模様。
- BFCLv2, UC Berkeley, 2024.08
単一呼び出し、複数定義されている中から適切なツールを呼び出すことで済むようなユースケースの場合は検討の余地があると思われる。ただし細かいreasoning_effortやverbosity等のパラメータ設定が記述されていないように見えるので、その辺はどうなんだろうか。
CODA: Coding LM via Diffusion Adaption, Chen+, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #DiffusionModel #Coding #SmallModel #OpenSource Issue Date: 2025-10-05 Comment
元ポスト:
HF:
https://huggingface.co/Salesforce/CoDA-v0-Instruct
cc-by-nc-4.0
Ming-UniVision: Joint Image Understanding and Generation via a Unified Continuous Tokenizer, inclusionAI, 2025.10
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #LanguageModel #UMM Issue Date: 2025-10-03 Comment
HF: https://huggingface.co/inclusionAI/Ming-UniVision-16B-A3B
元ポスト:
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation, inclusionAI, 2025.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #SpeechProcessing #Blog #Editing Issue Date: 2025-10-03 Comment
元ポスト:
Ming-Omniの後継モデルで、スピーチに特化して書き起こし、理解、編集などができるモデル
HF: https://huggingface.co/inclusionAI/Ming-UniAudio-16B-A3B
公式ポスト:
IBM Granite 4.0: hyper-efficient, high performance hybrid models for enterprise, IBM, 2025.10
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Transformer #LongContext #SmallModel #SSM (StateSpaceModel) Issue Date: 2025-10-02 Comment
元ポスト:
Mamba2とtransformerのハイブリッドモデルで、比率は9:1とMamba2ブロックが多めらしい。Mamba2の恩恵によりlokg-context時のメモリ使用量が70パーセント削減されるとのこと。
Apriel-1.5-15b-Thinker, ServiceNow-AI, 2025.09
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #Reasoning #SmallModel #VisionLanguageModel Issue Date: 2025-10-01 Comment
元ポスト:
Artificial Analysisによるベンチマーキングでは現状<20BでSoTAなReasoningモデルな模様。
MIT License
公式ポスト:
Nvidiaによるポスト:
GLM-4.6: Advanced Agentic, Reasoning and Coding Capabilies, Zhipu AI, 2025.09
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #MultiModal #MoE(Mixture-of-Experts) #read-later #VisionLanguageModel #One-Line Notes Issue Date: 2025-09-30 Comment
元ポスト:
続報:
Artificial Intelligenceによる評価:
OpenWeightモデルの中でトップレベルのベンチスコア
HFにてモデルが公開された模様。ベンチマークのスコアを見て思ったが、106BA12Bのモデルと9Bモデルのスコア差がベンチマークによっては小さいので、場合によってはSLMの方でtest time scacingを効かせた方が、時間的な制約がきつい場合は現実的には高い性能が出るのでは?
InternVL3.5-Flash, OpenGVLab, 2025.09
Paper/Blog Link My Issue
#Article #ComputerVision #Reasoning #VisionLanguageModel Issue Date: 2025-09-29 Comment
元ポスト:
Ring-1T-preview, inclusionAI, 2025.09
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Reasoning Issue Date: 2025-09-29 Comment
元ポスト:
DeepSeek-V3.2-Exp: Boosting Long-Context Efficiency with DeepSeek Sparse Attention, DeepSeek-AI, 2025.09
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Attention #Reference Collection #Sparse #SparseAttention Issue Date: 2025-09-29 Comment
元ポスト:
DeepSeek Sparse Attentionポイント解説:
解説:
DSA図解:
ポイント解説:
公式ポスト:
DSAについては
- [Paper Note] DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models, DeepSeek-AI+, arXiv'25, 2025.12
参照のこと。
