TrainingFramework (22) — 1/1
[Paper Note] VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use, Dongfu Jiang+, arXiv'25, 2025.08
Paper/Blog Link My Issue
#EfficiencyImprovement #Tools #NLP #LanguageModel #ReinforcementLearning #PostTraining #Asynchronous Issue Date: 2025-09-03 GPT Summary- VerlToolは、強化学習におけるツール統合の課題を解決するための統一的かつモジュラーなフレームワークを提供する。主な貢献は、互換性の確保、標準化されたAPIによるツール管理、非同期実行による速度向上、競争力のあるパフォーマンス評価である。これにより、マルチターンのインタラクションを形式化し、様々なタスクにおいて専門的なシステムと同等の結果を達成する。開発のオーバーヘッドを削減し、スケーラブルな基盤を提供する。コードはオープンソースで公開されている。 Comment
github: https://github.com/TIGER-AI-Lab/verl-tool
元ポスト:
[Paper Note] Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism, Mohammad Shoeybi+, arXiv'19, 2019.09
Paper/Blog Link My Issue
#Pretraining #Tools #NLP #LanguageModel #Supervised-FineTuning (SFT) #ReinforcementLearning #PostTraining #Selected Papers/Blogs #Parallelism #One-Line Notes #DistributedLearning Issue Date: 2026-09-29 GPT Summary- 大規模Transformerモデルを学習するため、ネイティブPyTorchに少数の通信操作を追加する層内モデル並列化手法を提案。パイプライン型並列化と併用可能で、512 GPU上で最大83億パラメータのモデルを学習し、15.1ペタFLOPSと76%のスケーリング効率を達成。GPT-2型・BERT型モデルで既存のSOTAを上回り、WikiText103、LAMBADA、RACEで高い性能を実現。 Comment
Megatron-LMがメモされていなかったので追加。代表的なLLM・マルチモーダルモデル学習のためのフレームワーク
事前学習、継続事前学習だけでなく
- Swallow: LLaMA-2 日本語継続事前学習モデル, Kazuki Fujii, 2023.12
以下のような事後学習フレームワークのバックエンドとしても利用される:
- Nemo-RL, Nvidia, 2025.05
- verl: Volcano Engine Reinforcement Learning for LLMs, ByteDance Seed Team, 2025.04
- slime, THUDM & Zhihu, 2025.09
- ms-swiftによるMegatron-LMベースのQwen3のファインチューニング, Aratako, 2025.05
Olmo-core: Building blocks for OLMo modeling and training, AI2, 2024.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #MoE(Mixture-of-Experts) #Initial Impression Notes #Author Thread-Post Issue Date: 2026-10-02 Comment
元ポスト:
pytorchによる大規模な分散学習を実現するためのフレームワークで、trillion級のMoEモデルが学習できるような機能を追加したv3.0.0がリリースされたとのことである。
Halo: Frontier-Lab Training for Everyone, White Circle, 2026.09
Paper/Blog Link My Issue
#Article #Multi #ComputerVision #EfficiencyImprovement #Pretraining #Tools #NLP #LanguageModel #ReinforcementLearning #AIAgents #MultiModal #Blog #mid-training #DPO #PostTraining #Parallelism #VisionLanguageModel #Asynchronous #Author Thread-Post Issue Date: 2026-09-29 Comment
元ポスト:
Miles: Enterprise-Grade Reinforcement Learning for Large-Scale Model Post-Training, RadixArk, 2026.08
Paper/Blog Link My Issue
#Article #Tools #NLP #LanguageModel #ReinforcementLearning #PEFT(Adaptor/LoRA) #PostTraining #Asynchronous #Author Thread-Post Issue Date: 2026-08-19 Comment
元ポスト:
🦋 Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning, Jian Hu and Molt Contributors, 2026.07
Paper/Blog Link My Issue
#Article #Tools #NLP #LanguageModel #ReinforcementLearning #AIAgents #PostTraining #One-Line Notes Issue Date: 2026-07-19 Comment
元ポスト:
ハイパフォーマンス、かつminimalな実装のRLライブラリで、研究用のハックがしやすいフレームワーク。verlの1/9, Slimeの1/3の行数で実現されているとのこと。
verl:
- verl: Volcano Engine Reinforcement Learning for LLMs, ByteDance Seed Team, 2025.04
Slime:
- slime, THUDM & Zhihu, 2025.09
Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel, nvidia, 2026.06
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Library #MoE(Mixture-of-Experts) #PostTraining #Selected Papers/Blogs #Finetuning #reading #One-Line Notes #EfficientEvaluation Issue Date: 2026-07-05 Comment
わずか数行を追加するだけで、MoEモデルをマルチGPU環境でFinetuningする際のスループットが約3--4倍、メモリ使用量が30%程度削減されるライブラリ。既存のQwen, Nemotron, GPT-OSS, DeepSeek V3などの一般的なMoEアーキテクチャに対して、最適化済みの実装が提供されているとのこと。
repository: https://github.com/NVIDIA-NeMo/Automodel
AReaL: A Large-Scale Asynchronous Reinforcement Learning System, inclusionAI, 2026.03
Paper/Blog Link My Issue
#Article #Tools #NLP #LanguageModel #ReinforcementLearning #Reasoning #read-later #Asynchronous Issue Date: 2026-03-05 Comment
元ポスト:
Accelerating Diffusion Models with an Open, Plug-and-Play Offering, Nvidia, 2026.01
Paper/Blog Link My Issue
#Article #ComputerVision #EfficiencyImprovement #Tools #NLP #DiffusionModel #TextToImageGeneration #Distillation #PostTraining #2D (Image) #Editing #3D (Video) #TextToVideoGeneration #ImageToTextGeneration Issue Date: 2026-01-29 Comment
元ポスト:
self forcingも実装されている
- [Paper Note] Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion, Xun Huang+, NeurIPS'25
OpenTinker Democratizing Agentic Reinforcement Learning as a Service, Zhu+, University of Illinois Urbana-Champaign, 2025.12
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #Tools #NLP #LanguageModel #ReinforcementLearning #Blog #PostTraining #KeyPoint Notes Issue Date: 2025-12-22 Comment
元ポスト:
code: https://github.com/open-tinker/OpenTinker
関連:
- verl: Volcano Engine Reinforcement Learning for LLMs, ByteDance Seed Team, 2025.04
- Tinker is a training API for {developers, builders, researchers}, THINKING MACHINES, 2025.10
Tinkerに着想を得てクライアントとサーバを分離した設計になっており、バックエンド側のGPUクラスタでサーバを一度起動するだけでクライアント側がスケジューラにジョブを送ればRLが実行される(ローカルにGPUは不要)。クライアント側はRLを実施したい環境のみをローカルで定義しコンフィグをロードしfitを呼び出すだけ。verlよりもよりも手間が省けているらしい。
リポジトリを見る限りは、verlをRLのコアエンジンとして使ってる模様。
Introducing torchforge – a PyTorch native library for scalable RL post-training and agentic development, PyTorch team at Meta, 2025.10
Paper/Blog Link My Issue
#Article #NLP #Library #ReinforcementLearning #AIAgents #Blog #Selected Papers/Blogs Issue Date: 2025-10-25 Comment
元ポスト:
slime, THUDM & Zhihu, 2025.09
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #Library #ReinforcementLearning #AIAgents #PostTraining #Selected Papers/Blogs #Asynchronous Issue Date: 2025-09-02 Comment
元ポスト:
GLM-4.5のRL学習に利用されたフレームワーク
- [Paper Note] GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models, GLM-4. 5 Team+, arXiv'25
RLinf: Reinforcement Learning Infrastructure for Agentic AI, RLinf, 2025.09
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #Library #ReinforcementLearning #PostTraining #Robotics #VisionLanguageActionModel #EmbodiedAI Issue Date: 2025-09-01 Comment
元ポスト:
rLLM, Agentica, 2025.06
Paper/Blog Link My Issue
#Article #NLP #Library #ReinforcementLearning #AIAgents #PostTraining #Initial Impression Notes Issue Date: 2025-07-04 Comment
>rLLM is an open-source framework for post-training language agents via reinforcement learning. With rLLM, you can easily build their custom agents and environments, train them with reinforcement learning, and deploy them for real-world workloads.
なるほど。
バックボーンにはverlが採用されており、シンプルかつ統一的なインタフェースでカスタムエージェントが学習できる模様?
https://rllm-project.readthedocs.io/en/latest/#key-features
元ポスト:
関連:
- verl: Volcano Engine Reinforcement Learning for LLMs, ByteDance Seed Team, 2025.04
v0.2がリリースされ、任意のagentia programの学習がサポートされた模様(マルチエージェントや複雑なワークフローに基づくものなど):
Nemo-RL, Nvidia, 2025.05
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #Library #ReinforcementLearning #PostTraining Issue Date: 2025-06-25
verl: Volcano Engine Reinforcement Learning for LLMs, ByteDance Seed Team, 2025.04
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Library #ReinforcementLearning #python #Selected Papers/Blogs #One-Line Notes #Reference Collection Issue Date: 2025-05-16 Comment
SoTAなRLアルゴリズムを数行のコードで実装可能で、Sequence Parallelismがサポートされているので長い系列を扱える。FSDP, Megatron-LM,vLLM,SGLangなどとシームレスに統合できるっぽい?
注意点(超重要):
inference backend(ブログ中ではvLLM, SGLangなどを仮定。ロールアウトに利用する)とtrainingのbackend(モデルを学習するフレームワーク, FSDPなどを仮定する)のミスマッチによってトークンの生起確率に差が生じ、ポリシーの更新がうまくいかなくなる。
- 論文では語られないLLM開発において重要なこと Swallow Projectを通して, Kazuki Fujii, NLPコロキウム, 2025.07
でも言われているように、ライブラリにはバグがあるのが普通なのね、、、。
Unsloth, unslothai, 2024.07
Paper/Blog Link My Issue
#Article #EfficiencyImprovement #NLP #LanguageModel #Library #Supervised-FineTuning (SFT) #InstructionTuning #PEFT(Adaptor/LoRA) #PostTraining #Selected Papers/Blogs #One-Line Notes Issue Date: 2024-10-08 Comment
single-GPUで、LLMのLoRA/QLoRAを高速/省メモリに実行できるライブラリ
現在でも鉄板
repeng
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #LanguageModel #Library #Alignment #TextualInversion #KeyPoint Notes Issue Date: 2024-03-21 Comment
LLMの出力のスタイルを数百個の事例だけで学習しチューニングできるライブラリ。promptで指定するのとは異なり、数値でスタイルの強さを指定することが可能らしい(元ツイート)。画像生成分野におけるTextual Inversionと同じ技術とのこと。
Textual Inversionとは、少量のサンプルを用いて、テキストエンコーダ部分に新たな「単語」を追加し、単語と対応する画像を用いてパラメータを更新することで、prompt中で「単語」を利用した場合に学習した画像のスタイルやオブジェクト(オリジナルの学習データに存在しなくても可)を生成できるようにする技術、らしい。
Huggiegface:
https://huggingface.co/docs/diffusers/training/text_inversion
(参考)GPTに質問した際のログ:
https://chat.openai.com/share/e4558c44-ce09-417f-9c77-6f3855e583fa
元ツイート:
LLaMA-Factory, hiyouga, 2023.07
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Library #Supervised-FineTuning (SFT) #PostTraining #One-Line Notes Issue Date: 2023-11-14 Comment
簡単に利用できるLLaMAのfinetuning frameworkとのこと。
元ツイート:
LLaMAベースなモデルなら色々対応している模様
trl_trlx
Paper/Blog Link My Issue
#Article #NLP #LanguageModel #Library #Alignment #Supervised-FineTuning (SFT) #ReinforcementLearning #PostTraining #One-Line Notes #Reference Collection #needs-revision Issue Date: 2023-07-23 Comment
関連:
- TRL - 強化学習によるLLMの学習のためのライブラリ, npaka, 2023.06
-
https://note.com/npaka/n/nbb974324d6e1
関連:
- trlを使って日本語LLMをSFTからRLHFまで一通り学習させてみる, AI SHIFT, 2023.07
-
https://www.ai-shift.co.jp/techblog/3583
Auto train advanced, HuggingFace, 2023.07
Paper/Blog Link My Issue
#Article #MachineLearning #Tools #LanguageModel #Supervised-FineTuning (SFT) #Blog #Repository #PEFT(Adaptor/LoRA) #One-Line Notes #needs-revision Issue Date: 2023-07-11 Comment
Hugging Face Hub上の任意のLLMに対して、localのカスタムトレーニングデータを使ってfinetuningがワンラインでできる。
peftも使える。
現在はもうメンテナンスされていないようだ。
LM Flow, OptimalScale, 2023.06
Paper/Blog Link My Issue
#Article #MachineLearning #Tools #LanguageModel #Supervised-FineTuning (SFT) #FoundationModel #One-Line Notes #needs-revision Issue Date: 2023-06-26 Comment
一般的なFoundation Modelのファインチューニングと推論を簡素化する拡張可能なツールキット。継続的なpretragning, instruction tuning, parameter efficientなファインチューニング,alignment tuning,大規模モデルの推論などさまざまな機能をサポート。
