TextToAudio
MiniMax-H3, MiniMaxAI, 2026.08
Paper/Blog Link My Issue
#Article #ComputerVision #NLP #Transformer #SpeechProcessing #OpenWeight #VideoGeneration/Understandings #UMM #Omni #TextToVideoGeneration #Author Thread-Post Issue Date: 2026-08-09 Comment
元ポスト:
Video Arenaと呼ばれるベンチマーク (Image-to-Video, Image-to-Video) でOpenWeightモデルでSoTA、全体で2位:
FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence, Black Forest Labs, 2026.07
Paper/Blog Link My Issue
#Article #ComputerVision #MultiModal #TextToImageGeneration #Blog #Proprietary #VideoGeneration/Understandings #Editing #UMM #One-Line Notes #ImageSynthesis #TextToVideoGeneration #WorldActionModel #Author Thread-Post Issue Date: 2026-07-24 Comment
元ポスト:
モデルは将来的にオープンになるようである
Bark, Suno-AI, 2023.04
Paper/Blog Link My Issue
#Article #NLP #Library #SpeechProcessing #One-Line Notes Issue Date: 2023-05-04 Comment
テキストプロンプトで音声生成ができるモデル。MIT License
