MiniMax推文
NVIDIA SANA 团队在 MiniMax H3 上的 Sol Engine 工作实现重大突破
NVIDIA SANA 团队在 MiniMax H3 上应用 Sol Engine,通过将生成过程分为 4 步低分辨率草图和 3 步目标分辨率细化,并采用 Sol-Attn 和 TAE 解码器,将单块 GB200 上 10 秒 768p 视频的延迟从 414 秒降至 14.93 秒,提速 27.7 倍。这一突破使单节点每月可服务 37.8 万条视频,GPU 利润率超 97%,推动高质量 AI 视频从异步批处理转向近实时交互式基础设施。
译文
The real breakthrough in the NVIDIA SANA team’s Sol Engine work on MiniMax H3! By splitting generation into a 4-step low-res H3 draft and a 3-step LTX refinement pass at target resolution with Sol-Attn, they’ve crushed 10s 768p latency on a single GB200 from 414s down to 14.93s (27.7x speedup). Replacing heavy VAE decodes with TAEH3/TAEHV while holding the latents stable for refinement is a masterclass in co-designing sampling topology with hardware kernel acceleration. When inference latency collapses this dramatically, unit economics fundamentally shift: a single node can suddenly serve 378K videos a month at 97%+ GPU margins. This is how high-fidelity AI video moves from asynchronous batch rendering to near-instant, interactive infrastructure. Huge respect to the team for setting a new engineering bar for our open-weights ecosystem! 🫡🩵
The real breakthrough in the NVIDIA SANA team’s Sol Engine work on MiniMax H3! By splitting generation into a 4-step low-res H3 draft and a 3-step LTX refinement pass at target resolution with Sol-Attn, they’ve crushed 10s 768p latency on a single GB200 from 414s down to 14.93s (27.7x speedup). Replacing heavy VAE decodes with TAEH3/TAEHV while holding the latents stable for refinement is a masterclass in co-designing sampling topology with hardware kernel acceleration. When inference latency collapses this dramatically, unit economics fundamentally shift: a single node can suddenly serve 378K videos a month at 97%+ GPU margins. This is how high-fidelity AI video moves from asynchronous batch rendering to near-instant, interactive infrastructure. Huge respect to the team for setting a new engineering bar for our open-weights ecosystem! 🫡🩵