Artificial Analysis推文
NVIDIA Groq 3 LPX 推理机架实测:Gemma 4 31B 输出速度达 3431 tokens/s
Artificial Analysis 通过 NVIDIA 提供的私有演示端点,对 Groq 3 LPX 推理机架进行了基准测试。在 10k 和 100k 输入序列长度下,Gemma 4 31B 的中位输出速度约为 3400 tokens/s,创下该模型在 100k 上下文下的最高记录。NVIDIA 宣布该机架已全面投产,将于今年晚些时候投入运营。
译文
Artificial Analysis 通过 NVIDIA Groq 3 LPX 的私有演示端点测得 Gemma 4 31B 的生成速度为 3,431 tokens/s——这是我们在私有或公共端点上测得的 Gemma 4 31B 在 100k 上下文下的最高输出速度。@nvidia 今日宣布,Groq 3 LPX“交互式 AI 推理”机架已全面投产,并将于今年晚些时候投入运营。NVIDIA 为我们提供了该机架私有部署的早期访问权限,用于运行 Gemma 4 31B 以进行基准测试。我们运行了用于无服务器 API 提供商基准测试的标准 1k、10k 和 100k 输入序列长度提示。在 10k 和 100k 输入序列长度下,该端点均提供了约 3,400 tokens/s 的中位输出速度(基于 50 次顺序单并发请求测得)。在测试的 10k 至 100k 输入序列长度之间,输出速度保持稳定,表明其具备稳健的长上下文推理性能。祝贺 @nvidia 团队成功发布!
Artificial Analysis has measured 3,431 tokens/s on Gemma 4 31B via a private demonstration endpoint for NVIDIA Groq 3 LPX - the highest 100k-context output speed we have measured for Gemma 4 31B on either a private or public endpoint @nvidia announced today that the Groq 3 LPX ‘interactive AI inference’ rack is in full-scale production and will enter operation later this year. NVIDIA granted us early access to a private deployment of the rack serving Gemma 4 31B for benchmarking purposes We ran the standard 1k, 10k and 100k input sequence length prompts that we use for benchmarking serverless API providers. At both the 10k and 100k input sequence lengths, the endpoint delivered a median output speed of approx. 3,400 tokens/s, measured over 50 sequential (single-concurrency) requests. Output speed was maintained between the 10k and 100k input sequence lengths tested, indicating robust long-context inference performance Congratulations to the @nvidia team on the launch!