Artificial Analysis推文
推出语音智能体竞技场:在真实场景中评估语音到语音模型
Hugging Face 发布了语音智能体竞技场,用于在真实世界场景中评估语音到语音模型。该竞技场通过人类参与者与模型进行实时对话,比较模型的对话偏好和任务成功率。初步结果显示,Gemini 3.1 Flash 在偏好上领先,而 Grok Voice 在任务成功率上最高。该平台旨在提供更贴近实际使用的评估,并计划扩展更多模型和场景。
译文
我们隆重推出全新的 Speech Agent Arena,用于在真实场景中评估语音到语音模型,以分析对话偏好与任务成功率。现有的语音到语音基准测试涵盖了推理、模拟智能体任务,以及轮次切换和打断处理等对话动态。Speech Agent Arena 则让人类在完成真实世界任务时对模型和级联系统进行比较,衡量对话偏好和工具使用成功率。这使我们能够提供更贴近实际应用的评估,揭示用户最喜欢与哪些模型对话,以及这些模型在多大程度上有效支持了用户的请求。 Speech Agent Arena 与任务成功率概览 人类参与者会在同一指定场景下对两个隐藏的语音到语音模型进行比较,场景分为 15 个智能体场景(需要调用工具的任务,例如点外卖)和 20 个非智能体场景(无需调用工具的任务,例如询问营业时间)。在与每个模型分别进行实时对话后,参与者会选择他们更偏好的模型,这些成对投票用于拟合偏好 Elo 评分。对于智能体场景,任务成功率是指在符合条件的对话中(即参与者无偏离或无法验证的情况),模型通过正确的最终工具调用完成所请求操作的比例。除“新患者牙科预约”示例外,其余场景的模型提示词、工具架构和参与者指令目前均保持私有,以减少过拟合。此外,Speech Agent Arena 目前使用经过筛选的合格付费第三方参与者来执行和评估智能体交互。 主要结果: ➤ Arena 偏好 Elo:@GoogleAI Gemini 3.1 Flash Live Preview - Minimal 以 1,046 Elo 领先,其次是 Gemini 3.1 Flash Live Preview - High(1,014)、@OpenAI GPT-Realtime-1.5(1,000)、GPT Realtime(2025 年 8 月版)(944),以及 @ElevenLabs Agents(默认级联系统,由 Scribe v2 Realtime / GPT-4o Mini / Eleven v3 组成,使用预注册工具架构)(937)。在已审阅的对话中,高偏好模型往往响应更快、语音更自然,并且产生更少的不自然声音或音频伪影。 ➤ 任务成功率:@SpaceXAI Grok Voice Think Fast 2.0 High 以 94.7% 领先,其次是 @OpenAI GPT-Realtime-2.1 High(91.5%)、@ElevenLabs Agents(默认级联系统)(90.5%)和 GPT-Realtime-2(High)(89.8%),GPT Realtime(2025 年 8 月版)与 GPT-Realtime-2.1 Minimal 并列 89.4%。Gemini 3.1 Flash Live Preview - Minimal 在整体偏好上以 1,046 Elo 领先,但任务成功率仅为 74.6%,这表明受偏好的对话并不总能带来成功的任务完成——有些对话听起来似乎已完成了所请求的操作,但实际所需的最终工具调用并未成功。 我们正在持续扩展对原生和级联语音到语音系统的覆盖范围,并欢迎在添加更多模型、提供商和场景时提供反馈。详见下方 ⬇️
Announcing our new Speech Agent Arena, evaluating Speech to Speech models on real-world scenarios to analyze conversational preference and task success rate Existing Speech to Speech benchmarks cover reasoning, simulated agentic tasks, and conversational dynamics such as turn-taking and interruption handling. The Speech Agent Arena compares models and cascaded systems as humans complete real-world tasks, measuring conversational preference and successful tool use. This allows us to provide an evaluation which closer reflects real-world use, offering insight into which models users most prefer speaking with and how effectively those models support their requests. Overview of the Speech Agent Arena and Task Success Rate Human participants compare two hidden Speech to Speech models on the same assigned scenario, one of 15 agentic scenarios (tasks requiring tool calling, such as ordering takeout) or 20 non-agentic scenarios (tasks without tool calling, such as asking about opening hours). After separate live conversations with each model, participants select which they preferred, with these pairwise votes used to fit a Preference Elo score. For agentic scenarios, Task Success Rate is the share of eligible conversations (no participant deviations or unverifiable cases) where the model completed the requested action through the correct final tool call or calls. For all but a New Patient Dental Booking example, scenario model prompts, tool schemas and participant instructions are currently private to reduce overfitting. Additionally, the Speech Agent Arena currently uses a qualified pool of paid, screened third-party participants to conduct and evaluate agent interactions. Key results: ➤ Arena Preference Elo: @GoogleAI Gemini 3.1 Flash Live Preview - Minimal leads at 1,046 Elo, followed by Gemini 3.1 Flash Live Preview - High at 1,014, @OpenAI GPT-Realtime-1.5 at 1,000, GPT Realtime (Aug '25) at 944, and @ElevenLabs Agents (default cascaded system of Scribe v2 Realtime / GPT-4o Mini / Eleven v3, with pre-registered tool schema) at 937. In reviewed conversations, highly preferred models tended to respond quickly, sound more natural and produce fewer unnatural sounds or audio artifacts ➤ Task Success Rate: @SpaceXAI Grok Voice Think Fast 2.0 High leads at 94.7%, followed by @OpenAI GPT-Realtime-2.1 High at 91.5%, @ElevenLabs Agents (Default Cascaded System) at 90.5%, and GPT-Realtime-2 (High) at 89.8%, with GPT Realtime (Aug '25) and GPT-Realtime-2.1 Minimal tied at 89.4%. Gemini 3.1 Flash Live Preview - Minimal leads overall preference at 1,046 Elo but records a 74.6% Task Success Rate, showing that a preferred conversation does not always result in successful task completion - some conversations can sound as though the requested action was completed even when the required final tool call was unsuccessful We are continuing to expand our coverage of native and cascaded Speech to Speech systems, and welcome feedback as we add more models, providers and scenarios. See more details below ⬇️
