Artificial Analysis推文
Agnes AI 发布 Agnes 2.5 Pro Beta,智能指数提升 9 分
新加坡 AI 实验室 Agnes AI 发布了 Agnes 2.5 Pro Beta,在 Artificial Analysis 智能指数上得分 49,较 Alpha 版提升 9 分,主要得益于智能体能力的显著增强,但输出 token 消耗约为前代的两倍。该模型在智能体指数上从 25 跃升至 44,接近 Gemini 3.7 Flash,同时在前沿推理评测中也有小幅提升。模型支持 1M 上下文,定价为每百万输入/输出/缓存命中 token 0.10/0.30/0.01 美元,现可通过其 API 使用。
译文
Agnes AI的Agnes 2.5 Pro Beta在Artificial Analysis Intelligence Index上获得49分,较Agnes 2.5 Pro Alpha提升9分,主要得益于智能体能力的大幅提升,但输出token消耗约为前代的2倍。Agnes AI(@agnesai_sapiens)是一家总部位于新加坡的AI实验室,训练全模态基础模型,并通过免费的omni-modal API提供模型服务。在Intelligence Index上获得49分后,Agnes 2.5 Pro Beta将Agnes从中游位置推至前沿邻近梯队,仅次于Gemini 3.5 Flash(high,52)和GPT-5.6 Luna(max,52),领先于MiniMax-M3(45)。主要结果: ➤ Agnes 2.5 Pro Beta在Artificial Analysis Intelligence Index上获得49分,较Agnes 2.5 Pro Alpha(40)跃升9分。这一成绩使其仅次于Gemini 3.5 Flash(high,52)和GPT-5.6 Luna(max,52),领先于MiniMax-M3(45)。 ➤ 智能体能力是此次跃升的主要驱动力。Artificial Analysis Agentic Index从25升至44,仅次于Gemini 3.7 Flash(high,45),领先于Gemini 3.5 Flash(high,40)和MiniMax-M3(36)。τ³-Banking从12%几乎翻三倍至36%,GDPval-AA v2的Elo从1171升至1456(人类基线为1000)。 ➤ 前沿推理评估的提升相对温和。Humanity's Last Exam从34%升至38%,GPQA Diamond从88%升至91%,CritPt从11%升至16%。 ➤ AA-Omniscience从-25提升至-11主要来自弃答策略,而非准确率提升。Agnes 2.5 Pro Beta仅尝试回答45%的问题(Agnes 2.5 Pro Alpha为94%),幻觉率从88%降至33%,但AA-Omniscience Accuracy从33%减半至17%。 ➤ 智能水平的提升需要消耗约为前代2倍的token。Agnes 2.5 Pro Beta在每项Intelligence Index任务中使用50k输出token,是Agnes 2.5 Pro Alpha(24k)的两倍以上。 其他模型详情: ➤ 上下文窗口:1M token ➤ 最大输出token:65k ➤ 输入模态:文本和图像 ➤ 定价:每100万输入/输出/缓存命中token分别为$0.10 / $0.30 / $0.01 ➤ 可用性:Agnes AI第一方API
Agnes AI's Agnes 2.5 Pro Beta scores 49 on the Artificial Analysis Intelligence Index, up 9 points from Agnes 2.5 Pro Alpha, driven by large agentic gains but using ~2x the output tokens Agnes AI (@agnesai_sapiens) is a Singapore-based AI lab that trains full-modality foundation models and offers them through a free omni-modal API. At 49 on the Intelligence Index, Agnes 2.5 Pro Beta moves Agnes from mid-pack to the frontier-adjacent tier, just behind Gemini 3.5 Flash (high, 52) and GPT-5.6 Luna (max, 52) and ahead of MiniMax-M3 (45). Key results: ➤ Agnes 2.5 Pro Beta scores 49 on the Artificial Analysis Intelligence Index, a 9-point jump from Agnes 2.5 Pro Alpha (40). This places it just behind Gemini 3.5 Flash (high, 52) and GPT-5.6 Luna (max, 52), and ahead of MiniMax-M3 (45). ➤ Agentic capabilities drive the jump. The Artificial Analysis Agentic Index rises from 25 to 44, just behind Gemini 3.7 Flash (high, 45) and ahead of Gemini 3.5 Flash (high, 40) and MiniMax-M3 (36). τ³-Banking nearly triples from 12% to 36%, and GDPval-AA v2 rises from an Elo of 1171 to 1456 against a human baseline of 1000. ➤ Frontier reasoning evaluations improves more modestly. Humanity's Last Exam rises from 34% to 38%, GPQA Diamond from 88% to 91%, and CritPt from 11% to 16%. ➤ The AA-Omniscience improvement from -25 to -11 comes from abstention, not increased accuracy. Agnes 2.5 Pro Beta attempts only 45% of questions against Agnes 2.5 Pro Alpha's 94%, cutting the hallucination rate from 88% to 33%, but AA-Omniscience Accuracy halves from 33% to 17%. ➤ The intelligence gain required ~2x as many tokens from its predecessor. Agnes 2.5 Pro Beta uses 50k output tokens per Intelligence Index task, more than double Agnes 2.5 Pro Alpha's 24k. Additional model details: ➤ Context window: 1M tokens ➤ Max output tokens: 65k ➤ Input modalities: Text and image ➤ Pricing: $0.10 / $0.30 / $0.01 per 1M input / output / cache hit tokens ➤ Availability: Agnes AI first-party API
