Artificial Analysis推文
推理实测活动:同一模型在不同端点表现差异显著
在旧金山的 Inference, Measured 活动中,我们发布了端点准确率指数,通过自托管权重作为参考,对多个提供商的端点进行相同评估,得分从73%到100%不等。量化、KV缓存压缩和上下文限制都会影响质量,即使定价页未显示。
译文
又一个美妙的夜晚,我们在旧金山举办了最新一期 Inference, Measured 活动。通过关于 serverless 推理、Endpoint Accuracy Index 和 AA-AgentPerf 的对话,我们探讨了为什么同一个模型并不总是同一个产品。我们全新的 Endpoint Accuracy Index 揭示了另一层差异。我们自行托管发布的权重作为 100% 参考基准,使用相同的三项评估对每个提供商的端点进行测量,并将结果与该参考基准进行对比评分。结果范围从 73% 到 100% 不等。量化、KV-cache 压缩和上下文限制都可能影响质量,即使这些并未显示在定价页面上。欢迎访问我们的网站查看评估详情和完整结果。感谢每一位加入我们并参与讨论性能、速度和成本之间权衡的朋友。
Another fantastic evening at our latest Inference, Measured event in San Francisco. We explored how the same model is not always the same product through conversations on serverless inference, Endpoint Accuracy Index, and AA-AgentPerf Our new Endpoint Accuracy Index reveals another layer. We self-host the released weights as a 100% reference, measure each provider endpoint using the same three evaluations, and score the results relative to that reference. The results range from 73% to 100%. Quantization, KV-cache compression, and context limits can all affect the quality, even when they don't appear on the pricing page. Visit our website to explore the evaluations and full results. Thank you to everyone who joined us and contributed to the conversation about measuring the tradeoffs between performance, speed, and cost.



