Artificial Analysis推文
手机端小模型评测:iPhone 17 Pro 与 Galaxy S26 Ultra 上的智能与推理性能
Artificial Analysis 与 Liquid AI 合作,推出针对手机端小模型的智能与推理性能评测,覆盖 iPhone 17 Pro 和 Galaxy S26 Ultra 上的多款 4-bit 或更低精度模型。评测结合智能评估与推理性能,提供端到端生成时间、输出速度、峰值内存等指标。初步结果显示,Nanbeige4.2-3B 和 LFM2.5-2.6B 在 16K 上下文限制下并列智能得分第一,但 LFM2.5-2.6B 效率更高。评测应用 Pipette 已免费开放下载。
译文
在 iPhone 17 Pro 上应该运行哪个模型?我们宣布推出针对移动设备端小模型的全新智能与推理测试:独立衡量小模型在典型端侧任务中的能力,以及它们在热门手机上的表现——与 @liquidai 合作。我们已与 @liquidai 合作推出移动设备推理基准测试,覆盖 iPhone 17 Pro 和 Galaxy S26 Ultra 上采用 4-bit 或更低精度的一系列模型。我们将手机级智能评估结果与 @liquidai 的推理性能基准结合发布,为用户和开发者提供小模型在手机上表现的全景视图。推理基准测试在受控环境中使用 Liquid AI 的推理基准测试软件进行,该软件已经过 Artificial Analysis 审查,并在 Liquid AI 的 GitHub 上开源。结果涵盖端到端生成时间、输出速度、峰值内存占用及其他指标。推理基准测试应用“Pipette”可在 iOS 和 Android 上免费下载,用户可以在自己的设备上测试多种模型。 我们根据每个模型在五项评估中的平均得分对手机级模型智能进行排名,这五项评估针对这些模型在实际任务型工作中的表现而选定:BFCL、IFBench、AA-Omniscience、GPQA Diamond 和 MATH-500。这些评估由 Artificial Analysis 使用我们独立的方法论运行。默认情况下,我们在每项评估中将模型限制为 16K 上下文,这代表了移动设备上缺乏足够内存来承载大规模 KV cache 的现实。这导致一些智能但冗长的模型相对得分下降——它们并非为手机内存对 token 使用的限制而设计。 我们将便携设备类别定义为量化后(含 KV cache,8K 上下文)适配 8 GB 内存的模型。我们预计智能和推理基准都会随着新模型、新设备、新推理框架和量化技术的发布而不断演进。我们的页面将持续更新这些新增内容。 初步结果: ➤ Nanbeige4.2-3B 和 LFM2.5-2.6B 以 63 分的平均评估得分并列第一(16K 上下文限制),领先于 Ornith-1.0-9B 的 62 分和 Qwen3.5 9B(Reasoning)的 61 分。LFM2.5-2.6B 以更高效率取得该得分:在 iPhone 17 Pro 上,它回答一个标准的 1,024-token 提示仅需 8.0 秒,占用 2.3 GB 内存;而 Nanbeige4.2-3B 需要 21.4 秒和 4.0 GB,两款 9B 模型则需要 25 秒以上和 6.9 GB。 ➤ 16K 上下文限制塑造了排行榜:例如,Qwen3.5 9B(Reasoning)在基准测试集上单次运行中输出 74.5M tokens,其中 29% 的生成触达 16K 上限,最终排名第四。当上限提升至 64K 时,Ling 3.0 Tiny 以 66 分位居第一,领先于 Nanbeige4.2-3B(65 分)、Qwen3.5 9B(Reasoning,64 分)和 LFM2.5 2.6B(64 分)。但 64K 窗口无法适配手机内存,而且在 iPhone 上以 55 tokens/s 的输出速度生成 64K tokens 可能意味着 20 分钟以上的等待和大量电量消耗。这就是我们主要结果以 16K 为上限的原因,但我们也发布了一组以 64K 为上限的结果,以及另一组以一分钟生成时间为上限的结果。 ➤ 速度-智能帕累托前沿很短:在 iPhone 17 Pro 上,有六款模型在智能和速度两方面均未被超越:LFM2.5-230M(27 分,0.9 秒)、MiniCPM5-1B(45 分,2.9 秒)、LFM2.5-8B-A1B(58 分,5.7 秒)、Ling 3.0 Tiny(59 分,5.7 秒)、LFM2.5-2.6B(63 分,8.0 秒)和 Nanbeige4.2-3B(63 分,21.4 秒)。LFM2.5-8B-A1B 和 Ling 3.0 Tiny 是混合专家模型,每个 token 仅激活约 1B 参数,这就是它们能以 8B 级权重在 6 秒内完成回答的原因。 ➤ 领先模型各有优势:Nanbeige4.2-3B 最为均衡(BFCL 76%、MATH-500 96%、GPQA Diamond 67%);Qwen3.5 9B(Non-reasoning)是最强的工具调用者(BFCL 77%)和科学推理者(GPQA Diamond 79%);LFM2.5-2.6B 在 iPhone 上所有已测模型中指令遵循能力最佳(IFBench 59%),MATH-500 超过 90%,并且在 AA-Omniscience 上的幻觉率远低于其他模型(非幻觉率 79%,而 Nanbeige4.2-3B 为 33%,Qwen3.5 9B(Reasoning)为 24%)。更多详情见下方帖子 ⬇️
Which model should you run on the iPhone 17 Pro? Announcing our new intelligence and inference testing for small models on mobile devices: independent measurement of how capable small models are in typical on-device tasks, and how they perform on popular phones - in partnership with @liquidai We have partnered with @liquidai to deliver mobile device inference benchmarking, covering a range of models in 4-bit or lower precision on the iPhone 17 Pro and Galaxy S26 Ultra We’re publishing our phone-scale intelligence evaluation results in combination with @liquidai's inference performance benchmarks to give users and developers a holistic view of how small models are performing on phones Inference benchmarking is conducted in a controlled environment using Liquid AI’s inference benchmarking software, which Artificial Analysis has examined and is open-sourced on Liquid AI’s GitHub. Results cover end-to-end generation time, output speed, peak memory usage and other metrics. The inference benchmarking app, ‘Pipette’, is available to download for free on iOS and Android, allowing users to test a variety of models on their own devices We are ranking phone-scale model intelligence based on each model’s average score in five evaluations chosen for the task-based work these models do in practice: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. These evaluations are run by Artificial Analysis using our independent methodology. By default, we limit models to 16K context on each of these evaluations, representing the lack of memory space for significant KV cache on mobile devices. This leads to some intelligent but verbose models dipping in relative score - they were not designed for the constraints that phone memory imposes on token use We are defining our portable device category as including models that fit within 8 GB of memory after quantization, including KV cache, at 8K context We expect both the intelligence and inference benchmarks to evolve over time, as new models, devices, inference frameworks and quantization techniques are released. Our pages will remain up to date with these new additions Initial results: ➤ Nanbeige4.2-3B and LFM2.5-2.6B share the top average evaluation score at 63 (with a 16K context limit), ahead of Ornith-1.0-9B at 62 and Qwen3.5 9B (Reasoning) at 61. LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0s using 2.3 GB of memory, against 21.4s and 4.0 GB for Nanbeige4.2-3B and 25+ seconds and 6.9 GB for the two 9B models ➤ The 16K context limit shapes the leaderboard: as an example, Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set, hitting the 16K limit on 29% of its generations and landing in fourth place overall. With the limit raised to 64K, Ling 3.0 Tiny takes first place with a score of 66, ahead of Nanbeige4.2-3B (65), Qwen3.5 9B (Reasoning, 64), and LFM2.5 2.6B (64). But a 64K window does not fit in mobile phone memory, and at 55 output tokens/s on an iPhone, generating 64K tokens could mean a 20+ minute wait and a lot of battery use. This is why our primary results are capped at 16K, but we're also publishing a set of results capped at 64K, and another capped at one minute of generation time ➤ The speed-intelligence Pareto frontier is short: six models are unbeaten on both intelligence and speed on an iPhone 17 Pro: LFM2.5-230M (27 at 0.9s), MiniCPM5-1B (45 at 2.9s), LFM2.5-8B-A1B (58 at 5.7s), Ling 3.0 Tiny (59 at 5.7s), LFM2.5-2.6B (63 at 8.0s) and Nanbeige4.2-3B (63 at 21.4s). LFM2.5-8B-A1B and Ling 3.0 Tiny are mixture-of-experts models that activate ~1B parameters per token, which is how they answer in under 6s with 8B-class weights ➤ Leading models have opposite strengths: Nanbeige4.2-3B is the most balanced (76% on BFCL, 96% on MATH-500, 67% on GPQA Diamond); Qwen3.5 9B (Non-reasoning) is the strongest tool caller (77% on BFCL) and scientific reasoner (79% on GPQA Diamond); LFM2.5-2.6B follows instructions best of any model measured on the iPhone (59% on IFBench), clears 90% on MATH-500 and hallucinates far less on AA-Omniscience (79% non-hallucination, against 33% for Nanbeige4.2-3B and 24% for Qwen3.5 9B (Reasoning)) More details below in thread ⬇️
