← 返回信息流

arXiv cs.AI论文

SparKV:面向高效端侧大模型推理的开销感知 KV Cache 加载方法

arxiv.org作者:Hongyao Liu, Liuqun Zhai, Junyi Wang, Zhengru Fang, Jingshu Chen, Jun Huang论文AI评分:70/100

针对端侧大模型推理中预填充阶段因处理完整输入上下文导致的高成本问题,提出 SparKV 框架。该方法结合云端 KV 流式传输与端侧计算,通过建模单个 KV 块的成本,自适应决策各块的加载策略,旨在优化有限硬件资源下的推理效率。

这篇是正式发表的长论文,站内提供中文解读,全文请到原文阅读 PDF。

Abstract:Efficient inference for on-device Large Language Models (LLMs) remains challenging due to limited hardware resources and the high cost of the prefill stage, which processes the full input context to construct Key-Value (KV) caches. We present SparKV, an adaptive KV loading framework that combines cloud-based KV streaming with on-device computation. SparKV models the cost of individual KV chunks and decides whether each chunk should be streamed or computed locally, while overlapping the two execution paths to reduce latency. To handle fluctuations in wireless connectivity and edge resource availability, SparKV further refines offline-generated schedules at runtime to rebalance communication and computation costs. Experiments across diverse datasets, LLMs, and edge devices show that SparKV reduces Time-to-First-Token by 1.3$x-5.1x with negligible impact on response quality, while lowering per-request energy consumption by 1.5x to 3.3x, demonstrating its robustness and practicality for real-world on-device deployment.

阅读原文