Jerry Liu推文
编码代理在长文档提取任务中展现出成本效益优势
在模式引导的复杂文档提取任务中,编码代理(如Claude Code和Codex)在长文档上展现出良好的成本/准确率平衡。短文档上,专用OCR工具成本更低且准确率相当或更高;长文档上,编码代理更接近帕累托最优。该结果来自ParseBench论文附录D。
译文
我们在围绕模式引导的复杂文档抽取任务📑中观察到一个有趣的性质:对于较长的文档,编码智能体框架在成本/准确率方面是很好的基线。我们测试了Claude Code和Codex,以及专门的OCR工具(包括LlamaParse)和原始VLM。 * 在短文档上,专门的OCR工具通常只需编码智能体成本的一小部分,且准确率相当或更高 * 在长文档上,编码智能体则更接近成本/准确率的帕累托曲线(见下图) 这是一个有趣的结果,虽然最终并不令人意外。复杂文档抽取是一项专门的推理任务,而编码智能体实际上是通用的推理框架。在处理长文档时,编码智能体有更多空间使用各种工具来搜索文档片段,而不是将整个文档加载到上下文中。它们还可以利用提示缓存来降低总token成本,即使多步推理在扩展。另一方面,它们确实会产生一定的基础token消耗,对于短文档而言,与专门的抽取器相比,这种消耗显得浪费。 这张具体的图位于我们ParseBench论文的附录D中,欢迎查看!ArXiv: https://arxiv.org/pdf/2607.29677 ExtractBench: https://www.extractbench.ai/
One of the interesting properties we’ve observed around schema-guided, complex document extraction tasks 📑 is that coding agent harnesses are good baselines (in terms of cost/accuracy) for longer documents. We tested Claude Code and Codex, along with specialized OCR tools (including LlamaParse) and raw VLMs. * On short documents, specialized OCR tools are generally a fraction of the cost of coding agents, with equivalent or higher accuracy * On longer documents, coding agents are a bit closer to the cost/accuracy Pareto curve (see bottom graph) It’s an interesting result, though ultimately not surprising. Complex document extraction is a specialized reasoning task, and coding agents are effectively generalized reasoning harnesses. Over long documents, coding agents have more room to use a variety of tools to search snippets of the document instead of loading the entire document into context. They can also make use of prompt caching to reduce total token cost even as it expands multi-step reasoning. On the flip side, they do generate a baseline degree of token usage that proves to be wasteful for shorter docs compared to specialized extractors. This specific graph is in our Appendix D in the ParseBench paper, come check it out! ArXiv: https://arxiv.org/pdf/2607.29677 ExtractBench: https://www.extractbench.ai/
