Jerry Liu推文
当前智能体框架的RAG趋势:两遍文档处理
作者指出当前智能体框架(如Codex、Cowork)处理文档时采用两遍法:先用轻量解析工具快速过滤,再用VLM对相关页面进行精细OCR。但现成工具存在成本高、精度不足等问题,因此推荐自家LiteParse和LlamaParse作为替代方案。
译文
当前智能体框架(Codex、Cowork)的最新RAG趋势是,对文档进行两轮处理,以解决在文档数据室中的知识工作任务:1️⃣ 快速轻量的一轮,通常使用免费/开源文档解析工具。这可以低成本地运行在10-100-1000个文件上,使智能体随后能够进行检索(例如grep、语义检索)以找到相关的上下文子集。2️⃣ 一轮“即时”的基于VLM的处理。一旦智能体找到相关的上下文页面,它会对文档截图,并调用自身的VLM(或编写代码)来解析这些页面。仅对大量临时客户文件转储使用基于VLM的OCR工具的问题在于,速度慢且成本高。进行即时VLM OCR使智能体能够低成本地筛选数据,同时仍能保持任务所需上下文的准确性。智能体框架默认使用现成工具进行两轮文档处理:第一轮使用pdf2text,第二轮使用自身(Opus 5)。请参阅下方视频,其中Cowork处理大量PDF,以回答关于Kimi k3论文中基准图表的问题。这些智能体提供的“开箱即用”文档处理的主要问题包括:* Opus 5并非最适合OCR的VLM。它在规模化时也过于昂贵,且缺乏接地能力。* 像pypdf、pdf2text这样的开源工具作为第一轮可能不够通用。* 智能体会编写大量一次性代码来重写OCR工具本应开箱即用提供的内容,例如图表处理、边界框、置信度分数,导致成本和速度增加。我们在@llama_index中拥有所有工具,可帮助任何智能体以更高准确性和更低成本进行两轮文档处理。1️⃣ 我们为第一轮提供liteparse——一个用Rust编写的免费/开源解析器,比其他开源解析器更快/更准确,支持50多种文档类型。2️⃣ 我们为第二轮提供LlamaParse——一个智能体文档引擎,使用VLM+框架在各种文档解析和提取任务中实现准确性和成本方面的SOTA。它可以从任何智能体框架中作为MCP或技能调用。它接受页码作为输入,因此智能体可以选择对文档子集运行LlamaParse,而不是整个文档,作为“放大”轮次。欢迎查看!LiteParse: https://github.com/run-llama/liteparse LlamaParse: https://cloud.llamaindex.ai/ 所有相关文档,包括MCP,均在此处:https://developers.llamaindex.ai/llamaparse/for-agents/mcp/
The latest RAG trend for the current agent harnesses (Codex, Cowork) is to do two passes of document processing to solve a knowledge work task over a data room of documents: 1️⃣ A fast and light pass, oftentimes using a free/OSS doc parsing tool. This can be cheaply run across 10-100-1k’s of files, and enables the agent to then do retrieval (e.g. grep, semantic) to find relevant subsets of context. 2️⃣ A “just-in-time” VLM-based pass. Once the agent finds the relevant pages of context, it will screenshot the documents can call its own VLM (or write code) to dissect the pages. The issue with only using VLM-based OCR tools over massive ad-hoc customer file dumps is that it’s slow and expensive. Doing JIT VLM OCR allows the agent to filter through the data cheaply, but still preserve accuracy for the context that’s needed for the task. The agent harnesses do two-pass document processing by default using off the shelf-tools: pdf2text as the first pass, and using itself (Opus 5) as the second pass. See the below video where Cowork runs over a bunch of PDFs to answer a question about a benchmark graph in the Kimi k3 paper. The main issues here with the “out of the box” doc processing these agents offer are: * Opus 5 is not the best VLM for OCR. It is also way too expensive at scale and lacks grounding * The OSS tools like pypdf, pdf2text, may not be versatile enough as the first pass. * The agent will write a lot of throwaway code to rewrite things an OCR tool would’ve provided out of the box, like chart processing, bounding boxes, confidence scores, leading to increased cost and speed. We have all the tools within @llama_index to help any agent do two-pass document processing with higher accuracy and lower cost. 1️⃣ We have liteparse for the first pass - a free/OSS parser written in Rust that’s faster/more accurate than other OSS parsers, and supports 50+ document types 2️⃣ We have LlamaParse for the second pass - an agentic document engine that uses VLMs+harnesses to achieve SOTA in accuracy and cost across various doc parsing and extraction tasks. It can be called from any agent harness as an MCP or skill. It takes in page numbers as input, so that the agent can choose to run LlamaParse over a subset of the doc instead of the full doc as a “zoom-in” pass. Come check it out! LiteParse: https://github.com/run-llama/liteparse LlamaParse: https://cloud.llamaindex.ai/ All the relevant docs, including MCP, are here: https://developers.llamaindex.ai/llamaparse/for-agents/mcp/