Jerry Liu推文
LlamaParse 推出原生智能体电子表格提取功能
LlamaIndex 在 LlamaParse 中新增了原生智能体电子表格提取功能。该功能针对电子表格与 PDF 等文档格式的差异,采用代码解释器而非 OCR 方式,通过调优的智能体引擎(模型+框架)实现大规模模式引导的提取,可将资产负债表等密集表格转换为干净的结构化字段。用户可在 API 文档中集成,或在 UI 的 agentic plus 模式中启用电子表格选项。
译文
我们已将原生的智能体电子表格提取功能 📊 引入 LlamaParse。电子表格与 PDF(或任何其他文档格式)截然不同。它们可以横跨任意多的行/列,无法保证表格结构,并且信息可能跨多个工作表关联。处理电子表格的原生方式通常是通过代码解释器,而非 OCR。我们引入了一个经过调优的智能体引擎(模型+框架),能够支持大规模、基于模式引导的电子表格提取。您可以将诸如资产负债表这类密集表格提取为清晰的结构化字段。查看文档了解如何集成到 API:https://developers.llamaindex.ai/llamaparse/extract/guides/configuring-extract/#spreadsheet-mode 在界面上,选择“agentic plus”,并在高级选项部分查看“spreadsheet options”。在此注册 LlamaParse!http://cloud.llamaindex.ai/
We've introduced native, agentic spreadsheet extraction 📊 into LlamaParse. Spreadsheets are a wildly different format from PDFs (or any other document format). They can span arbitrarily many rows/columns, lack any guarantees on tabular structure, and can have information linked across multiple sheets. The native ways to deal with spreadsheets is usually through a code interpeter as opposed to OCR. We've introduced a tuned agentic engine (model+harness) that is equipped for large-scale schema guided extraction from spreadsheets. You can extract out dense sheets like balance sheets into clean structured fields. Check out the docs for how to integrate into the API: https://developers.llamaindex.ai/llamaparse/extract/guides/configuring-extract/#spreadsheet-mode On the UI, select "agentic plus", and see "spreadsheet options" in the advanced options section. Signup for LlamaParse here! http://cloud.llamaindex.ai/
