Docling 阅读器
Docling 可将 PDF、DOCX、HTML 及其他文档格式提取为富文本表示(包括布局、表格等),并能导出为 Markdown 或 JSON 格式。
本笔记本中展示的 Docling 阅读器和 Docling 节点解析器可无缝将 Docling 集成到 LlamaIndex 中,使您能够:
- 在您的LLM应用中轻松快速地使用各种文档类型,以及
- 利用 Docling 的丰富格式实现高级的、文档原生的基础功能。
- 👉 为获得最佳转换速度,请尽可能使用GPU加速;例如,如果在Colab上运行,请启用GPU运行时。
- Notebook 使用 HuggingFace 的推理 API;如需增加 LLM 配额,可通过环境变量
HF_TOKEN提供令牌。 - 可按如下所示安装依赖项(
--no-warn-conflicts专为 Colab 预置的 Python 环境设计;如需更严格的使用可自行移除):
%pip install -q --progress-bar off --no-warn-conflicts llama-index-core llama-index-readers-docling llama-index-node-parser-docling llama-index-embeddings-huggingface llama-index-llms-huggingface-api llama-index-readers-file python-dotenv我们现在可以定义主要参数:
from llama_index.embeddings.huggingface import HuggingFaceEmbeddingfrom llama_index.llms.huggingface_api import HuggingFaceInferenceAPIimport osfrom dotenv import load_dotenv
def get_env_from_colab_or_os(key): try: from google.colab import userdata
try: return userdata.get(key) except userdata.SecretNotFoundError: pass except ImportError: pass return os.getenv(key)
load_dotenv()EMBED_MODEL = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")GEN_MODEL = HuggingFaceInferenceAPI( token=get_env_from_colab_or_os("HF_TOKEN"), model_name="mistralai/Mixtral-8x7B-Instruct-v0.1",)SOURCE = "https://arxiv.org/pdf/2408.09869" # Docling Technical ReportQUERY = "Which are the main AI models in Docling?"使用Markdown导出
Section titled “Using Markdown export”要创建一个简单的RAG流水线,我们可以:
- 定义一个
DoclingReader,默认导出为Markdown格式,并且 - 对于这些基于Markdown的文档,使用标准节点解析器,例如
MarkdownNodeParser
from llama_index.core import VectorStoreIndexfrom llama_index.core.node_parser import MarkdownNodeParserfrom llama_index.readers.docling import DoclingReader
reader = DoclingReader()node_parser = MarkdownNodeParser()
index = VectorStoreIndex.from_documents( documents=reader.load_data(SOURCE), transformations=[node_parser], embed_model=EMBED_MODEL,)result = index.as_query_engine(llm=GEN_MODEL).query(QUERY)print(f"Q: {QUERY}\nA: {result.response.strip()}\n\nSources:")display([(n.text, n.metadata) for n in result.source_nodes])Q: Which are the main AI models in Docling?A: 1. A layout analysis model, an accurate object-detector for page elements. 2. TableFormer, a state-of-the-art table structure recognition model.
Sources:
[('3.2 AI models\n\nAs part of Docling, we initially release two highly capable AI models to the open-source community, which have been developed and published recently by our team. The first model is a layout analysis model, an accurate object-detector for page elements [13]. The second model is TableFormer [12, 9], a state-of-the-art table structure recognition model. We provide the pre-trained weights (hosted on huggingface) and a separate package for the inference code as docling-ibm-models . Both models are also powering the open-access deepsearch-experience, our cloud-native service for knowledge exploration tasks.', {'Header_2': '3.2 AI models'}), ("5 Applications\n\nThanks to the high-quality, richly structured document conversion achieved by Docling, its output qualifies for numerous downstream applications. For example, Docling can provide a base for detailed enterprise document search, passage retrieval or classification use-cases, or support knowledge extraction pipelines, allowing specific treatment of different structures in the document, such as tables, figures, section structure or references. For popular generative AI application patterns, such as retrieval-augmented generation (RAG), we provide quackling , an open-source package which capitalizes on Docling's feature-rich document output to enable document-native optimized vector embedding and chunking. It plugs in seamlessly with LLM frameworks such as LlamaIndex [8]. Since Docling is fast, stable and cheap to run, it also makes for an excellent choice to build document-derived datasets. With its powerful table structure recognition, it provides significant benefit to automated knowledge-base construction [11, 10]. Docling is also integrated within the open IBM data prep kit [6], which implements scalable data transforms to build large-scale multi-modal training datasets.", {'Header_2': '5 Applications'})]使用Docling格式
Section titled “Using Docling format”为了利用 Docling 丰富的原生格式,我们:
- 创建一个具有JSON导出类型的
DoclingReader,并且 - 使用一个
DoclingNodeParser来正确解析该Docling格式。
请注意,现在来源信息还包含了文档级别的定位信息(例如页码或边界框信息):
from llama_index.node_parser.docling import DoclingNodeParser
reader = DoclingReader(export_type=DoclingReader.ExportType.JSON)node_parser = DoclingNodeParser()
index = VectorStoreIndex.from_documents( documents=reader.load_data(SOURCE), transformations=[node_parser], embed_model=EMBED_MODEL,)result = index.as_query_engine(llm=GEN_MODEL).query(QUERY)print(f"Q: {QUERY}\nA: {result.response.strip()}\n\nSources:")display([(n.text, n.metadata) for n in result.source_nodes])Q: Which are the main AI models in Docling?A: The main AI models in Docling are a layout analysis model and TableFormer. The layout analysis model is an accurate object-detector for page elements, and TableFormer is a state-of-the-art table structure recognition model.
Sources:
[('As part of Docling, we initially release two highly capable AI models to the open-source community, which have been developed and published recently by our team. The first model is a layout analysis model, an accurate object-detector for page elements [13]. The second model is TableFormer [12, 9], a state-of-the-art table structure recognition model. We provide the pre-trained weights (hosted on huggingface) and a separate package for the inference code as docling-ibm-models . Both models are also powering the open-access deepsearch-experience, our cloud-native service for knowledge exploration tasks.', {'schema_name': 'docling_core.transforms.chunker.DocMeta', 'version': '1.0.0', 'doc_items': [{'self_ref': '#/texts/34', 'parent': {'$ref': '#/body'}, 'children': [], 'label': 'text', 'prov': [{'page_no': 3, 'bbox': {'l': 107.07593536376953, 't': 406.1695251464844, 'r': 504.1148681640625, 'b': 330.2677307128906, 'coord_origin': 'BOTTOMLEFT'}, 'charspan': [0, 608]}]}], 'headings': ['3.2 AI models'], 'origin': {'mimetype': 'application/pdf', 'binary_hash': 14981478401387673002, 'filename': '2408.09869v3.pdf'}}), ('With Docling , we open-source a very capable and efficient document conversion tool which builds on the powerful, specialized AI models and datasets for layout analysis and table structure recognition we developed and presented in the recent past [12, 13, 9]. Docling is designed as a simple, self-contained python library with permissive license, running entirely locally on commodity hardware. Its code architecture allows for easy extensibility and addition of new features and models.', {'schema_name': 'docling_core.transforms.chunker.DocMeta', 'version': '1.0.0', 'doc_items': [{'self_ref': '#/texts/9', 'parent': {'$ref': '#/body'}, 'children': [], 'label': 'text', 'prov': [{'page_no': 1, 'bbox': {'l': 107.0031967163086, 't': 136.7283935546875, 'r': 504.04998779296875, 'b': 83.30133056640625, 'coord_origin': 'BOTTOMLEFT'}, 'charspan': [0, 488]}]}], 'headings': ['1 Introduction'], 'origin': {'mimetype': 'application/pdf', 'binary_hash': 14981478401387673002, 'filename': '2408.09869v3.pdf'}})]为了演示这种使用模式,我们首先设置一个测试文档目录。
from pathlib import Pathfrom tempfile import mkdtempimport requests
tmp_dir_path = Path(mkdtemp())r = requests.get(SOURCE)with open(tmp_dir_path / f"{Path(SOURCE).name}.pdf", "wb") as out_file: out_file.write(r.content)使用上述任意变体中的 reader 和 node_parser 定义后,与 SimpleDirectoryReader 的用法如下所示:
from llama_index.core import SimpleDirectoryReader
dir_reader = SimpleDirectoryReader( input_dir=tmp_dir_path, file_extractor={".pdf": reader},)
index = VectorStoreIndex.from_documents( documents=dir_reader.load_data(SOURCE), transformations=[node_parser], embed_model=EMBED_MODEL,)result = index.as_query_engine(llm=GEN_MODEL).query(QUERY)print(f"Q: {QUERY}\nA: {result.response.strip()}\n\nSources:")display([(n.text, n.metadata) for n in result.source_nodes])Q: Which are the main AI models in Docling?A: The main AI models in Docling are a layout analysis model and TableFormer. The layout analysis model is an accurate object-detector for page elements, and TableFormer is a state-of-the-art table structure recognition model.
Sources:
[('As part of Docling, we initially release two highly capable AI models to the open-source community, which have been developed and published recently by our team. The first model is a layout analysis model, an accurate object-detector for page elements [13]. The second model is TableFormer [12, 9], a state-of-the-art table structure recognition model. We provide the pre-trained weights (hosted on huggingface) and a separate package for the inference code as docling-ibm-models . Both models are also powering the open-access deepsearch-experience, our cloud-native service for knowledge exploration tasks.', {'file_path': '/var/folders/76/4wwfs06x6835kcwj4186c0nc0000gn/T/tmpgwz4gpzx/2408.09869.pdf', 'file_name': '2408.09869.pdf', 'file_type': 'application/pdf', 'file_size': 5566574, 'creation_date': '2024-10-24', 'last_modified_date': '2024-10-24', 'schema_name': 'docling_core.transforms.chunker.DocMeta', 'version': '1.0.0', 'doc_items': [{'self_ref': '#/texts/34', 'parent': {'$ref': '#/body'}, 'children': [], 'label': 'text', 'prov': [{'page_no': 3, 'bbox': {'l': 107.07593536376953, 't': 406.1695251464844, 'r': 504.1148681640625, 'b': 330.2677307128906, 'coord_origin': 'BOTTOMLEFT'}, 'charspan': [0, 608]}]}], 'headings': ['3.2 AI models'], 'origin': {'mimetype': 'application/pdf', 'binary_hash': 14981478401387673002, 'filename': '2408.09869.pdf'}}), ('With Docling , we open-source a very capable and efficient document conversion tool which builds on the powerful, specialized AI models and datasets for layout analysis and table structure recognition we developed and presented in the recent past [12, 13, 9]. Docling is designed as a simple, self-contained python library with permissive license, running entirely locally on commodity hardware. Its code architecture allows for easy extensibility and addition of new features and models.', {'file_path': '/var/folders/76/4wwfs06x6835kcwj4186c0nc0000gn/T/tmpgwz4gpzx/2408.09869.pdf', 'file_name': '2408.09869.pdf', 'file_type': 'application/pdf', 'file_size': 5566574, 'creation_date': '2024-10-24', 'last_modified_date': '2024-10-24', 'schema_name': 'docling_core.transforms.chunker.DocMeta', 'version': '1.0.0', 'doc_items': [{'self_ref': '#/texts/9', 'parent': {'$ref': '#/body'}, 'children': [], 'label': 'text', 'prov': [{'page_no': 1, 'bbox': {'l': 107.0031967163086, 't': 136.7283935546875, 'r': 504.04998779296875, 'b': 83.30133056640625, 'coord_origin': 'BOTTOMLEFT'}, 'charspan': [0, 488]}]}], 'headings': ['1 Introduction'], 'origin': {'mimetype': 'application/pdf', 'binary_hash': 14981478401387673002, 'filename': '2408.09869.pdf'}})]