结构化分层检索
在多个文档上做好RAG很困难。通用框架在接收到用户查询时,首先选择相关文档,然后再选择其中的内容。
但选择文档可能很困难——我们如何根据用户查询的不同属性动态选择文档?
在本笔记本中,我们将向您展示我们的多文档RAG架构:
- 将每个文档表示为一个简洁的元数据字典,包含不同属性:提取的摘要以及结构化元数据。
- 将此元数据字典作为筛选条件存储在向量数据库中。
- 给定用户查询,首先执行自动检索 - 推断相关的语义查询和筛选条件集以查询此数据(有效结合文本转SQL和语义搜索)。
%pip install llama-index-readers-github%pip install llama-index-vector-stores-weaviate%pip install llama-index-llms-openai!pip install llama-index llama-hub在本节中,我们将加载 LlamaIndex 的 GitHub 问题。
import nest_asyncio
nest_asyncio.apply()import os
os.environ["GITHUB_TOKEN"] = "ghp_..."os.environ["OPENAI_API_KEY"] = "sk-..."import os
from llama_index.readers.github import ( GitHubRepositoryIssuesReader, GitHubIssuesClient,)
github_client = GitHubIssuesClient()loader = GitHubRepositoryIssuesReader( github_client, owner="run-llama", repo="llama_index", verbose=True,)
orig_docs = loader.load_data()
limit = 100
docs = []for idx, doc in enumerate(orig_docs): doc.metadata["index_id"] = int(doc.id_) if idx >= limit: break docs.append(doc)Found 100 issues in the repo page 1Resulted in 100 documentsFound 100 issues in the repo page 2Resulted in 200 documentsFound 100 issues in the repo page 3Resulted in 300 documentsFound 64 issues in the repo page 4Resulted in 364 documentsNo more issues found, stoppingimport weaviate
# cloudauth_config = weaviate.AuthApiKey( api_key="XRa15cDIkYRT7AkrpqT6jLfE4wropK1c1TGk")client = weaviate.Client( "https://llama-index-test-v0oggsoz.weaviate.network", auth_client_secret=auth_config,)
class_name = "LlamaIndex_docs"# optional: delete schemaclient.schema.delete_class(class_name)from llama_index.vector_stores.weaviate import WeaviateVectorStorefrom llama_index.core import VectorStoreIndex, StorageContext
vector_store = WeaviateVectorStore( weaviate_client=client, index_name=class_name)storage_context = StorageContext.from_defaults(vector_store=vector_store)doc_index = VectorStoreIndex.from_documents( docs, storage_context=storage_context)from llama_index.core import SummaryIndexfrom llama_index.core.async_utils import run_jobsfrom llama_index.llms.openai import OpenAIfrom llama_index.core.schema import IndexNodefrom llama_index.core.vector_stores import ( FilterOperator, MetadataFilter, MetadataFilters,)
async def aprocess_doc(doc, include_summary: bool = True): """Process doc.""" metadata = doc.metadata
date_tokens = metadata["created_at"].split("T")[0].split("-") year = int(date_tokens[0]) month = int(date_tokens[1]) day = int(date_tokens[2])
assignee = ( "" if "assignee" not in doc.metadata else doc.metadata["assignee"] ) size = "" if len(doc.metadata["labels"]) > 0: size_arr = [l for l in doc.metadata["labels"] if "size:" in l] size = size_arr[0].split(":")[1] if len(size_arr) > 0 else "" new_metadata = { "state": metadata["state"], "year": year, "month": month, "day": day, "assignee": assignee, "size": size, }
# now extract out summary summary_index = SummaryIndex.from_documents([doc]) query_str = "Give a one-sentence concise summary of this issue." query_engine = summary_index.as_query_engine( llm=OpenAI(model="gpt-3.5-turbo") ) summary_txt = await query_engine.aquery(query_str) summary_txt = str(summary_txt)
index_id = doc.metadata["index_id"] # filter for the specific doc id filters = MetadataFilters( filters=[ MetadataFilter( key="index_id", operator=FilterOperator.EQ, value=int(index_id) ), ] )
# create an index node using the summary text index_node = IndexNode( text=summary_txt, metadata=new_metadata, obj=doc_index.as_retriever(filters=filters), index_id=doc.id_, )
return index_node
async def aprocess_docs(docs): """Process metadata on docs."""
index_nodes = [] tasks = [] for doc in docs: task = aprocess_doc(doc) tasks.append(task)
index_nodes = await run_jobs(tasks, show_progress=True, workers=3)
return index_nodesindex_nodes = await aprocess_docs(docs) 1%| | 1/100 [00:00<00:55, 1.78it/s]/home/loganm/llama_index_proper/llama_index/.venv/lib/python3.11/site-packages/openai/_resource.py:38: ResourceWarning: unclosed <socket.socket fd=71, family=2, type=1, proto=6, laddr=('172.25.21.0', 40832), raddr=('104.18.7.192', 443)> self._delete = client.deleteResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=73 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=71 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback 12%|█▏ | 12/100 [00:04<00:31, 2.79it/s]/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=76 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=77 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=78 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/llama_index_proper/llama_index/.venv/lib/python3.11/site-packages/openai/resources/chat/completions.py:1337: ResourceWarning: unclosed <socket.socket fd=81, family=2, type=1, proto=6, laddr=('172.25.21.0', 40848), raddr=('104.18.7.192', 443)> completions.create,ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=81 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=82 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=83 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=84 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback 21%|██ | 21/100 [00:06<00:22, 3.58it/s]/home/loganm/llama_index_proper/llama_index/.venv/lib/python3.11/site-packages/openai/_resource.py:34: ResourceWarning: unclosed <socket.socket fd=81, family=2, type=1, proto=6, laddr=('172.25.21.0', 40866), raddr=('104.18.7.192', 443)> self._get = client.getResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/llama_index_proper/llama_index/.venv/lib/python3.11/site-packages/openai/_resource.py:34: ResourceWarning: unclosed <socket.socket fd=82, family=2, type=1, proto=6, laddr=('172.25.21.0', 40868), raddr=('104.18.7.192', 443)> self._get = client.getResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=86 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback 38%|███▊ | 38/100 [00:12<00:24, 2.54it/s]/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=90 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=92 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/llama_index_proper/llama_index/.venv/lib/python3.11/site-packages/openai/_resource.py:34: ResourceWarning: unclosed <socket.socket fd=94, family=2, type=1, proto=6, laddr=('172.25.21.0', 40912), raddr=('104.18.7.192', 443)> self._get = client.getResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=94 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback 50%|█████ | 50/100 [00:17<00:19, 2.51it/s]/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=95 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=96 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=97 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback 73%|███████▎ | 73/100 [00:24<00:07, 3.42it/s]/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=101 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=102 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback 82%|████████▏ | 82/100 [00:27<00:06, 2.94it/s]/home/loganm/miniconda3/envs/llama_index/lib/python3.11/functools.py:76: ResourceWarning: unclosed <socket.socket fd=102, family=2, type=1, proto=6, laddr=('172.25.21.0', 40998), raddr=('104.18.7.192', 443)> return partial(update_wrapper, wrapped=wrapped,ResourceWarning: Enable tracemalloc to get the object allocation traceback 92%|█████████▏| 92/100 [00:32<00:03, 2.15it/s]/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=106 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback/home/loganm/miniconda3/envs/llama_index/lib/python3.11/asyncio/selector_events.py:835: ResourceWarning: unclosed transport <_SelectorSocketTransport fd=111 read=idle write=<idle, bufsize=0>> _warn(f"unclosed transport {self!r}", ResourceWarning, source=self)ResourceWarning: Enable tracemalloc to get the object allocation traceback100%|██████████| 100/100 [00:36<00:00, 2.71it/s]index_nodes[5].metadata{'state': 'open', 'year': 2024, 'month': 1, 'day': 13, 'assignee': '', 'size': 'XL'}我们将汇总后的元数据以及原始文档都加载到向量数据库中。
- 汇总元数据: 这部分内容将存入
LlamaIndex_auto集合中。 - 原始文档: 这些内容将存入
LlamaIndex_docs集合中。
通过同时存储汇总的元数据以及原始文档,我们可以执行结构化的分层检索策略。
我们加载到一个支持自动检索的向量数据库中。
这将被放入 LlamaIndex_auto
import weaviate
# cloudauth_config = weaviate.AuthApiKey( api_key="XRa15cDIkYRT7AkrpqT6jLfE4wropK1c1TGk")client = weaviate.Client( "https://llama-index-test-v0oggsoz.weaviate.network", auth_client_secret=auth_config,)
class_name = "LlamaIndex_auto"# optional: delete schemaclient.schema.delete_class(class_name)from llama_index.vector_stores.weaviate import WeaviateVectorStorefrom llama_index.core import VectorStoreIndex, StorageContext
vector_store_auto = WeaviateVectorStore( weaviate_client=client, index_name=class_name)storage_context_auto = StorageContext.from_defaults( vector_store=vector_store_auto)# Since "index_nodes" are concise summaries, we can directly feed them as objects into VectorStoreIndexindex = VectorStoreIndex( objects=index_nodes, storage_context=storage_context_auto)在本节中,我们将设置自动检索器。我们需要执行以下几个步骤。
- 定义模式: 定义向量数据库模式(例如元数据字段)。当LLM决定推断哪些元数据过滤器时,这些信息将被放入LLM输入提示中。
- 实例化 VectorIndexAutoRetriever 类:这会在我们汇总的元数据索引之上创建一个检索器,并将定义的模式作为输入。
- 定义一个包装检索器:这使我们能够将每个节点后处理成一个
IndexNode,其中包含链接回源文档的索引ID。这将允许我们在下一节中进行递归检索(这依赖于链接到下游检索器/查询引擎/其他节点的IndexNode对象)。注意:我们正在改进这个抽象概念。
运行此检索器将基于我们的文本摘要和顶级 IndeNode 对象的元数据进行检索。然后,将使用它们底层的检索器从特定的 GitHub 问题中检索内容。
from llama_index.core.vector_stores import MetadataInfo, VectorStoreInfo
vector_store_info = VectorStoreInfo( content_info="Github Issues", metadata_info=[ MetadataInfo( name="state", description="Whether the issue is `open` or `closed`", type="string", ), MetadataInfo( name="year", description="The year issue was created", type="integer", ), MetadataInfo( name="month", description="The month issue was created", type="integer", ), MetadataInfo( name="day", description="The day issue was created", type="integer", ), MetadataInfo( name="assignee", description="The assignee of the ticket", type="string", ), MetadataInfo( name="size", description="How big the issue is (XS, S, M, L, XL, XXL)", type="string", ), ],)2. 实例化 VectorIndexAutoRetriever
Section titled “2. Instantiate VectorIndexAutoRetriever”from llama_index.core.retrievers import VectorIndexAutoRetriever
retriever = VectorIndexAutoRetriever( index, vector_store_info=vector_store_info, similarity_top_k=2, empty_query_top_k=10, # if only metadata filters are specified, this is the limit verbose=True,)现在我们可以开始从 Github Issues 中检索相关上下文了!
为了完成RAG管道的设置,我们将把递归检索器与我们的RetrieverQueryEngine结合起来,除了检索到的节点外,还会生成响应。
from llama_index.core import QueryBundle
nodes = retriever.retrieve(QueryBundle("Tell me about some issues on 01/11"))Using query str: issuesUsing filters: [('day', '==', '11'), ('month', '==', '01')][1;3;38;2;11;159;203mRetrieval entering 9995: VectorIndexRetriever[0m[1;3;38;2;237;90;200mRetrieving from object VectorIndexRetriever with query issues[0m[1;3;38;2;11;159;203mRetrieval entering 9985: VectorIndexRetriever[0m[1;3;38;2;237;90;200mRetrieving from object VectorIndexRetriever with query issues[0m结果是相关文档中的源文本块。
让我们查看附加到源数据块上的日期(原始元数据中已存在)。
print(f"Number of source nodes: {len(nodes)}")nodes[0].node.metadataNumber of source nodes: 2
{'state': 'open', 'created_at': '2024-01-11T20:37:34Z', 'url': 'https://api.github.com/repos/run-llama/llama_index/issues/9995', 'source': 'https://github.com/run-llama/llama_index/pull/9995', 'labels': ['size:XXL'], 'index_id': 9995}接入 RetrieverQueryEngine
Section titled “Plug into RetrieverQueryEngine”我们接入检索查询引擎来合成结果。
from llama_index.core.query_engine import RetrieverQueryEnginefrom llama_index.llms.openai import OpenAI
llm = OpenAI(model="gpt-3.5-turbo")
query_engine = RetrieverQueryEngine.from_args(retriever, llm=llm)response = query_engine.query("Tell me about some issues on 01/11")Using query str: issuesUsing filters: [('day', '==', '11'), ('month', '==', '01')][1;3;38;2;11;159;203mRetrieval entering 9995: VectorIndexRetriever[0m[1;3;38;2;237;90;200mRetrieving from object VectorIndexRetriever with query issues[0m[1;3;38;2;11;159;203mRetrieval entering 9985: VectorIndexRetriever[0m[1;3;38;2;237;90;200mRetrieving from object VectorIndexRetriever with query issues[0mprint(str(response))There are two issues that were created on 01/11. The first issue is related to ensuring backwards compatibility with the new Pinecone client version bifurcation. The second issue is a feature request to implement the Language Agent Tree Search (LATS) agent in llama-index.response = query_engine.query( "Tell me about some open issues related to agents")Using query str: agentsUsing filters: [('state', '==', 'open')][1;3;38;2;11;159;203mRetrieval entering 10058: VectorIndexRetriever[0m[1;3;38;2;237;90;200mRetrieving from object VectorIndexRetriever with query agents[0m[1;3;38;2;11;159;203mRetrieval entering 9899: VectorIndexRetriever[0m[1;3;38;2;237;90;200mRetrieving from object VectorIndexRetriever with query agents[0mprint(str(response))There are two open issues related to agents. One issue is about adding context for agents, updating a stale link, and adding a notebook to demo a react agent with context. The other issue is a feature request for parallelism when using the top agent from a multi-document agent while comparing multiple documents.这向您展示了如何在文档摘要之上创建一个结构化检索层,使您能够根据用户查询动态拉取相关文档。
您可能会注意到这与我们的多文档智能体之间的相似性。两种架构都旨在实现强大的多文档检索。
本笔记本的目标是展示如何在多文档环境中应用结构化查询。您实际上也可以将此自动检索算法应用于我们的多智能体设置。多智能体设置主要侧重于在跨文档和每个文档中添加智能体推理,允许使用思维链进行多部分查询。