Upstash 向量存储
我们将了解如何使用 LlamaIndex 与 Upstash Vector 进行交互!
! pip install -q llama-index upstash-vectorfrom llama_index.core import VectorStoreIndex, SimpleDirectoryReaderfrom llama_index.core.vector_stores import UpstashVectorStorefrom llama_index.core import StorageContextimport textwrapimport openai# Setup the OpenAI APIopenai.api_key = "sk-..."# Download data! mkdir -p 'data/paul_graham/'! wget 'https://raw.githubusercontent.com/run-llama/llama_index/main/docs/examples/data/paul_graham/paul_graham_essay.txt' -O 'data/paul_graham/paul_graham_essay.txt'--2024-02-03 20:04:25-- https://raw.githubusercontent.com/run-llama/llama_index/main/docs/examples/data/paul_graham/paul_graham_essay.txtResolving raw.githubusercontent.com (raw.githubusercontent.com)... 185.199.108.133, 185.199.109.133, 185.199.110.133, ...Connecting to raw.githubusercontent.com (raw.githubusercontent.com)|185.199.108.133|:443... connected.HTTP request sent, awaiting response... 200 OKLength: 75042 (73K) [text/plain]Saving to: ‘data/paul_graham/paul_graham_essay.txt’
data/paul_graham/pa 100%[===================>] 73.28K --.-KB/s in 0.01s
2024-02-03 20:04:25 (5.96 MB/s) - ‘data/paul_graham/paul_graham_essay.txt’ saved [75042/75042]现在,我们可以使用LlamaIndex的SimpleDirectoryReader加载文档
documents = SimpleDirectoryReader("./data/paul_graham/").load_data()
print("# Documents:", len(documents))# Documents: 1要在 Upstash 上创建索引,请访问 https://console.upstash.com/vector,创建一个具有 1536 维度和 Cosine 距离度量的索引。复制下方的 URL 和令牌
vector_store = UpstashVectorStore(url="https://...", token="...")
storage_context = StorageContext.from_defaults(vector_store=vector_store)index = VectorStoreIndex.from_documents( documents, storage_context=storage_context)现在我们已经成功创建了一个索引,并用论文中的向量填充了它!数据需要一点时间来建立索引,之后就可以进行查询了。
query_engine = index.as_query_engine()res1 = query_engine.query("What did the author learn?")print(textwrap.fill(str(res1), 100))
print("\n")
res2 = query_engine.query("What is the author's opinion on startups?")print(textwrap.fill(str(res2), 100))The author learned that the study of philosophy in college did not live up to their expectations.They found that other fields took up most of the space of ideas, leaving little room for what theyperceived as the ultimate truths that philosophy was supposed to explore. As a result, they decidedto switch to studying AI.
The author's opinion on startups is that they are in need of help and support, especially in thebeginning stages. The author believes that founders of startups are often helpless and face variouschallenges, such as getting incorporated and understanding the intricacies of running a company. Theauthor's investment firm, Y Combinator, aims to provide seed funding and comprehensive support tostartups, offering them the guidance and resources they need to succeed.你可以传递 MetadataFilters 和你的 VectorStoreQuery 来筛选从Upstash向量存储返回的节点。
import os
from llama_index.vector_stores.upstash import UpstashVectorStorefrom llama_index.core.vector_stores.types import ( MetadataFilter, MetadataFilters, FilterOperator,)
vector_store = UpstashVectorStore( url=os.environ.get("UPSTASH_VECTOR_URL") or "", token=os.environ.get("UPSTASH_VECTOR_TOKEN") or "",)
index = VectorStoreIndex.from_vector_store(vector_store=vector_store)
filters = MetadataFilters( filters=[ MetadataFilter( key="author", value="Marie Curie", operator=FilterOperator.EQ ) ],)
retriever = index.as_retriever(filters=filters)
retriever.retrieve("What is inception about?")我们还可以将多个 MetadataFilters 与 AND 或 OR 条件结合使用
from llama_index.core.vector_stores import FilterOperator, FilterCondition
filters = MetadataFilters( filters=[ MetadataFilter( key="theme", value=["Fiction", "Horror"], operator=FilterOperator.IN, ), MetadataFilter(key="year", value=1997, operator=FilterOperator.GT), ], condition=FilterCondition.AND,)
retriever = index.as_retriever(filters=filters)retriever.retrieve("Harry Potter?")