MyScale向量存储
在本笔记本中,我们将快速演示如何使用MyScaleVectorStore。
如果您在 Colab 上打开这个笔记本,您可能需要安装 LlamaIndex 🦙。
%pip install llama-index-vector-stores-myscale!pip install llama-index创建 MyScale 客户端
Section titled “Creating a MyScale Client”import loggingimport sys
logging.basicConfig(stream=sys.stdout, level=logging.INFO)logging.getLogger().addHandler(logging.StreamHandler(stream=sys.stdout))from os import environimport clickhouse_connect
environ["OPENAI_API_KEY"] = "sk-*"
# initialize clientclient = clickhouse_connect.get_client( host="YOUR_CLUSTER_HOST", port=8443, username="YOUR_USERNAME", password="YOUR_CLUSTER_PASSWORD",)加载文档,使用MyScaleVectorStore构建并存储向量存储索引
Section titled “Load documents, build and store the VectorStoreIndex with MyScaleVectorStore”这里我们将使用一套保罗·格雷厄姆的文章作为文本来源,将其转化为嵌入向量,存储在MyScaleVectorStore中,并通过查询为大语言模型问答循环寻找上下文。
from llama_index.core import VectorStoreIndex, SimpleDirectoryReaderfrom llama_index.vector_stores.myscale import MyScaleVectorStorefrom IPython.display import Markdown, display# load documentsdocuments = SimpleDirectoryReader("../data/paul_graham").load_data()print("Document ID:", documents[0].doc_id)print("Number of Documents: ", len(documents))Document ID: a5f2737c-ed18-4e5d-ab9a-75955edb816dNumber of Documents: 1下载数据
!mkdir -p 'data/paul_graham/'!wget 'https://raw.githubusercontent.com/run-llama/llama_index/main/docs/examples/data/paul_graham/paul_graham_essay.txt' -O 'data/paul_graham/paul_graham_essay.txt'您可以使用 SimpleDirectoryReader 单独处理您的文件:
loader = SimpleDirectoryReader("./data/paul_graham/")documents = loader.load_data()for file in loader.input_files: print(file) # Here is where you would do any preprocessing../data/paul_graham/paul_graham_essay.txt# initialize with metadata filter and store indexesfrom llama_index.core import StorageContext
for document in documents: document.metadata = {"user_id": "123", "favorite_color": "blue"}vector_store = MyScaleVectorStore(myscale_client=client)storage_context = StorageContext.from_defaults(vector_store=vector_store)index = VectorStoreIndex.from_documents( documents, storage_context=storage_context)现在 MyScale 向量存储支持过滤搜索和混合搜索
import textwrap
from llama_index.core.vector_stores import ExactMatchFilter, MetadataFilters
# set Logging to DEBUG for more detailed outputsquery_engine = index.as_query_engine( filters=MetadataFilters( filters=[ ExactMatchFilter(key="user_id", value="123"), ] ), similarity_top_k=2, vector_store_query_mode="hybrid",)response = query_engine.query("What did the author learn?")print(textwrap.fill(str(response), 100))for document in documents: index.delete_ref_doc(document.doc_id)