Deep Lake 向量存储快速入门
Deep Lake 可以通过 pip 进行安装。
%pip install llama-index-vector-stores-deeplake!pip install llama-index!pip install deeplake接下来,让我们导入所需的模块并设置必要的环境变量:
import osimport textwrap
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader, Documentfrom llama_index.vector_stores.deeplake import DeepLakeVectorStore
os.environ["OPENAI_API_KEY"] = "sk-********************************"os.environ["ACTIVELOOP_TOKEN"] = "********************************"我们将把保罗·格雷厄姆的一篇文章嵌入并存储在本地的Deep Lake向量存储中。首先,我们将数据下载到名为data/paul_graham的目录中
import urllib.request
urllib.request.urlretrieve( "https://raw.githubusercontent.com/run-llama/llama_index/main/docs/examples/data/paul_graham/paul_graham_essay.txt", "data/paul_graham/paul_graham_essay.txt",)我们现在可以从源数据文件创建文档。
# load documentsdocuments = SimpleDirectoryReader("./data/paul_graham/").load_data()print( "Document ID:", documents[0].doc_id, "Document Hash:", documents[0].hash,)Document ID: a98b6686-e666-41a9-a0bc-b79f0d666bde Document Hash: beaa54b3e9cea641e91e6975d2207af4f4200f4b2d629725d688f272372ce5bb最后,让我们创建 Deep Lake 向量存储库并用数据填充它。我们使用默认的张量配置,该配置创建包含 text (str)、metadata(json)、id (str, auto-populated)、embedding (float32) 的张量。在此处了解更多关于张量可定制性的信息。
from llama_index.core import StorageContext
dataset_path = "./dataset/paul_graham"
# Create an index over the documentsvector_store = DeepLakeVectorStore(dataset_path=dataset_path, overwrite=True)storage_context = StorageContext.from_defaults(vector_store=vector_store)index = VectorStoreIndex.from_documents( documents, storage_context=storage_context)Uploading data to deeplake dataset.
100%|██████████| 22/22 [00:00<00:00, 684.80it/s]
Dataset(path='./dataset/paul_graham', tensors=['text', 'metadata', 'embedding', 'id'])
tensor htype shape dtype compression ------- ------- ------- ------- ------- text text (22, 1) str None metadata json (22, 1) str None embedding embedding (22, 1536) float32 None id text (22, 1) str NoneDeep Lake 提供高度灵活的向量搜索和混合搜索选项 这些教程中详细讨论。在本快速入门中,我们展示一个使用默认选项的简单示例。
query_engine = index.as_query_engine()response = query_engine.query( "What did the author learn?",)print(textwrap.fill(str(response), 100)) The author learned that working on things that are not prestigious can be a good thing, as it canlead to discovering something real and avoiding the wrong track. The author also learned thatignorance can be beneficial, as it can lead to discovering something new and unexpected. The authoralso learned the importance of working hard, even at the parts of the job they don't like, in orderto set an example for others. The author also learned the value of unsolicited advice, as it can bebeneficial in unexpected ways, such as when Robert Morris suggested that the author should make sureY Combinator wasn't the last cool thing they did.response = query_engine.query("What was a hard moment for the author?")print(textwrap.fill(str(response), 100))The author experienced a hard moment when one of his programs on the IBM 1401 computer did notterminate. This was a social as well as a technical error, as the data center manager's expressionmade clear.query_engine = index.as_query_engine()response = query_engine.query("What was a hard moment for the author?")print(textwrap.fill(str(response), 100))The author experienced a hard moment when one of his programs on the IBM 1401 computer did notterminate. This was a social as well as a technical error, as the data center manager's expressionmade clear.要查找要删除的文档ID,您可以直接查询底层的deeplake数据集
import deeplake
ds = deeplake.load(dataset_path)
idx = ds.id[0].numpy().tolist()idx./dataset/paul_graham loaded successfully.
['42f8220e-673d-4c65-884d-5a48a1a15b03']index.delete(idx[0])