Chroma
Chroma 是一个专注于开发者生产力和幸福感的AI原生开源向量数据库。Chroma基于Apache 2.0许可证发布。
Chroma是完全类型化、完全测试和完全文档化的。
使用以下命令安装 Chroma:
pip install chromadbChroma 可在多种模式下运行。以下展示了与 LlamaIndex 集成的各种模式示例。
in-memory- 在 Python 脚本或 Jupyter 笔记本中in-memory with persistence- 在脚本或笔记本中,并保存/加载到磁盘in a docker container- 作为在您本地机器或云端运行的服务器
与任何其他数据库一样,您可以:
.add.get.update.upsert.delete.peek- 并且
.query执行相似性搜索。
查看完整文档请访问文档。
在这个基础示例中,我们获取Paul Graham的文章,将其分割成多个片段,使用开源嵌入模型进行嵌入处理,加载到Chroma中,然后进行查询。
如果您在 Colab 上打开这个笔记本,您可能需要安装 LlamaIndex 🦙。
%pip install llama-index-vector-stores-chroma%pip install llama-index-embeddings-huggingface!pip install llama-index创建 Chroma 索引
Section titled “Creating a Chroma Index”# !pip install llama-index chromadb --quiet# !pip install chromadb# !pip install sentence-transformers# !pip install pydantic==1.10.11# importfrom llama_index.core import VectorStoreIndex, SimpleDirectoryReaderfrom llama_index.vector_stores.chroma import ChromaVectorStorefrom llama_index.core import StorageContextfrom llama_index.embeddings.huggingface import HuggingFaceEmbeddingfrom IPython.display import Markdown, displayimport chromadb# set up OpenAIimport osimport getpass
os.environ["OPENAI_API_KEY"] = getpass.getpass("OpenAI API Key:")import openai
openai.api_key = os.environ["OPENAI_API_KEY"]下载数据
!mkdir -p 'data/paul_graham/'!wget 'https://raw.githubusercontent.com/run-llama/llama_index/main/docs/examples/data/paul_graham/paul_graham_essay.txt' -O 'data/paul_graham/paul_graham_essay.txt'# create client and a new collectionchroma_client = chromadb.EphemeralClient()chroma_collection = chroma_client.create_collection("quickstart")
# define embedding functionembed_model = HuggingFaceEmbedding(model_name="BAAI/bge-base-en-v1.5")
# load documentsdocuments = SimpleDirectoryReader("./data/paul_graham/").load_data()
# set up ChromaVectorStore and load in datavector_store = ChromaVectorStore(chroma_collection=chroma_collection)storage_context = StorageContext.from_defaults(vector_store=vector_store)index = VectorStoreIndex.from_documents( documents, storage_context=storage_context, embed_model=embed_model)
# Query Dataquery_engine = index.as_query_engine()response = query_engine.query("What did the author do growing up?")display(Markdown(f"<b>{response}</b>"))/Users/loganmarkewich/llama_index/llama-index/lib/python3.9/site-packages/tqdm/auto.py:21: TqdmWarning: IProgress not found. Please update jupyter and ipywidgets. See https://ipywidgets.readthedocs.io/en/stable/user_install.html from .autonotebook import tqdm as notebook_tqdm/Users/loganmarkewich/llama_index/llama-index/lib/python3.9/site-packages/bitsandbytes/cextension.py:34: UserWarning: The installed version of bitsandbytes was compiled without GPU support. 8-bit optimizers, 8-bit multiplication, and GPU quantization are unavailable. warn("The installed version of bitsandbytes was compiled without GPU support. "
'NoneType' object has no attribute 'cadam32bit_grad_fp32'作者从小就开始从事写作和编程。他们撰写短篇小说,并尝试在IBM 1401计算机上编写程序。后来,他们获得了一台微型计算机,并开始更广泛地进行编程。
扩展之前的示例,如果您想要保存到磁盘,只需初始化 Chroma 客户端并传入您希望数据保存的目录路径。
Caution: Chroma会尽最大努力自动将数据保存到磁盘,但多个内存中的客户端可能会相互覆盖工作内容。最佳实践是,在任何给定时间,每个路径只运行一个客户端。
# save to disk
db = chromadb.PersistentClient(path="./chroma_db")chroma_collection = db.get_or_create_collection("quickstart")vector_store = ChromaVectorStore(chroma_collection=chroma_collection)storage_context = StorageContext.from_defaults(vector_store=vector_store)
index = VectorStoreIndex.from_documents( documents, storage_context=storage_context, embed_model=embed_model)
# load from diskdb2 = chromadb.PersistentClient(path="./chroma_db")chroma_collection = db2.get_or_create_collection("quickstart")vector_store = ChromaVectorStore(chroma_collection=chroma_collection)index = VectorStoreIndex.from_vector_store( vector_store, embed_model=embed_model,)
# Query Data from the persisted indexquery_engine = index.as_query_engine()response = query_engine.query("What did the author do growing up?")display(Markdown(f"<b>{response}</b>"))作者从小就开始从事写作和编程。他们撰写短篇故事,并尝试在IBM 1401计算机上编写程序。后来,他们获得了一台微型计算机,并开始编写游戏和文字处理软件。
基础示例(使用Docker容器)
Section titled “Basic Example (using the Docker Container)”你也可以在Docker容器中单独运行Chroma服务器,创建一个客户端连接它,然后将其传递给LlamaIndex。
以下是克隆、构建和运行 Docker 镜像的方法:
git clone git@github.com:chroma-core/chroma.gitdocker-compose up -d --build# create the chroma client and add our dataimport chromadb
remote_db = chromadb.HttpClient()chroma_collection = remote_db.get_or_create_collection("quickstart")vector_store = ChromaVectorStore(chroma_collection=chroma_collection)storage_context = StorageContext.from_defaults(vector_store=vector_store)
index = VectorStoreIndex.from_documents( documents, storage_context=storage_context, embed_model=embed_model)# Query Data from the Chroma Docker indexquery_engine = index.as_query_engine()response = query_engine.query("What did the author do growing up?")display(Markdown(f"<b>{response}</b>"))在构建实际应用程序的过程中,您不仅需要添加数据,还需要更新和删除数据。
Chroma让用户提供ids来简化此处的簿记工作。ids可以是文件名,也可以是类似filename_paragraphNumber的组合哈希值等。
以下是一个基础示例,展示如何进行各种操作:
doc_to_update = chroma_collection.get(limit=1)doc_to_update["metadatas"][0] = { **doc_to_update["metadatas"][0], **{"author": "Paul Graham"},}chroma_collection.update( ids=[doc_to_update["ids"][0]], metadatas=[doc_to_update["metadatas"][0]])updated_doc = chroma_collection.get(limit=1)print(updated_doc["metadatas"][0])
# delete the last documentprint("count before", chroma_collection.count())chroma_collection.delete(ids=[doc_to_update["ids"][0]])print("count after", chroma_collection.count()){'_node_content': '{"id_": "be08c8bc-f43e-4a71-ba64-e525921a8319", "embedding": null, "metadata": {}, "excluded_embed_metadata_keys": [], "excluded_llm_metadata_keys": [], "relationships": {"1": {"node_id": "2cbecdbb-0840-48b2-8151-00119da0995b", "node_type": null, "metadata": {}, "hash": "4c702b4df575421e1d1af4b1fd50511b226e0c9863dbfffeccb8b689b8448f35"}, "3": {"node_id": "6a75604a-fa76-4193-8f52-c72a7b18b154", "node_type": null, "metadata": {}, "hash": "d6c408ee1fbca650fb669214e6f32ffe363b658201d31c204e85a72edb71772f"}}, "hash": "b4d0b960aa09e693f9dc0d50ef46a3d0bf5a8fb3ac9f3e4bcf438e326d17e0d8", "text": "", "start_char_idx": 0, "end_char_idx": 4050, "text_template": "{metadata_str}\\n\\n{content}", "metadata_template": "{key}: {value}", "metadata_seperator": "\\n"}', 'author': 'Paul Graham', 'doc_id': '2cbecdbb-0840-48b2-8151-00119da0995b', 'document_id': '2cbecdbb-0840-48b2-8151-00119da0995b', 'ref_doc_id': '2cbecdbb-0840-48b2-8151-00119da0995b'}count before 20count after 19