使用LlamaIndex和Cleanlab构建可信赖的RAG系统
大型语言模型偶尔会产生错误的答案幻觉,特别是在其训练数据中未充分支持的问题上。虽然各组织正在采用检索增强生成(RAG)技术为大型语言模型提供专有数据支持,但错误的RAG响应仍然是一个问题。
本教程展示如何构建可信赖的RAG应用程序:使用Cleanlab为每个LLM响应的可信度评分,并通过评估特定RAG组件来诊断为何响应不可信赖。
由最先进的不确定性估计技术驱动,Cleanlab可信度评分帮助您自动捕获任何LLM应用中的错误响应。可信度评分实时进行,无需任何数据标注或模型训练工作。Cleanlab为特定RAG组件(如检索到的上下文)提供额外的实时评估,帮助您追溯为何RAG响应会出现错误。Cleanlab让您轻松预防RAG应用中的不准确响应,避免失去用户的信任。
本教程需要:
- Cleanlab API 密钥:在 tlm.cleanlab.ai/ 注册获取免费密钥
- OpenAI API 密钥:用于向大语言模型发送补全请求
首先安装所需的依赖项。
%pip install llama-index cleanlab-tlmimport os, refrom typing import List, ClassVarimport pandas as pd
from llama_index.llms.openai import OpenAIfrom llama_index.embeddings.openai import OpenAIEmbedding
from cleanlab_tlm import TrustworthyRAG, Eval, get_default_evals使用其API密钥初始化OpenAI客户端。
os.environ["OPENAI_API_KEY"] = "<your-openai-api-key>"
llm = OpenAI(model="gpt-4o-mini")embed_model = OpenAIEmbedding(embed_batch_size=10)现在,我们使用默认配置初始化 Cleanlab 客户端。通过调整可选配置,您可以获得更好的检测精度和延迟。
os.environ["CLEANLAB_TLM_API_KEY"] = "<your-cleanlab-api-key"
trustworthy_rag = ( TrustworthyRAG()) # Optional configurations can improve accuracy/latency本教程使用英伟达2024财年第一季度财报作为示例数据源,用于填充RAG应用的知识库。
!wget -nc 'https://cleanlab-public.s3.amazonaws.com/Datasets/NVIDIA_Financial_Results_Q1_FY2024.md'!mkdir -p ./data!mv NVIDIA_Financial_Results_Q1_FY2024.md data/--2025-05-07 16:13:28-- https://cleanlab-public.s3.amazonaws.com/Datasets/NVIDIA_Financial_Results_Q1_FY2024.mdResolving cleanlab-public.s3.amazonaws.com (cleanlab-public.s3.amazonaws.com)... 54.231.236.193, 16.182.70.65, 52.217.14.204, ...Connecting to cleanlab-public.s3.amazonaws.com (cleanlab-public.s3.amazonaws.com)|54.231.236.193|:443... connected.HTTP request sent, awaiting response... 200 OKLength: 7379 (7.2K) [binary/octet-stream]Saving to: ‘NVIDIA_Financial_Results_Q1_FY2024.md’
NVIDIA_Financial_Re 100%[===================>] 7.21K --.-KB/s in 0s
2025-05-07 16:13:28 (97.7 MB/s) - ‘NVIDIA_Financial_Results_Q1_FY2024.md’ saved [7379/7379]with open( "data/NVIDIA_Financial_Results_Q1_FY2024.md", "r", encoding="utf-8") as file: data = file.read()
print(data[:200])# NVIDIA Announces Financial Results for First Quarter Fiscal 2024
NVIDIA (NASDAQ: NVDA) today reported revenue for the first quarter ended April 30, 2023, of $7.19 billion, down 13% from a year ago现在让我们使用LlamaIndex构建一个简单的RAG流水线。我们已经为LLM和嵌入模型初始化了OpenAI API。
from llama_index.core import Settings, VectorStoreIndex, SimpleDirectoryReader
Settings.llm = llmSettings.embed_model = embed_model加载数据并创建索引 + 查询引擎
Section titled “Load Data and Create Index + Query Engine”让我们基于刚才提取的文档创建一个索引。本教程中我们沿用LlamaIndex的默认索引设置。
documents = SimpleDirectoryReader("data").load_data()# Optional step since we're loading just one data filefor doc in documents: doc.excluded_llm_metadata_keys.append( "file_path" ) # file_path wouldn't be a useful metadata to add to LLM's context since our datasource contains just 1 fileindex = VectorStoreIndex.from_documents(documents)生成的索引用于为数据上的查询引擎提供支持。
query_engine = index.as_query_engine()请注意,Cleanlab 对用于 RAG 的索引和查询引擎是不可知的,并且与您为系统这些组件所做的任何选择兼容。
此外,您可以直接在现有的自定义RAG流程中使用Cleanlab(使用任何其他LLM生成器,无论是否支持流式处理)。
Cleanlab只需要发送给您的LLM的提示(包括系统指令、检索到的上下文、用户查询等)以及生成的响应。
我们定义一个事件处理器,用于存储LlamaIndex发送给LLM的提示词。更多详细信息请参考插桩文档。
from llama_index.core.instrumentation import get_dispatcherfrom llama_index.core.instrumentation.events import BaseEventfrom llama_index.core.instrumentation.event_handlers import BaseEventHandlerfrom llama_index.core.instrumentation.events.llm import LLMPredictStartEvent
class PromptEventHandler(BaseEventHandler): events: ClassVar[List[BaseEvent]] = [] PROMPT_TEMPLATE: str = ""
@classmethod def class_name(cls) -> str: return "PromptEventHandler"
def handle(self, event) -> None: if isinstance(event, LLMPredictStartEvent): self.PROMPT_TEMPLATE = event.template.default_template.template self.events.append(event)
# Root dispatcherroot_dispatcher = get_dispatcher()
# Register event handlerevent_handler = PromptEventHandler()root_dispatcher.add_event_handler(event_handler)对于每个查询,我们可以从event_handler.PROMPT_TEMPLATE获取提示。让我们看看它的实际效果。
使用我们的RAG应用程序
Section titled “Use our RAG application”现在向量数据库已加载文本块及其对应的嵌入向量,我们可以开始查询它以回答问题。
query = "What was NVIDIA's total revenue in the first quarter of fiscal 2024?"
response = query_engine.query(query)print(response)NVIDIA's total revenue in the first quarter of fiscal 2024 was $7.19 billion.这个回答对于我们的简单查询确实是正确的。让我们看看 LlamaIndex 为此查询检索到的文档片段,从中我们可以轻松验证这个回答是正确的。
def get_retrieved_context(response, print_chunks=False): if isinstance(response, list): texts = [node.text for node in response] else: texts = [src.node.text for src in response.source_nodes]
if print_chunks: for idx, text in enumerate(texts): print(f"--- Chunk {idx + 1} ---\n{text[:200]}...") return "\n".join(texts)context_str = get_retrieved_context(response, True)--- Chunk 1 ---# NVIDIA Announces Financial Results for First Quarter Fiscal 2024
NVIDIA (NASDAQ: NVDA) today reported revenue for the first quarter ended April 30, 2023, of $7.19 billion, down 13% from a year ago ...--- Chunk 2 ---- **Gross Margins**: GAAP and non-GAAP gross margins are expected to be 68.6% and 70.0%, respectively, plus or minus 50 basis points.- **Operating Expenses**: GAAP and non-GAAP operating expenses are...使用 Cleanlab 添加信任层
Section titled “Add a Trust Layer with Cleanlab”让我们添加一个检测层来实时标记不可靠的RAG响应。TrustworthyRAG运行Cleanlab最先进的不确定性估计器——可信语言模型,以提供可信度评分,表明您的RAG响应正确的总体置信度。
为了诊断为何响应不可信,TrustworthyRAG 可以对特定 RAG 组件运行额外评估。让我们看看它默认运行的评估内容:
default_evals = get_default_evals()for eval in default_evals: print(f"{eval.name}")context_sufficiencyresponse_groundednessresponse_helpfulnessquery_ease每个评估返回一个介于0-1之间的分数(越高越好),用于评估您RAG系统的不同方面:
-
context_sufficiency: 评估检索到的上下文是否包含足够信息来完整回答查询。低分表示上下文中缺少关键信息(可能是由于检索效果不佳或文档缺失所致)。
-
response_groundedness: 评估回复中陈述的主张/信息是否明确得到所提供上下文的支持。
-
response_helpfulness: 评估回复是否以有帮助的方式尝试回答用户查询。
-
query_ease: 评估用户查询是否看起来易于AI系统正确处理。复杂、模糊、棘手或带有不满情绪的查询会获得较低分数。
要运行 TrustworthyRAG,我们需要发送给大语言模型的提示,其中包含系统消息、检索到的文本块、用户查询以及大语言模型的响应。 上述定义的事件处理器提供了这个提示。 让我们定义一个辅助函数来运行 Cleanlab 的检测。
# Helper function to run real-time Evalsdef get_eval(query, response, event_handler, evaluator): # Get context used by LLM to generate response context = get_retrieved_context(response) # Get prompt template used to build the prompt pt = event_handler.PROMPT_TEMPLATE # Build prompt full_prompt = pt.format(context_str=context, query_str=query)
eval_result = evaluator.score( query=query, context=context, response=response.response, prompt=full_prompt, ) # Evaluate the response using TrustworthyRAG print("### Evaluation results:") for metric, value in eval_result.items(): print(f"{metric}: {value['score']}")
# Helper function run end-to-end RAGdef get_answer(query, evaluator=trustworthy_rag, event_handler=event_handler): response = query_engine.query(query)
print( f"### Query:\n{query}\n\n### Trimmed Context:\n{get_retrieved_context(response)[:300]}..." ) print(f"\n### Generated response:\n{response.response}\n")
get_eval(query, response, event_handler, evaluator)get_eval(query, response, event_handler, trustworthy_rag)### Evaluation results:trustworthiness: 1.0context_sufficiency: 0.9975124377856721response_groundedness: 0.9975124378045552response_helpfulness: 0.9975124367363073query_ease: 0.9975071027792313分析: 较高的 trustworthiness_score 表明此响应非常可信,即非虚构且很可能正确。此处检索到的上下文足以回答此查询,这通过较高的 context_sufficiency 得分反映出来。较高的 query_ease 得分也表明这是一个简单的查询。
现在让我们运行一个具有挑战性的查询,该查询无法通过我们RAG应用程序知识库中唯一的文档来回答。
get_answer( "How does the report explain why NVIDIA's Gaming revenue decreased year over year?")### Query:How does the report explain why NVIDIA's Gaming revenue decreased year over year?
### Trimmed Context:# NVIDIA Announces Financial Results for First Quarter Fiscal 2024
NVIDIA (NASDAQ: NVDA) today reported revenue for the first quarter ended April 30, 2023, of $7.19 billion, down 13% from a year ago and up 19% from the previous quarter.
- **Quarterly revenue** of $7.19 billion, up 19% from the pre...
### Generated response:The report indicates that NVIDIA's Gaming revenue decreased year over year by 38%, which is attributed to a combination of factors, although specific reasons are not detailed. The context highlights that the revenue for the first quarter was $2.24 billion, down from the previous year, while it did show an increase of 22% from the previous quarter. This suggests that while there may have been a seasonal or cyclical recovery, the overall year-over-year decline reflects challenges in the gaming segment during that period.
### Evaluation results:trustworthiness: 0.8018049078305449context_sufficiency: 0.26134514055082803response_groundedness: 0.8147481620994604response_helpfulness: 0.28647897539109127query_ease: 0.952132218665045分析:生成器大语言模型通过提供可靠响应来避免推测,这体现在较高的trustworthiness_score值上。较低的context_sufficiency分数反映出检索上下文存在不足,而较低的response_helpfulness值表明该响应实际上并未解答用户的查询。
让我们看看我们的RAG系统如何回应另一个具有挑战性的问题。
get_answer( "How much did Nvidia's revenue decrease this quarter vs last quarter, in dollars?")### Query:How much did Nvidia's revenue decrease this quarter vs last quarter, in dollars?
### Trimmed Context:# NVIDIA Announces Financial Results for First Quarter Fiscal 2024
NVIDIA (NASDAQ: NVDA) today reported revenue for the first quarter ended April 30, 2023, of $7.19 billion, down 13% from a year ago and up 19% from the previous quarter.
- **Quarterly revenue** of $7.19 billion, up 19% from the pre...
### Generated response:NVIDIA's revenue decreased by $1.10 billion this quarter compared to the last quarter.
### Evaluation results:trustworthiness: 0.572441384819641context_sufficiency: 0.9974990573223977response_groundedness: 0.006136548076912901response_helpfulness: 0.997512230771839query_ease: 0.8018484929561781分析: 生成的回复错误地声称英伟达本季度收入下降,而实际上引用的报告指出季度环比增长19%。
Cleanlab的低可信度评分帮助我们实时自动捕获这个不正确的RAG响应!为了追溯这个响应不可信的根本原因,我们看到response_groundedness评分较低,这表明我们的LLM模型因编造此虚假信息而受到指责。
让我们再试一个:
get_answer( "If NVIDIA's Data Center segment maintains its Q1 FY2024 quarter-over-quarter growth rate for the next four quarters, what would be its projected annual revenue?")### Query:If NVIDIA's Data Center segment maintains its Q1 FY2024 quarter-over-quarter growth rate for the next four quarters, what would be its projected annual revenue?
### Trimmed Context:# NVIDIA Announces Financial Results for First Quarter Fiscal 2024
NVIDIA (NASDAQ: NVDA) today reported revenue for the first quarter ended April 30, 2023, of $7.19 billion, down 13% from a year ago and up 19% from the previous quarter.
- **Quarterly revenue** of $7.19 billion, up 19% from the pre...
### Generated response:If NVIDIA's Data Center segment maintains its quarter-over-quarter growth rate of 18% from Q1 FY2024 for the next four quarters, the projected revenue for the next four quarters can be calculated as follows:
1. Q1 FY2024 revenue: $4.28 billion2. Q2 FY2024 projected revenue: $4.28 billion * 1.18 = $5.05 billion3. Q3 FY2024 projected revenue: $5.05 billion * 1.18 = $5.95 billion4. Q4 FY2024 projected revenue: $5.95 billion * 1.18 = $7.02 billion5. Q5 FY2024 projected revenue: $7.02 billion * 1.18 = $8.27 billion
Now, summing these revenues for the five quarters (including Q1 FY2024):
- Total projected revenue = $4.28 billion + $5.05 billion + $5.95 billion + $7.02 billion + $8.27 billion = $30.57 billion
Therefore, the projected annual revenue for the Data Center segment would be approximately $30.57 billion.
### Evaluation results:trustworthiness: 0.23124932848015411context_sufficiency: 0.9299227307108295response_groundedness: 0.31247206392894905response_helpfulness: 0.9975055879546202query_ease: 0.7724662723193096分析: 在审阅生成的回复时,我们发现它高估(汇总了第一季度的财务数据)了预期收入。Cleanlab再次通过其较低的trustworthiness_score值帮助我们自动捕获这个错误回复。根据额外的评估结果,该问题的根本原因似乎仍然是LLM模型未能将其回复建立在检索到的上下文基础上。
您还可以指定自定义评估以衡量特定标准,并将其与默认评估相结合,对您的RAG系统进行全面/定制化评估。
例如,以下是如何创建并运行一个自定义评估,用于检查生成响应的简洁性。
conciseness_eval = Eval( name="response_conciseness", criteria="Evaluate whether the Generated response is concise and to the point without unnecessary verbosity or repetition. A good response should be brief but comprehensive, covering all necessary information without extra words or redundant explanations.", response_identifier="Generated Response",)
# Combine default evals with a custom evalcombined_evals = get_default_evals() + [conciseness_eval]
# Initialize TrustworthyRAG with combined evalscombined_trustworthy_rag = TrustworthyRAG(evals=combined_evals)get_answer( "What significant transitions did Jensen comment on?", evaluator=combined_trustworthy_rag,)### Query:What significant transitions did Jensen comment on?
### Trimmed Context:# NVIDIA Announces Financial Results for First Quarter Fiscal 2024
NVIDIA (NASDAQ: NVDA) today reported revenue for the first quarter ended April 30, 2023, of $7.19 billion, down 13% from a year ago and up 19% from the previous quarter.
- **Quarterly revenue** of $7.19 billion, up 19% from the pre...
### Generated response:Jensen Huang commented on the significant transitions the computer industry is undergoing, particularly in the areas of accelerated computing and generative AI.
### Evaluation results:trustworthiness: 0.9810004109697261context_sufficiency: 0.9902170786836257response_groundedness: 0.9975123614036665response_helpfulness: 0.9420916924086002query_ease: 0.5334109647649754response_conciseness: 0.842668665703559将您的LLM替换为Cleanlab的
Section titled “Replace your LLM with Cleanlab’s”除了评估已从您的LLM生成的响应外,Cleanlab还可以同时生成并评估响应(使用众多支持的模型之一)。
您可以通过调用trustworthy_rag.generate(query=query, context=context, prompt=full_prompt)来实现此功能
这将替换您RAG系统中的自有LLM,可能更加便捷/准确/快速。
让我们将我们的OpenAI LLM替换为调用Cleanlab的端点:
query = "How much did Nvidia's revenue decrease this quarter vs last quarter, in dollars?"relevant_chunks = query_engine.retrieve(query)context = get_retrieved_context(relevant_chunks)print(f"### Query:\n{query}\n\n### Trimmed Context:\n{context[:300]}")
pt = event_handler.PROMPT_TEMPLATEfull_prompt = pt.format(context_str=context, query_str=query)
result = trustworthy_rag.generate( query=query, context=context, prompt=full_prompt)print(f"\n### Generated Response:\n{result['response']}\n")print("### Evaluation Scores:")for metric, value in result.items(): if metric != "response": print(f"{metric}: {value['score']}")### Query:How much did Nvidia's revenue decrease this quarter vs last quarter, in dollars?
### Trimmed Context:# NVIDIA Announces Financial Results for First Quarter Fiscal 2024
NVIDIA (NASDAQ: NVDA) today reported revenue for the first quarter ended April 30, 2023, of $7.19 billion, down 13% from a year ago and up 19% from the previous quarter.
- **Quarterly revenue** of $7.19 billion, up 19% from the pre
### Generated Response:NVIDIA's revenue for the first quarter of fiscal 2024 was $7.19 billion, and for the previous quarter (Q4 FY23), it was $6.05 billion. Therefore, the revenue increased by $1.14 billion from the previous quarter, not decreased.
So, the revenue did not decrease this quarter vs last quarter; it actually increased by $1.14 billion.
### Evaluation Scores:trustworthiness: 0.6810414232214796context_sufficiency: 0.9974887437375295response_groundedness: 0.9975116791816968response_helpfulness: 0.3293002430120912query_ease: 0.33275910932109172虽然实现一个能准确回答任何可能问题的RAG应用仍然困难,但你可以轻松使用Cleanlab部署一个可信赖的RAG应用,该应用至少会标记可能不准确的答案。了解更多关于可调整以提升准确性/延迟的可选配置,请参阅Cleanlab文档。