Pandas查询引擎
本指南向您展示如何使用我们的 PandasQueryEngine:通过大型语言模型将自然语言转换为Pandas Python代码。
PandasQueryEngine 的输入是一个 Pandas 数据框,输出是一个响应。LLM 会推断需要执行的数据框操作以获取结果。
警告: 此工具为LLM提供对 eval 函数的访问权限。
在运行此工具的机器上可能执行任意代码。
虽然对代码进行了一定程度的过滤,但不建议在生产环境中使用此工具,
除非有严格沙箱或虚拟机保护。
如果您在 Colab 上打开这个笔记本,您可能需要安装 LlamaIndex 🦙。
!pip install llama-index llama-index-experimentalimport loggingimport sysfrom IPython.display import Markdown, display
import pandas as pdfrom llama_index.experimental.query_engine import PandasQueryEngine
logging.basicConfig(stream=sys.stdout, level=logging.INFO)logging.getLogger().addHandler(logging.StreamHandler(stream=sys.stdout))让我们从一个玩具数据框开始
Section titled “Let’s start on a Toy DataFrame”这里让我们加载一个非常简单的包含城市和人口对的数据框,并对其运行 PandasQueryEngine。
通过设置 verbose=True 我们可以查看中间生成的指令。
# Test on some sample datadf = pd.DataFrame( { "city": ["Toronto", "Tokyo", "Berlin"], "population": [2930000, 13960000, 3645000], })query_engine = PandasQueryEngine(df=df, verbose=True)response = query_engine.query( "What is the city with the highest population?",)INFO:httpx:HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"> Pandas Instructions:```df['city'][df['population'].idxmax()]```> Pandas Output: Tokyodisplay(Markdown(f"<b>{response}</b>"))东京
# get pandas python instructionsprint(response.metadata["pandas_instruction_str"])df['city'][df['population'].idxmax()]我们还可以采取使用大型语言模型来合成响应的步骤。
query_engine = PandasQueryEngine(df=df, verbose=True, synthesize_response=True)response = query_engine.query( "What is the city with the highest population? Give both the city and population",)print(str(response))INFO:httpx:HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"> Pandas Instructions:```df.loc[df['population'].idxmax()]```> Pandas Output: city Tokyopopulation 13960000Name: 1, dtype: objectINFO:httpx:HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"The city with the highest population is Tokyo, with a population of 13,960,000.泰坦尼克号数据集是机器学习入门中最受欢迎的表格数据集之一 来源:https://www.kaggle.com/c/titanic
!wget 'https://raw.githubusercontent.com/jerryjliu/llama_index/main/docs/examples/data/csv/titanic_train.csv' -O 'titanic_train.csv'--2024-01-13 17:45:15-- https://raw.githubusercontent.com/jerryjliu/llama_index/main/docs/examples/data/csv/titanic_train.csvResolving raw.githubusercontent.com (raw.githubusercontent.com)... 2606:50c0:8003::154, 2606:50c0:8002::154, 2606:50c0:8001::154, ...Connecting to raw.githubusercontent.com (raw.githubusercontent.com)|2606:50c0:8003::154|:443... connected.HTTP request sent, awaiting response... 200 OKLength: 57726 (56K) [text/plain]Saving to: ‘titanic_train.csv’
titanic_train.csv 100%[===================>] 56.37K --.-KB/s in 0.009s
2024-01-13 17:45:15 (6.45 MB/s) - ‘titanic_train.csv’ saved [57726/57726]df = pd.read_csv("./titanic_train.csv")query_engine = PandasQueryEngine(df=df, verbose=True)response = query_engine.query( "What is the correlation between survival and age?",)INFO:httpx:HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"> Pandas Instructions:```df['survived'].corr(df['age'])```> Pandas Output: -0.07722109457217755display(Markdown(f"<b>{response}</b>"))-0.07722109457217755
# get pandas python instructionsprint(response.metadata["pandas_instruction_str"])df['survived'].corr(df['age'])让我们来看看这些提示!
from llama_index.core import PromptTemplatequery_engine = PandasQueryEngine(df=df, verbose=True)prompts = query_engine.get_prompts()print(prompts["pandas_prompt"].template)You are working with a pandas dataframe in Python.The name of the dataframe is `df`.This is the result of `print(df.head())`:{df_str}
Follow these instructions:{instruction_str}Query: {query_str}
Expression:print(prompts["response_synthesis_prompt"].template)Given an input question, synthesize a response from the query results.Query: {query_str}
Pandas Instructions (optional):{pandas_instructions}
Pandas Output: {pandas_output}
Response:您也可以更新提示词:
new_prompt = PromptTemplate( """\You are working with a pandas dataframe in Python.The name of the dataframe is `df`.This is the result of `print(df.head())`:{df_str}
Follow these instructions:{instruction_str}Query: {query_str}
Expression: """)
query_engine.update_prompts({"pandas_prompt": new_prompt})这是指令字符串(您可以通过在初始化时传入 instruction_str 来自定义)
instruction_str = """\1. Convert the query to executable Python code using Pandas.2. The final line of code should be a Python expression that can be called with the `eval()` function.3. The code should represent a solution to the query.4. PRINT ONLY THE EXPRESSION.5. Do not quote the expression."""如果你想学习使用我们的查询管道语法和上述提示组件来构建自己的Pandas查询引擎,请查看以下教程。