文件
Docx阅读器 #
基类:EventBaseReader
Docx 解析器。
workflows/handler.py 中的源代码llama_index/readers/file/docs/base.py
103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None, fs: Optional[AbstractFileSystem] = None) -> List[文档]
解析文件。
workflows/handler.py 中的源代码llama_index/readers/file/docs/base.py
106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 | |
HWP阅读器 #
基类:EventBaseReader
Hwp 解析器。
workflows/handler.py 中的源代码llama_index/readers/file/docs/base.py
136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None, fs: Optional[AbstractFileSystem] = None) -> List[文档]
从Hwp文件中加载数据并提取表格。
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
file
|
Path
|
Hwp 文件的路径。 |
required |
返回:
| 类型 | 描述 |
|---|---|
List[文档]
|
文档列表 |
workflows/handler.py 中的源代码llama_index/readers/file/docs/base.py
148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 | |
PDF阅读器 #
基类:EventBaseReader
PDF解析器。
workflows/handler.py 中的源代码llama_index/readers/file/docs/base.py
28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 | |
load_data #
load_data(file: Union[Path, PurePosixPath], extra_info: Optional[Dict] = None, fs: Optional[AbstractFileSystem] = None) -> List[文档]
解析文件。
workflows/handler.py 中的源代码llama_index/readers/file/docs/base.py
37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 | |
电子书阅读器 #
基类:EventBaseReader
Epub 解析器。
workflows/handler.py 中的源代码llama_index/readers/file/epub/base.py
18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None, fs: Optional[AbstractFileSystem] = None) -> List[文档]
解析文件。
workflows/handler.py 中的源代码llama_index/readers/file/epub/base.py
21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 | |
扁平读取器 #
基类:EventBaseReader
扁平化读取器。
从文件中提取原始文本并将文件类型保存在元数据中
workflows/handler.py 中的源代码llama_index/readers/file/flat/base.py
12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None, fs: Optional[AbstractFileSystem] = None) -> List[文档]
将文件解析为字符串。
workflows/handler.py 中的源代码llama_index/readers/file/flat/base.py
34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 | |
HTML标签读取器 #
基类:EventBaseReader
读取HTML文件并使用BeautifulSoup从特定标签中提取文本。
默认情况下,从 <section> 标签读取文本。
workflows/handler.py 中的源代码llama_index/readers/file/html/base.py
11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 | |
图像读取器 #
基类:EventBaseReader
图像解析器。
使用 DONUT 或 pytesseract 从图像中提取文本。
workflows/handler.py 中的源代码llama_index/readers/file/image/base.py
19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None, fs: Optional[AbstractFileSystem] = None) -> List[文档]
解析文件。
workflows/handler.py 中的源代码llama_index/readers/file/image/base.py
75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 | |
图像标题读取器 #
基类:EventBaseReader
图像解析器。
使用 Blip 为图像添加标题。
workflows/handler.py 中的源代码llama_index/readers/file/image_caption/base.py
9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None) -> List[文档]
解析文件。
workflows/handler.py 中的源代码llama_index/readers/file/image_caption/base.py
59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 | |
图像表格图表读取器 #
基类:EventBaseReader
图像解析器。
从图表或图形中提取表格数据。
workflows/handler.py 中的源代码llama_index/readers/file/image_deplot/base.py
8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None) -> List[文档]
解析文件。
workflows/handler.py 中的源代码llama_index/readers/file/image_deplot/base.py
57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 | |
图像视觉大语言模型阅读器 #
基类:EventBaseReader
图像解析器。
使用 Blip2(一种类似于 GPT4 的多模态视觉大语言模型)为图像添加标题。
workflows/handler.py 中的源代码llama_index/readers/file/image_vision_llm/base.py
9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None) -> List[文档]
解析文件。
workflows/handler.py 中的源代码llama_index/readers/file/image_vision_llm/base.py
64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 | |
IPYNB阅读器 #
基类:EventBaseReader
图像解析器。
workflows/handler.py 中的源代码llama_index/readers/file/ipynb/base.py
10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None, fs: Optional[AbstractFileSystem] = None) -> List[文档]
解析文件。
workflows/handler.py 中的源代码llama_index/readers/file/ipynb/base.py
22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 | |
Markdown阅读器 #
基类:EventBaseReader
Markdown 解析器。
从markdown文件中提取文本。 返回字典,其中键为标题,值为标题之间的文本内容。
workflows/handler.py 中的源代码llama_index/readers/file/markdown/base.py
16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 | |
markdown_to_tups #
markdown_to_tups(markdown_text: str) -> List[Tuple[Optional[str], str]]
将Markdown文件转换为包含标题和文本的元组列表。
workflows/handler.py 中的源代码llama_index/readers/file/markdown/base.py
39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 | |
remove_images #
remove_images(content: str) -> str
移除Markdown内容中的图像但保留描述。
workflows/handler.py 中的源代码llama_index/readers/file/markdown/base.py
103 104 105 106 | |
remove_hyperlinks #
remove_hyperlinks(content: str) -> str
移除Markdown内容中的超链接。
workflows/handler.py 中的源代码llama_index/readers/file/markdown/base.py
108 109 110 111 | |
parse_tups #
parse_tups(filepath: str, errors: str = 'ignore', fs: Optional[AbstractFileSystem] = None) -> List[Tuple[Optional[str], str]]
将文件解析为元组。
workflows/handler.py 中的源代码llama_index/readers/file/markdown/base.py
117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 | |
load_data #
load_data(file: str, extra_info: Optional[Dict] = None, fs: Optional[AbstractFileSystem] = None) -> List[文档]
将文件解析为字符串。
workflows/handler.py 中的源代码llama_index/readers/file/markdown/base.py
133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 | |
Mbox阅读器 #
基类:EventBaseReader
Mbox 解析器。
从邮箱文件中提取消息。 返回包含每条消息的日期、主题、发送者、接收者和内容的字符串。
workflows/handler.py 中的源代码llama_index/readers/file/mbox/base.py
19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None, fs: Optional[AbstractFileSystem] = None) -> List[文档]
将文件解析为字符串。
workflows/handler.py 中的源代码llama_index/readers/file/mbox/base.py
56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 | |
分页CSV读取器 #
基类:EventBaseReader
分页CSV解析器。
以LLM友好的格式在单独文档中显示每一行。
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
encoding
|
str
|
用于打开文件的编码。 默认为 utf-8。 |
'utf-8'
|
workflows/handler.py 中的源代码llama_index/readers/file/paged_csv/base.py
15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None, delimiter: str = ',', quotechar: str = '"') -> List[文档]
解析文件。
workflows/handler.py 中的源代码llama_index/readers/file/paged_csv/base.py
32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 | |
PyMuPDF阅读器 #
基类:EventBaseReader
使用 PyMuPDF 库读取 PDF 文件。
workflows/handler.py 中的源代码llama_index/readers/file/pymu_pdf/base.py
10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 | |
load_data #
load_data(file_path: Union[Path, str], metadata: bool = True, extra_info: Optional[Dict] = None) -> List[文档]
从PDF文件加载文档列表,并接受字典格式的额外信息。
workflows/handler.py 中的源代码llama_index/readers/file/pymu_pdf/base.py
13 14 15 16 17 18 19 20 | |
加载 #
load(file_path: Union[Path, str], metadata: bool = True, extra_info: Optional[Dict] = None) -> List[文档]
从PDF文件加载文档列表,并接受字典格式的额外信息。
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
file_path
|
Union[Path, str]
|
PDF文件的路径(接受字符串或Path类型)。 |
required |
metadata
|
bool
|
是否包含元数据。默认为 True。 |
True
|
extra_info
|
Optional[Dict]
|
与每个文档相关的额外信息,以字典格式表示。默认为 None。 |
None
|
引发:
| 类型 | 描述 |
|---|---|
TypeError
|
如果 extra_info 不是字典类型。 |
TypeError
|
如果 file_path 不是字符串或路径。 |
返回:
| 类型 | 描述 |
|---|---|
List[文档]
|
List[Document]: 文档列表。 |
workflows/handler.py 中的源代码llama_index/readers/file/pymu_pdf/base.py
22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 | |
RTF阅读器 #
基类:EventBaseReader
RTF(富文本格式)阅读器。读取 rtf 文件并转换为文档。
workflows/handler.py 中的源代码llama_index/readers/file/rtf/base.py
10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 | |
load_data #
load_data(input_file: Union[Path, str], extra_info: Optional[Dict[str, Any]] = None, **load_kwargs: Any) -> List[文档]
从RTF文件加载数据。
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
input_file
|
Path | str
|
RTF 文件的路径。 |
required |
extra_info
|
Dict[str, Any]
|
RTF 文件的路径。 |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[文档]
|
List[Document]: 文档列表。 |
workflows/handler.py 中的源代码llama_index/readers/file/rtf/base.py
13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 | |
PPTX阅读器 #
基类:EventBaseReader
增强版 PowerPoint 解析器。
提取文本、表格、图表、演讲者备注,并可选择为图像添加标题。 支持多线程处理和基于LLM的内容整合。 每张幻灯片始终返回一个文档。
workflows/handler.py 中的源代码llama_index/readers/file/slides/base.py
32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 | |
load_data #
load_data(file: Union[str, Path], extra_info: Optional[Dict] = None, fs: Optional[AbstractFileSystem] = None) -> List[文档]
使用增强内容提取功能解析 PowerPoint 文件。
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
file
|
Union[str, Path]
|
PowerPoint 文件路径 |
required |
extra_info
|
Optional[Dict]
|
要包含的额外元数据 |
None
|
fs
|
Optional[AbstractFileSystem]
|
用于读取的文件系统 |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[文档]
|
文档列表(每页幻灯片一个) |
workflows/handler.py 中的源代码llama_index/readers/file/slides/base.py
79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 | |
extract_with_validation #
extract_with_validation(file_path: str, extract_images: bool = True, context_consolidation_with_llm: bool = False, fs: Optional[AbstractFileSystem] = None) -> Dict[str, Any]
从PowerPoint文件中提取内容,支持验证和多线程处理。
workflows/handler.py 中的源代码llama_index/readers/file/slides/base.py
153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 | |
CSV读取器 #
基类:EventBaseReader
CSV解析器。
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
concat_rows
|
bool
|
是否将所有行合并为一个文档。 如果设为False,将为每一行创建一个文档。 默认为True。 |
True
|
workflows/handler.py 中的源代码llama_index/readers/file/tabular/base.py
18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None) -> List[文档]
解析文件。
返回:
| 类型 | 描述 |
|---|---|
List[文档]
|
Union[str, List[str]]: 一个字符串或字符串列表。 |
workflows/handler.py 中的源代码llama_index/readers/file/tabular/base.py
34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 | |
PandasCSV读取器 #
基类:EventBaseReader
基于Pandas的CSV解析器。
使用Pandas read_csv函数中的分隔符检测功能解析CSV文件。
如需特殊参数,请使用pandas_config字典。
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
concat_rows
|
bool
|
是否将所有行合并为一个文档。 如果设为False,将为每一行创建一个文档。 默认为True。 |
True
|
col_joiner
|
str
|
用于每行连接列的分隔符。 默认设置为 ", "。 |
', '
|
row_joiner
|
str
|
用于连接每行的分隔符。
仅在 |
'\n'
|
pandas_config
|
dict
|
|
{}
|
workflows/handler.py 中的源代码llama_index/readers/file/tabular/base.py
64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None, fs: Optional[AbstractFileSystem] = None) -> List[文档]
解析文件。
workflows/handler.py 中的源代码llama_index/readers/file/tabular/base.py
107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 | |
PandasExcel读取器 #
基类:EventBaseReader
自定义Excel解析器,在每行中包含标题名称。
使用 Pandas 的 read_excel 函数解析 Excel 文件,但会将每行格式化为包含标题名称,例如:"姓名: joao, 职位: analyst"。
首行(标题)不会包含在生成的文档中。
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
concat_rows
|
bool
|
决定是否将所有行合并为一个文档。 如果设为False,则每行创建一个文档。 默认为True。 |
True
|
sheet_name
|
str | int | None
|
默认为 None,表示读取所有工作表。 或者,传递一个字符串或整数来指定要读取的工作表。 |
None
|
field_separator
|
str
|
用于分隔每个字段的字符或字符串。默认值:", "。 |
', '
|
key_value_separator
|
str
|
用于分隔键与值的字符或字符串。默认值:": "。 |
': '
|
pandas_config
|
dict
|
|
{}
|
workflows/handler.py 中的源代码llama_index/readers/file/tabular/base.py
136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None, fs: Optional[AbstractFileSystem] = None) -> List[文档]
解析文件。
workflows/handler.py 中的源代码llama_index/readers/file/tabular/base.py
177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 | |
非结构化读取器 #
基类:EventBaseReader
适用于多种文件的通用非结构化文本读取器。
workflows/handler.py 中的源代码llama_index/readers/file/unstructured/base.py
24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 | |
from_api
classmethod
#
from_api(api_key: str, url: str = None)
设置服务器URL和API密钥。
workflows/handler.py 中的源代码llama_index/readers/file/unstructured/base.py
71 72 73 74 | |
load_data #
load_data(file: Optional[Path] = None, unstructured_kwargs: Optional[Dict] = None, document_kwargs: Optional[Dict] = None, extra_info: Optional[Dict] = None, split_documents: Optional[bool] = False, excluded_metadata_keys: Optional[List[str]] = None) -> List[文档]
使用 Unstructured.io 加载数据。
根据配置情况,如果设置了url或use_api为True, 它将通过API调用解析文件,否则在本地进行解析。 如果split_documents为True,extra_info会被返回的元数据扩展。
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
file
|
Optional[Path]
|
要加载文件的路径。 |
None
|
unstructured_kwargs
|
Optional[Dict]
|
用于非结构化分区的附加参数。 |
None
|
document_kwargs
|
Optional[Dict]
|
文档创建的附加参数。 |
None
|
extra_info
|
Optional[Dict]
|
要添加到文档元数据中的额外信息。 |
None
|
split_documents
|
Optional[bool]
|
是否拆分文档。 |
False
|
excluded_metadata_keys
|
Optional[List[str]]
|
要从元数据中排除的键。 |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[文档]
|
List[Document]: 已解析文档的列表。 |
workflows/handler.py 中的源代码llama_index/readers/file/unstructured/base.py
76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 | |
视频音频读取器 #
基类:EventBaseReader
视频音频解析器。
从视频/音频文件的转录文本中提取文字。
workflows/handler.py 中的源代码llama_index/readers/file/video_audio/base.py
19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None, fs: Optional[AbstractFileSystem] = None) -> List[文档]
解析文件。
workflows/handler.py 中的源代码llama_index/readers/file/video_audio/base.py
45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 | |
XML阅读器 #
基类:EventBaseReader
XML 读取器。
读取XML文档,提供有助于推断节点间关系的选项。
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
tree_level_split
|
int
|
从XML树的哪个层级开始分割文档, |
0
|
workflows/handler.py 中的源代码llama_index/readers/file/xml/base.py
42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 | |
load_data #
load_data(file: Path, extra_info: Optional[Dict] = None) -> List[文档]
从输入文件加载数据。
参数:
| 名称 | 类型 | 描述 | 默认 |
|---|---|---|---|
file
|
Path
|
输入文件的路径。 |
required |
extra_info
|
Optional[Dict]
|
附加信息。默认为无。 |
None
|
返回:
| 类型 | 描述 |
|---|---|
List[文档]
|
List[Document]: 文档列表。 |
workflows/handler.py 中的源代码llama_index/readers/file/xml/base.py
83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 | |
选项: 成员:- CSVReader - DocxReader - EpubReader - FlatReader - HTMLTagReader - HWPReader - IPYNBReader - ImageCaptionReader - ImageReader - ImageTabularChartReader - ImageVisionLLMReader - MarkdownReader - MboxReader - PDFReader - PagedCSVReader - PandasCSVReader - PandasExcelReader - PptxReader - PyMuPDFReader - RTFReader - UnstructuredReader - VideoAudioReader - XMLReader