如何使用 PyMuPDFM 将 PDF 转换为 Markdown 及其评估-Python教程-PHP中文网

首页

后端开发

Python教程

如何使用 PyMuPDFM 将 PDF 转换为 Markdown 及其评估

Linda Hamilton

Oct 07, 2024 pm 06:12 PM

PyMuPDF4LLM 是一个旨在将 PDF 转换为 Markdown 格式的库。在这里，我将分享我测试这个库的经验。

安装

首先使用以下命令安装库：


pip install pymupdf4llm

用法

基本用法非常简单，只需三行代码即可将 PDF 转换为 Markdown：


import pymupdf4llm
md_text = pymupdf4llm.to_markdown("input.pdf")
print(md_text)

您可以指定参数来调整内容的提取方式。

按页提取文本

默认情况下，整个 PDF 会转换为单个文本输出。但是，您可以通过指定 page_chunks=True 逐页提取文本。


md_text = pymupdf4llm.to_markdown("input.pdf", page_chunks=True)

提取图像

要将图像提取为文件，请使用 write_images=True 选项：


md_text = pymupdf4llm.to_markdown("input.pdf", write_images=True)

也可以使用base64编码直接在Markdown中嵌入图像：


md_text = pymupdf4llm.to_markdown("input.pdf", embed_images=True)

转换结果评估

为了进行测试，使用了具有不同 Markdown 元素的各种 PDF。

How to Convert PDFs to Markdown Using PyMuPDFM and Its Evaluation

标头转换

标题已正确转换为 Markdown 格式。这是结果的一部分：


# Sample Markdown Guide

This is a sample markdown file that includes various features for quick reference.

## 1. Headers

...

## 3. Lists

粗体和斜体文本

粗体和斜体格式也已正确转换：


**Bold: **Bold Text****

_Italic: *Italic Text*_

**_Bold and Italic: ***Bold and Italic***_**

列表转换

第一级有序列表转换没有问题，但嵌套列表和无序列表转换不准确。

How to Convert PDFs to Markdown Using PyMuPDFM and Its Evaluation


## 3. Lists

### Unordered List

Item 1

Item 2

Sub-item 1

Sub-item 2

### Ordered List

1. First item

2. Second item

1. Sub-item A

2. Sub-item B

链接转换

提取了链接的URL，但包含链接的整行变成了超链接，偏离了原始格式。

How to Convert PDFs to Markdown Using PyMuPDFM and Its Evaluation


## 4. Links and Images

[You can add links using [Link Text](URL).](https://www.example.com/)

图像提取

默认情况下不会提取图像，但可以使用 write_images=True 将图像保存在本地。


md_text = pymupdf4llm.to_markdown("input.pdf", write_images=True)

然后在 Markdown 中引用保存的图像，如下所示：


<p>### Image Example</p>

<p>![](input.pdf-1-0.png)</p>

表转换

没有垂直边框的简单表格无法准确转换（可能是因为不明确的列边界导致表格被视为纯文本）。

How to Convert PDFs to Markdown Using PyMuPDFM and Its Evaluation


<p>## 5. Tables</p>

<p>**Column 1** **Column 2** **Column 3**</p>

<p>Row 1 Data A Data B</p>

<p>Row 2 Data C Data D</p>

代码转换

代码块已正确转换，但语言规范（例如 python）未保留。内联代码转换也存在问题。

How to Convert PDFs to Markdown Using PyMuPDFM and Its Evaluation


<p>## 6. Code</p>

<p>### Inline Code</p>

<p>Use backticks for inline code: print("Hello, world!")</p>

<p>### Code Block</p>

<p>Use triple backticks for code blocks:</p>

<p>```<br>
def greet(name):<br>
  return f"Hello, {name}!"<br>
print(greet("Markdown"))<br>
```</p>

多行文本

对于多行文本，换行符将按照原始 PDF 中的显示方式保留。

How to Convert PDFs to Markdown Using PyMuPDFM and Its Evaluation


<p>Markdown is a lightweight and versatile markup language favored by developers, writers, and bloggers alike</p>

<p>due to its simplicity in formatting text, enabling users to create readable and well-structured documents—</p>

<p>whether for documentation, blog posts, or articles—without the complexity of HTML, while also offering the</p>

<p>ability to convert content seamlessly into other formats like HTML, PDF, and even slideshows, making it an</p>

<p>ideal choice for projects that require both clarity and flexibility in presentation.</p>

结论

尽管在准确转换列表和链接方面存在挑战，PyMuPDF4LLM 是将 PDF 转换为 Markdown 的有用工具。它可以在本地工作，无需外部语言模型，适合无法访问互联网的环境。

以上是如何使用 PyMuPDFM 将 PDF 转换为 Markdown 及其评估的详细内容。更多信息请关注PHP中文网其他相关文章！

声明

本文内容由网友自发贡献，版权归原作者所有，本站不承担相应法律责任。如您发现有涉嫌抄袭侵权的内容，请联系admin@php.cn

列表和阵列之间的选择如何影响涉及大型数据集的Python应用程序的整体性能？May 03, 2025 am 12:11 AM

ForhandlinglargedatasetsinPython,useNumPyarraysforbetterperformance.1)NumPyarraysarememory-efficientandfasterfornumericaloperations.2)Avoidunnecessarytypeconversions.3)Leveragevectorizationforreducedtimecomplexity.4)Managememoryusagewithefficientdata

说明如何将内存分配给Python中的列表与数组。May 03, 2025 am 12:10 AM

Inpython，ListSusedynamicMemoryAllocationWithOver-Asalose，而alenumpyArraySallaySallocateFixedMemory.1）listssallocatemoremoremoremorythanneededinentientary上，respizeTized.2）numpyarsallaysallaysallocateAllocateAllocateAlcocateExactMemoryForements，OfferingPrediCtableSageButlessemageButlesseflextlessibility。

您如何在Python数组中指定元素的数据类型？May 03, 2025 am 12:06 AM

Inpython，YouCansspecthedatatAtatatPeyFelemereModeRernSpant.1）Usenpynernrump.1）Usenpynyp.dloatp.dloatp.ploatm64，formor professisconsiscontrolatatypes。

什么是Numpy，为什么对于Python中的数值计算很重要？May 03, 2025 am 12:03 AM

NumPyisessentialfornumericalcomputinginPythonduetoitsspeed,memoryefficiency,andcomprehensivemathematicalfunctions.1)It'sfastbecauseitperformsoperationsinC.2)NumPyarraysaremorememory-efficientthanPythonlists.3)Itoffersawiderangeofmathematicaloperation

讨论'连续内存分配”的概念及其对数组的重要性。May 03, 2025 am 12:01 AM

Contiguousmemoryallocationiscrucialforarraysbecauseitallowsforefficientandfastelementaccess.1)Itenablesconstanttimeaccess,O(1),duetodirectaddresscalculation.2)Itimprovescacheefficiencybyallowingmultipleelementfetchespercacheline.3)Itsimplifiesmemorym

您如何切成python列表？May 02, 2025 am 12:14 AM

SlicingaPythonlistisdoneusingthesyntaxlist[start:stop:step].Here'showitworks:1)Startistheindexofthefirstelementtoinclude.2)Stopistheindexofthefirstelementtoexclude.3)Stepistheincrementbetweenelements.It'susefulforextractingportionsoflistsandcanuseneg

在Numpy阵列上可以执行哪些常见操作？May 02, 2025 am 12:09 AM

numpyallowsforvariousoperationsonArrays：1）basicarithmeticlikeaddition，减法，乘法和division; 2）evationAperationssuchasmatrixmultiplication; 3）element-wiseOperations wiseOperationswithOutexpliitloops; 4）

Python的数据分析中如何使用阵列？May 02, 2025 am 12:09 AM

Arresinpython，尤其是Throughnumpyandpandas，weessentialFordataAnalysis，offeringSpeedAndeffied.1）NumpyArseNable efflaysenable efficefliceHandlingAtaSetSetSetSetSetSetSetSetSetSetSetsetSetSetSetSetsopplexoperationslikemovingaverages.2）

See all articles