yii框架需借助smalot/pdfparser等第三方库实现pdf分页文本提取,通过封装服务类逐页解析并返回纯文本,注意中文支持、大文件处理、扫描件ocr及安全性校验。

Yii 框架本身不直接支持 PDF 内容读取和分页提取,需借助第三方 PHP PDF 解析库(如 Smalot/PdfParser 或 setasign/Fpdi + tcpdf 组合),再在 Yii 应用中封装调用。以下以轻量、纯 PHP、支持中文(需字体配置)的 Smalot/PdfParser 为例,给出在 Yii2 中分页提取 PDF 文本内容的完整实现方案。
安装 PdfParser 扩展
在项目根目录执行:
composer require smalot/pdfparser
安装后即可在控制器或服务类中 use 相关类。
创建 PDF 提取服务类(推荐)
在 common/services/ 或 console/services/ 下新建 PdfTextExtractor.php:
<?php namespace common\services;
<p>use Smalot\PdfParser\Parser;<p>class PdfTextExtractor
{
/**</p>
逐页提取 PDF 的纯文本内容
@param string $filePath 本地 PDF 文件路径(绝对路径)
-
@return array 每页文本组成的索引数组,键为页码(从 1 开始) */ public function extractByPage(string $filePath): array { if (!file_exists($filePath) || pathinfo($filePath, PATHINFO_EXTENSION) !== 'pdf') { throw new \InvalidArgumentException('Invalid PDF file path.'); }
$parser = new Parser(); $pdf = $parser->parseFile($filePath);
$pages = $pdf->getPages(); $result = [];
foreach ($pages as $index => $page) { // 获取当前页原始文本(已自动处理换行、空格合并等基础清理) $text = trim($page->getText()); $result[$index + 1] = $text ?: ''; }
return $result; } }
在控制器中调用分页提取
例如在 frontend/controllers/DocumentController.php 中:
use common\services\PdfTextExtractor;
use yii\web\Controller;
use yii\web\UploadedFile;
<p>public function actionExtractPdf()
{
$model = new UploadForm(); // 假设你有上传表单模型</p><pre class="brush:php;toolbar:false;">if ($model->load(Yii::$app->request->post())) {
$file = UploadedFile::getInstance($model, 'pdf_file');
if ($file && $file->extension === 'pdf') {
$tempPath = Yii::getAlias('@runtime/uploads/') . $file->baseName . '_' . time() . '.' . $file->extension;
$file->saveAs($tempPath);
try {
$extractor = new PdfTextExtractor();
$pages = $extractor->extractByPage($tempPath);
// 示例:返回第 1 页和第 2 页内容
$response = [
'success' => true,
'pages' => [
'page_1' => mb_strlen($pages[1] ?? '') > 200
? mb_substr($pages[1], 0, 200) . '...'
: $pages[1] ?? '',
'page_2' => $pages[2] ?? '',
],
'total_pages' => count($pages),
];
} catch (\Exception $e) {
$response = ['success' => false, 'error' => $e->getMessage()];
}
// 清理临时文件
@unlink($tempPath);
return $this->asJson($response);
}
}
return $this->render('upload');}
注意事项与常见问题
-
中文支持:Smalot/PdfParser 对含中文字体的 PDF 支持有限;若提取为空或乱码,说明 PDF 使用了未嵌入字体或 CID 字体。可尝试用
pdftotext(Poppler 工具)替代,需服务器安装:exec("pdftotext -layout -f 1 -l 5 {$filePath} -", $output)。 -
大文件处理:避免一次性加载超大 PDF(如 >100MB),建议加内存限制或分批处理页码范围(
getPages($start, $limit)需自行扩展 Parser)。 -
非文本 PDF(扫描件):该方案仅适用于可复制文本的 PDF。扫描件需先 OCR(如 Tesseract),Yii 中可调用 shell 命令或集成
thiagoalessio/tesseract_ocr。 - 安全性:务必校验上传文件 MIME 类型和扩展名,禁用用户可控的文件路径拼接,防止 LFI 或任意文件读取。











