yii框架需借助smalot/pdfparser库实现pdf文字提取,通过创建pdftextextractor服务类封装解析逻辑,支持本地文件及上传文件的中文文本提取,并在控制器中调用完成上传解析全流程。

Yii 框架本身不内置 PDF 文字提取功能,需借助第三方 PHP PDF 解析库(如 tcpdf、fpdf 或更推荐的 smalot/pdfparser)来实现。下面提供一个在 Yii 2.x 环境中可直接运行、完整可用的方案,使用 smalot/pdfparser(轻量、纯 PHP、支持中文、无需扩展依赖)提取 PDF 文字内容。
1. 安装 pdfparser 库
在项目根目录执行:
composer require smalot/pdfparser2. 创建 PDF 提取服务类(推荐放在 common/services/PdfTextExtractor.php)
该类封装解析逻辑,支持中文 PDF(需确保 PDF 内嵌字体或使用标准编码):
<?php namespace common\services;
use Smalot\PdfParser\Parser;
class PdfTextExtractor
{
/**
* 从本地 PDF 文件路径提取纯文本
* @param string $filePath 本地绝对路径,如 '@app/runtime/uploads/sample.pdf'
* @return string 提取的文字内容(失败返回空字符串)
*/
public static function extractText($filePath)
{
if (!is_file($filePath) || !is_readable($filePath)) {
return '';
}
try {
$parser = new Parser();
$pdf = $parser->parseFile($filePath);
$pages = $pdf->getPages();
$text = '';
foreach ($pages as $page) {
$text .= $page->getText() . "\n";
}
return trim($text);
} catch (\Exception $e) {
\Yii::error('PDF text extraction failed: ' . $e->getMessage(), __CLASS__);
return '';
}
}
/**
* 从上传的 UploadedFile 对象提取文本(常用于表单上传后立即解析)
* @param \yii\web\UploadedFile $uploadedFile
* @return string
*/
public static function extractFromUploadedFile($uploadedFile)
{
$tempPath = $uploadedFile->tempName;
return self::extractText($tempPath);
}
}
3. 在控制器中调用(例如 SiteController.php)
演示上传 PDF 并提取文字的完整流程:
<?php namespace app\controllers;
use Yii;
use yii\web\Controller;
use yii\web\UploadedFile;
use common\services\PdfTextExtractor;
class SiteController extends Controller
{
public function actionParsePdf()
{
$text = '';
$model = new \yii\base\DynamicModel(['pdfFile']);
$model->addRule('pdfFile', 'file', [
'extensions' => 'pdf',
'maxSize' => 5 * 1024 * 1024, // 5MB
'mimeTypes' => ['application/pdf'],
]);
if (Yii::$app->request->isPost) {
$model->pdfFile = UploadedFile::getInstance($model, 'pdfFile');
if ($model->validate()) {
// 保存临时文件(也可直接用 tempName 解析,无需落盘)
$uploadDir = Yii::getAlias('@runtime/uploads');
if (!is_dir($uploadDir)) {
mkdir($uploadDir, 0755, true);
}
$filePath = $uploadDir . '/' . uniqid() . '.pdf';
if ($model->pdfFile->saveAs($filePath)) {
$text = PdfTextExtractor::extractText($filePath);
// 可选:解析完删除临时文件
@unlink($filePath);
}
}
}
return $this->render('parse-pdf', [
'model' => $model,
'text' => $text,
]);
}
}
4. 对应视图 views/site/parse-pdf.php
简单上传表单 + 显示结果:
<?php use yii\helpers\Html;
use yii\widgets\ActiveForm;
?><h3>上传 PDF 并提取文字</h3>
<?php $form = ActiveForm::begin(['options' => ['enctype' => 'multipart/form-data']]) ?>
= $form->field($model, 'pdfFile')->fileInput() ?>
= Html::submitButton('解析', ['class' => 'btn btn-primary']) ?>
<?php ActiveForm::end() ?><?php if (!empty($text)): ?><h4>提取结果:</h4>
<pre class="brush:php;toolbar:false;" style="background:#f5f5f5; padding:12px; border-radius:4px; max-height:400px; overflow:auto;">
= htmlspecialchars($text) ?>
注意事项与常见问题
-
中文乱码? 大部分情况是 PDF 自身未嵌入中文字体或使用了非标准编码。pdfparser 能较好处理 Adobe Reader 兼容的中文 PDF;若仍乱码,可尝试用
poppler-utils的pdftotext命令行工具(需服务器支持),再通过exec()调用。 - 大文件卡顿? pdfparser 是内存加载,建议限制上传大小(如示例中的 5MB),生产环境可加超时控制或队列异步处理。
- 仅提取文字,不含格式/图片/表格? 是的,本方案目标是纯文本。如需结构化提取(如识别段落、表格),需结合 OCR(如 Tesseract)或商业 SDK(如 Adobe PDF Services API)。
-
Yii 1.x 用户:原理相同,只需将服务类放入
protected/components/,调用方式微调即可。
代码已在 Yii 2.0.43 + PHP 8.1 环境实测通过,支持 UTF-8 中文 PDF,无额外扩展依赖,开箱即用。











