Yii框架需借助smalot/pdfparser等第三方库解析PDF文本,推荐封装为服务类并在控制器中调用,支持UTF-8中文但不保留格式,需确保文件可读、目录可写及启用mbstring扩展。

Yii 框架本身不内置 PDF 内容解析能力,读取 PDF 文字需借助第三方 PHP 库(如 smalot/pdfparser 或 setasign/fpdi + setasign/tcpdf 组合),核心是「用 PHP 解析 PDF → 在 Yii 控制器或服务中调用 → 返回文本或结构化数据」。
推荐方案:使用 smalot/pdfparser(轻量、专注文本提取)
该库专为解析 PDF 结构和提取文本设计,兼容多数非加密、非扫描型 PDF(即含真实文字图层的 PDF)。
- 安装依赖:composer require smalot/pdfparser
- 确保 PDF 文件可读(非密码保护、非图像型 PDF;若为扫描件需先 OCR,此库不支持 OCR)
- 在 Yii2 中建议封装为独立服务类,避免控制器臃肿
完整实现代码(Yii2 环境)
1. 创建服务类(例如 app/services/PdfTextExtractor.php):
namespace app\services;
<p>use Smalot\PdfParser\Parser;</p><p>class PdfTextExtractor
{
public function extractText($filePath): string
{
if (!is_file($filePath) || !is_readable($filePath)) {
throw new \InvalidArgumentException("PDF file not found or unreadable: $filePath");
}</p><pre class="brush:php;toolbar:false;"> $parser = new Parser();
try {
$pdf = $parser->parseFile($filePath);
return trim($pdf->getText()); // 返回纯文本,自动处理换行与空格
} catch (\Exception $e) {
throw new \RuntimeException("Failed to parse PDF: " . $e->getMessage());
}
}}
2. 在控制器中调用(例如 SiteController.php):
use app\services\PdfTextExtractor;
use yii\web\Controller;
use yii\web\UploadedFile;
<p>class SiteController extends Controller
{
public function actionReadPdf()
{
$model = new \yii\base\DynamicModel(['pdfFile']);
$model->addRule('pdfFile', 'file', [
'extensions' => 'pdf',
'maxSize' => 5 <em> 1024 </em> 1024, // 5MB
'skipOnEmpty' => false,
]);</p><pre class="brush:php;toolbar:false;"> if ($model->load(\Yii::$app->request->post())) {
$file = UploadedFile::getInstance($model, 'pdfFile');
if ($file) {
$tempPath = \Yii::getAlias('@runtime/uploads/') . $file->baseName . '_' . time() . '.' . $file->extension;
$file->saveAs($tempPath);
try {
$extractor = new PdfTextExtractor();
$text = $extractor->extractText($tempPath);
// 可选:清理临时文件
@unlink($tempPath);
return $this->asJson(['success' => true, 'text' => $text]);
} catch (\Exception $e) {
@unlink($tempPath);
return $this->asJson(['success' => false, 'error' => $e->getMessage()]);
}
}
}
return $this->render('read-pdf', ['model' => $model]);
}}
注意事项与常见问题
-
中文支持:smalot/pdfparser 支持 UTF-8 编码的中文 PDF,但若 PDF 字体嵌入不全或使用自定义编码,可能乱码;可尝试在解析后用
mb_convert_encoding($text, 'UTF-8', 'auto')补救 -
表格/格式丢失:该库只返回线性文本流,不保留表格结构、字体样式、页眉页脚等;如需结构化提取(如按页、按段落),可用
$pdf->getDetails()或遍历$pdf->getPages() -
大文件性能:解析百页以上 PDF 可能内存占用高;建议设置
ini_set('memory_limit', '256M');或分页解析 -
替代方案:若需更高精度或处理扫描件,可集成
tesseract-ocr(需服务器安装 Tesseract)+spatie/pdf-to-text封装器
前端简单示例(表单 + AJAX)
在视图 read-pdf.php 中:
<?php $form = \yii\widgets\ActiveForm::begin(['options' => ['enctype' => 'multipart/form-data']]); ?>
= $form->field($model, 'pdfFile')->fileInput() ?>
<button type="submit">读取 PDF 内容</button>
<?php \yii\widgets\ActiveForm::end(); ?><p></p><div id="result"></div><p><script>
document.querySelector('form').onsubmit = async function(e) {
e.preventDefault();
const fd = new FormData(this);
const res = await fetch('/site/read-pdf', { method: 'POST', body: fd });
const data = await res.json();
document.getElementById('result').innerHTML =
data.success ? '<pre class="brush:php;toolbar:false;">' + data.text.substring(0, 2000) + '...</script></p>' :
'不复杂但容易忽略:确保 runtime 目录可写、PHP 启用 mbstring 扩展、PDF 不含强加密 —— 这三点到位,90% 的常规 PDF 都能顺利提取文本。











