Code to export web pages to Word documents in PHP_PHP tutorial-PHP Tutorial-php.cn

Home

Backend Development

PHP Tutorial

Code to export web pages to Word documents in PHP_PHP tutorial

WBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWB

Jul 21, 2016 pm 03:18 PM

docphpwordgenerallyforcodeuseCanExportdocumentmethodyeshaveofWeb page

Generally, there are two ways to export doc documents. One is to use com and install it on the server as an extension library of PHP, then create a com and call its methods. A server with office installed can call a com called word.application to generate a word document. However, I do not recommend this method because the execution efficiency is relatively low (I tested it and found that when executing the code, the server will actually Open a word client). The ideal com should have no interface and perform data conversion in the background, so the effect will be better, but these extensions generally require charges.

The second method is to use PHP to write the content of our doc document directly into a file with the suffix doc. Using this method does not require relying on third-party extensions, and the execution efficiency is higher.

Word itself is still very powerful. It can open files in html format and retain the format. Even if the suffix is doc, it can still open it normally. This provides us with convenience. But there is a problem. The pictures in the HTML format file have only one address, and the real pictures are saved elsewhere. That is to say, if the HTML format is written into the doc, the doc will not be able to contain the pictures. So how do we create a doc document containing images? We can use the mht format which is very close to html.

The mht format is very similar to html, except that in the mht format, externally linked files, such as images, Javascript, and CSS, will be encoded and stored in base64. Therefore, a single mht file can save all the resources in a web page. Of course, its size will be larger than that of html.

Can the mht format be recognized by word? I saved a web page as mht, then changed the suffix to doc, and then opened it with word. OK, word can also recognize mht files and can display pictures.

Okay, now that doc can recognize mht, the next step is to consider how to put pictures into mht. Since the address of the image in the html code is written in the src attribute of the img tag, as long as the src attribute value in the html code is extracted, the image address can be obtained. Of course, it is possible that what you get is a relative path. It doesn't matter. Just add the prefix of the URL and change it to an absolute path. With the image address, we can obtain the specific content of the image file through the file_get_content function, then call the base64_encode function to encode the file content into base64 encoding, and finally insert it into the appropriate location of the mht file.

Finally, we have two ways to send the file to the client. One is to first generate a doc document on the server side, and then record the address of the doc document. Finally, through header("location: xx.doc"); allows the client to download this doc. Another method is to directly send an html request, modify the header part of the HTML protocol, set its content-type to application/doc, set content-disposition to attachment, followed by the file name. After sending the html protocol, directly The file content is sent to the client, and the client can also be downloaded to the doc document.

Implementation

Through the above introduction of principles, I believe everyone should have a preliminary understanding of the implementation process. Below I will give an export function. This function can Export the HTML code into an mht document. There are 3 parameters, the last 2 of which are optional parameters
content: HTML code to be converted
absolutePath: If the image addresses in the HTML code are all relative paths, then This parameter is the absolute path missing in the HTML code.
isEraseLink: Whether to remove hyperlinks in HTML code
The return value is the file content of mht. You can save it as a file with the suffix doc through file_put_content
The main function of this function is actually to analyze HTML All image addresses in the code and download them one by one. After obtaining the content of the image, call the MhtFileMaker class to add the image to the mht file. The specific adding details are encapsulated in the MhtFileMaker class.

Copy code The code is as follows:

/**
* Get word document content based on HTML code
* Create a document that is essentially mht. This function will analyze the file content and download the image resources in the page from a remote location
* This function depends on the class MhtFileMaker
* This function will analyze the img tag and extract the attribute value of src. However, the attribute value of src must be surrounded by quotes, otherwise it cannot be extracted.
*
* @param string $content HTML content
* @param string $absolutePath The absolute path of the web page. If the image path in the HTML content is a relative path, you need to fill in this parameter to let the function automatically fill it into an absolute path. This parameter needs to end with /
* @param bool $isEraseLink Whether to remove links in HTML content
*/
function getWordDocument( $content , $absolutePath = "" , $isEraseLink = true )
{
$mht = new MhtFileMaker();
if ($isEraseLink)
$content = preg_replace('/(s*.*?s*)/i' , '$1' , $content); //Remove the link
$images = array();
$files = array();
$matches = array();
//This algorithm requires the attributes after src Values must be enclosed in quotes
if ( preg_match_all('/ Code to export web pages to Word documents in PHP_PHP tutorial

Code to export web pages to Word documents in PHP_PHP tutorial

/i',$content ,$matches ) )
{
$arrPath = $matches[1];
for ( $i=0;$i{
$path = $arrPath[$i];
$imgPath = trim( $path );
if ( $imgPath != "" )
{
$ files[] = $imgPath;
if( substr($imgPath,0,7) == 'http://')
{
//Absolute link, without prefix
}
else
{
$imgPath = $absolutePath.$imgPath;
}
$images[] = $imgPath;
}
}
}
$ mht->AddContents("tmp.html",$mht->GetMimeType("tmp.html"),$content);
for ( $i=0;$i{
$image = $images[$i];
if ( @fopen($image , 'r') )
{
$imgcontent = @file_get_contents( $image );
if ( $content )
$mht->AddContents($files[$i],$mht->GetMimeType($image),$imgcontent);
}
else
{
echo "file:".$image." not exist!
";
}
}
return $mht->GetFile();
}

Usage:

Copy code The code is as follows:

 
$fileContent = getWordDocument($content,"http://www.yoursite.com/Music/etc/"); 
$fp = fopen("test.doc", 'w'); 
fwrite($fp, $fileContent); 
fclose($fp); 

Among them, the $content variable should be the HTML source code, and the following link should be the URL address that can fill in the relative path of the image in the HTML code
Note that before using this function, you need to include the class MhtFileMaker. This class can help us generate Mht documents.

Copy code The code is as follows:

 
/*********************************************************************** 
Class: Mht File Maker 
Version: 1.2 beta 
Date: 02/11/2007 
Author: Wudi  
Description: The class can make .mht file. 
***********************************************************************/ 
class MhtFileMaker{ 
var $config = array(); 
var $headers = array(); 
var $headers_exists = array(); 
var $files = array(); 
var $boundary; 
var $dir_base; 
var $page_first; 
function MhtFile($config = array()){ 
} 
function SetHeader($header){ 
$this->headers[] = $header; 
$key = strtolower(substr($header, 0, strpos($header, ':'))); 
$this->headers_exists[$key] = TRUE; 
} 
function SetFrom($from){ 
$this->SetHeader("From: $from"); 
} 
function SetSubject($subject){ 
$this->SetHeader("Subject: $subject"); 
} 
function SetDate($date = NULL, $istimestamp = FALSE){ 
if ($date == NULL) { 
$date = time(); 
} 
if ($istimestamp == TRUE) { 
$date = date('D, d M Y H:i:s O', $date); 
} 
$this->SetHeader("Date: $date"); 
} 
function SetBoundary($boundary = NULL){ 
if ($boundary == NULL) { 
$this->boundary = '--' . strtoupper(md5(mt_rand())) . '_MULTIPART_MIXED'; 
} else { 
$this->boundary = $boundary; 
} 
} 
function SetBaseDir($dir){ 
$this->dir_base = str_replace("\", "/", realpath($dir)); 
} 
function SetFirstPage($filename){ 
$this->page_first = str_replace("\", "/", realpath("{$this->dir_base}/$filename")); 
} 
function AutoAddFiles(){ 
if (!isset($this->page_first)) { 
exit ('Not set the first page.'); 
} 
$filepath = str_replace($this->dir_base, '', $this->page_first); 
$filepath = 'http://mhtfile' . $filepath; 
$this->AddFile($this->page_first, $filepath, NULL); 
$this->AddDir($this->dir_base); 
} 
function AddDir($dir){ 
$handle_dir = opendir($dir); 
while ($filename = readdir($handle_dir)) { 
if (($filename!='.') && ($filename!='..') && ("$dir/$filename"!=$this->page_first)) { 
if (is_dir("$dir/$filename")) { 
$this->AddDir("$dir/$filename"); 
} elseif (is_file("$dir/$filename")) { 
$filepath = str_replace($this->dir_base, '', "$dir/$filename"); 
$filepath = 'http://mhtfile' . $filepath; 
$this->AddFile("$dir/$filename", $filepath, NULL); 
} 
} 
} 
closedir($handle_dir); 
} 
function AddFile($filename, $filepath = NULL, $encoding = NULL){ 
if ($filepath == NULL) { 
$filepath = $filename; 
} 
$mimetype = $this->GetMimeType($filename); 
$filecont = file_get_contents($filename); 
$this->AddContents($filepath, $mimetype, $filecont, $encoding); 
} 
function AddContents($filepath, $mimetype, $filecont, $encoding = NULL){ 
if ($encoding == NULL) { 
$filecont = chunk_split(base64_encode($filecont), 76); 
$encoding = 'base64'; 
} 
$this->files[] = array('filepath' => $filepath, 
'mimetype' => $mimetype, 
'filecont' => $filecont, 
'encoding' => $encoding); 
} 
function CheckHeaders(){ 
if (!array_key_exists('date', $this->headers_exists)) { 
$this->SetDate(NULL, TRUE); 
} 
if ($this->boundary == NULL) { 
$this->SetBoundary(); 
} 
} 
function CheckFiles(){ 
if (count($this->files) == 0) { 
return FALSE; 
} else { 
return TRUE; 
} 
} 
function GetFile(){ 
$this->CheckHeaders(); 
if (!$this->CheckFiles()) { 
exit ('No file was added.'); 
} 
$contents = implode("rn", $this->headers); 
$contents .= "rn"; 
$contents .= "MIME-Version: 1.0rn"; 
$contents .= "Content-Type: multipart/related;rn"; 
$contents .= "tboundary="{$this->boundary}";rn"; 
$contents .= "ttype="" . $this->files[0]['mimetype'] . ""rn"; 
$contents .= "X-MimeOLE: Produced By Mht File Maker v1.0 betarn"; 
$contents .= "rn"; 
$contents .= "This is a multi-part message in MIME format.rn"; 
$contents .= "rn"; 
foreach ($this->files as $file) { 
$contents .= "--{$this->boundary}rn"; 
$contents .= "Content-Type: $file[mimetype]rn"; 
$contents .= "Content-Transfer-Encoding: $file[encoding]rn"; 
$contents .= "Content-Location: $file[filepath]rn"; 
$contents .= "rn"; 
$contents .= $file['filecont']; 
$contents .= "rn"; 
} 
$contents .= "--{$this->boundary}--rn"; 
return $contents; 
} 
function MakeFile($filename){ 
$contents = $this->GetFile(); 
$fp = fopen($filename, 'w'); 
fwrite($fp, $contents); 
fclose($fp); 
}
function GetMimeType($filename){ 
$pathinfo = pathinfo($filename); 
switch ($pathinfo['extension']) { 
case 'htm': $mimetype = 'text/html'; break; 
case 'html': $mimetype = 'text/html'; break; 
case 'txt': $mimetype = 'text/plain'; break; 
case 'cgi': $mimetype = 'text/plain'; break; 
case 'php': $mimetype = 'text/plain'; break; 
case 'css': $mimetype = 'text/css'; break; 
case 'jpg': $mimetype = 'image/jpeg'; break; 
case 'jpeg': $mimetype = 'image/jpeg'; break; 
case 'jpe': $mimetype = 'image/jpeg'; break; 
case 'gif': $mimetype = 'image/gif'; break; 
case 'png': $mimetype = 'image/png'; break; 
default: $mimetype = 'application/octet-stream'; break; 
} 
return $mimetype; 
} 
} 
?> 

Statement

The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn

PHP Dependency Injection Container: A Quick StartMay 13, 2025 am 12:11 AM

APHPDependencyInjectionContainerisatoolthatmanagesclassdependencies,enhancingcodemodularity,testability,andmaintainability.Itactsasacentralhubforcreatingandinjectingdependencies,thusreducingtightcouplingandeasingunittesting.

Dependency Injection vs. Service Locator in PHPMay 13, 2025 am 12:10 AM

Select DependencyInjection (DI) for large applications, ServiceLocator is suitable for small projects or prototypes. 1) DI improves the testability and modularity of the code through constructor injection. 2) ServiceLocator obtains services through center registration, which is convenient but may lead to an increase in code coupling.

PHP performance optimization strategies.May 13, 2025 am 12:06 AM

PHPapplicationscanbeoptimizedforspeedandefficiencyby:1)enablingopcacheinphp.ini,2)usingpreparedstatementswithPDOfordatabasequeries,3)replacingloopswitharray_filterandarray_mapfordataprocessing,4)configuringNginxasareverseproxy,5)implementingcachingwi

PHP Email Validation: Ensuring Emails Are Sent CorrectlyMay 13, 2025 am 12:06 AM

PHPemailvalidationinvolvesthreesteps:1)Formatvalidationusingregularexpressionstochecktheemailformat;2)DNSvalidationtoensurethedomainhasavalidMXrecord;3)SMTPvalidation,themostthoroughmethod,whichchecksifthemailboxexistsbyconnectingtotheSMTPserver.Impl

How to make PHP applications fasterMay 12, 2025 am 12:12 AM

TomakePHPapplicationsfaster,followthesesteps:1)UseOpcodeCachinglikeOPcachetostoreprecompiledscriptbytecode.2)MinimizeDatabaseQueriesbyusingquerycachingandefficientindexing.3)LeveragePHP7 Featuresforbettercodeefficiency.4)ImplementCachingStrategiessuc

PHP Performance Optimization Checklist: Improve Speed NowMay 12, 2025 am 12:07 AM

ToimprovePHPapplicationspeed,followthesesteps:1)EnableopcodecachingwithAPCutoreducescriptexecutiontime.2)ImplementdatabasequerycachingusingPDOtominimizedatabasehits.3)UseHTTP/2tomultiplexrequestsandreduceconnectionoverhead.4)Limitsessionusagebyclosin

PHP Dependency Injection: Improve Code TestabilityMay 12, 2025 am 12:03 AM

Dependency injection (DI) significantly improves the testability of PHP code by explicitly transitive dependencies. 1) DI decoupling classes and specific implementations make testing and maintenance more flexible. 2) Among the three types, the constructor injects explicit expression dependencies to keep the state consistent. 3) Use DI containers to manage complex dependencies to improve code quality and development efficiency.

PHP Performance Optimization: Database Query OptimizationMay 12, 2025 am 12:02 AM

DatabasequeryoptimizationinPHPinvolvesseveralstrategiestoenhanceperformance.1)Selectonlynecessarycolumnstoreducedatatransfer.2)Useindexingtospeedupdataretrieval.3)Implementquerycachingtostoreresultsoffrequentqueries.4)Utilizepreparedstatementsforeffi

See all articles

Hot AI Tools

Undresser.AI Undress

AI-powered app for creating realistic nude photos

AI Clothes Remover

Online AI tool for removing clothes from photos.

Undress AI Tool

Undress images for free

Clothoff.io

AI clothes remover

Video Face Swap

Swap faces in any video effortlessly with our completely free AI face swap tool!

Hot Article

Roblox: Grow A Garden - Complete Mutation Guide

3 weeks agoByDDD

How to fix KB5055612 fails to install in Windows 10?

3 weeks agoByDDD

Roblox: Bubble Gum Simulator Infinity - How To Get And Use Royal Keys

3 weeks agoBy尊渡假赌尊渡假赌尊渡假赌

Mandragora: Whispers Of The Witch Tree - How To Unlock The Grappling Hook

3 weeks agoBy尊渡假赌尊渡假赌尊渡假赌

Nordhold: Fusion System, Explained

3 weeks agoBy尊渡假赌尊渡假赌尊渡假赌

Hot Tools

PhpStorm Mac version

The latest (2018.2.1) professional PHP integrated development tool

DVWA

Damn Vulnerable Web App (DVWA) is a PHP/MySQL web application that is very vulnerable. Its main goals are to be an aid for security professionals to test their skills and tools in a legal environment, to help web developers better understand the process of securing web applications, and to help teachers/students teach/learn in a classroom environment Web application security. The goal of DVWA is to practice some of the most common web vulnerabilities through a simple and straightforward interface, with varying degrees of difficulty. Please note that this software

SublimeText3 Chinese version

Chinese version, very easy to use

SecLists

SecLists is the ultimate security tester's companion. It is a collection of various types of lists that are frequently used during security assessments, all in one place. SecLists helps make security testing more efficient and productive by conveniently providing all the lists a security tester might need. List types include usernames, passwords, URLs, fuzzing payloads, sensitive data patterns, web shells, and more. The tester can simply pull this repository onto a new test machine and he will have access to every type of list he needs.

Dreamweaver Mac version

Visual web development tools

Hot Topics

1668

1426

1328

1273

1256