Example of extracting the body content of a web page with php

Home

Backend Development

PHP Tutorial

Example of extracting the body content of a web page with php_PHP tutorial

WBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWB

Jul 13, 2016 am 10:12 AM

phpexamplecontentextracttextofWeb pageidentifydifficulty

Example of php extracting the text content of a web page

Because the difficulty lies in how to identify and retain the article part of the web page, delete other useless information, and make it universal, You cannot formulate collection rules based on the target station like Locomotive, because there are various web pages in the search engine results.

To retrieve the data of a page and how to match the text part, Zheng Xiao came up with an idea on the way to get off work:

　1. Extract the body tag part –> Remove all links –> Remove all scripts and comments –> Remove all blank tags (including those that do not contain Chinese characters) –> Get the results.

　2. Directly match non-linked Chinese parts that match the div, p, h tags???

There will still be a lot of other redundant information, such as bottom information, etc. . How to do it? I wonder if anyone has any ideas or suggestions?

This class is an algorithm for extracting the text part of a web page implemented in PHP found on the Internet. Zheng Xiao also tested it locally and the accuracy rate is very high.

The code is as follows

class Readability {
//The name of the mark bit that saves the judgment result
const ATTR_CONTENT_SCORE = "contentScore";

//The DOM parsing class currently only supports UTF-8 encoding
const DOM_DEFAULT_CHARSET = "utf-8";

//The content displayed when the judgment fails
const MESSAGE_CAN_NOT_GET = "Readability was unable to parse this page for content.";

// DOM parsing class (built-in in PHP5)
protected $DOM = null;

//Source code that needs to be parsed
protected $source = "";

//List of parent elements of chapters
private $parentNodes = array();

// Tags that need to be deleted
// Note: added extra tags from http://www.45it.net

private $junkTags = Array("style", "form", "iframe", "script", "button", "input", "textarea",
"noscript", "select", "option", "object", "applet", "basefont",
"bgsound", "blink", "canvas", "command", "menu", "nav", "datalist",
"embed", "frame", "frameset", "keygen", "label", "marquee", "link");

//Attributes that need to be deleted
private $junkAttrs = Array("style", "class", "onclick", "onmouseover", "align", "border", "margin");

/**
* Constructor
* @param $input_char encoding of string. Default utf-8, can be omitted
*/
function __construct($source, $input_char = "utf-8") {
$this->source = $source;

// The DOM parsing class can only handle characters in UTF-8 format
$source = mb_convert_encoding($source, 'HTML-ENTITIES', $input_char);

// Preprocess HTML tags, remove redundant tags, etc.
$source = $this->preparSource($source);

// 生成 DOM 解析类
$this->DOM = new DOMDocument('1.0', $input_char);
try {
//libxml_use_internal_errors(true);
// 会有些错误信息，不过不要紧 :^)
if (!@$this->DOM->loadHTML(''.$source)) {
throw new Exception("Parse HTML Error!");
}

foreach ($this->DOM->childNodes as $item) {
if ($item->nodeType == XML_PI_NODE) {
$this->DOM->removeChild($item); // remove hack
}
}

// insert proper
$this->DOM->encoding = Readability::DOM_DEFAULT_CHARSET;
} catch (Exception $e) {
// ...
}
}

/**
* Preprocess HTML tags so that they can be accurately processed by DOM parsing classes
*
* @return String
*/
private function preparSource($string) {
// 剔除多余的 HTML 编码标记，避免解析出错
preg_match("/charset=([＼w|＼-]+);?/", $string, $match);
if (isset($match[1])) {
$string = preg_replace("/charset=([＼w|＼-]+);?/", "", $string, 1);
}

// Replace all doubled-up
tags with

tags, and remove fonts.
$string = preg_replace("/
[ ＼r＼n＼s]*
/i", "

", $string);
$string = preg_replace("/]*>/i", "", $string);

// @see https://github.com/feelinglucky/php-readability/issues/7
// - from http://stackoverflow.com/questions/7130867/remove-script-tag-from-html-content
$string = preg_replace("#<script>(.*?)</script>#is", "", $string);

return trim($string);
}

/**
* Remove all $TagName tags in DOM elements
*
* @return DOMDocument
*/
private function removeJunkTag($RootNode, $TagName) {

$Tags = $RootNode->getElementsByTagName($TagName);

//Note: always index 0, because removing a tag removes it from the results as well.
while($Tag = $Tags->item(0)){
$parentNode = $Tag->parentNode;
$parentNode->removeChild($Tag);
}

return $RootNode;

}

/**
* Remove all unnecessary attributes from the element
*/
private function removeJunkAttr($RootNode, $Attr) {
$Tags = $RootNode->getElementsByTagName("*");

$i = 0;
while($Tag = $Tags->item($i++)) {
$Tag->removeAttribute($Attr);
}

return $RootNode;
}

/**
* Get the box model of the main content of the page based on the rating
* The determination algorithm comes from: http://code.google.com/p/arc90labs-readability/
* This is forwarded by Zheng Xiao’s blog
* @return DOMNode
*/
private function getTopBox() {
// 获得页面所有的章节
$allParagraphs = $this->DOM->getElementsByTagName("p");

// Study all the paragraphs and find the chunk that has the best score.
// A score is determined by things like: Number of

's, commas, special classes, etc.
$i = 0;
while($paragraph = $allParagraphs->item($i++)) {
$parentNode = $paragraph->parentNode;
$contentScore = intval($parentNode->getAttribute(Readability::ATTR_CONTENT_SCORE));
$className = $parentNode->getAttribute("class");
$id = $parentNode->getAttribute("id");

// Look for a special classname
if (preg_match("/(comment|meta|footer|footnote)/i", $className)) {
$contentScore -= 50;
} else if(preg_match(
"/((^|＼＼s)(post|hentry|entry[-]?(content|text|body)?|article[-]?(content|text|body)?)(＼＼s|$))/i",
$className)) {
$contentScore += 25;
}

// Look for a special ID
if (preg_match("/(comment|meta|footer|footnote)/i", $id)) {
$contentScore -= 50;
} else if (preg_match(
"/^(post|hentry|entry[-]?(content|text|body)?|article[-]?(content|text|body)?)$/i",
$id)) {
$contentScore += 25;
}

// Add a point for the paragraph found
// Add points for any commas within this paragraph
if (strlen($paragraph->nodeValue) > 10) {
$contentScore += strlen($paragraph->nodeValue);
}

// 保存父元素的判定得分
$parentNode->setAttribute(Readability::ATTR_CONTENT_SCORE, $contentScore);

// 保存章节的父元素，以便下次快速获取
array_push($this->parentNodes, $parentNode);
}

$topBox = null;

// Assignment from index for performance.
// See http://www.peachpit.com/articles/article.aspx?p=31567&seqNum=5
for ($i = 0, $len = sizeof($this->parentNodes); $i $parentNode = $this->parentNodes[$i];
$contentScore = intval($parentNode->getAttribute(Readability::ATTR_CONTENT_SCORE));
$orgContentScore = intval($topBox ? $topBox->getAttribute(Readability::ATTR_CONTENT_SCORE) : 0);

if ($contentScore && $contentScore > $orgContentScore) {
$topBox = $parentNode;
}
}

// At this time, $topBox should be the main element of the page content that has been determined
return $topBox;
}

/**
* Get HTML page title
*
* @return String
*/
public function getTitle() {
$split_point = ' - ';
$titleNodes = $this->DOM->getElementsByTagName("title");

if ($titleNodes->length
&& $titleNode = $titleNodes->item(0)) {
// @see http://stackoverflow.com/questions/717328/how-to-explode-string-right-to-left
$title = trim($titleNode->nodeValue);
$result = array_map('strrev', explode($split_point, strrev($title)));
return sizeof($result) > 1 ? array_pop($result) : $title;
}

return null;
}

/**
* Get Leading Image Url
*
* @return String
*/
public function getLeadImageUrl($node) {
$images = $node->getElementsByTagName("img");

if ($images->length && $leadImage = $images->item(0)) {
return $leadImage->getAttribute("src");
}

return null;
}

/**
* Get the main content of the page (content after Readability)
*
* @return Array
*/
public function getContent() {
if (!$this->DOM) return false;

//Get page title
$ContentTitle = $this->getTitle();

// Get the main content of the page
$ContentBox = $this->getTopBox();

//Check if we found a suitable top-box.
if($ContentBox === null)
throw new RuntimeException(Readability::MESSAGE_CAN_NOT_GET);

// Copy content to new DOMDocument
$Target = new DOMDocument;
$Target->appendChild($Target->importNode($ContentBox, true));

//Delete unnecessary tags
foreach ($this->junkTags as $tag) {
$Target = $this->removeJunkTag($Target, $tag);
}

//Delete unnecessary attributes
foreach ($this->junkAttrs as $attr) {
$Target = $this->removeJunkAttr($Target, $attr);
}

$content = mb_convert_encoding($Target->saveHTML(), Readability::DOM_DEFAULT_CHARSET, "HTML-ENTITIES");

//Multiple data, returned in the form of array
return Array(
'lead_image_url' => $this->getLeadImageUrl($Target),
'word_count' => mb_strlen(strip_tags($content), Readability::DOM_DEFAULT_CHARSET),
'title' => $ContentTitle ? $ContentTitle : null,
'content' => $content
);
}

function __destruct() { }
}

It is also very simple to use. When instantiating, pass in the html source code and corresponding encoding of the web page, and then directly call its getContent method to return the extracted text part. The extracted article may also contain a small number of links. , you can modify it later

Statement

The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn

PHP and Python: Different Paradigms ExplainedApr 18, 2025 am 12:26 AM

PHP is mainly procedural programming, but also supports object-oriented programming (OOP); Python supports a variety of paradigms, including OOP, functional and procedural programming. PHP is suitable for web development, and Python is suitable for a variety of applications such as data analysis and machine learning.

PHP and Python: A Deep Dive into Their HistoryApr 18, 2025 am 12:25 AM

PHP originated in 1994 and was developed by RasmusLerdorf. It was originally used to track website visitors and gradually evolved into a server-side scripting language and was widely used in web development. Python was developed by Guidovan Rossum in the late 1980s and was first released in 1991. It emphasizes code readability and simplicity, and is suitable for scientific computing, data analysis and other fields.

Choosing Between PHP and Python: A GuideApr 18, 2025 am 12:24 AM

PHP is suitable for web development and rapid prototyping, and Python is suitable for data science and machine learning. 1.PHP is used for dynamic web development, with simple syntax and suitable for rapid development. 2. Python has concise syntax, is suitable for multiple fields, and has a strong library ecosystem.

PHP and Frameworks: Modernizing the LanguageApr 18, 2025 am 12:14 AM

PHP remains important in the modernization process because it supports a large number of websites and applications and adapts to development needs through frameworks. 1.PHP7 improves performance and introduces new features. 2. Modern frameworks such as Laravel, Symfony and CodeIgniter simplify development and improve code quality. 3. Performance optimization and best practices further improve application efficiency.

PHP's Impact: Web Development and BeyondApr 18, 2025 am 12:10 AM

PHPhassignificantlyimpactedwebdevelopmentandextendsbeyondit.1)ItpowersmajorplatformslikeWordPressandexcelsindatabaseinteractions.2)PHP'sadaptabilityallowsittoscaleforlargeapplicationsusingframeworkslikeLaravel.3)Beyondweb,PHPisusedincommand-linescrip

How does PHP type hinting work, including scalar types, return types, union types, and nullable types?Apr 17, 2025 am 12:25 AM

PHP type prompts to improve code quality and readability. 1) Scalar type tips: Since PHP7.0, basic data types are allowed to be specified in function parameters, such as int, float, etc. 2) Return type prompt: Ensure the consistency of the function return value type. 3) Union type prompt: Since PHP8.0, multiple types are allowed to be specified in function parameters or return values. 4) Nullable type prompt: Allows to include null values and handle functions that may return null values.

How does PHP handle object cloning (clone keyword) and the __clone magic method?Apr 17, 2025 am 12:24 AM

In PHP, use the clone keyword to create a copy of the object and customize the cloning behavior through the \_\_clone magic method. 1. Use the clone keyword to make a shallow copy, cloning the object's properties but not the object's properties. 2. The \_\_clone method can deeply copy nested objects to avoid shallow copying problems. 3. Pay attention to avoid circular references and performance problems in cloning, and optimize cloning operations to improve efficiency.

PHP vs. Python: Use Cases and ApplicationsApr 17, 2025 am 12:23 AM

PHP is suitable for web development and content management systems, and Python is suitable for data science, machine learning and automation scripts. 1.PHP performs well in building fast and scalable websites and applications and is commonly used in CMS such as WordPress. 2. Python has performed outstandingly in the fields of data science and machine learning, with rich libraries such as NumPy and TensorFlow.

See all articles

Hot AI Tools

Undresser.AI Undress

AI-powered app for creating realistic nude photos

AI Clothes Remover

Online AI tool for removing clothes from photos.

Undress AI Tool

Undress images for free

Clothoff.io

AI clothes remover

AI Hentai Generator

Generate AI Hentai for free.

Hot Article

R.E.P.O. Energy Crystals Explained and What They Do (Yellow Crystal)

1 months agoBy尊渡假赌尊渡假赌尊渡假赌

R.E.P.O. Best Graphic Settings

1 months agoBy尊渡假赌尊渡假赌尊渡假赌

Assassin's Creed Shadows: Seashell Riddle Solution

3 weeks agoByDDD

What's New in Windows 11 KB5054979 & How to Fix Update Issues

2 weeks agoByDDD

Will R.E.P.O. Have Crossplay?

1 months agoBy尊渡假赌尊渡假赌尊渡假赌

Hot Tools

Safe Exam Browser

Safe Exam Browser is a secure browser environment for taking online exams securely. This software turns any computer into a secure workstation. It controls access to any utility and prevents students from using unauthorized resources.

WebStorm Mac version

Useful JavaScript development tools

SAP NetWeaver Server Adapter for Eclipse

Integrate Eclipse with SAP NetWeaver application server.

MinGW - Minimalist GNU for Windows

This project is in the process of being migrated to osdn.net/projects/mingw, you can continue to follow us there. MinGW: A native Windows port of the GNU Compiler Collection (GCC), freely distributable import libraries and header files for building native Windows applications; includes extensions to the MSVC runtime to support C99 functionality. All MinGW software can run on 64-bit Windows platforms.

Atom editor mac version download

The most popular open source editor

Hot Topics

Where is the login entrance for gmail email?

7554

CakePHP Tutorial

1382

What is the format of the account name of steam

win11 activation key permanent

nyt connections hints and answers