PHP抽取网页标题并剔除不相关的seo关键字

首页

后端开发

php教程

PHP抽取网页标题并剔除不相关的seo关键字_PHP教程

WBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWB

Jul 13, 2016 pm 05:44 PM

phpseo关键字在场景我们抽取标题的网页

场景描述:

过往我们在抽取网页标题的时候，都会直接抽取之间的内容. 但实际情况是这样,例如javaeye 的一篇文章 http://www.iteye.com/news/21643 , 的内容为 "10年软件开发教会我最重要的10件事 - 非技术 - ITeye资讯", 但实际引用中我们期望的标题应该为 "10年软件开发教会我最重要的10件事". 所以标题后面堆砌了很多不相关的关键字(应该是为了 seo 吧). 所以我们希望过滤掉这些关键字. 有下面的方法可以参考:

1. 查找 h1 等标签.(分析sina news 一些网站之后, 觉得不可行,会有很多干扰)

2. 从全文去标题后，将之间的内容切割(按 _ | -)为 a1,a2,a3,a4，然后从最长的词组a3开始从全文查找. 如果查找成功,那么开始向左边迭代查询 a2,a1,直到查询失败为止。左侧失败后，再继续向右迭代，同理. (这里我采用的是这种方法)

Php代码
/**
* @author pqcc
* @date: 2011-06-18
* Description: 给定一个网页内容，提取网页的标题. 提取的标题不包括 seo 关键字.
* e.g: 一篇新闻标题的从直接抽取结果为 "大学英语四六级本周六开考 909万人参考_新浪教育_新浪网", * 但我们希望的结果是:"大学英语四六级本周六开考 909万人参考". * 适用范围: 文章最终页标题的提取, 不包括专题页等. */ class TitlePurify{ private $matches_preg = [-_s|—]; function getTitle($contents){/*{{{*/ $preg = "/<title>]*>([w| ||W]*?)/i";
 preg_match($preg, $contents, $matches);
 if(count($matches) return "标题抽取失败";
 }
 $title = $matches[1];
 return $this->trimTitle($title, $contents);
 }/*}}}*/

 function trimMeta($contents){/*{{{*/
 // 首先去除内容, <meta> 内容. $preg = "/<title>]*>([w| ||W]*?)/i";
 $contents = preg_replace($preg, , $contents);
 $preg = "/]*>/i";
 $contents = preg_replace($preg, , $contents);
 return $contents;
 }/*}}}*/

 // 获取长度最长的 item 所处的index.
 function getMaxIndex($titles){/*{{{*/
 $maxItemIndex = 0;
 $maxLength = 0;
 $loop = 0;
 foreach($titles as $item){
 if(strlen($item)>$maxLength){
 $maxLength = strlen($item);
 $maxItemIndex = $loop;
 }
 $loop++;
 }
 return $maxItemIndex;
 }/*}}}*/

 function trim($title, $titles, $contents, $maxItemIndex){/*{{{*/
 //@todo : 此处可优化contents
 // 如果查找成功. result = tempTitle.
 $tempTitle = $titles[$maxItemIndex];
 $result = $tempTitle;
 $count = count($titles);
 // while 从当前index 向左进行迭代(直到到达第一个或者匹配失败才中止).
 $leftIndex = $maxItemIndex-1;
 while(true && $leftIndex>=0){
 // tempTitle+左一个.
 preg_match("/({$this->matches_preg}+{$tempTitle})/i", $title, $matches);
 if(count($matches)>1){
 // temp 用于匹配失败后,进行回滚.
 $temp = $titles[$leftIndex] . $matches[1];
 $tempTitle = $titles[$leftIndex] . $matches[1];
 // 继续拿着 tempTitle 去匹配.
 preg_match("/$tempTitle/i", $contents, $matches);
 // 如果查找失败....
 if(count($matches) $tempTitle = $temp;
 break;
 }else{
 $result = $tempTitle;
 }
 }else{ // 正常情况下, 不会出现该情况.
 break;
 }
 $leftIndex--;&

声明

本文内容由网友自发贡献，版权归原作者所有，本站不承担相应法律责任。如您发现有涉嫌抄袭侵权的内容，请联系admin@php.cn

PHP电子邮件：分步发送指南May 09, 2025 am 12:14 AM

phpisusedforsendendemailsduetoitsignegrationwithservermailservicesand andexternalsmtpproviders，自动化notifications andMarketingCampaigns.1）设置设置yourphpenvironcormentswironmentswithaweberswithawebserverserverserverandphp，确保themailfunctionisenabled.2）useabasicscruct

如何通过PHP发送电子邮件：示例和代码May 09, 2025 am 12:13 AM

发送电子邮件的最佳方法是使用PHPMailer库。1)使用mail()函数简单但不可靠，可能导致邮件进入垃圾邮件或无法送达。2)PHPMailer提供更好的控制和可靠性，支持HTML邮件、附件和SMTP认证。3)确保正确配置SMTP设置并使用加密（如STARTTLS或SSL/TLS）以增强安全性。4)对于大量邮件，考虑使用邮件队列系统来优化性能。

高级PHP电子邮件：自定义标题和功能May 09, 2025 am 12:13 AM

CustomHeadersheadersandAdvancedFeaturesInphpeMailenHanceFunctionalityAndreliability.1）CustomHeadersheadersheadersaddmetadatatatatataatafortrackingandCategorization.2）htmlemailsallowformattingandttinganditive.3）attachmentscanmentscanmentscanbesmentscanbestmentscanbesentscanbesentingslibrarieslibrarieslibrariesliblarikelikephpmailer.4）smtppapapairatienticationaltication enterticationallimpr

使用PHP和SMTP发送电子邮件的指南May 09, 2025 am 12:06 AM

使用PHP和SMTP发送邮件可以通过PHPMailer库实现。1)安装并配置PHPMailer，2)设置SMTP服务器细节，3)定义邮件内容，4)发送邮件并处理错误。使用此方法可以确保邮件的可靠性和安全性。

使用PHP发送电子邮件的最佳方法是什么？May 08, 2025 am 12:21 AM

ThebestapproachforsendingemailsinPHPisusingthePHPMailerlibraryduetoitsreliability,featurerichness,andeaseofuse.PHPMailersupportsSMTP,providesdetailederrorhandling,allowssendingHTMLandplaintextemails,supportsattachments,andenhancessecurity.Foroptimalu

PHP中依赖注入的最佳实践May 08, 2025 am 12:21 AM

使用依赖注入(DI)的原因是它促进了代码的松耦合、可测试性和可维护性。1)使用构造函数注入依赖，2)避免使用服务定位器，3)利用依赖注入容器管理依赖，4)通过注入依赖提高测试性，5)避免过度注入依赖，6)考虑DI对性能的影响。

PHP性能调整技巧和技巧May 08, 2025 am 12:20 AM

phperformancetuningiscialbecapeitenhancesspeedandeffice，whatevitalforwebapplications.1）cachingwithapcureduccureducesdatabaseloadprovesrovesponsemetimes.2）优化

PHP电子邮件安全性：发送电子邮件的最佳实践May 08, 2025 am 12:16 AM

ThebestpracticesforsendingemailssecurelyinPHPinclude:1)UsingsecureconfigurationswithSMTPandSTARTTLSencryption,2)Validatingandsanitizinginputstopreventinjectionattacks,3)EncryptingsensitivedatawithinemailsusingOpenSSL,4)Properlyhandlingemailheaderstoa

See all articles