Maison >développement back-end >Tutoriel Python >python爬虫beta版之抓取知乎单页面

python爬虫beta版之抓取知乎单页面

高洛峰original: 2016-12-02 16:51:451879parcourir

鉴于之前用python写爬虫，帮运营人员抓取过京东的商品品牌以及分类，这次也是用python来搞简单的抓取单页面版，后期再补充哈。

#-*- coding: UTF-8 -*- 
import requests
import sys
from bs4 import BeautifulSoup

#－－－－－－知乎答案收集－－－－－－－－－－

#获取网页body里的内容
def get_content(url , data = None):
    header={
        &#39;Accept&#39;: &#39;text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8&#39;,
        &#39;Accept-Encoding&#39;: &#39;gzip, deflate, sdch&#39;,
        &#39;Accept-Language&#39;: &#39;zh-CN,zh;q=0.8&#39;,
        &#39;Connection&#39;: &#39;keep-alive&#39;,
        &#39;User-Agent&#39;: &#39;Mozilla/5.0 (Windows NT 6.3; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/43.0.235&#39;
    }

    req = requests.get(url, headers=header)
    req.encoding = &#39;utf-8&#39;
    bs = BeautifulSoup(req.text, "html.parser")  # 创建BeautifulSoup对象
    body = bs.body # 获取body部分
    return body

#获取问题标题
def get_title(html_text):
     data = html_text.find(&#39;span&#39;, {&#39;class&#39;: &#39;zm-editable-content&#39;})
     return data.string.encode(&#39;utf-8&#39;)

#获取问题内容
def get_question_content(html_text):
     data = html_text.find(&#39;div&#39;, {&#39;class&#39;: &#39;zm-editable-content&#39;})
     if data.string is None:
         out = &#39;&#39;;
         for datastring in data.strings:
             out = out + datastring.encode(&#39;utf-8&#39;)
         print &#39;内容：\n&#39; + out
     else:
         print &#39;内容：\n&#39; + data.string.encode(&#39;utf-8&#39;)

#获取点赞数
def get_answer_agree(body):
    agree = body.find(&#39;span&#39;,{&#39;class&#39;: &#39;count&#39;})
    print &#39;点赞数：&#39; + agree.string.encode(&#39;utf-8&#39;) + &#39;\n&#39;

#获取答案
def get_response(html_text):
     response = html_text.find_all(&#39;div&#39;, {&#39;class&#39;: &#39;zh-summary summary clearfix&#39;})
     for index in range(len(response)):
         #获取标签
         answerhref = response[index].find(&#39;a&#39;, {&#39;class&#39;: &#39;toggle-expand&#39;})
         if not(answerhref[&#39;href&#39;].startswith(&#39;javascript&#39;)):
             url = &#39;http://www.zhihu.com/&#39; + answerhref[&#39;href&#39;]
             print url
             body = get_content(url)
             get_answer_agree(body)
             answer = body.find(&#39;div&#39;, {&#39;class&#39;: &#39;zm-editable-content clearfix&#39;})
             if answer.string is None:
                 out = &#39;&#39;;
                 for datastring in answer.strings:
                     out = out + &#39;\n&#39; + datastring.encode(&#39;utf-8&#39;)
                 print out
             else:
                 print answer.string.encode(&#39;utf-8&#39;)


html_text = get_content(&#39;https://www.zhihu.com/question/43879769&#39;)
title = get_title(html_text)
print "标题：\n" + title + &#39;\n&#39;
questiondata = get_question_content(html_text)
print &#39;\n&#39;
data = get_response(html_text)

输出结果：

Déclaration：

Le contenu de cet article est volontairement contribué par les internautes et les droits d'auteur appartiennent à l'auteur original. Ce site n'assume aucune responsabilité légale correspondante. Si vous trouvez un contenu suspecté de plagiat ou de contrefaçon, veuillez contacter admin@php.cn

Article précédent：python中round(x,[n])的使用Article suivant：python 数据类型 ---字符串

Articles Liés

Voir plus