搜尋
首頁後端開發Python教學如何写个爬虫程序扒下知乎某个回答所有点赞用户名单?

问一个人知乎账号想fo一下,结果她不告诉我,现在想想有点奇怪,有点好奇,好在我知道几个她点赞过的问题,想用社会工程学的方法筛选下,找出她的知乎账号。(匿了没法邀请,算了她应该不会来这个区)

回复内容:

每个回答的div里面都有一个叫 data-aid="12345678"的东西,
然后根据, www.zhihu.com/answer/12345678/voters_profile?&offset=10
这个json数据连接分析所有点赞的id和个人连接就行, 每10的点赞人数为一个json连接.

刚刚试了一下,需要登陆之后才能得到完整的数据, 登陆知乎可以参考我写的博客.
python模拟登陆知乎


比如我这个回答的data-aid = '22229844'
如何写个爬虫程序扒下知乎某个回答所有点赞用户名单?
<span class="c">#!/usr/bin/env python</span>
<span class="c"># -*- coding: utf-8 -*-</span>

<span class="kn">import</span> <span class="nn">requests</span>
<span class="kn">from</span> <span class="nn">bs4</span> <span class="kn">import</span> <span class="n">BeautifulSoup</span>
<span class="kn">import</span> <span class="nn">time</span>
<span class="kn">import</span> <span class="nn">json</span>
<span class="kn">import</span> <span class="nn">os</span>
<span class="kn">import</span> <span class="nn">sys</span>

<span class="n">url</span> <span class="o">=</span> <span class="s">'http://www.zhihu.com'</span>
<span class="n">loginURL</span> <span class="o">=</span> <span class="s">'http://www.zhihu.com/login/email'</span>

<span class="n">headers</span> <span class="o">=</span> <span class="p">{</span>
    <span class="s">"User-Agent"</span><span class="p">:</span> <span class="s">'Mozilla/5.0 (Macintosh; Intel Mac OS X 10.10; rv:41.0) Gecko/20100101 Firefox/41.0'</span><span class="p">,</span>
    <span class="s">"Referer"</span><span class="p">:</span> <span class="s">"http://www.zhihu.com/"</span><span class="p">,</span>
    <span class="s">'Host'</span><span class="p">:</span> <span class="s">'www.zhihu.com'</span><span class="p">,</span>
<span class="p">}</span>

<span class="n">data</span> <span class="o">=</span> <span class="p">{</span>
    <span class="s">'email'</span><span class="p">:</span> <span class="s">'</span>
<span class="n">xxxxx</span><span class="nd">@gmail.com</span><span class="s">',</span>
    <span class="s">'password'</span><span class="p">:</span> <span class="s">'</span>
<span class="n">xxxxxxx</span><span class="s">',</span>
    <span class="s">'rememberme'</span><span class="p">:</span> <span class="s">"true"</span><span class="p">,</span>
<span class="p">}</span>

<span class="n">s</span> <span class="o">=</span> <span class="n">requests</span><span class="o">.</span><span class="n">session</span><span class="p">()</span>
<span class="c"># 如果成功登陆过,用保存的cookies登录</span>
<span class="k">if</span> <span class="n">os</span><span class="o">.</span><span class="n">path</span><span class="o">.</span><span class="n">exists</span><span class="p">(</span><span class="s">'cookiefile'</span><span class="p">):</span>
    <span class="k">with</span> <span class="nb">open</span><span class="p">(</span><span class="s">'cookiefile'</span><span class="p">)</span> <span class="k">as</span> <span class="n">f</span><span class="p">:</span>
        <span class="n">cookie</span> <span class="o">=</span> <span class="n">json</span><span class="o">.</span><span class="n">load</span><span class="p">(</span><span class="n">f</span><span class="p">)</span>
    <span class="n">s</span><span class="o">.</span><span class="n">cookies</span><span class="o">.</span><span class="n">update</span><span class="p">(</span><span class="n">cookie</span><span class="p">)</span>
    <span class="n">req1</span> <span class="o">=</span> <span class="n">s</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="n">url</span><span class="p">,</span> <span class="n">headers</span><span class="o">=</span><span class="n">headers</span><span class="p">)</span>
    <span class="k">with</span> <span class="nb">open</span><span class="p">(</span><span class="s">'zhihu.html'</span><span class="p">,</span> <span class="s">'w'</span><span class="p">)</span> <span class="k">as</span> <span class="n">f</span><span class="p">:</span>
        <span class="n">f</span><span class="o">.</span><span class="n">write</span><span class="p">(</span><span class="n">req1</span><span class="o">.</span><span class="n">content</span><span class="p">)</span>
<span class="c"># 第一次需要手动输入验证码登录</span>
<span class="k">else</span><span class="p">:</span>
    <span class="n">req</span> <span class="o">=</span> <span class="n">s</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="n">url</span><span class="p">,</span> <span class="n">headers</span><span class="o">=</span><span class="n">headers</span><span class="p">)</span>
    <span class="k">print</span> <span class="n">req</span>

    <span class="n">soup</span> <span class="o">=</span> <span class="n">BeautifulSoup</span><span class="p">(</span><span class="n">req</span><span class="o">.</span><span class="n">text</span><span class="p">,</span> <span class="s">"html.parser"</span><span class="p">)</span>
    <span class="n">xsrf</span> <span class="o">=</span> <span class="n">soup</span><span class="o">.</span><span class="n">find</span><span class="p">(</span><span class="s">'input'</span><span class="p">,</span> <span class="p">{</span><span class="s">'name'</span><span class="p">:</span> <span class="s">'_xsrf'</span><span class="p">,</span> <span class="s">'type'</span><span class="p">:</span> <span class="s">'hidden'</span><span class="p">})</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s">'value'</span><span class="p">)</span>

    <span class="n">data</span><span class="p">[</span><span class="s">'_xsrf'</span><span class="p">]</span> <span class="o">=</span> <span class="n">xsrf</span>

    <span class="n">timestamp</span> <span class="o">=</span> <span class="nb">int</span><span class="p">(</span><span class="n">time</span><span class="o">.</span><span class="n">time</span><span class="p">()</span> <span class="o">*</span> <span class="mi">1000</span><span class="p">)</span>
    <span class="n">captchaURL</span> <span class="o">=</span> <span class="s">'http://www.zhihu.com/captcha.gif?='</span> <span class="o">+</span> <span class="nb">str</span><span class="p">(</span><span class="n">timestamp</span><span class="p">)</span>
    <span class="k">print</span> <span class="n">captchaURL</span>

    <span class="k">with</span> <span class="nb">open</span><span class="p">(</span><span class="s">'zhihucaptcha.gif'</span><span class="p">,</span> <span class="s">'wb'</span><span class="p">)</span> <span class="k">as</span> <span class="n">f</span><span class="p">:</span>
        <span class="n">captchaREQ</span> <span class="o">=</span> <span class="n">s</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="n">captchaURL</span><span class="p">)</span>
        <span class="n">f</span><span class="o">.</span><span class="n">write</span><span class="p">(</span><span class="n">captchaREQ</span><span class="o">.</span><span class="n">content</span><span class="p">)</span>
    <span class="n">loginCaptcha</span> <span class="o">=</span> <span class="nb">raw_input</span><span class="p">(</span><span class="s">'input captcha:</span><span class="se">\n</span><span class="s">'</span><span class="p">)</span><span class="o">.</span><span class="n">strip</span><span class="p">()</span>
    <span class="n">data</span><span class="p">[</span><span class="s">'captcha'</span><span class="p">]</span> <span class="o">=</span> <span class="n">loginCaptcha</span>
    <span class="c"># print data</span>
    <span class="n">loginREQ</span> <span class="o">=</span> <span class="n">s</span><span class="o">.</span><span class="n">post</span><span class="p">(</span><span class="n">loginURL</span><span class="p">,</span>  <span class="n">headers</span><span class="o">=</span><span class="n">headers</span><span class="p">,</span> <span class="n">data</span><span class="o">=</span><span class="n">data</span><span class="p">)</span>
    <span class="c"># print loginREQ.url</span>
    <span class="c"># print s.cookies.get_dict()</span>
    <span class="k">if</span> <span class="ow">not</span> <span class="n">loginREQ</span><span class="o">.</span><span class="n">json</span><span class="p">()[</span><span class="s">'r'</span><span class="p">]:</span>
        <span class="c"># print loginREQ.json()</span>
        <span class="k">with</span> <span class="nb">open</span><span class="p">(</span><span class="s">'cookiefile'</span><span class="p">,</span> <span class="s">'wb'</span><span class="p">)</span> <span class="k">as</span> <span class="n">f</span><span class="p">:</span>
            <span class="n">json</span><span class="o">.</span><span class="n">dump</span><span class="p">(</span><span class="n">s</span><span class="o">.</span><span class="n">cookies</span><span class="o">.</span><span class="n">get_dict</span><span class="p">(),</span> <span class="n">f</span><span class="p">)</span>
    <span class="k">else</span><span class="p">:</span>
        <span class="k">print</span> <span class="s">'login failed, try again!'</span>
        <span class="n">sys</span><span class="o">.</span><span class="n">exit</span><span class="p">(</span><span class="mi">1</span><span class="p">)</span>

<span class="c"># 以http://www.zhihu.com/question/27621722/answer/48820436这个大神的399各赞为例子.</span>
<span class="n">zanBaseURL</span> <span class="o">=</span> <span class="s">'http://www.zhihu.com/answer/22229844/voters_profile?&offset={0}'</span>
<span class="n">page</span> <span class="o">=</span> <span class="mi">0</span>
<span class="n">count</span> <span class="o">=</span> <span class="mi">0</span>
<span class="k">while</span> <span class="mi">1</span><span class="p">:</span>
    <span class="n">zanURL</span> <span class="o">=</span> <span class="n">zanBaseURL</span><span class="o">.</span><span class="n">format</span><span class="p">(</span><span class="nb">str</span><span class="p">(</span><span class="n">page</span><span class="p">))</span>
    <span class="n">page</span> <span class="o">+=</span> <span class="mi">10</span>
    <span class="n">zanREQ</span> <span class="o">=</span> <span class="n">s</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="n">zanURL</span><span class="p">,</span> <span class="n">headers</span><span class="o">=</span><span class="n">headers</span><span class="p">)</span>
    <span class="n">zanData</span> <span class="o">=</span> <span class="n">zanREQ</span><span class="o">.</span><span class="n">json</span><span class="p">()[</span><span class="s">'payload'</span><span class="p">]</span>
    <span class="k">if</span> <span class="ow">not</span> <span class="n">zanData</span><span class="p">:</span>
        <span class="k">break</span>
    <span class="k">for</span> <span class="n">item</span> <span class="ow">in</span> <span class="n">zanData</span><span class="p">:</span>
        <span class="c"># print item</span>
        <span class="n">zanSoup</span> <span class="o">=</span> <span class="n">BeautifulSoup</span><span class="p">(</span><span class="n">item</span><span class="p">,</span> <span class="s">"html.parser"</span><span class="p">)</span>
        <span class="n">zanInfo</span> <span class="o">=</span> <span class="n">zanSoup</span><span class="o">.</span><span class="n">find</span><span class="p">(</span><span class="s">'a'</span><span class="p">,</span> <span class="p">{</span><span class="s">'target'</span><span class="p">:</span> <span class="s">"_blank"</span><span class="p">,</span> <span class="s">'class'</span><span class="p">:</span> <span class="s">'zg-link'</span><span class="p">})</span>
        <span class="k">if</span> <span class="n">zanInfo</span><span class="p">:</span>
            <span class="k">print</span> <span class="s">'nickname:'</span><span class="p">,</span> <span class="n">zanInfo</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s">'title'</span><span class="p">),</span>  <span class="s">'    '</span><span class="p">,</span>
            <span class="k">print</span> <span class="s">'person_url:'</span><span class="p">,</span> <span class="n">zanInfo</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s">'href'</span><span class="p">)</span>
        <span class="k">else</span><span class="p">:</span>
            <span class="n">anonymous</span> <span class="o">=</span> <span class="n">zanSoup</span><span class="o">.</span><span class="n">find</span><span class="p">(</span>
                <span class="s">'img'</span><span class="p">,</span> <span class="p">{</span><span class="s">'title'</span><span class="p">:</span> <span class="bp">True</span><span class="p">,</span> <span class="s">'class'</span><span class="p">:</span> <span class="s">"zm-item-img-avatar"</span><span class="p">})</span>
            <span class="k">print</span> <span class="s">'nickname:'</span><span class="p">,</span> <span class="n">anonymous</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s">'title'</span><span class="p">)</span>

        <span class="n">count</span> <span class="o">+=</span> <span class="mi">1</span>
    <span class="k">print</span> <span class="n">count</span>
这里有个 Python 3 的项目 Zhihu-py3 7sDream/zhihu-py3 · GitHub

封装了知乎爬虫的各方面需求,比如获取用户信息,获取问题信息,获取答案信息,之类的……当然也包括点赞用户啥的,虽然是单线程同步式 但是平常用用还是可以滴。

这里是她的文档:Welcome to zhihu-py3’s documentation!

欢迎 Star 以及 Fork 或者贡献代码。

===

获取点赞用户灰常简单 大概就这样

<span class="kn">from</span> <span class="nn">zhihu</span> <span class="kn">import</span> <span class="n">ZhihuClient</span>

<span class="n">client</span> <span class="o">=</span> <span class="n">ZhihuClient</span><span class="p">(</span><span class="s">'cookies.json'</span><span class="p">)</span>

<span class="n">url</span> <span class="o">=</span> <span class="s">'http://www.zhihu.com/question/36338520/answer/67029821'</span>
<span class="n">answer</span> <span class="o">=</span> <span class="n">client</span><span class="o">.</span><span class="n">answer</span><span class="p">(</span><span class="n">url</span><span class="p">)</span>

<span class="k">print</span><span class="p">(</span><span class="s">'问题:{0}'</span><span class="o">.</span><span class="n">format</span><span class="p">(</span><span class="n">answer</span><span class="o">.</span><span class="n">question</span><span class="o">.</span><span class="n">title</span><span class="p">))</span>
<span class="k">print</span><span class="p">(</span><span class="s">'答主:{0}'</span><span class="o">.</span><span class="n">format</span><span class="p">(</span><span class="n">answer</span><span class="o">.</span><span class="n">author</span><span class="o">.</span><span class="n">name</span><span class="p">))</span>
<span class="k">print</span><span class="p">(</span><span class="s">'此答案共有{0}人点赞:</span><span class="se">\n</span><span class="s">'</span><span class="o">.</span><span class="n">format</span><span class="p">(</span><span class="n">answer</span><span class="o">.</span><span class="n">upvote_num</span><span class="p">))</span>

<span class="k">for</span> <span class="n">upvoter</span> <span class="ow">in</span> <span class="n">answer</span><span class="o">.</span><span class="n">upvoters</span><span class="p">:</span>
    <span class="k">print</span><span class="p">(</span><span class="n">upvoter</span><span class="o">.</span><span class="n">name</span><span class="p">,</span> <span class="n">upvoter</span><span class="o">.</span><span class="n">url</span><span class="p">)</span>
看到第一名的答案中的评论,补充一下如何发现aid这个关键特征的思路:

一句话概述:对人工操作时发送的HTTP Request/Response进行分析,找出关键定位特征。
工具:firebug

1. 点击 任意一个答案页面下面的超链接 等人赞同 发现会向类似于这样的
zhihu.com/answer/222298
URL发送数据。
从这个URL的格式上已经很容易猜到这就是给答案22229844 进行投票的投票者资料了,一看服务器返回的Response (一段JSON数据)也能说明这一点。那么只要我们可以向这个URL发送一段GET请求就能知道投票者了。剩下的就是要解决怎么找出这个URL的问题,也就是找到这个22229844 。
2. 既然知道当点击 等人赞同 会触发一段Ajax向这个URL发送请求,那这个22229844 要么在DOM中存储了,要么是计算出来的。既然如此,在DOM中搜索22229844这个字符串,很轻松就能找到这样的一个div:
<span class="nt"><div</span> <span class="na">data-copyable=</span><span class="s">"1"</span> <span class="na">data-isowner=</span><span class="s">"0"</span> <span class="na">data-helpful=</span><span class="s">"1"</span> <span class="na">data-deleted=</span><span class="s">"0"</span> <span class="na">data-created=</span><span class="s">"1444404675"</span> <span class="na">data-collapsed=</span><span class="s">"0"</span> <span class="na">data-atoken=</span><span class="s">"67029821"</span> <span class="na">data-author=</span><span class="s">"洛克"</span> <span class="na">data-qtoken=</span><span class="s">"36338520"</span> <span class="na">data-aid=</span><span class="s">"22229844"</span> <span class="na">itemtype=</span><span class="s">"http://schema.org/Answer"</span> <span class="na">itemscope=</span><span class="s">""</span> <span class="na">itemprop=</span><span class="s">"topAnswer"</span> <span class="na">class=</span><span class="s">"zm-item-answer"</span> <span class="na">tabindex=</span><span class="s">"-1"</span><span class="nt">></span>
我好奇的是,你说的社会工程学是啥?
据我所知,一般所谓社会工程学就是黑客的骗术,凭借已知信息骗取信任拿到自己要的信息,但是核心就是骗。
你现在是知道她点了哪个答案的赞,跟社会工程学有什么关系呢?
你是想说你知道她点过的多个答案,准备从同时赞过这些答案的人中找到她?
运气好可能一下子就找出来了,运气不好恐怕一堆候选人等着你。关键看你知道她赞过几个答案了。
论技术的话,我觉得用不着python写js在控制台跑就好了 找轮子哥 他有源码 轮子哥有爬取用户自动分析性别颜值值得关注程度的源码
陳述
本文內容由網友自願投稿,版權歸原作者所有。本站不承擔相應的法律責任。如發現涉嫌抄襲或侵權的內容,請聯絡admin@php.cn
Python和時間:充分利用您的學習時間Python和時間:充分利用您的學習時間Apr 14, 2025 am 12:02 AM

要在有限的時間內最大化學習Python的效率,可以使用Python的datetime、time和schedule模塊。 1.datetime模塊用於記錄和規劃學習時間。 2.time模塊幫助設置學習和休息時間。 3.schedule模塊自動化安排每週學習任務。

Python:遊戲,Guis等Python:遊戲,Guis等Apr 13, 2025 am 12:14 AM

Python在遊戲和GUI開發中表現出色。 1)遊戲開發使用Pygame,提供繪圖、音頻等功能,適合創建2D遊戲。 2)GUI開發可選擇Tkinter或PyQt,Tkinter簡單易用,PyQt功能豐富,適合專業開發。

Python vs.C:申請和用例Python vs.C:申請和用例Apr 12, 2025 am 12:01 AM

Python适合数据科学、Web开发和自动化任务,而C 适用于系统编程、游戏开发和嵌入式系统。Python以简洁和强大的生态系统著称,C 则以高性能和底层控制能力闻名。

2小時的Python計劃:一種現實的方法2小時的Python計劃:一種現實的方法Apr 11, 2025 am 12:04 AM

2小時內可以學會Python的基本編程概念和技能。 1.學習變量和數據類型,2.掌握控制流(條件語句和循環),3.理解函數的定義和使用,4.通過簡單示例和代碼片段快速上手Python編程。

Python:探索其主要應用程序Python:探索其主要應用程序Apr 10, 2025 am 09:41 AM

Python在web開發、數據科學、機器學習、自動化和腳本編寫等領域有廣泛應用。 1)在web開發中,Django和Flask框架簡化了開發過程。 2)數據科學和機器學習領域,NumPy、Pandas、Scikit-learn和TensorFlow庫提供了強大支持。 3)自動化和腳本編寫方面,Python適用於自動化測試和系統管理等任務。

您可以在2小時內學到多少python?您可以在2小時內學到多少python?Apr 09, 2025 pm 04:33 PM

兩小時內可以學到Python的基礎知識。 1.學習變量和數據類型,2.掌握控制結構如if語句和循環,3.了解函數的定義和使用。這些將幫助你開始編寫簡單的Python程序。

如何在10小時內通過項目和問題驅動的方式教計算機小白編程基礎?如何在10小時內通過項目和問題驅動的方式教計算機小白編程基礎?Apr 02, 2025 am 07:18 AM

如何在10小時內教計算機小白編程基礎?如果你只有10個小時來教計算機小白一些編程知識,你會選擇教些什麼�...

如何在使用 Fiddler Everywhere 進行中間人讀取時避免被瀏覽器檢測到?如何在使用 Fiddler Everywhere 進行中間人讀取時避免被瀏覽器檢測到?Apr 02, 2025 am 07:15 AM

使用FiddlerEverywhere進行中間人讀取時如何避免被檢測到當你使用FiddlerEverywhere...

See all articles

熱AI工具

Undresser.AI Undress

Undresser.AI Undress

人工智慧驅動的應用程序,用於創建逼真的裸體照片

AI Clothes Remover

AI Clothes Remover

用於從照片中去除衣服的線上人工智慧工具。

Undress AI Tool

Undress AI Tool

免費脫衣圖片

Clothoff.io

Clothoff.io

AI脫衣器

AI Hentai Generator

AI Hentai Generator

免費產生 AI 無盡。

熱門文章

R.E.P.O.能量晶體解釋及其做什麼(黃色晶體)
3 週前By尊渡假赌尊渡假赌尊渡假赌
R.E.P.O.最佳圖形設置
3 週前By尊渡假赌尊渡假赌尊渡假赌
R.E.P.O.如果您聽不到任何人,如何修復音頻
3 週前By尊渡假赌尊渡假赌尊渡假赌
WWE 2K25:如何解鎖Myrise中的所有內容
4 週前By尊渡假赌尊渡假赌尊渡假赌

熱工具

ZendStudio 13.5.1 Mac

ZendStudio 13.5.1 Mac

強大的PHP整合開發環境

SublimeText3 英文版

SublimeText3 英文版

推薦:為Win版本,支援程式碼提示!

DVWA

DVWA

Damn Vulnerable Web App (DVWA) 是一個PHP/MySQL的Web應用程序,非常容易受到攻擊。它的主要目標是成為安全專業人員在合法環境中測試自己的技能和工具的輔助工具,幫助Web開發人員更好地理解保護網路應用程式的過程,並幫助教師/學生在課堂環境中教授/學習Web應用程式安全性。 DVWA的目標是透過簡單直接的介面練習一些最常見的Web漏洞,難度各不相同。請注意,該軟體中

SublimeText3漢化版

SublimeText3漢化版

中文版,非常好用

EditPlus 中文破解版

EditPlus 中文破解版

體積小,語法高亮,不支援程式碼提示功能