首页 >后端开发 >Python教程 >分享十种py3爬取网页资源的方法

分享十种py3爬取网页资源的方法

Y2J原创: 2017-05-11 11:04:142986浏览

这两天学习了python3实现抓取网页资源的方法，发现了很多种方法，所以，今天添加一点小笔记。

1、最简单

import urllib.request
response = urllib.request.urlopen(&#39;http://python.org/&#39;)
html = response.read()

2、使用 Request

import urllib.request
 
req = urllib.request.Request(&#39;http://python.org/&#39;)
response = urllib.request.urlopen(req)
the_page = response.read()

3、发送数据

#! /usr/bin/env python3
 
import urllib.parse
import urllib.request
 
url = &#39;http://localhost/login.php&#39;
user_agent = &#39;Mozilla/4.0 (compatible; MSIE 5.5; Windows NT)&#39;
values = {
     &#39;act&#39; : &#39;login&#39;,
     &#39;login[email]&#39; : &#39;yzhang@i9i8.com&#39;,
     &#39;login[password]&#39; : &#39;123456&#39;
     }
 
data = urllib.parse.urlencode(values)
req = urllib.request.Request(url, data)
req.add_header(&#39;Referer&#39;, &#39;http://www.python.org/&#39;)
response = urllib.request.urlopen(req)
the_page = response.read()
 
print(the_page.decode("utf8"))

4、发送数据和header

#! /usr/bin/env python3
 
import urllib.parse
import urllib.request
 
url = &#39;http://localhost/login.php&#39;
user_agent = &#39;Mozilla/4.0 (compatible; MSIE 5.5; Windows NT)&#39;
values = {
     &#39;act&#39; : &#39;login&#39;,
     &#39;login[email]&#39; : &#39;yzhang@i9i8.com&#39;,
     &#39;login[password]&#39; : &#39;123456&#39;
     }
headers = { &#39;User-Agent&#39; : user_agent }
 
data = urllib.parse.urlencode(values)
req = urllib.request.Request(url, data, headers)
response = urllib.request.urlopen(req)
the_page = response.read()
 
print(the_page.decode("utf8"))

5、http 错误

#! /usr/bin/env python3
 
import urllib.request
 
req = urllib.request.Request(&#39;http://www.python.org/fish.html&#39;)
try:
  urllib.request.urlopen(req)
except urllib.error.HTTPError as e:
  print(e.code)
  print(e.read().decode("utf8"))

6、异常处理1

#! /usr/bin/env python3
 
from urllib.request import Request, urlopen
from urllib.error import URLError, HTTPError
req = Request("http://twitter.com/")
try:
  response = urlopen(req)
except HTTPError as e:
  print(&#39;The server couldn\&#39;t fulfill the request.&#39;)
  print(&#39;Error code: &#39;, e.code)
except URLError as e:
  print(&#39;We failed to reach a server.&#39;)
  print(&#39;Reason: &#39;, e.reason)
else:
  print("good!")
  print(response.read().decode("utf8"))

7、异常处理2

#! /usr/bin/env python3
 
from urllib.request import Request, urlopen
from urllib.error import URLError
req = Request("http://twitter.com/")
try:
  response = urlopen(req)
except URLError as e:
  if hasattr(e, &#39;reason&#39;):
    print(&#39;We failed to reach a server.&#39;)
    print(&#39;Reason: &#39;, e.reason)
  elif hasattr(e, &#39;code&#39;):
    print(&#39;The server couldn\&#39;t fulfill the request.&#39;)
    print(&#39;Error code: &#39;, e.code)
else:
  print("good!")
  print(response.read().decode("utf8"))

8、HTTP 认证

#! /usr/bin/env python3
 
import urllib.request
 
# create a password manager
password_mgr = urllib.request.HTTPPasswordMgrWithDefaultRealm()
 
# Add the username and password.
# If we knew the realm, we could use it instead of None.
top_level_url = "https://cms.tetx.com/"
password_mgr.add_password(None, top_level_url, &#39;yzhang&#39;, &#39;cccddd&#39;)
 
handler = urllib.request.HTTPBasicAuthHandler(password_mgr)
 
# create "opener" (OpenerDirector instance)
opener = urllib.request.build_opener(handler)
 
# use the opener to fetch a URL
a_url = "https://cms.tetx.com/"
x = opener.open(a_url)
print(x.read())
 
# Install the opener.
# Now all calls to urllib.request.urlopen use our opener.
urllib.request.install_opener(opener)
 
a = urllib.request.urlopen(a_url).read().decode(&#39;utf8&#39;)
print(a)

9、使用代理

#! /usr/bin/env python3
 
import urllib.request
 
proxy_support = urllib.request.ProxyHandler({&#39;sock5&#39;: &#39;localhost:1080&#39;})
opener = urllib.request.build_opener(proxy_support)
urllib.request.install_opener(opener)

 
a = urllib.request.urlopen("http://g.cn").read().decode("utf8")
print(a)

10、超时

#! /usr/bin/env python3
 
import socket
import urllib.request
 
# timeout in seconds
timeout = 2
socket.setdefaulttimeout(timeout)
 
# this call to urllib.request.urlopen now uses the default timeout
# we have set in the socket module
req = urllib.request.Request(&#39;http://twitter.com/&#39;)
a = urllib.request.urlopen(req).read()
print(a)

【相关推荐】

1. Python免费视频教程

2. Python学习手册

3. 马哥教育python基础语法全讲解视频

以上是分享十种py3爬取网页资源的方法的详细内容。更多信息请关注PHP中文网其他相关文章！

声明：

本文内容由网友自发贡献，版权归原作者所有，本站不承担相应法律责任。如您发现有涉嫌抄袭侵权的内容，请联系admin@php.cn

上一篇：Pycharm设置代码风格的图文教程下一篇：详解Python搭建Django项目的全过程

查看更多