Home >Backend Development >Python Tutorial >A course that teaches you to use python to crawl w3shcool and save it to local code examples

A course that teaches you to use python to crawl w3shcool and save it to local code examples

Y2JOriginal: 2017-04-27 11:42:032284browse

This article mainly introduces the method analysis of python crawling the JQuery course of w3shcool and saving it locally. Has very good reference value. Let’s take a look with the editor below

I have been busy looking for a job recently. In my spare time, I also found some reptile projects to practice my skills and write code. I know that I am a rookie, but I need to practice more. Shushan has Road work is the path. If you have any testing pits, can you introduce them to me? Automation, functions, and interfaces can all be done.

First of all, we clearly understand our needs. Many students want to see some technologies when they have nothing to do. For example, I want to see the syntax of JQuery, but I don’t have the Internet now, and I don’t have e-books on my mobile phone. Really It makes us uncomfortable, so don't worry, I'm here to meet your needs. First of all, your need is to obtain the syntax of JQuery, then I see this need, I have a website that responds, so let's go on Go analyze this website. www.w3school.com.cn/jquery/jquery_syntax.asp This is the syntax URL, http://www.w3school.com.cn/jquery/jquery_intro.asp This is the introduction URL, then we got a lot of URL analysis , our www.w3school.com.cn/jquery is the same, so let’s analyze how to get these in the interface. We can see that there is a corresponding target bar on the right, so let’s analyze it

Let’s take a look at these links. We can splice these links together with http://www.w3school.com.cn. Then form our new url,

upload the code

import urllib.request
from bs4 import BeautifulSoup 
import time
def head():
 headers={
 &#39;User-Agent&#39;:&#39;Mozilla/5.0 (Windows NT 6.1; WOW64; rv:52.0) Gecko/20100101 Firefox/52.0&#39;
 }
 return headers
def parse_url(url):
 hea=head()
 resposne=urllib.request.Request(url,headers=hea)
 html=urllib.request.urlopen(resposne).read().decode(&#39;gb2312&#39;)
 return html
def url_s():
 url=&#39;http://www.w3school.com.cn/jquery/index.asp&#39;
 html=parse_url(url)
 soup=BeautifulSoup(html)
 me=soup.find_all(id=&#39;course&#39;)
 m_url_text=[]
 m_url=[]
 for link in me:
  m_url_text.append(link.text)
  m=link.find_all(&#39;a&#39;)
  for i in m:
   m_url.append(i.get(&#39;href&#39;))
 for i in m_url_text:
  h=i.encode(&#39;utf-8&#39;).decode(&#39;utf-8&#39;)
  m_url_text=h.split(&#39;\n&#39;)
 return m_url,m_url_text

so that we can use the url_s function to get all our links.

[&#39;/jquery/index.asp&#39;, &#39;/jquery/jquery_intro.asp&#39;, &#39;/jquery/jquery_install.asp&#39;, &#39;/jquery/jquery_syntax.asp&#39;, &#39;/jquery/jquery_selectors.asp&#39;, &#39;/jquery/jquery_events.asp&#39;, &#39;/jquery/jquery_hide_show.asp&#39;, &#39;/jquery/jquery_fade.asp&#39;, &#39;/jquery/jquery_slide.asp&#39;, &#39;/jquery/jquery_animate.asp&#39;, &#39;/jquery/jquery_stop.asp&#39;, &#39;/jquery/jquery_callback.asp&#39;, &#39;/jquery/jquery_chaining.asp&#39;, &#39;/jquery/jquery_dom_get.asp&#39;, &#39;/jquery/jquery_dom_set.asp&#39;, &#39;/jquery/jquery_dom_add.asp&#39;, &#39;/jquery/jquery_dom_remove.asp&#39;, &#39;/jquery/jquery_css_classes.asp&#39;, &#39;/jquery/jquery_css.asp&#39;, &#39;/jquery/jquery_dimensions.asp&#39;, &#39;/jquery/jquery_traversing.asp&#39;, &#39;/jquery/jquery_traversing_ancestors.asp&#39;, &#39;/jquery/jquery_traversing_descendants.asp&#39;, &#39;/jquery/jquery_traversing_siblings.asp&#39;, &#39;/jquery/jquery_traversing_filtering.asp&#39;, &#39;/jquery/jquery_ajax_intro.asp&#39;, &#39;/jquery/jquery_ajax_load.asp&#39;, &#39;/jquery/jquery_ajax_get_post.asp&#39;, &#39;/jquery/jquery_noconflict.asp&#39;, &#39;/jquery/jquery_examples.asp&#39;, &#39;/jquery/jquery_quiz.asp&#39;, &#39;/jquery/jquery_reference.asp&#39;, &#39;/jquery/jquery_ref_selectors.asp&#39;, &#39;/jquery/jquery_ref_events.asp&#39;, &#39;/jquery/jquery_ref_effects.asp&#39;, &#39;/jquery/jquery_ref_manipulation.asp&#39;, &#39;/jquery/jquery_ref_attributes.asp&#39;, &#39;/jquery/jquery_ref_css.asp&#39;, &#39;/jquery/jquery_ref_ajax.asp&#39;, &#39;/jquery/jquery_ref_traversing.asp&#39;, &#39;/jquery/jquery_ref_data.asp&#39;, &#39;/jquery/jquery_ref_dom_element_methods.asp&#39;, &#39;/jquery/jquery_ref_core.asp&#39;, &#39;/jquery/jquery_ref_prop.asp&#39;], [&#39;jQuery 教程&#39;, &#39;&#39;, &#39;jQuery 教程&#39;, &#39;jQuery 简介&#39;, &#39;jQuery 安装&#39;, &#39;jQuery 语法&#39;, &#39;jQuery 选择器&#39;, &#39;jQuery 事件&#39;, &#39;&#39;, &#39;jQuery 效果&#39;, &#39;&#39;, &#39;jQuery 隐藏/显示&#39;, &#39;jQuery 淡入淡出&#39;, &#39;jQuery 滑动&#39;, &#39;jQuery 动画&#39;, &#39;jQuery stop()&#39;, &#39;jQuery Callback&#39;, &#39;jQuery Chaining&#39;, &#39;&#39;, &#39;jQuery HTML&#39;, &#39;&#39;, &#39;jQuery 获取&#39;, &#39;jQuery 设置&#39;, &#39;jQuery 添加&#39;, &#39;jQuery 删除&#39;, &#39;jQuery CSS 类&#39;, &#39;jQuery css()&#39;, &#39;jQuery 尺寸&#39;, &#39;&#39;, &#39;jQuery 遍历&#39;, &#39;&#39;, &#39;jQuery 遍历&#39;, &#39;jQuery 祖先&#39;, &#39;jQuery 后代&#39;, &#39;jQuery 同胞&#39;, &#39;jQuery 过滤&#39;, &#39;&#39;, &#39;jQuery AJAX&#39;, &#39;&#39;, &#39;jQuery AJAX 简介&#39;, &#39;jQuery 加载&#39;, &#39;jQuery Get/Post&#39;, &#39;&#39;, &#39;jQuery 杂项&#39;, &#39;&#39;, &#39;jQuery noConflict()&#39;, &#39;&#39;, &#39;jQuery 实例&#39;, &#39;&#39;, &#39;jQuery 实例&#39;, &#39;jQuery 测验&#39;, &#39;&#39;, &#39;jQuery 参考手册&#39;, &#39;&#39;, &#39;jQuery 参考手册&#39;, &#39;jQuery 选择器&#39;, &#39;jQuery 事件&#39;, &#39;jQuery 效果&#39;, &#39;jQuery 文档操作&#39;, &#39;jQuery 属性操作&#39;, &#39;jQuery CSS 操作&#39;, &#39;jQuery Ajax&#39;, &#39;jQuery 遍历&#39;, &#39;jQuery 数据&#39;, &#39;jQuery DOM 元素&#39;, &#39;jQuery 核心&#39;, &#39;jQuery 属性&#39;, &#39;&#39;, &#39;&#39;])

This is the name of all links and the corresponding grammar modules of the corresponding links. Then our next step is to splice urls, using str splicing

 [&#39;http://www.w3school.com.cn//jquery/index.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_intro.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_install.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_syntax.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_selectors.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_events.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_hide_show.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_fade.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_slide.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_animate.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_stop.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_callback.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_chaining.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_dom_get.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_dom_set.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_dom_add.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_dom_remove.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_css_classes.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_css.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_dimensions.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_traversing.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_traversing_ancestors.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_traversing_descendants.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_traversing_siblings.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_traversing_filtering.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ajax_intro.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ajax_load.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ajax_get_post.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_noconflict.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_examples.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_quiz.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_reference.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ref_selectors.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ref_events.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ref_effects.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ref_manipulation.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ref_attributes.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ref_css.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ref_ajax.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ref_traversing.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ref_data.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ref_dom_element_methods.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ref_core.asp&#39;, &#39;http://www.w3school.com.cn//jquery/jquery_ref_prop.asp&#39;]

Then we have all these urls, then let’s analyze the article text.

Analysis can show that all our texts are in an id=maincontent, then we directly parse the id=maincontent tag in each interface, obtain the response text document, and save it.

So all our code is as follows:

import urllib.request
from bs4 import BeautifulSoup 
import time
def head():
 headers={
 &#39;User-Agent&#39;:&#39;Mozilla/5.0 (Windows NT 6.1; WOW64; rv:52.0) Gecko/20100101 Firefox/52.0&#39;
 }
 return headers
def parse_url(url):
 hea=head()
 resposne=urllib.request.Request(url,headers=hea)
 html=urllib.request.urlopen(resposne).read().decode(&#39;gb2312&#39;)
 return html
def url_s():
 url=&#39;http://www.w3school.com.cn/jquery/index.asp&#39;
 html=parse_url(url)
 soup=BeautifulSoup(html)
 me=soup.find_all(id=&#39;course&#39;)
 m_url_text=[]
 m_url=[]
 for link in me:
  m_url_text.append(link.text)
  m=link.find_all(&#39;a&#39;)
  for i in m:
   m_url.append(i.get(&#39;href&#39;))
 for i in m_url_text:
  h=i.encode(&#39;utf-8&#39;).decode(&#39;utf-8&#39;)
  m_url_text=h.split(&#39;\n&#39;)
 return m_url,m_url_text
def xml():
 url,url_text=url_s()
 url_jque=[]
 for link in url:
  url_jque.append('http://www.w3school.com.cn/'+link)
 return url_jque
def xiazai():
 urls=xml()
 i=0
 for url in urls:
  html=parse_url(url)
  soup=BeautifulSoup(html)
  me=soup.find_all(id='maincontent')
  with open(r'%s.txt'%i,'wb') as f:
   for h in me:
    f.write(h.text.encode('utf-8'))
    print(i)
  i+=1
if __name__ == '__main__':
 xiazai()

import urllib.request
from bs4 import BeautifulSoup 
import time
def head():
 headers={
 &#39;User-Agent&#39;:&#39;Mozilla/5.0 (Windows NT 6.1; WOW64; rv:52.0) Gecko/20100101 Firefox/52.0&#39;
 }
 return headers
def parse_url(url):
 hea=head()
 resposne=urllib.request.Request(url,headers=hea)
 html=urllib.request.urlopen(resposne).read().decode(&#39;gb2312&#39;)
 return html
def url_s():
 url=&#39;http://www.w3school.com.cn/jquery/index.asp&#39;
 html=parse_url(url)
 soup=BeautifulSoup(html)
 me=soup.find_all(id=&#39;course&#39;)
 m_url_text=[]
 m_url=[]
 for link in me:
  m_url_text.append(link.text)
  m=link.find_all(&#39;a&#39;)
  for i in m:
   m_url.append(i.get(&#39;href&#39;))
 for i in m_url_text:
  h=i.encode(&#39;utf-8&#39;).decode(&#39;utf-8&#39;)
  m_url_text=h.split(&#39;\n&#39;)
 return m_url,m_url_text

def xml():
 url,url_text=url_s()
 url_jque=[]
 for link in url:
  url_jque.append('http://www.w3school.com.cn/'+link)
 return url_jque
def xiazai():
 urls=xml()
 i=0
 for url in urls:
  html=parse_url(url)
  soup=BeautifulSoup(html)
  me=soup.find_all(id='maincontent')
  with open(r'%s.txt'%i,'wb') as f:
   for h in me:
    f.write(h.text.encode('utf-8'))
    print(i)
  i+=1
if __name__ == '__main__':
 xiazai()

Result

##Now, our crawling work is completed, and the rest It’s just minor repairs and minor changes, but we should have completed all the major content.

In fact, Python’s crawler is still very simple. As long as we can analyze the elements of the website and find out the common terms of all elements, we can analyze and solve our problems very well

The above is the detailed content of A course that teaches you to use python to crawl w3shcool and save it to local code examples. For more information, please follow other related articles on the PHP Chinese website!

Statement：

The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn

Previous article：Detailed introduction to how to use Naive Bayes algorithm in pythonNext article：Detailed introduction to how to use Naive Bayes algorithm in python

See more

A course that teaches you to use python to crawl w3shcool and save it to local code examples

Related articles