Solution to garbled code problem in Python web crawler-Python Tutorial-php.cn

Home

Backend Development

Python Tutorial

Solution to garbled code problem in Python web crawler

高洛峰

Feb 11, 2017 pm 01:13 PM

pythonGarbled charactersWeb Crawler

This article mainly introduces in detail the solution to the problem of garbled characters in Python web crawlers. It has certain reference value. Interested friends can refer to it.

There are many ways to solve the problem of garbled characters in crawlers. Various problems, here are not only Chinese garbled characters, encoding conversion, but also some garbled characters such as Japanese, Korean, Russian, Tibetan, etc., because the solutions are the same, so they are explained here.

The reason why the web crawler appears garbled code

The encoding format of the source web page is inconsistent with the encoding format after crawling.
If the source web page is a byte stream encoded by gbk, and after we grab it, the program directly uses utf-8 to encode and output it to the storage file. This will inevitably cause garbled code when the source web page is encoded and captured. When the program directly uses the processing encoding to be consistent, there will be no garbled characters; at this time, if the character encoding is unified, there will be no garbled characters.

Pay attention to the distinction

Source network code A,
code B used directly by the program,
code C for unified conversion of characters.

Solution to garbled codes

Determine the code A of the source web page, code A is often in three positions in the web page

1.Content-Type of http header
The site that obtains the server header can use it to tell the browser some information about the page content. The Content-Type entry is written as "text/html; charset=utf-8".

2.meta charset

3. Document definition in the web page header

<script> 
if(document.charset){ 
 alert(document.charset+"!!!!"); 
 document.charset = &#39;GBK&#39;; 
 alert(document.charset); 
} 
else if(document.characterSet){ 
 alert(document.characterSet+"????"); 
 document.characterSet = &#39;GBK&#39;; 
 alert(document.characterSet); 
}</script>

When obtaining the source web page encoding, judge these three in order Part of the data is enough, from front to back, and the same is true for priority.
There is no encoding information among the above three. Generally, third-party web page encoding intelligent identification tools such as chardet are used for

Installation: pip install chardet

Python chardet character encoding judgment

Using chardet can easily realize encoding detection of strings/files. Although HTML pages have charset tags, they are sometimes incorrect. Then chardet can help us a lot.
chardet example

import urllib 
rawdata = urllib.urlopen('http://www.php.cn/').read() 
import chardet 
chardet.detect(rawdata) 
{'confidence': 0.99, 'encoding': 'GB2312'}

chardet can directly use the detect function to detect the encoding of the given character. The function return value is a dictionary with two elements, one is the detection credibility, and the other is the detected encoding.

How to deal with Chinese character encoding in the process of developing your own crawler?
The following are all for python2.7. If not processed, all the collected characters will be garbled. The solution is Process html into unified utf-8 encoding and encounter windows-1252 encoding, which belongs to the chardet encoding recognition training that has not been completed

import chardet 
a='abc' 
type(a) 
str 
chardet.detect(a) 
{'confidence': 1.0, 'encoding': 'ascii'} 
 
 
a ="我" 
chardet.detect(a) 
{'confidence': 0.73, 'encoding': 'windows-1252'} 
a.decode('windows-1252') 
u'\xe6\u02c6\u2018' 
chardet.detect(a.decode('windows-1252').encode('utf-8')) 
type(a.decode('windows-1252')) 
unicode 
type(a.decode('windows-1252').encode('utf-8')) 
str 
chardet.detect(a.decode('windows-1252').encode('utf-8')) 
{'confidence': 0.87625, 'encoding': 'utf-8'} 
 
 
a ="我是中国人" 
type(a) 
str 
{'confidence': 0.9690625, 'encoding': 'utf-8'} 
chardet.detect(a) 
# -*- coding:utf-8 -*- 
import chardet 
import urllib2 
#抓取网页html 
html = urllib2.urlopen('http://www.jb51.net/').read() 
print html 
mychar=chardet.detect(html) 
print mychar 
bianma=mychar['encoding'] 
if bianma == 'utf-8' or bianma == 'UTF-8': 
 html=html.decode('utf-8','ignore').encode('utf-8') 
else: 
 html =html.decode('gb2312','ignore').encode('utf-8') 
print html 
print chardet.detect(html)

python code file Encoding
py file defaults to ASCII encoding. When Chinese is displayed, a conversion from ASCII to the system default encoding will occur. At this time, an error will occur: SyntaxError: Non-ASCII character. It is necessary to add encoding instructions in the first line of the code file:

# -*- coding:utf-8 -*- 
 
print '中文'

The string input directly as above is encoded according to the code file'utf -8' to process
If unicode encoding is used, the following method is used:

s1 = u'Chinese' #u means using unicode encoding to store information

decode is a method that any string has to convert the string into unicode format. The parameter indicates the encoding format of the source string.
encode is also a method that any string has, converting the string into the format specified by the parameter.

The above is the entire content of this article. I hope it will be helpful to everyone's learning. I also hope that everyone will support the PHP Chinese website.

For more related articles on solutions to garbled code problems in Python web crawlers, please pay attention to the PHP Chinese website!

Statement

The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn

The Main Purpose of Python: Flexibility and Ease of UseApr 17, 2025 am 12:14 AM

Python's flexibility is reflected in multi-paradigm support and dynamic type systems, while ease of use comes from a simple syntax and rich standard library. 1. Flexibility: Supports object-oriented, functional and procedural programming, and dynamic type systems improve development efficiency. 2. Ease of use: The grammar is close to natural language, the standard library covers a wide range of functions, and simplifies the development process.

Python: The Power of Versatile ProgrammingApr 17, 2025 am 12:09 AM

Python is highly favored for its simplicity and power, suitable for all needs from beginners to advanced developers. Its versatility is reflected in: 1) Easy to learn and use, simple syntax; 2) Rich libraries and frameworks, such as NumPy, Pandas, etc.; 3) Cross-platform support, which can be run on a variety of operating systems; 4) Suitable for scripting and automation tasks to improve work efficiency.

Learning Python in 2 Hours a Day: A Practical GuideApr 17, 2025 am 12:05 AM

Yes, learn Python in two hours a day. 1. Develop a reasonable study plan, 2. Select the right learning resources, 3. Consolidate the knowledge learned through practice. These steps can help you master Python in a short time.

Python vs. C : Pros and Cons for DevelopersApr 17, 2025 am 12:04 AM

Python is suitable for rapid development and data processing, while C is suitable for high performance and underlying control. 1) Python is easy to use, with concise syntax, and is suitable for data science and web development. 2) C has high performance and accurate control, and is often used in gaming and system programming.

Python: Time Commitment and Learning PaceApr 17, 2025 am 12:03 AM

The time required to learn Python varies from person to person, mainly influenced by previous programming experience, learning motivation, learning resources and methods, and learning rhythm. Set realistic learning goals and learn best through practical projects.

Python: Automation, Scripting, and Task ManagementApr 16, 2025 am 12:14 AM

Python excels in automation, scripting, and task management. 1) Automation: File backup is realized through standard libraries such as os and shutil. 2) Script writing: Use the psutil library to monitor system resources. 3) Task management: Use the schedule library to schedule tasks. Python's ease of use and rich library support makes it the preferred tool in these areas.

Python and Time: Making the Most of Your Study TimeApr 14, 2025 am 12:02 AM

To maximize the efficiency of learning Python in a limited time, you can use Python's datetime, time, and schedule modules. 1. The datetime module is used to record and plan learning time. 2. The time module helps to set study and rest time. 3. The schedule module automatically arranges weekly learning tasks.

Python: Games, GUIs, and MoreApr 13, 2025 am 12:14 AM

Python excels in gaming and GUI development. 1) Game development uses Pygame, providing drawing, audio and other functions, which are suitable for creating 2D games. 2) GUI development can choose Tkinter or PyQt. Tkinter is simple and easy to use, PyQt has rich functions and is suitable for professional development.

See all articles

Hot AI Tools

Undresser.AI Undress

AI-powered app for creating realistic nude photos

AI Clothes Remover

Online AI tool for removing clothes from photos.

Undress AI Tool

Undress images for free

Clothoff.io

AI clothes remover

AI Hentai Generator

Generate AI Hentai for free.

Hot Article

R.E.P.O. Energy Crystals Explained and What They Do (Yellow Crystal)

1 months agoBy尊渡假赌尊渡假赌尊渡假赌

R.E.P.O. Best Graphic Settings

1 months agoBy尊渡假赌尊渡假赌尊渡假赌

Assassin's Creed Shadows: Seashell Riddle Solution

2 weeks agoByDDD

R.E.P.O. How to Fix Audio if You Can't Hear Anyone

1 months agoBy尊渡假赌尊渡假赌尊渡假赌

R.E.P.O. Chat Commands and How to Use Them

1 months agoBy尊渡假赌尊渡假赌尊渡假赌

Hot Tools

EditPlus Chinese cracked version

Small size, syntax highlighting, does not support code prompt function

mPDF

mPDF is a PHP library that can generate PDF files from UTF-8 encoded HTML. The original author, Ian Back, wrote mPDF to output PDF files "on the fly" from his website and handle different languages. It is slower than original scripts like HTML2FPDF and produces larger files when using Unicode fonts, but supports CSS styles etc. and has a lot of enhancements. Supports almost all languages, including RTL (Arabic and Hebrew) and CJK (Chinese, Japanese and Korean). Supports nested block-level elements (such as P, DIV),