Analyzing Chinese coding issues in PHP development,

Home

Backend Development

PHP Tutorial

Analyzing Chinese coding issues in PHP development, _PHP tutorial

WBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWB

Jul 12, 2016 am 08:51 AM

phpChinesedevelopofcodingparsequestion

Analysis of Chinese coding issues in PHP development,

In fact, Chinese coding in PHP development is not as complicated as imagined. Although there are no fixed rules for locating and solving problems, each Each operating environment is different, but the underlying principles are the same.

Understanding the knowledge of character sets is the basis for solving character problems.

The problem of Chinese encoding in PHP programming has troubled many people. The reason for this problem is actually very simple. Each country (or region) stipulates the character encoding set for computer information exchange, such as the extended ASCII code of the United States. China's GB2312-80, Japan's JIS, etc. As the basis for information processing in this country/region, character encoding sets play an important role in unifying encoding. Character encoding sets are divided into two categories according to length: SBCS (single-byte character set) and DBCS (double-byte character set). In early software (especially operating systems), in order to solve the computer processing of local character information, various localized versions (L10N) appeared. In order to differentiate, concepts such as LANG and Codepage were introduced. However, due to the overlapping code ranges of various local character sets, it is difficult to exchange information with each other; the cost of independent maintenance of each localized version of the software is high. Therefore, it is necessary to extract the commonalities in localization work and process them consistently to minimize special localization processing content. This is also called internationalization (118N). Various language information is further standardized as Locale information. The underlying character set processed became Unicode, which contains almost all glyphs.

Currently, most of the core character processing of software with international features is based on Unicode. When the software is running, the corresponding local character encoding settings are determined according to the locale/Lang/Codepage settings at that time, and local characters are processed accordingly. . During the processing, it is necessary to convert between Unicode and local character sets, or even between two different local character sets with Unicode as an intermediate. This method is further extended in the network environment, and any character information at both ends of the network also needs to be converted into acceptable content according to the character set settings.

Character set encoding issues in databases
Popular relational database systems all support database character set encoding, which means that you can specify its own character set settings when creating a database, and the database data is stored in the specified encoding. . When an application accesses data, there will be character set encoding conversion at entry and exit. For Chinese data, the database character encoding setting should ensure the integrity of the data. GB2312, GBK, UTF-8, etc. are all optional database character set encodings; of course we can also choose ISO8859-1 (8-bit), but we have to use

Before writing data with a program, first split a 16-bit Chinese character or Unicode into two 8-bit characters. After reading the data, you also need to merge the two bytes and identify the SBCS characters. Therefore We do not recommend using ISO8859-1 as the database character set encoding. Not only does this not make full use of the character set encoding support of the database itself, but it also increases the complexity of programming. When programming, you can first use the management functions provided by the database management system to check whether the Chinese data is correct.

Before querying the database, the PHP program first executes mysql_query("SET NAMES xxxx"); where xxxx is the encoding of your web page (charset=xxxx). If charset=utf8 in the web page, then xxxx=utf8, if charset in the web page =gb2312, then xxxx=gb2312. Almost all WEB programs have a common code to connect to the database, which is placed in a file. In this file, just add mysql_query("SET NAMES xxxx").

SET NAMES Shows what character set is used in the SQL statement sent by the client. Therefore, the SET NAMES 'utf-8' statement tells the server that "future messages from this client will use the character set utf-8." It also specifies the character set for the results that the server sends back to the client (for example, if you use a SELECT statement, it indicates what character set is used for the column values).

Commonly used techniques when locating problems
Locating Chinese encoding problems usually uses the stupidest and most effective method - printing the internal code of the string after processing by the program you think is suspicious. By printing the internal code of a string, you can find out when Chinese characters are converted to Unicode, when Unicode is converted back to Chinese internal code, when one Chinese character becomes two Unicode characters, when a Chinese string is converted to A string of question marks, when were the high bits of the Chinese character string cut off...

Using appropriate sample strings can also help distinguish the type of problem. For example: "aaah aa?@aa" and other Chinese and English character strings with both GB and GBK characteristic characters. Generally speaking, English characters will not be distorted no matter how they are converted or processed (if you encounter them, you can try to increase the length of consecutive English letters).

Solving garbled code problems in various applications

1) Use the tag to set the page encoding
The purpose of this tag is to declare what the client's browser uses The character set encoding displays this page, xxx can be GB2312, GBK, UTF-8 (different from MySQL, which is UTF8), etc. Therefore, most pages can use this method to tell the browser what encoding to use when displaying this page, so as to avoid encoding errors and garbled characters. But sometimes we will find that this sentence still doesn't work. No matter which xxx is, the browser always uses the same encoding. I will talk about this later.

Please note that belongs to HTML information and is just a statement, which only indicates that the server has passed the HTML information to the browser.

2) header("content-type:text/html; charset=xxx");
The function of this function header() is to send the information in the brackets to the http header. If the content in the brackets is as mentioned in the article, the function is basically the same as the label. If you compare the first one, you will find that the characters are similar. But the difference is that if there is this function, the browser will always use the xxx encoding you requested and will never be disobedient, so this function is very useful. Why is this? Then we have to talk about the difference between http header and HTML information:

The http header is a string sent by the server before sending HTML information to the browser using the http protocol. The tag belongs to HTML information, so the content sent by header() reaches the browser first. The popular point is that the priority of header() is higher than (I don’t know if this can be said). If a php page has both header("content-type:text/html;charset=xxx") and header("content-type:text/html;charset=xxx"), the browser will only recognize the former http header and not the meta. Of course, this function can only be used within php pages.

There is also a question left, why does the former definitely work, but the latter sometimes does not work? This is the reason why we want to talk about Apache next.

3) AddDefaultCharset
In the conf folder in the Apache root directory, there is the entire Apache configuration document httpd.conf.

Open httpd.conf with a text editor. Line 708 (different versions may be different) contains AddDefaultCharset xxx, where xxx is the encoding name. The meaning of this line of code: Set the character set in the http header of the web page file in the entire server to your default xxx character set. Having this line is equivalent to adding a line of header("content-type: text/html; charset=xxx") to each file. Now you can understand why the browser always uses gb2312 even though the setting is utf-8.

If there is header("content-type:text/html; charset=xxx") in the web page, the default character set will be changed to the character set you set, so this function will always be useful. If you add a "#" in front of AddDefaultCharset xxx, comment out this sentence, and the page does not contain header("content-type..."), then it is the meta tag's turn to take effect.

The priority order of the above is listed below:
.. header("content-type:text/html; charset=xxx")
.. AddDefaultCharset xxx
..

If you are a web programmer, it is recommended to add a header ("content-type: text/html; charset=xxx") to each of your pages, so as to ensure that it can be displayed correctly on any server. Portability is also relatively strong.

4) default_charset configuration in php.ini:
default_charset = "gb2312" in php.ini defines the default language character set of php. It is generally recommended to comment out this line and let the browser automatically select the language based on the charset in the web page header instead of making a mandatory requirement, so that web services in multiple languages can be provided on the same server.

For questions about php encoding, you can also refer to the following articles:

Analysis of PHP string encoding issues
Two methods for PHP to determine character encoding
A function that automatically detects the encoding in the content and converts it
PHP code for converting GB2312 and UTF8 encoding
http ://www.cnblogs.com/GarfieldTom/archive/2012/11/02/2750776.html
PHP Big5 Utf-8 GB2312 encoding conversion solution
php encoding, garbled code problem

Conclusion
In fact, Chinese coding in PHP development is not as complicated as imagined. Although there are no fixed rules for positioning and solving problems, and various operating environments are also different, the principles behind it are the same. Understanding the knowledge of character sets is the basis for solving character problems. However, with the changes in the Chinese character set, not only PHP programming, but also problems in Chinese information processing will still exist for some time.

Statement

The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn

Python解析XML中的特殊字符和转义序列Aug 08, 2023 pm 12:46 PM

Python解析XML中的特殊字符和转义序列XML（eXtensibleMarkupLanguage）是一种常用的数据交换格式，用于在不同系统之间传输和存储数据。在处理XML文件时，经常会遇到包含特殊字符和转义序列的情况，这可能会导致解析错误或者误解数据。因此，在使用Python解析XML文件时，我们需要了解如何处理这些特殊字符和转义序列。一、特殊字符和

Python编程解析百度地图API文档中的坐标转换功能Aug 01, 2023 am 08:57 AM

Python编程解析百度地图API文档中的坐标转换功能导读：随着互联网的快速发展，地图定位功能已经成为现代人生活中不可或缺的一部分。而百度地图作为国内最受欢迎的地图服务之一，提供了一系列的API供开发者使用。本文将通过Python编程，解析百度地图API文档中的坐标转换功能，并给出相应的代码示例。一、引言在开发中，我们有时会涉及到坐标的转换问题。百度地图AP

使用Python解析SOAP消息Aug 08, 2023 am 09:27 AM

使用Python解析SOAP消息SOAP（SimpleObjectAccessProtocol）是一种基于XML的远程过程调用（RPC）协议，用于在网络上不同的应用程序之间进行通信。Python提供了许多库和工具来处理SOAP消息，其中最常用的是suds库。suds是Python的一个SOAP客户端库，可以用于解析和生成SOAP消息。它提供了一种简单而

PHP8.0中的XML解析库May 14, 2023 am 08:19 AM

随着PHP8.0的发布，许多新特性都被引入和更新了，其中包括XML解析库。PHP8.0中的XML解析库提供了更快的解析速度和更好的可读性，这对于PHP开发者来说是一个重要的提升。在本文中，我们将探讨PHP8.0中的XML解析库的新特性以及如何使用它。什么是XML解析库？XML解析库是一种软件库，用于解析和处理XML文档。XML是一种用于将数据存储为结构化文档

使用Python解析带有命名空间的XML文档Aug 09, 2023 pm 04:25 PM

使用Python解析带有命名空间的XML文档XML是一种常用的数据交换格式，能够适应各种应用场景。在处理XML文档时，有时会遇到带有命名空间（namespace）的情况。命名空间可以防止不同XML文档中元素名的冲突，提高了XML的灵活性和可扩展性。本文将介绍如何使用Python解析带有命名空间的XML文档，并给出相应的代码示例。首先，我们需要导入xml.et

PHP中的HTTP Basic鉴权方法解析及应用Aug 06, 2023 am 08:16 AM

PHP中的HTTPBasic鉴权方法解析及应用HTTPBasic鉴权是一种简单但常用的身份验证方法，它通过在HTTP请求头中添加用户名和密码的Base64编码字符串进行身份验证。本文将介绍HTTPBasic鉴权的原理和使用方法，并提供PHP代码示例供读者参考。一、HTTPBasic鉴权原理HTTPBasic鉴权的原理非常简单，当客户端发送一个请求时

PHP 爬虫实战之获取网页源码和内容解析Jun 13, 2023 am 10:46 AM

PHP爬虫是一种自动化获取网页信息的程序，它可以获取网页代码、抓取数据并存储到本地或数据库中。使用爬虫可以快速获取大量的数据，为后续的数据分析和处理提供巨大的帮助。本文将介绍如何使用PHP实现一个简单的爬虫，以获取网页源码和内容解析。一、获取网页源码在开始之前，我们应该先了解一下HTTP协议和HTML的基本结构。HTTP是HyperText

PHP中的单点登录（SSO）鉴权方法解析Aug 08, 2023 am 09:21 AM

PHP中的单点登录（SSO）鉴权方法解析引言：随着互联网的发展，用户通常要同时访问多个网站进行各种操作。为了提高用户体验，单点登录（SingleSign-On，简称SSO）应运而生。本文将探讨PHP中的SSO鉴权方法，并提供相应的代码示例。一、什么是单点登录（SSO）？单点登录（SSO）是一种集中化认证的方法，在多个应用系统中，用户只需要登录一次，就能访问

See all articles