search
HomeBackend DevelopmentPHP TutorialHow to solve the Chinese garbled problem in PHP? _PHP Tutorial

The problem of Chinese encoding in PHP programming has troubled many people. The reason for this problem is actually very simple. Each country (or region) stipulates the character encoding set for computer information exchange, such as the American extension ASCII code, China's GB2312-80, Japan's JIS, etc. As the basis for information processing in this country/region, character encoding sets play an important role in unifying encoding. Character encoding sets are divided into two categories according to length: SBCS (single-byte character set) and DBCS (double-byte character set). In early software (especially operating systems), in order to solve the computer processing of local character information, various localized versions (L10N) appeared. In order to differentiate, concepts such as LANG and Codepage were introduced. However, due to the overlapping code ranges of various local character sets, it is difficult to exchange information with each other; the cost of independent maintenance of each localized version of the software is high. Therefore, it is necessary to extract the commonalities in localization work and process them consistently to minimize special localization processing content. This is also called internationalization (118N). Various language information is further standardized as Locale information. The underlying character set processed became Unicode, which contains almost all glyphs.

Currently, most of the core character processing of software with international features is based on Unicode. When the software is running, the corresponding local character encoding settings are determined according to the locale/Lang/Codepage settings at that time, and local characters are processed accordingly. . During the processing, it is necessary to convert between Unicode and local character sets, or even between two different local character sets with Unicode as an intermediate. This method is further extended in the network environment, and any character information at both ends of the network also needs to be converted into acceptable content according to the character set settings.

Character set encoding problem in database

Popular relational database systems all support database character set encoding, which means that you can specify its own character set settings when creating a database, and the data in the database is stored in the specified encoding. When an application accesses data, there will be character set encoding conversion at entry and exit. For Chinese data, the database character encoding setting should ensure the integrity of the data. GB2312, GBK, UTF-8, etc. are all optional database character set encodings; of course we can also choose ISO8859-1 (8-bit), but we have to split a 16-bit Chinese character or Unicode before the application writes data. Divide it into two 8-bit characters. After reading the data, you need to merge the two bytes and identify the SBCS characters. Therefore, we do not recommend using ISO8859-1 as the database character set encoding. Not only does this not make full use of the character set encoding support of the database itself, but it also increases the complexity of programming. When programming, you can first use the management functions provided by the database management system to check whether the Chinese data is correct.

Before querying the database, the PHP program first executes mysql_query ("SET NAMES xxxx"); where xxxx is the encoding of your web page (charset=xxxx), if charset=utf8 in the web page, then xxxx=utf8, if charset in the web page =gb2312, then xxxx=gb2312. Almost all WEB programs have a common code for connecting to the database, which is placed in a file. In this file, just add mysql_query ("SET NAMES xxxx").

SET NAMES Shows what character set is used in the SQL statement sent by the client. Therefore, the SET NAMES 'utf-8' statement tells the server that "future messages from this client will use the character set utf-8." It also specifies the character set for the results that the server sends back to the client (for example, if you use a SELECT statement, it indicates what character set is used for the column values).

Commonly used techniques when locating problems

To locate Chinese encoding problems, the stupidest and most effective method is usually used - printing the internal code of the string after processing by the program you think is suspicious. By printing the internal code of a string, you can find out when Chinese characters are converted to Unicode, when Unicode is converted back to Chinese internal code, when one Chinese character becomes two Unicode characters, when a Chinese string is converted to A string of question marks, when the high bits of the Chinese string were truncated.

Using appropriate sample strings can also help distinguish the type of problem. For example: "aaah aa?@aa" and other Chinese and English character strings with both GB and GBK characteristic characters. Generally speaking, English characters will not be distorted no matter how they are converted or processed (if you encounter them, you can try to increase the length of consecutive English letters).

Solving garbled code problems in various applications

  1. Use tags to set page encoding
  2. The purpose of this tag is to declare what character set encoding the client's browser uses to display the page. xxx can be GB2312, GBK, UTF-8 (different from MySQL, which is UTF8), etc. Therefore, most pages can use this method to tell the browser what encoding to use when displaying this page, so as to avoid encoding errors and garbled characters. But sometimes we will find that this sentence still doesn't work. No matter which xxx is, the browser always uses the same encoding. I will talk about this later.

    Please note that it belongs to HTML information and is just a statement, which only indicates that the server has passed the HTML information to the browser.

  3. header("content-type:text/html; charset=xxx");
  4. The function header() is to send the information in the brackets to the http header. If the content in the brackets is as mentioned in the article, the function is basically the same as the label. If you compare the first one, you will find that the characters are similar. But the difference is that if there is this function, the browser will always use the xxx encoding you requested and will never be disobedient, so this function is very useful. Why is this? Then we have to talk about the difference between http header and HTML information:

    The http header is a string sent by the server before sending HTML information to the browser using the http protocol. The tag belongs to HTML information, so the content sent by header() reaches the browser first. The popular point is that header() has a higher priority than header() (I don’t know if I can say this). If a php page has both header ("content-type:text/html;charset=xxx") and header ("content-type:text/html;charset=xxx"), the browser will only recognize the former http header and not the meta. Of course, this function can only be used within php pages.

    There is also a question left, why does the former definitely work, but the latter sometimes does not work? This is the reason why we want to talk about Apache next.

  5. AddDefaultCharset
  6. In the conf folder in the Apache root directory, there is the entire Apache configuration document httpd.conf.

    Open httpd.conf with a text editor. Line 708 (different versions may be different) contains AddDefaultCharset xxx, where xxx is the encoding name. The meaning of this line of code: Set the character set in the http header of the web page file in the entire server to your default xxx character set. Having this line is equivalent to adding a header line ("content-type: text/html; charset=xxx") to each file. Now you can understand why the browser always uses gb2312 even though it is set to utf-8.

    If there is a header ("content-type: text/html; charset=xxx") in the web page, the default character set will be changed to the character set you set, so this function will always be useful. If you add a "#" in front of AddDefaultCharset xxx, comment out this sentence, and the page does not contain header ("content-type..."), then it is the meta tag's turn to take effect.

    The above is listed below in order of priority:

    .. header("content-type:text/html; charset=xxx")

    .. AddDefaultCharset xxx

    ..

    If you are a web programmer, it is recommended to add a header ("content-type: text/html; charset=xxx") to each of your pages, so as to ensure that it can be displayed correctly on any server. Portability is also relatively strong.

  7. default_charset configuration in php.ini
  8. default_charset = "gb2312" in php.ini defines the default language character set of php. It is generally recommended to comment out this line and let the browser automatically select the language based on the charset in the web page header instead of making a mandatory requirement, so that web services in multiple languages ​​can be provided on the same server.

In fact, Chinese coding in PHP development is not as complicated as imagined. Although there are no fixed rules for locating and solving problems, and various operating environments are also different, the principles behind it are the same. Understanding the knowledge of character sets is the basis for solving character problems. However, with the changes in the Chinese character set, not only PHP programming, but also problems in Chinese information processing will still exist for some time.

www.bkjia.comtruehttp: //www.bkjia.com/PHPjc/752529.htmlTechArticleThe problem of Chinese encoding in PHP programming has troubled many people. The reason for this problem is actually very simple. Every country (or region) all stipulate the character encoding for computer information exchange...
Statement
The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn
php怎么把负数转为正整数php怎么把负数转为正整数Apr 19, 2022 pm 08:59 PM

php把负数转为正整数的方法:1、使用abs()函数将负数转为正数,使用intval()函数对正数取整,转为正整数,语法“intval(abs($number))”;2、利用“~”位运算符将负数取反加一,语法“~$number + 1”。

php怎么实现几秒后执行一个函数php怎么实现几秒后执行一个函数Apr 24, 2022 pm 01:12 PM

实现方法:1、使用“sleep(延迟秒数)”语句,可延迟执行函数若干秒;2、使用“time_nanosleep(延迟秒数,延迟纳秒数)”语句,可延迟执行函数若干秒和纳秒;3、使用“time_sleep_until(time()+7)”语句。

php怎么除以100保留两位小数php怎么除以100保留两位小数Apr 22, 2022 pm 06:23 PM

php除以100保留两位小数的方法:1、利用“/”运算符进行除法运算,语法“数值 / 100”;2、使用“number_format(除法结果, 2)”或“sprintf("%.2f",除法结果)”语句进行四舍五入的处理值,并保留两位小数。

php怎么根据年月日判断是一年的第几天php怎么根据年月日判断是一年的第几天Apr 22, 2022 pm 05:02 PM

判断方法:1、使用“strtotime("年-月-日")”语句将给定的年月日转换为时间戳格式;2、用“date("z",时间戳)+1”语句计算指定时间戳是一年的第几天。date()返回的天数是从0开始计算的,因此真实天数需要在此基础上加1。

php怎么判断有没有小数点php怎么判断有没有小数点Apr 20, 2022 pm 08:12 PM

php判断有没有小数点的方法:1、使用“strpos(数字字符串,'.')”语法,如果返回小数点在字符串中第一次出现的位置,则有小数点;2、使用“strrpos(数字字符串,'.')”语句,如果返回小数点在字符串中最后一次出现的位置,则有。

php怎么替换nbsp空格符php怎么替换nbsp空格符Apr 24, 2022 pm 02:55 PM

方法:1、用“str_replace(" ","其他字符",$str)”语句,可将nbsp符替换为其他字符;2、用“preg_replace("/(\s|\&nbsp\;||\xc2\xa0)/","其他字符",$str)”语句。

php字符串有没有下标php字符串有没有下标Apr 24, 2022 am 11:49 AM

php字符串有下标。在PHP中,下标不仅可以应用于数组和对象,还可应用于字符串,利用字符串的下标和中括号“[]”可以访问指定索引位置的字符,并对该字符进行读写,语法“字符串名[下标值]”;字符串的下标值(索引值)只能是整数类型,起始值为0。

php怎么设置implode没有分隔符php怎么设置implode没有分隔符Apr 18, 2022 pm 05:39 PM

在PHP中,可以利用implode()函数的第一个参数来设置没有分隔符,该函数的第一个参数用于规定数组元素之间放置的内容,默认是空字符串,也可将第一个参数设置为空,语法为“implode(数组)”或者“implode("",数组)”。

See all articles

Hot AI Tools

Undresser.AI Undress

Undresser.AI Undress

AI-powered app for creating realistic nude photos

AI Clothes Remover

AI Clothes Remover

Online AI tool for removing clothes from photos.

Undress AI Tool

Undress AI Tool

Undress images for free

Clothoff.io

Clothoff.io

AI clothes remover

AI Hentai Generator

AI Hentai Generator

Generate AI Hentai for free.

Hot Article

R.E.P.O. Energy Crystals Explained and What They Do (Yellow Crystal)
2 weeks agoBy尊渡假赌尊渡假赌尊渡假赌
Repo: How To Revive Teammates
4 weeks agoBy尊渡假赌尊渡假赌尊渡假赌
Hello Kitty Island Adventure: How To Get Giant Seeds
4 weeks agoBy尊渡假赌尊渡假赌尊渡假赌

Hot Tools

Safe Exam Browser

Safe Exam Browser

Safe Exam Browser is a secure browser environment for taking online exams securely. This software turns any computer into a secure workstation. It controls access to any utility and prevents students from using unauthorized resources.

ZendStudio 13.5.1 Mac

ZendStudio 13.5.1 Mac

Powerful PHP integrated development environment

SublimeText3 English version

SublimeText3 English version

Recommended: Win version, supports code prompts!

Zend Studio 13.0.1

Zend Studio 13.0.1

Powerful PHP integrated development environment

Dreamweaver CS6

Dreamweaver CS6

Visual web development tools