What is the difference between utf8_unicode_ci and utf8_general

Home

Database

Mysql Tutorial

What is the difference between utf8_unicode_ci and utf8_general_ci in Mysql?

不言

Mar 27, 2019 am 10:04 AM

mysql

The content of this article is about the difference between utf8_unicode_ci and utf8_general_ci in Mysql? It has certain reference value. Friends in need can refer to it. I hope it will be helpful to you.

What is the difference between utf8_general_ci and utf8_unicode_ci in Mysql? In programming languages, unicode is usually used to process Chinese characters to prevent garbled characters. So in MySQL, why does everyone use utf8_general_ci instead of utf8_unicode_ci?

After using it for so long, I found that I didn’t even know the difference between utf_bin and utf_general_ci. .
ci is case insensitive, that is, "case insensitive", a and A will be treated as the same in character judgment;
bin is binary, a and A will be treated differently.
For example, if you Run:
SELECT * FROM table WHERE txt = 'a'
Then you will not find the line with txt = 'A' in utf8_bin, but utf8_general_ci can.
utf8_general_ci is not case-sensitive. You will use this when registering your username and email address.
utf8_general_cs is case-sensitive. If this is used for username and email, there will be adverse consequences.
utf8_bin: String Each string is compiled and stored with binary data. It is case-sensitive and can store binary content

1. Official document description
The following is an excerpt from the Mysql 5.1 Chinese manual about utf8_unicode_ci and utf8_general_ci:

Currently, the utf8_unicode_ci collation rule only partially supports the Unicode collation rule algorithm. Some characters are still not supported. Also, combined tokens are not fully supported. This mainly affects some minority languages in Vietnam and Russia, such as: Udmurt, Tatar, Bashkir and Mari.

The most important feature of utf8_unicode_ci is to support expansion, that is, when a letter is regarded as equal to other letter combinations. For example, 'ß' is equivalent to 'ss' in German and some other languages.

utf8_general_ci is a legacy collation rule and does not support extensions. It is only capable of character-by-character comparisons. This means that comparisons made by the utf8_general_ci collation are fast, but less accurate than those using the utf8_unicode_ci collation).

For example, using the two collation rules utf8_general_ci and utf8_unicode_ci the following comparisons are equal:
Ä = A
Ö = O
Ü = U

Between the two collation rules The difference is that for utf8_general_ci the following equation holds:
ß = s

However, for utf8_unicode_ci the following equation holds:
ß = ss

For one language only When sorting using utf8_unicode_ci does not work well, the utf8 character set collation rules related to the specific language are implemented. For example, for German and French, utf8_unicode_ci works just fine, so there is no need to create special utf8 collation rules for these two languages.

utf8_general_ci also works with German and French, except that 'ß' equals 's' instead of 'ss'. If your application can accept this, you should use utf8_general_ci because it is fast. Otherwise, use utf8_unicode_ci since it is more accurate.

If you want to use gb2312 encoding, it is recommended that you use latin1 as the default character set of the data table, so that you can directly insert data in the command line tool in Chinese and display it directly. Do not use gb2312 Or gbk and other character sets. If you are worried about query sorting and other issues, you can use binary attribute constraints, for example:

create table my_table ( name varchar(20) binary not null default &#39;&#39;)type=myisam default charset latin1;

2. Brief summary
utf8_unicode_ci and utf8_general_ci for Chinese and English There is no real difference.
utf8_general_ci The proofreading speed is fast, but the accuracy is slightly worse.
utf8_unicode_ci has high accuracy, but the proofing speed is slightly slower.

If your application is in German, French or Russian, please be sure to use utf8_unicode_ci. Generally, it is enough to use utf8_general_ci, and no problem has been found so far. . .

3. Detailed summary

1. For a language, only when the utf8_unicode_ci sorting is not done well, the utf8 character set correction related to the specific language will be performed. rule. For example, for German and French, utf8_unicode_ci works just fine, so there is no need to create special utf8 collation rules for these two languages.
2. utf8_general_ci is also applicable to German and French, except that '?' is equal to 's', not 'ss'. If your application can accept this, you should use utf8_general_ci because it is fast. Otherwise, use utf8_unicode_ci since it is more accurate.

Use one sentence to summarize the above paragraph: utf8_unicode_ci is more accurate, and utf8_general_ci is faster. Under normal circumstances, the accuracy of utf8_general_ci is enough for our use. After I read many program source codes, I found that most of them also use utf8_general_ci, so when creating a new database, generally choose utf8_general_ci.

4. How to use UTF8 in MySQL5.0
Add the following parameters in my.cnf

[mysqld]
init_connect=&#39;SET NAMES utf8′
default-character-set=utf8
default-collation = utf8_general_ci

Execute query mysql> show variables; Related as follows:

character_set_client | utf8 
character_set_connection | utf8 
character_set_database | utf8 
character_set_results | utf8 
character_set_server | utf8 
character_set_system | utf8

collation_connection | utf8_general_ci 
collation_database | utf8_general_ci 
collation_server | utf8_general_ci

Personal opinion, for the use of databases, utf8 - general is accurate enough, and compared with utf8 - unicode, it has an advantage in speed, so you can use it with confidence

附1：旧数据升级办法

以原来的字符集为latin1为例，升级成为utf8的字符集。原来的表: old_table (default charset=latin1)，新表：new_table(default charset=utf8)。

第一步：导出旧数据

mysqldump --default-character-set=latin1 -hlocalhost -uroot -B my_db --tables old_table > old.sql

第二步：转换编码(类似unix/linux环境下)

iconv -t utf-8 -f gb2312 -c old.sql > new.sql

或者可以去掉 -f 参数，让iconv自动判断原来的字符集

iconv -t utf-8 -c old.sql > new.sql

在这里，假定原来的数据默认是gb2312编码。

第三步：导入

修改old.sql，在插入/更新语句开始之前，增加一条sql语句： "SET NAMES utf8;"，保存。

mysql -hlocalhost -uroot my_db < new.sql

大功告成！！

附2：支持查看utf8字符集的MySQL客户端有
1.) MySQL-Front，据说这个项目已经被MySQL AB勒令停止了，不知为何，如果国内还有不少破解版可以下载（不代表我推荐使用破解版 :-P）。
2.) Navicat，另一款非常不错的MySQL客户端，汉化版刚出来，还邀请我试用过，总的来说还是不错的，不过也需要付费。
3.) PhpMyAdmin，开源的php项目，非常好。
4.) Linux下的终端工具（Linux terminal），把终端的字符集设置为utf8，连接到MySQL之后，执行 SET NAMES UTF8; 也能读写utf8数据了。

本篇文章到这里就已经全部结束了，更多其他精彩内容可以关注PHP中文网的MySQL视频教程栏目！

The above is the detailed content of What is the difference between utf8_unicode_ci and utf8_general_ci in Mysql?. For more information, please follow other related articles on the PHP Chinese website!

Statement

This article is reproduced at:脚本之家. If there is any infringement, please contact admin@php.cn delete

Explain the InnoDB Buffer Pool and its importance for performance.Apr 19, 2025 am 12:24 AM

InnoDBBufferPool reduces disk I/O by caching data and indexing pages, improving database performance. Its working principle includes: 1. Data reading: Read data from BufferPool; 2. Data writing: After modifying the data, write to BufferPool and refresh it to disk regularly; 3. Cache management: Use the LRU algorithm to manage cache pages; 4. Reading mechanism: Load adjacent data pages in advance. By sizing the BufferPool and using multiple instances, database performance can be optimized.

MySQL vs. Other Programming Languages: A ComparisonApr 19, 2025 am 12:22 AM

Compared with other programming languages, MySQL is mainly used to store and manage data, while other languages such as Python, Java, and C are used for logical processing and application development. MySQL is known for its high performance, scalability and cross-platform support, suitable for data management needs, while other languages have advantages in their respective fields such as data analytics, enterprise applications, and system programming.

Learning MySQL: A Step-by-Step Guide for New UsersApr 19, 2025 am 12:19 AM

MySQL is worth learning because it is a powerful open source database management system suitable for data storage, management and analysis. 1) MySQL is a relational database that uses SQL to operate data and is suitable for structured data management. 2) The SQL language is the key to interacting with MySQL and supports CRUD operations. 3) The working principle of MySQL includes client/server architecture, storage engine and query optimizer. 4) Basic usage includes creating databases and tables, and advanced usage involves joining tables using JOIN. 5) Common errors include syntax errors and permission issues, and debugging skills include checking syntax and using EXPLAIN commands. 6) Performance optimization involves the use of indexes, optimization of SQL statements and regular maintenance of databases.

MySQL: Essential Skills for Beginners to MasterApr 18, 2025 am 12:24 AM

MySQL is suitable for beginners to learn database skills. 1. Install MySQL server and client tools. 2. Understand basic SQL queries, such as SELECT. 3. Master data operations: create tables, insert, update, and delete data. 4. Learn advanced skills: subquery and window functions. 5. Debugging and optimization: Check syntax, use indexes, avoid SELECT*, and use LIMIT.

MySQL: Structured Data and Relational DatabasesApr 18, 2025 am 12:22 AM

MySQL efficiently manages structured data through table structure and SQL query, and implements inter-table relationships through foreign keys. 1. Define the data format and type when creating a table. 2. Use foreign keys to establish relationships between tables. 3. Improve performance through indexing and query optimization. 4. Regularly backup and monitor databases to ensure data security and performance optimization.

MySQL: Key Features and Capabilities ExplainedApr 18, 2025 am 12:17 AM

MySQL is an open source relational database management system that is widely used in Web development. Its key features include: 1. Supports multiple storage engines, such as InnoDB and MyISAM, suitable for different scenarios; 2. Provides master-slave replication functions to facilitate load balancing and data backup; 3. Improve query efficiency through query optimization and index use.

The Purpose of SQL: Interacting with MySQL DatabasesApr 18, 2025 am 12:12 AM

SQL is used to interact with MySQL database to realize data addition, deletion, modification, inspection and database design. 1) SQL performs data operations through SELECT, INSERT, UPDATE, DELETE statements; 2) Use CREATE, ALTER, DROP statements for database design and management; 3) Complex queries and data analysis are implemented through SQL to improve business decision-making efficiency.

MySQL for Beginners: Getting Started with Database ManagementApr 18, 2025 am 12:10 AM

The basic operations of MySQL include creating databases, tables, and using SQL to perform CRUD operations on data. 1. Create a database: CREATEDATABASEmy_first_db; 2. Create a table: CREATETABLEbooks(idINTAUTO_INCREMENTPRIMARYKEY, titleVARCHAR(100)NOTNULL, authorVARCHAR(100)NOTNULL, published_yearINT); 3. Insert data: INSERTINTObooks(title, author, published_year)VA

See all articles

Hot AI Tools

Undresser.AI Undress

AI-powered app for creating realistic nude photos

AI Clothes Remover

Online AI tool for removing clothes from photos.

Undress AI Tool

Undress images for free

Clothoff.io

AI clothes remover

Video Face Swap

Swap faces in any video effortlessly with our completely free AI face swap tool!

Hot Article

Assassin's Creed Shadows: Seashell Riddle Solution

3 weeks agoByDDD

What's New in Windows 11 KB5054979 & How to Fix Update Issues

2 weeks agoByDDD

Where to find the Crane Control Keycard in Atomfall

3 weeks agoByDDD

Assassin's Creed Shadows - How To Find The Blacksmith And Unlock Weapon And Armour Customisation

1 months agoByDDD

Roblox: Dead Rails - How To Complete Every Challenge

3 weeks agoByDDD

Hot Tools

SublimeText3 English version

Recommended: Win version, supports code prompts!

mPDF

mPDF is a PHP library that can generate PDF files from UTF-8 encoded HTML. The original author, Ian Back, wrote mPDF to output PDF files "on the fly" from his website and handle different languages. It is slower than original scripts like HTML2FPDF and produces larger files when using Unicode fonts, but supports CSS styles etc. and has a lot of enhancements. Supports almost all languages, including RTL (Arabic and Hebrew) and CJK (Chinese, Japanese and Korean). Supports nested block-level elements (such as P, DIV),

SublimeText3 Mac version

God-level code editing software (SublimeText3)

MinGW - Minimalist GNU for Windows

This project is in the process of being migrated to osdn.net/projects/mingw, you can continue to follow us there. MinGW: A native Windows port of the GNU Compiler Collection (GCC), freely distributable import libraries and header files for building native Windows applications; includes extensions to the MSVC runtime to support C99 functionality. All MinGW software can run on 64-bit Windows platforms.

Atom editor mac version download

The most popular open source editor

Hot Topics

Where is the login entrance for gmail email?

7637

CakePHP Tutorial

1391

What is the format of the account name of steam

win11 activation key permanent

nyt connections hints and answers

150