1. Understanding character encoding:
1. The default encoding of String in Java is UTF-8, which can be obtained using the following statement: Charset.defaultCharset( );
2. Under the Windows operating system, the default encoding of text files is ANSI, which is GBK for Chinese Windows. For example, if we use the Notepad program to create a new text document, its default character encoding is ANSI.
3. Text text documents have four encoding options: ANSI, Unicode (including Unicode Big Endian and Unicode Little Endian), UTF-8, UTF-16
4, so we read txt files may sometimes not know their encoding format, so a program needs to be used to dynamically determine the encoding of the txt file.
ANSI : No format definition, for Chinese operating systems it is GBK or GB2312
UTF-8 : The first three bytes are: 0xE59B9E (UTF-8), 0xEFBBBF (UTF-8 inclusive BOM)
UTF-16: The first two bytes are: 0xFEFF
Unicode: The first two bytes are: 0xFFFE
For example: Unicode documents start with 0xFFFE, use The program just takes out the first few bytes and makes a judgment.
5. Correspondence between Java encoding and Text encoding:
Java reads Text files. If the encoding format does not match, garbled characters will appear. Therefore, you need to set the correct character encoding when reading text files. The encoding format of Text documents is written in the file header. In the program, the encoding format of the file needs to be parsed first. After obtaining the encoding format, reading the file in this format will avoid garbled characters.
Free online video tutorial recommendation: java learning
2. For example:
There is a text file: test.txt
Test code:
/** * 文件名:CharsetCodeTest.java * 功能描述:文件字符编码测试 */ import java.io.*; public class CharsetCodeTest { public static void main(String[] args) throws Exception { String filePath = "test.txt"; String content = readTxt(filePath); System.out.println(content); } public static String readTxt(String path) { StringBuilder content = new StringBuilder(""); try { String fileCharsetName = getFileCharsetName(path); System.out.println("文件的编码格式为:"+fileCharsetName); InputStream is = new FileInputStream(path); InputStreamReader isr = new InputStreamReader(is, fileCharsetName); BufferedReader br = new BufferedReader(isr); String str = ""; boolean isFirst = true; while (null != (str = br.readLine())) { if (!isFirst) content.append(System.lineSeparator()); //System.getProperty("line.separator"); else isFirst = false; content.append(str); } br.close(); } catch (Exception e) { e.printStackTrace(); System.err.println("读取文件:" + path + "失败!"); } return content.toString(); } public static String getFileCharsetName(String fileName) throws IOException { InputStream inputStream = new FileInputStream(fileName); byte[] head = new byte[3]; inputStream.read(head); String charsetName = "GBK";//或GB2312,即ANSI if (head[0] == -1 && head[1] == -2 ) //0xFFFE charsetName = "UTF-16"; else if (head[0] == -2 && head[1] == -1 ) //0xFEFF charsetName = "Unicode";//包含两种编码格式:UCS2-Big-Endian和UCS2-Little-Endian else if(head[0]==-27 && head[1]==-101 && head[2] ==-98) charsetName = "UTF-8"; //UTF-8(不含BOM) else if(head[0]==-17 && head[1]==-69 && head[2] ==-65) charsetName = "UTF-8"; //UTF-8-BOM inputStream.close(); //System.out.println(code); return charsetName; } }
Running results:
Recommended related articles and tutorials: Getting started with java
The above is the detailed content of Java implements obtaining the character encoding of a text file. For more information, please follow other related articles on the PHP Chinese website!