Java implements obtaining the character encoding of a text file-JavaBase-php.cn

Home

Java

JavaBase

Java implements obtaining the character encoding of a text file

王林

Dec 23, 2019 am 11:49 AM

javaCharacter Encodingaccomplishtext fileObtain

Java implements obtaining the character encoding of a text file

1. Understanding character encoding:

1. The default encoding of String in Java is UTF-8, which can be obtained using the following statement: Charset.defaultCharset( );

2. Under the Windows operating system, the default encoding of text files is ANSI, which is GBK for Chinese Windows. For example, if we use the Notepad program to create a new text document, its default character encoding is ANSI.

3. Text text documents have four encoding options: ANSI, Unicode (including Unicode Big Endian and Unicode Little Endian), UTF-8, UTF-16

4, so we read txt files may sometimes not know their encoding format, so a program needs to be used to dynamically determine the encoding of the txt file.

ANSI : No format definition, for Chinese operating systems it is GBK or GB2312

UTF-8 : The first three bytes are: 0xE59B9E (UTF-8), 0xEFBBBF (UTF-8 inclusive BOM)

UTF-16: The first two bytes are: 0xFEFF

Unicode: The first two bytes are: 0xFFFE

For example: Unicode documents start with 0xFFFE, use The program just takes out the first few bytes and makes a judgment.

5. Correspondence between Java encoding and Text encoding:

Java implements obtaining the character encoding of a text file

Java reads Text files. If the encoding format does not match, garbled characters will appear. Therefore, you need to set the correct character encoding when reading text files. The encoding format of Text documents is written in the file header. In the program, the encoding format of the file needs to be parsed first. After obtaining the encoding format, reading the file in this format will avoid garbled characters.

Free online video tutorial recommendation: java learning

2. For example:

There is a text file: test.txt

Java implements obtaining the character encoding of a text file

Test code:

/**
 * 文件名：CharsetCodeTest.java
 * 功能描述：文件字符编码测试
 */
 
import java.io.*;
 
public class CharsetCodeTest {
    public static void main(String[] args) throws Exception {
        String filePath = "test.txt";
        String content = readTxt(filePath);
        System.out.println(content);
    }
 
 
public static String readTxt(String path) {
        StringBuilder content = new StringBuilder("");
        try {
            String fileCharsetName = getFileCharsetName(path);
            System.out.println("文件的编码格式为："+fileCharsetName);
 
            InputStream is = new FileInputStream(path);
            InputStreamReader isr = new InputStreamReader(is, fileCharsetName);
            BufferedReader br = new BufferedReader(isr);
 
            String str = "";
            boolean isFirst = true;
            while (null != (str = br.readLine())) {
                if (!isFirst)
                    content.append(System.lineSeparator());
                    //System.getProperty("line.separator");
                else
                    isFirst = false;
                content.append(str);
            }
            br.close();
        } catch (Exception e) {
            e.printStackTrace();
            System.err.println("读取文件:" + path + "失败!");
        }
        return content.toString();
    }
 
 
    public static String getFileCharsetName(String fileName) throws IOException {
        InputStream inputStream = new FileInputStream(fileName);
        byte[] head = new byte[3];
        inputStream.read(head);
 
        String charsetName = "GBK";//或GB2312，即ANSI
        if (head[0] == -1 && head[1] == -2 ) //0xFFFE
            charsetName = "UTF-16";
        else if (head[0] == -2 && head[1] == -1 ) //0xFEFF
            charsetName = "Unicode";//包含两种编码格式：UCS2-Big-Endian和UCS2-Little-Endian
        else if(head[0]==-27 && head[1]==-101 && head[2] ==-98)
            charsetName = "UTF-8"; //UTF-8(不含BOM)
        else if(head[0]==-17 && head[1]==-69 && head[2] ==-65)
            charsetName = "UTF-8"; //UTF-8-BOM
 
        inputStream.close();
 
        //System.out.println(code);
        return charsetName;
    }
}

Running results:

Java implements obtaining the character encoding of a text file

Recommended related articles and tutorials: Getting started with java

The above is the detailed content of Java implements obtaining the character encoding of a text file. For more information, please follow other related articles on the PHP Chinese website!

Statement

This article is reproduced at:CSDN. If there is any infringement, please contact admin@php.cn delete

带你搞懂Java结构化数据处理开源库SPLMay 24, 2022 pm 01:34 PM

本篇文章给大家带来了关于java的相关知识，其中主要介绍了关于结构化数据处理开源库SPL的相关问题，下面就一起来看一下java下理想的结构化数据处理类库，希望对大家有帮助。

Java集合框架之PriorityQueue优先级队列Jun 09, 2022 am 11:47 AM

本篇文章给大家带来了关于java的相关知识，其中主要介绍了关于PriorityQueue优先级队列的相关知识，Java集合框架中提供了PriorityQueue和PriorityBlockingQueue两种类型的优先级队列，PriorityQueue是线程不安全的，PriorityBlockingQueue是线程安全的，下面一起来看一下，希望对大家有帮助。

完全掌握Java锁（图文解析）Jun 14, 2022 am 11:47 AM

本篇文章给大家带来了关于java的相关知识，其中主要介绍了关于java锁的相关问题，包括了独占锁、悲观锁、乐观锁、共享锁等等内容，下面一起来看一下，希望对大家有帮助。

一起聊聊Java多线程之线程安全问题Apr 21, 2022 pm 06:17 PM

本篇文章给大家带来了关于java的相关知识，其中主要介绍了关于多线程的相关问题，包括了线程安装、线程加锁与线程不安全的原因、线程安全的标准类等等内容，希望对大家有帮助。

Java基础归纳之枚举May 26, 2022 am 11:50 AM

本篇文章给大家带来了关于java的相关知识，其中主要介绍了关于枚举的相关问题，包括了枚举的基本操作、集合类对枚举的支持等等内容，下面一起来看一下，希望对大家有帮助。

详细解析Java的this和super关键字Apr 30, 2022 am 09:00 AM

本篇文章给大家带来了关于Java的相关知识，其中主要介绍了关于关键字中this和super的相关问题，以及他们的一些区别，下面一起来看一下，希望对大家有帮助。

java中封装是什么May 16, 2019 pm 06:08 PM

封装是一种信息隐藏技术，是指一种将抽象性函式接口的实现细节部分包装、隐藏起来的方法；封装可以被认为是一个保护屏障，防止指定类的代码和数据被外部类定义的代码随机访问。封装可以通过关键字private，protected和public实现。

Java数据结构之AVL树详解Jun 01, 2022 am 11:39 AM

本篇文章给大家带来了关于java的相关知识，其中主要介绍了关于平衡二叉树（AVL树）的相关知识，AVL树本质上是带了平衡功能的二叉查找树，下面一起来看一下，希望对大家有帮助。

See all articles

Hot AI Tools

Undresser.AI Undress

AI-powered app for creating realistic nude photos

AI Clothes Remover

Online AI tool for removing clothes from photos.

Undress AI Tool

Undress images for free

Clothoff.io

AI clothes remover

AI Hentai Generator

Generate AI Hentai for free.

Hot Article

R.E.P.O. Energy Crystals Explained and What They Do (Yellow Crystal)

2 weeks agoBy尊渡假赌尊渡假赌尊渡假赌

How Long Does It Take To Beat Split Fiction?

1 months agoByDDD

R.E.P.O. Save File Location: Where Is It & How to Protect It?

1 months agoByDDD

R.E.P.O. Best Graphic Settings

2 weeks agoBy尊渡假赌尊渡假赌尊渡假赌

Assassin's Creed Shadows: Seashell Riddle Solution

1 weeks agoByDDD

Hot Tools

SublimeText3 English version

Recommended: Win version, supports code prompts!

SAP NetWeaver Server Adapter for Eclipse

Integrate Eclipse with SAP NetWeaver application server.

WebStorm Mac version

Useful JavaScript development tools

SublimeText3 Linux new version

SublimeText3 Linux latest version

MinGW - Minimalist GNU for Windows

This project is in the process of being migrated to osdn.net/projects/mingw, you can continue to follow us there. MinGW: A native Windows port of the GNU Compiler Collection (GCC), freely distributable import libraries and header files for building native Windows applications; includes extensions to the MSVC runtime to support C99 functionality. All MinGW software can run on 64-bit Windows platforms.

Hot Topics

Where is the login entrance for gmail email?

7391

1630

1357

1268

1216