Home  >  Article  >  Java  >  Use java.nio.charset.CharsetDecoder to automatically identify character set methods

Use java.nio.charset.CharsetDecoder to automatically identify character set methods

高洛峰
高洛峰Original
2017-03-12 09:43:232314browse

This articleIntroductionUsing java.nio.charset.CharsetDecoder to automatically identify Character setmethod

Study on how to use java.nio.charset. The most effective way to automatically identify character sets found is to use the third-party

class library

jchardet. There is also cpdetector, which actually uses jchardet. I accidentally discovered that jdk's java.nio.charset.CharsetDecoder can be used to identify character sets. 1. Principle

Generally, two methods are used to construct InputStreamReader:

InputStreamReader reader = new InputStreamReader(in, charsetName);

or

InputStreamReader reader = new InputStreamReader(in, charset);

If the charset does not match, garbled characters will be output.

There is also a construction method, which is to use CharsetDecoder:

CharsetDecoder cd = charset.newDecoder();
InputStreamReader reader = new InputStreamReader(in, cd);

If there is no match at this time,

throws an exception

:

java.nio.charset.MalformedInputException: Input length = 1
    at java.nio.charset.CoderResult.throwException(CoderResult.java:277)
    at sun.nio.cs.StreamDecoder.implRead(StreamDecoder.java:338)
    at sun.nio.cs.StreamDecoder.read(StreamDecoder.java:177)
        ....

In this way, it can be used as character set detection. 2. Use of AutoCharsetReader

AutoCharsetReader is a class written based on the above principles and with reference to InputStreamReader.

Inherits

Reader , can be seen as Charset adaptive InputStreamReader.

AutoCharsetReader ar= new AutoCharsetReader(in);char c = ar.read();
...char[] cbuf = new char[2000];
ar.read(cbuf);
...
BufferedReader br = new BufferedReader(ar);
br.readLine();
...

Another example is Lucene's TextField that creates a full-text

index

requires a Reader parameter. You can use this class directly:

Field field = new TextField("content", new AutoCharsetReader(file));

After reading the file, you can get the charset of the file. Note, this is after reading.

Charset charset = ar.charset();

3. Alternative character set

Because of the method of multiple attempts to finalize the character set, so provide alternatives. The default alternative character set provided by the current code is as follows:

    private final static String[] _defaultCharsets = {        
            "US-ASCII",            "UTF-8",            "GB2312", 
            "BIG5",            "GBK",            "GB18030",                
            "UTF-16BE", 
            "UTF-16LE", 
            "UTF-16",            "UNICODE"};

also provides a method to change the alternative character set. For example:

AutoCharsetReader ar = new AutoCharsetReader(in).setCharset("ascii", "utf-8", "gbk");

The order will affect the detection results. For example, if GBK is before GB2312, the detection result can only be GBK, not GB2312, because GBK contains GB2312. 4. Only for character set detection

Can be used only for character set detection:

charset = AutoCharsetReader.quickDetect(file.toURI().toURL(), charsets);
or:
charset = AutoCharsetReader.deepDetect(file.toURI().toURL(), charsets, stops);

quickDetect only reads one character and is suitable for single character set files. For html, you may need to read it all to know the charset, so use deepDetect. The parameter charsets can be null

. If a set of files, the known possible character sets are "ascii", "utf-8", "gb2312", and "gbk", when it is detected that the character set of a file is "utf- 8" or "gbk", the results can be returned immediately without continuing to read the file. At this time, you can assign the stops parameter to {"utf-8", "gbk"}. If

null

, you need to read them all. 5. Others

#In order to improve efficiency, this class has a buffer. If the initial character set decoding fails, there is no need to re-read io . The buffer size defaults to 8192. The object can be

defined by itself when constructing the

buffer size. If the parameter is less than 16, set it to 16.

### ###

The above is the detailed content of Use java.nio.charset.CharsetDecoder to automatically identify character set methods. For more information, please follow other related articles on the PHP Chinese website!

Statement:
The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn