


Use java.nio.charset.CharsetDecoder to automatically identify character set methods
This articleIntroductionUsing java.nio.charset.CharsetDecoder to automatically identify Character setmethod
Study on how to use java.nio.charset. The most effective way to automatically identify character sets found is to use the third-party
class libraryjchardet. There is also cpdetector, which actually uses jchardet. I accidentally discovered that jdk's java.nio.charset.CharsetDecoder can be used to identify character sets. 1. Principle
Generally, two methods are used to construct InputStreamReader:
InputStreamReader reader = new InputStreamReader(in, charsetName);
or
InputStreamReader reader = new InputStreamReader(in, charset);
If the charset does not match, garbled characters will be output.
There is also a construction method, which is to use CharsetDecoder:
CharsetDecoder cd = charset.newDecoder(); InputStreamReader reader = new InputStreamReader(in, cd);
If there is no match at this time,
throws an exception:
java.nio.charset.MalformedInputException: Input length = 1 at java.nio.charset.CoderResult.throwException(CoderResult.java:277) at sun.nio.cs.StreamDecoder.implRead(StreamDecoder.java:338) at sun.nio.cs.StreamDecoder.read(StreamDecoder.java:177) ....
In this way, it can be used as character set detection. 2. Use of AutoCharsetReader
AutoCharsetReader is a class written based on the above principles and with reference to InputStreamReader.
InheritsReader , can be seen as Charset adaptive InputStreamReader.
AutoCharsetReader ar= new AutoCharsetReader(in);char c = ar.read(); ...char[] cbuf = new char[2000]; ar.read(cbuf); ... BufferedReader br = new BufferedReader(ar); br.readLine(); ...
Another example is Lucene's TextField that creates a full-text
indexrequires a Reader parameter. You can use this class directly:
Field field = new TextField("content", new AutoCharsetReader(file));
After reading the file, you can get the charset of the file. Note, this is after reading.
Charset charset = ar.charset();
3. Alternative character set
Because of the method of multiple attempts to finalize the character set, so provide alternatives. The default alternative character set provided by the current code is as follows:
private final static String[] _defaultCharsets = { "US-ASCII", "UTF-8", "GB2312", "BIG5", "GBK", "GB18030", "UTF-16BE", "UTF-16LE", "UTF-16", "UNICODE"};
also provides a method to change the alternative character set. For example:
AutoCharsetReader ar = new AutoCharsetReader(in).setCharset("ascii", "utf-8", "gbk");
The order will affect the detection results. For example, if GBK is before GB2312, the detection result can only be GBK, not GB2312, because GBK contains GB2312. 4. Only for character set detection
Can be used only for character set detection:
charset = AutoCharsetReader.quickDetect(file.toURI().toURL(), charsets); or: charset = AutoCharsetReader.deepDetect(file.toURI().toURL(), charsets, stops);
quickDetect only reads one character and is suitable for single character set files. For html, you may need to read it all to know the charset, so use deepDetect. The parameter charsets can be null
. If a set of files, the known possible character sets are "ascii", "utf-8", "gb2312", and "gbk", when it is detected that the character set of a file is "utf- 8" or "gbk", the results can be returned immediately without continuing to read the file. At this time, you can assign the stops parameter to {"utf-8", "gbk"}. If
null, you need to read them all. 5. Others
#In order to improve efficiency, this class has a buffer. If the initial character set decoding fails, there is no need to re-read io . The buffer size defaults to 8192. The object can be
defined by itself when constructing thebuffer size. If the parameter is less than 16, set it to 16.
The above is the detailed content of Use java.nio.charset.CharsetDecoder to automatically identify character set methods. For more information, please follow other related articles on the PHP Chinese website!

JVM works by converting Java code into machine code and managing resources. 1) Class loading: Load the .class file into memory. 2) Runtime data area: manage memory area. 3) Execution engine: interpret or compile execution bytecode. 4) Local method interface: interact with the operating system through JNI.

JVM enables Java to run across platforms. 1) JVM loads, validates and executes bytecode. 2) JVM's work includes class loading, bytecode verification, interpretation execution and memory management. 3) JVM supports advanced features such as dynamic class loading and reflection.

Java applications can run on different operating systems through the following steps: 1) Use File or Paths class to process file paths; 2) Set and obtain environment variables through System.getenv(); 3) Use Maven or Gradle to manage dependencies and test. Java's cross-platform capabilities rely on the JVM's abstraction layer, but still require manual handling of certain operating system-specific features.

Java requires specific configuration and tuning on different platforms. 1) Adjust JVM parameters, such as -Xms and -Xmx to set the heap size. 2) Choose the appropriate garbage collection strategy, such as ParallelGC or G1GC. 3) Configure the Native library to adapt to different platforms. These measures can enable Java applications to perform best in various environments.

OSGi,ApacheCommonsLang,JNA,andJVMoptionsareeffectiveforhandlingplatform-specificchallengesinJava.1)OSGimanagesdependenciesandisolatescomponents.2)ApacheCommonsLangprovidesutilityfunctions.3)JNAallowscallingnativecode.4)JVMoptionstweakapplicationbehav

JVMmanagesgarbagecollectionacrossplatformseffectivelybyusingagenerationalapproachandadaptingtoOSandhardwaredifferences.ItemploysvariouscollectorslikeSerial,Parallel,CMS,andG1,eachsuitedfordifferentscenarios.Performancecanbetunedwithflagslike-XX:NewRa

Java code can run on different operating systems without modification, because Java's "write once, run everywhere" philosophy is implemented by Java virtual machine (JVM). As the intermediary between the compiled Java bytecode and the operating system, the JVM translates the bytecode into specific machine instructions to ensure that the program can run independently on any platform with JVM installed.

The compilation and execution of Java programs achieve platform independence through bytecode and JVM. 1) Write Java source code and compile it into bytecode. 2) Use JVM to execute bytecode on any platform to ensure the code runs across platforms.


Hot AI Tools

Undresser.AI Undress
AI-powered app for creating realistic nude photos

AI Clothes Remover
Online AI tool for removing clothes from photos.

Undress AI Tool
Undress images for free

Clothoff.io
AI clothes remover

Video Face Swap
Swap faces in any video effortlessly with our completely free AI face swap tool!

Hot Article

Hot Tools

Atom editor mac version download
The most popular open source editor

mPDF
mPDF is a PHP library that can generate PDF files from UTF-8 encoded HTML. The original author, Ian Back, wrote mPDF to output PDF files "on the fly" from his website and handle different languages. It is slower than original scripts like HTML2FPDF and produces larger files when using Unicode fonts, but supports CSS styles etc. and has a lot of enhancements. Supports almost all languages, including RTL (Arabic and Hebrew) and CJK (Chinese, Japanese and Korean). Supports nested block-level elements (such as P, DIV),

Dreamweaver Mac version
Visual web development tools

SublimeText3 Linux new version
SublimeText3 Linux latest version

Dreamweaver CS6
Visual web development tools
