ML-based charset encoding detector. MojibusterEncodingDetector
runs structural detectors for UTF-32 / UTF-16 / UTF-8 and then a
33-class byte-bigram Naive Bayes classifier covering CJK multi-byte,
EBCDIC, DOS OEM, Cyrillic, Windows single-byte, ISO-8859, and Mac
encodings. This detector replaces ICU4J and juniversalchardet as the
default statistical encoding detector in Tika.
Latest Versions
2 versions โ| Version | Vulnerabilities | Usages | Date | |
|---|---|---|---|---|
4.0.x | 4.0.0-beta-1 |
10
| Jul 04, 2026 | |
| 4.0.0-alpha-1 |
5
| May 10, 2026 |