ML-based charset encoding detector. MojibusterEncodingDetector runs structural detectors for UTF-32 / UTF-16 / UTF-8 and then a 33-class byte-bigram Naive Bayes classifier covering CJK multi-byte, EBCDIC, DOS OEM, Cyrillic, Windows single-byte, ISO-8859, and Mac encodings. This detector replaces ICU4J and juniversalchardet as the default statistical encoding detector in Tika.

Latest Versions

2 versions โ†’
VersionVulnerabilitiesUsagesDate
4.0.x
4.0.0-beta-1
10
Jul 04, 2026
4.0.0-alpha-1
5
May 10, 2026
2 versions โ†’