Language-agnostic text quality scorer that discriminates between clean UTF-8 text and mojibake, reversed text, wrong-codec decodings, and other corruption forms. Provides a standalone "languageyness" score suitable for re-OCR triggering and charset-decoding arbitration. Runtime classes (JunkDetector, ScriptDetector, feature extractors) and bundled model resources live here. Training and evaluation CLI tools live in the tools subpackage.

Latest Versions

3 versions →
VersionVulnerabilitiesUsagesDate
4.0.x
4.0.0
11
Aug 23, 2026
4.0.0-beta-1
11
Jul 04, 2026
4.0.0-alpha-1
4
May 10, 2026
3 versions →