Language-agnostic text quality scorer that discriminates between clean UTF-8 text and
mojibake, reversed text, wrong-codec decodings, and other corruption forms.
Provides a standalone "languageyness" score suitable for re-OCR triggering and
charset-decoding arbitration.
Runtime classes (JunkDetector, ScriptDetector, feature extractors) and bundled model
resources live here. Training and evaluation CLI tools live in the tools subpackage.
Latest Versions
3 versions →| Version ▼ | Vulnerabilities | Usages | Date | |
|---|---|---|---|---|
4.0.x | 4.0.0 |
11
| Aug 23, 2026 | |
| 4.0.0-beta-1 |
11
| Jul 04, 2026 | ||
| 4.0.0-alpha-1 |
4
| May 10, 2026 |