Install & Compatibility
Where this runs
tested against v3.5.1 · pip install
no network on importno background threads
Install × environment matrix
Each cell = how many times install + import succeeded across repeated harness runs. Partial = flaky.
glibc = Debian/Ubuntu slim · musl = Alpine Linux
muslpy 3.10–3.910 runs
installs and imports cleanly · install 0.0s · import 0.091s · 18.7MB
glibcpy 3.10–3.910 runs
installs and imports cleanly · install 1.9s · import 0.075s · 19MB
17MB installed
● package 17MB
Code
Verified usage
Verified import paths — ran on the pinned version, not inferred.
from_bytes
✓ from charset_normalizer import from_bytes
Primary API for detecting encoding of a bytes/bytearray object. Returns a CharsetMatches container.
from_path
✓ from charset_normalizer import from_path
Primary API for detecting encoding of a file on disk; accepts str, bytes, or os.PathLike.
from_fp
✓ from charset_normalizer import from_fp
Primary API for detecting encoding from an already-open binary file pointer. Does NOT close the file pointer.
detect
✓ from charset_normalizer import detect
✗ import chardet; chardet.detect(...)
Legacy Chardet-compatible shim. Officially deprecated in favour of from_bytes; not planned for removal. Returns a dict with 'encoding', 'confidence', 'language'.
is_binary
✓ from charset_normalizer import is_binary
Utility to detect whether bytes/path/fp point to binary (non-text) content. Added in 3.3.x.
CharsetMatches
✓ from charset_normalizer.models import CharsetMatches
✗ from charset_normalizer import CharsetNormalizerMatches
Class was renamed from CharsetNormalizerMatches to CharsetMatches in 3.0. The old alias was removed.
CharsetMatch
✓ from charset_normalizer.models import CharsetMatch
✗ from charset_normalizer import CharsetNormalizerMatch
Renamed from CharsetNormalizerMatch in 3.0. Old alias removed.
Detect encoding of raw bytes, decode the content, and use the Chardet-compatible legacy shim.
from charset_normalizer import from_bytes, from_path, detect
# --- from raw bytes ---
raw = b'\xff\xfe' + 'Hello, world!'.encode('utf-16-le')
results = from_bytes(raw)
best = results.best()
if best is not None:
print('Encoding:', best.encoding) # e.g. 'utf_16'
print('Language:', best.language) # e.g. 'English' or ''
print('Decoded :', str(best)) # decoded unicode string
else:
print('Could not detect encoding (possibly binary data)')
# --- from a file path ---
# results2 = from_path('./data/sample.txt')
# print(str(results2.best()))
# --- Chardet-compatible legacy shim (deprecated but stable) ---
result = detect(raw)
print(result) # {'encoding': 'UTF-16', 'confidence': 1.0, 'language': ''}
if result['encoding']:
decoded = raw.decode(result['encoding'])
print('Legacy decoded:', decoded)
normalizer --version
Debug
Known issues
breakingClass aliases CharsetNormalizerMatch, CharsetNormalizerMatches, CharsetDetector, and CharsetDoctor were removed in 3.0. Code referencing these names will raise ImportError or AttributeError.fixReplace with CharsetMatch and CharsetMatches imported from charset_normalizer.models, or use the top-level from_bytes/from_path functions directly.
affects: <3.0
breakingPython 3.6 support was dropped in 3.1.0, and Python 3.5 support was dropped in 2.1.0. Installing 3.x on Python 3.6 is unsupported.fixPin charset-normalizer<3.1 for Python 3.6, or upgrade the Python interpreter.
affects: <3.1 for Python 3.6; <2.1 for Python 3.5
gotchadetect() is the legacy Chardet-compatible shim and is officially deprecated. It also lowers confidence automatically for small byte samples (3.4.3+), so results on short inputs may differ from Chardet.fixMigrate to from_bytes(...).best() for new code. Check best() for None before calling str() or accessing .encoding.
affects: >=3.0
gotchaFeeding truncated or incomplete multi-byte byte sequences (e.g. a partial UTF-16 or UTF-32 file) will likely produce incorrect or empty detection results. The library is not designed for streaming partial payloads.fixAlways pass the full byte sequence. Do not slice input for 'performance' — the library already samples internally (5 blocks of 512 bytes by default).
affects: all
gotchafrom_bytes/from_path return a CharsetMatches container, not a string or a single result. Calling str() directly on the container gives unexpected output. Always call .best() first, then check for None.fixUse: result = from_bytes(raw).best(); text = str(result) if result is not None else ''
affects: all
gotchaThe import name uses an underscore (charset_normalizer) but the PyPI/install name uses a hyphen (charset-normalizer). Using import charset-normalizer raises a SyntaxError.fixAlways use: from charset_normalizer import ...
affects: all
deprecatedInternal module charset_normalizer.assets was moved into charset_normalizer.constant in 3.3.x. Any code importing from charset_normalizer.assets directly will break on 3.3+.fixDo not import internal modules. Use only the public API: from_bytes, from_path, from_fp, detect, is_binary.
affects: >=3.3
Errors
Common errors & fixes
AttributeError: partially initialized module 'charset_normalizer' has no attribute 'md__mypyc' (most likely due to a circular import)
This error typically indicates a corrupted or incomplete installation of `charset-normalizer`, often due to file shadowing, stale `__pycache__` files, or issues within specific build environments like PyInstaller.
fixReinstall the package cleanly using `pip install --force-reinstall charset-normalizer` or, if using conda, `conda install -c conda-forge charset-normalizer` after uninstalling any existing version.
ModuleNotFoundError: No module named 'charset_normalizer'
The `charset-normalizer` package is not installed in the active Python environment or is not discoverable in the Python path.
fixInstall the package using `pip install charset-normalizer` or `conda install charset-normalizer` depending on your environment.
normalizer: command not found
The `normalizer` CLI tool, which comes with the `charset-normalizer` library, is not found in your system's PATH or was not installed correctly.
fixEnsure `charset-normalizer` is installed in an environment whose scripts directory is in your system's PATH, or run the tool using `python -m charset_normalizer`.
ImportError: cannot import name 'COMMON_SAFE_ASCII_CHARACTERS' from 'charset_normalizer.constant'
This usually points to a version incompatibility or a corrupted installation, often occurring when `charset-normalizer` is used alongside other libraries (like `transformers` or `chardet`) that expect a different internal structure or version.
fixCleanly uninstall both `charset-normalizer` and any directly dependent libraries (like `chardet` if present), then reinstall `charset-normalizer` and the dependent libraries to ensure compatible versions are used.
Upgrade
Version history
3.5.1latest on PyPI · released Aug 15, 2026
Audit
Dependencies
No dependency data recorded yet.