Registry / serialization / pysimdjson

pysimdjson

JSON →
library7.0.2pypypi✓ verified 31d ago

pysimdjson provides high-performance Python bindings for the simdjson C++ library, a SIMD-accelerated JSON parser. It offers both a compatibility API similar to Python's built-in `json` module and a native API for significantly faster parsing, especially when only parts of a JSON document are needed. The library is actively maintained, with the current version being 7.0.2.

pip install pysimdjson
INSTALL
IMPORT
SIG · PYSIMDJSON
P
pysimdjson
serializationpythonv7.0.2
Install
1.9s avg
Import
19ms
Disk
31MB
Pass rate
10/ 10
Env Coverage10 / 10
glibc
3.9–3.13
musl
3.9–3.13
Install & Compatibility
Where this runs
tested against v7.0.2 · pip install
no network on importno background threads
Install × environment matrix
Each cell = how many times install + import succeeded across repeated harness runs. Partial = flaky.
glibc = Debian/Ubuntu slim · musl = Alpine Linux
musl
py 3.10–3.95 runs
installs and imports cleanly · install 0.0s · import 0.022s · 33.5MB
glibc
py 3.10–3.95 runs
installs and imports cleanly · install 1.9s · import 0.016s · 33MB
31MB installed
● package 31MB
Code
Verified usage

Verified import paths — ran on the pinned version, not inferred.

Parser
✓ from simdjson import Parser
✗ from pysimdjson import Parser
loads
✓ from simdjson import loads
✗ from pysimdjson import loads
dumps
✓ from simdjson import dumps
✗ from pysimdjson import dumps

This quickstart demonstrates both the high-performance native API using `pysimdjson.Parser` for selective data extraction and the `pysimdjson.loads` function for full document parsing, similar to the standard `json` module. The native API is generally preferred for large documents to avoid unnecessary object materialization.

from pysimdjson import Parser json_data = b'{"name": "Alice", "age": 30, "city": "New York", "details": {"occupation": "Engineer", "hobbies": ["reading", "hiking"]}}' # Using the native Parser API for performance and partial loading parser = Parser() try: # Parsing bytes is generally fastest doc = parser.parse(json_data) # Accessing elements without fully materializing the document name = doc['name'].as_str() age = doc['age'].as_int() occupation = doc['details']['occupation'].as_str() first_hobby = doc['details']['hobbies'][0].as_str() print(f"Name: {name}, Age: {age}") print(f"Occupation: {occupation}, First Hobby: {first_hobby}") # Convert a subtree to a Python object if needed details_dict = doc['details'].as_dict() print(f"Details as dict: {details_dict}") except RuntimeError as e: print(f"Error during parsing or access: {e}") # For simple full document loading, compatible with json.loads from pysimdjson import loads full_python_obj = loads(json_data) print(f"Full Python object (loads): {full_python_obj}")
Debug
Known issues
breakingPython 3.5 and 3.6 support has been removed in prior major releases. Current versions (>=7.0.0) require Python 3.9 or newer. Ensure your environment meets the Python version requirement.
fix
Upgrade your Python interpreter to version 3.9 or higher.
affects: <7.0.0 (Python 3.5, 3.6)
gotchaFor optimal performance, especially with large JSON documents, avoid fully materializing the entire document into Python objects. Use the native `Parser` API with methods like `at_pointer()` or direct proxy access (e.g., `doc['key']`) to extract only the necessary parts.
fix
Utilize `pysimdjson.Parser().parse(data)` and navigate the `doc` object using dictionary-like or list-like access, then convert to Python primitives (e.g., `.as_str()`, `.as_int()`, `.as_dict()`) only for the data you need.
affects: All
gotchaWhen reusing a `pysimdjson.Parser` instance, ensure that no `Object` or `Array` proxies from a previously parsed document are still in scope. Calling `parse()` or `load()` on a parser while old proxies exist may lead to a `RuntimeError` due to memory management conflicts.
fix
Allow previously created `Object` and `Array` proxy objects to go out of scope or explicitly delete them before reusing a `Parser` instance for a new document. Consider creating a new `Parser` instance for each document if managing proxy lifetimes is complex.
affects: All
gotchapysimdjson primarily operates on `bytes` and assumes UTF-8 encoding. It does not provide options to specify alternative encodings, unlike the standard `json` module. Providing `str` will be slower due to internal encoding.
fix
Always provide JSON data as `bytes` (e.g., `b'{ "key": "value" }'`) to the parser for best performance and to avoid encoding issues.
affects: All
Upgrade
Version history
7.0.2latest on PyPI · released Jun 28, 2025
Audit
Dependencies
pythonrequiredRequired Python interpreter version.
C++11-capable compileroptionalRequired for building from source if pre-compiled binary wheels are not available for your platform.
Agent activity
9 hits · last 30 days
node
8
Resources
pysimdjson — pip install pysimdjson · libregistry