Skip to content

Speed up JSON string serialization - #844

Open
jonathan343 wants to merge 1 commit into
smithy-lang:developfrom
jonathan343:perf/json-string-escaping
Open

jonathan343 wants to merge 1 commit into
smithy-lang:developfrom
jonathan343:perf/json-string-escaping

Conversation

@jonathan343

Copy link
Copy Markdown
Contributor

Summary

Replaces our custom JSON string escaping code with Python's standard JSON string encoder, json.encoder.encode_basestring.

The JSON bytes stay the same for valid strings, including non-English text and emoji. Strings that cannot be encoded as UTF-8 still raise an error. They now fail before writing any part of that string, and the error details may change.

Performance testing

Compared develop with this branch on Python 3.12.14 using single-string inputs of different sizes. Each input is ASCII, so each input character is one byte before JSON escaping.

The old code makes one replacement for each double quote, backslash, or control character from U+0000 through U+001F. The test inputs contain an evenly spread mix of quotes, backslashes, newlines, tabs, and NUL characters, with roughly 0%, 1%, or 10% of characters needing escaping. The counts below match the original regex exactly.

Times are medians of 33 samples per case, measuring the full JSONCodec.serialize call. Input construction and correctness checks are outside the timed portion.

Input bytes Characters needing escaping Before After Speedup
1,024 0 2.7 µs 2.0 µs 1.4x
1,024 10 3.9 µs 2.2 µs 1.7x
1,024 102 13.2 µs 2.3 µs 5.8x
16,384 0 33.8 µs 24.1 µs 1.4x
16,384 163 50.9 µs 25.0 µs 2.0x
16,384 1,638 188.8 µs 25.7 µs 7.3x
262,144 0 530.7 µs 390.6 µs 1.4x
262,144 2,621 795.6 µs 393.3 µs 2.0x
262,144 26,214 3053.2 µs 404.9 µs 7.5x

Both versions produced identical JSON bytes for every input. These are synthetic string benchmarks, not measurements of AWS request latency.


By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.

Replace custom escaping with the standard library encoder. Keep Unicode output as UTF-8 and reject invalid strings before writing.
@jonathan343
jonathan343 requested a review from a team as a code owner October 10, 2026 05:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant