Repository navigation
Speed up JSON string serialization - #844
Open
jonathan343 wants to merge 1 commit into
Open
jonathan343 wants to merge 1 commit into
jonathan343 wants to merge 1 commit into
Conversation
Replace custom escaping with the standard library encoder. Keep Unicode output as UTF-8 and reject invalid strings before writing.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Replaces our custom JSON string escaping code with Python's standard JSON string encoder,
json.encoder.encode_basestring.The JSON bytes stay the same for valid strings, including non-English text and emoji. Strings that cannot be encoded as UTF-8 still raise an error. They now fail before writing any part of that string, and the error details may change.
Performance testing
Compared
developwith this branch on Python 3.12.14 using single-string inputs of different sizes. Each input is ASCII, so each input character is one byte before JSON escaping.The old code makes one replacement for each double quote, backslash, or control character from U+0000 through U+001F. The test inputs contain an evenly spread mix of quotes, backslashes, newlines, tabs, and NUL characters, with roughly 0%, 1%, or 10% of characters needing escaping. The counts below match the original regex exactly.
Times are medians of 33 samples per case, measuring the full
JSONCodec.serializecall. Input construction and correctness checks are outside the timed portion.Both versions produced identical JSON bytes for every input. These are synthetic string benchmarks, not measurements of AWS request latency.
By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.