Skip to content

Generate schemas as on-demand hydration from tuples - #841

Open
kuhe wants to merge 3 commits into
developfrom
kuhe/chore/schemas
Open

kuhe wants to merge 3 commits into
developfrom
kuhe/chore/schemas

Conversation

@kuhe

@kuhe kuhe commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

In the schemas.py file, generate schema declarations as a compact tuples instead of a constructor or factory invocation. This expands lazily to the Schema class object, but only on-use, rather than on import of the client.

The result is a reduction in the memory footprint of schemas (delayed init, compact standby mode), and faster startup.

Both the faster startup and the size-on-disk reduction help any caller, and usage on Lambda in particular.


Code example:

# compact tuple (new)
QUERY_INCOMPATIBLE_OPERATION_INPUT = (18, _s0, _s171, {_s72: _s73})

# Schema init (old)
QUERY_INCOMPATIBLE_OPERATION_INPUT = Schema(...)

effect on clients

client form whole-pkg .py size import+hydrate time peak mem (RSS Δ)
DDB baseline 3.08 MB 111 ms 14.3 MiB
compact 1.27 MB 104 ms 12.0 MiB
EC2 baseline 15.58 MB 736 ms 75.9 MiB
compact 11.83 MB 650 ms 62.8 MiB

picture of schemas.py only

client form size LOC exec time mem after load
DDB baseline 1.75 MB 31,821 44.9 ms 8.4 MiB
compact 0.16 MB 6,717 1.2 ms 0.8 MiB
EC2 baseline 3.75 MB 126,101 124.4 ms 27.8 MiB
compact 1.74 MB 87,510 12.1 ms 7.8 MiB

@kuhe
kuhe force-pushed the kuhe/chore/schemas branch 3 times, most recently from 40881a1 to 8ffafa1 Compare October 9, 2026 19:32
@kuhe
kuhe force-pushed the kuhe/chore/schemas branch from 8ffafa1 to 8dbd922 Compare October 9, 2026 19:55
@kuhe
kuhe marked this pull request as ready for review October 9, 2026 20:05
@kuhe
kuhe requested a review from a team as a code owner October 9, 2026 20:05
Comment on lines +54 to +57
reportUnknownVariableType = false
reportUnknownMemberType = false
reportUnknownArgumentType = false
reportUnknownLambdaType = false

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why are these needed and can we have more isolated suppressions (e.g. per-file)? Or does the entire package need them?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll look into fixing, since I'd rather not add suppressions

…wide

Replace the four reportUnknown* settings in the generated pyproject with
inline # pyright: ignore comments emitted only on the recursive/forward-ref
schema lines that pyright cannot type (the lambda member targets and the
assignment lines holding them). Keeps strict type checking everywhere else in
the generated client.
# recursive member target that resolves back to this tuple sees the in-progress
# object instead of recursing forever.
schema = Schema(id=id_, shape_type=shape_type, traits=_expand_traits(traits))
_cache[key] = (data, schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

another thread can get this schema before it's populated and cache an empty members_by_index

members = schema.members
for index, mname in enumerate(member_names):
target_ref, member_traits = member_targets[index]
target = _resolve(target_ref)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If this raises an error, a partial schema could stay cached

traits[tid] = Trait.new(tid)
bit >>= 1
i += 1
_traits_cache[indicator] = traits

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is shared across schemas- would it be possible to return a copy or cache something immutable?

# Mutate the published members dict in place so a recursive target that
# captured it by reference during the cycle break observes the full map.
members = schema.members
for index, mname in enumerate(member_names):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we add tests for recursion, failed hydration, shared traits, and concurrent access?


schema: Schema = field(repr=False)
"""The schema of the operation."""
static_schema: "StaticOperationSchema" = field(repr=False)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This makes callers use a private tuple format to construct APIOperation. Could we keep the existing constructor and add an internal lazy path?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants