Skip to content

vmem: Add stronger tlb invalidation barrier - #1876

Merged
syntactically merged 1 commit into
mainfrom
lm/extra-tlbi
Oct 3, 2026
Merged

syntactically merged 1 commit into
mainfrom
lm/extra-tlbi

Conversation

@syntactically

Copy link
Copy Markdown
Member

No description provided.

@hyperlight-gh-bot

This comment has been minimized.

@syntactically syntactically added kind/enhancement For PRs adding features, improving functionality, docs, tests, etc. ready-for-review PR is ready for (re-)review labels Oct 2, 2026
Copilot AI balanced review requested due to automatic review settings October 2, 2026 12:29

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unaligned ranges can leave stale permissions cached, and the new architecture-specific behavior lacks tests.

Review effort: Balanced
Findings: 1 High severity · 1 Medium severity · 2 Low severity

Open (4)
What changed in this PR

Adds a guest paging barrier for invalidating cached translations after permission downgrades.

Changes:

  • Exposes downgrade_in_place.
  • Uses invlpg on amd64.
  • Uses DSB, TLBI, and ISB on AArch64.
File Description
src/​hyperlight_guest_bin/​src/​paging.rs Exposes the barrier API.
src/​hyperlight_guest_bin/​src/​arch/​amd64/​paging.rs Adds amd64 invalidation.
src/​hyperlight_guest_bin/​src/​arch/​aarch64/​paging.rs Adds AArch64 invalidation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/hyperlight_guest_bin/src/paging.rs
Comment thread src/hyperlight_guest_bin/src/paging.rs
Comment thread src/hyperlight_guest_bin/src/arch/aarch64/paging.rs
Comment thread src/hyperlight_guest_bin/src/arch/amd64/paging.rs
@hyperlight-gh-bot

This comment has been minimized.

Base automatically changed from lm/modify-mappings to main October 2, 2026 15:45
@hyperlight-gh-bot

This comment has been minimized.

@github-actions github-actions Bot removed the ready-for-review PR is ready for (re-)review label Oct 2, 2026
Signed-off-by: Lucy Menon <168595099+syntactically@users.noreply.github.com>
@hyperlight-gh-bot

Copy link
Copy Markdown

Benchmark Results

Measured commit: 7d244de79713
Baseline commit: 43cf3539bd10

kvm / amd (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 772.31 ns (➖ 1.00x faster)
vec_bytes 585.00 ns (➖ 1.00x slower)
373.19 µs (➖ 1.00x faster)

payload_allocation

slot_pool_segmented
262144 531.27 ns (➖ 1.03x slower)
65536 141.69 ns (➖ 1.02x faster)

sandboxes

create_initialized_and_drop
medium 79.13 ms (➖ 1.01x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
7.76 ns (➖ 1.00x faster) 7.81 ns (➖ 1.01x slower) 7.73 ns (➖ 1.00x slower)

snapshot_files

load_snapshot_unverified
small 92.64 µs (➖ 1.00x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.26 µs (➖ 1.00x faster) 8.00 µs (➖ 1.10x slower)
65536 2.07 µs (➖ 1.01x slower) 1.93 µs (➖ 1.00x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 6.22 µs (➖ 1.02x faster) 6.22 µs (➖ 1.03x faster)
8192 1.07 µs (➖ 1.02x slower) 1.08 µs (➖ 1.02x slower)
262144 26.97 µs (➖ 1.01x faster)
kvm / intel (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 708.67 ns (➖ 1.06x faster)
vec_bytes 503.77 ns (➖ 1.08x faster)
653.10 µs (➖ 1.06x faster)

payload_allocation

slot_pool_segmented
262144 510.32 ns (➖ 1.12x faster)
65536 136.27 ns (➖ 1.13x faster)

sandboxes

create_initialized_and_drop
medium 73.73 ms (➖ 1.08x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
7.28 ns (➖ 1.04x faster) 6.92 ns (➖ 1.09x faster) 7.11 ns (➖ 1.08x faster)

snapshot_files

load_snapshot_unverified
small 44.45 µs (➖ 1.10x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.54 µs (➖ 1.13x faster) 7.59 µs (➖ 1.13x faster)
65536 2.12 µs (➖ 1.13x faster) 2.12 µs (➖ 1.12x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 7.13 µs (➖ 1.10x faster) 7.13 µs (➖ 1.12x faster)
8192 745.26 ns (➖ 1.12x faster) 727.91 ns (➖ 1.08x faster)
262144 29.59 µs (➖ 1.13x faster)
mshv3 / amd (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 953.48 ns (➖ 1.00x slower)
vec_bytes 704.25 ns (➖ 1.01x faster)
283.74 µs (➖ 1.26x faster)

payload_allocation

slot_pool_segmented
262144 739.96 ns (➖ 1.01x slower)
65536 193.49 ns (➖ 1.02x faster)

sandboxes

create_initialized_and_drop
medium 57.00 ms (➖ 1.02x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
10.59 ns (➖ 1.04x slower) 9.94 ns (➖ 1.03x slower) 9.69 ns (➖ 1.04x faster)

snapshot_files

load_snapshot_unverified
small 86.00 µs (➖ 1.02x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 10.16 µs (➖ 1.01x slower) 8.82 µs (➖ 1.01x faster)
65536 2.26 µs (➖ 1.00x slower) 2.48 µs (➖ 1.01x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 9.18 µs (➖ 1.11x slower) 8.11 µs (➖ 1.01x slower)
8192 1.31 µs (➖ 1.03x slower) 1.30 µs (➖ 1.03x faster)
262144 35.69 µs (➖ 1.02x faster)
mshv3 / intel (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 941.95 ns (➖ 1.01x slower)
vec_bytes 655.65 ns (➖ 1.01x faster)
765.30 µs (➖ 1.14x slower)

payload_allocation

slot_pool_segmented
262144 629.73 ns (➖ 1.02x faster)
65536 165.23 ns (➖ 1.00x slower)

sandboxes

create_initialized_and_drop
medium 61.52 ms (➖ 1.01x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
8.76 ns (➖ 1.00x faster) 8.75 ns (➖ 1.00x faster) 10.00 ns (➖ 1.05x slower)

snapshot_files

load_snapshot_unverified
small 43.85 µs (➖ 1.00x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.77 µs (➖ 1.00x slower) 7.74 µs (➖ 1.04x faster)
65536 2.22 µs (➖ 1.02x slower) 2.26 µs (➖ 1.01x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 7.50 µs (➖ 1.00x faster) 7.57 µs (➖ 1.00x slower)
8192 875.03 ns (➖ 1.02x slower) 877.32 ns (➖ 1.01x faster)
262144 35.25 µs (➖ 1.08x faster)
hyperv-ws2025 / amd (Windows) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.17 µs (➖ 1.04x slower)
vec_bytes 787.36 ns (➖ 1.21x faster)
2.26 ms (➖ 1.11x faster)

payload_allocation

slot_pool_segmented
262144 790.83 ns (➖ 1.02x faster)
65536 234.16 ns (➖ 1.02x slower)

sandboxes

create_initialized_and_drop
medium 88.23 ms (➖ 1.25x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
9.94 ns (➖ 1.39x faster) 10.50 ns (➖ 1.03x faster) 10.12 ns (➖ 1.04x faster)

snapshot_files

load_snapshot_unverified
small 802.40 µs (➖ 1.01x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 9.23 µs (➖ 1.06x slower) 9.16 µs (➖ 1.02x faster)
65536 2.41 µs (➖ 1.05x faster) 2.40 µs (➖ 1.13x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 9.22 µs (➖ 1.06x slower) 9.10 µs (➖ 1.05x slower)
8192 1.29 µs (➖ 1.11x faster) 1.31 µs (➖ 1.07x faster)
262144 46.09 µs (➖ 1.27x slower)
hyperv-ws2025 / intel (Windows) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.22 µs (➖ 1.01x slower)
vec_bytes 819.82 ns (➖ 1.05x slower)
3.13 ms (➖ 1.00x slower)

payload_allocation

slot_pool_segmented
262144 734.12 ns (➖ 1.00x faster)
65536 210.43 ns (➖ 1.01x faster)

sandboxes

create_initialized_and_drop
medium 108.38 ms (➖ 1.11x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
10.03 ns (➖ 1.00x slower) 10.12 ns (➖ 1.00x slower) 10.59 ns (➖ 1.00x faster)

snapshot_files

load_snapshot_unverified
small 614.64 µs (➖ 1.02x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.75 µs (➖ 1.02x slower) 7.69 µs (➖ 1.01x faster)
65536 2.33 µs (➖ 1.03x slower) 2.31 µs (➖ 1.01x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.50 µs (➖ 1.02x slower) 8.41 µs (➖ 1.00x slower)
8192 1.14 µs (➖ 1.11x faster) 1.23 µs (➖ 1.10x slower)
262144 50.18 µs (➖ 1.17x slower)

Reported by cargo ci bench-report --candidate run:37059659974 --baseline run:36945564806 --config-file bench_report.toml.

@syntactically
syntactically merged commit 7d244de into main Oct 3, 2026
60 checks passed
@syntactically
syntactically deleted the lm/extra-tlbi branch October 3, 2026 00:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kind/enhancement For PRs adding features, improving functionality, docs, tests, etc.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants