Skip to content

-mm=gc shared library: enable threads on the first call after the load; coroutine frames uncollectable - #524

Merged
ASDAlexander77 merged 3 commits into
mainfrom
gc-enable-threads-after-load
Oct 6, 2026
Merged

ASDAlexander77 merged 3 commits into
mainfrom
gc-enable-threads-after-load

Conversation

@ASDAlexander77

Copy link
Copy Markdown
Owner

Fixes #522. Fixes #523.

Follow-up to #521, plus a -mm=gc async bug that the new test found on Linux.

#522 and #523: threads are enabled on the first call after the load

#521 stopped __tslang_gc_enter from enabling the collector's threads for an already-registered thread, because on the loading thread that ran inside DllMain and deadlocked. That left two holes:

Now the first call that may enables threads, whichever thread makes it, before it goes on. A call may do this unless the thread holds the Windows loader lock: it is inside a DllMain, detected through PEB.LoaderLock. The offset (0x110 on x64) was checked against dt ntdll!_PEB in a dump of the original hang; x86 uses 0xA0. A call made during the load carries on as before, and the loading thread's next call, after the load, enables threads. No other thread can call in before the load finishes. Off Windows there is no such lock.

Coroutine frames are uncollectable

The coroutine lowering allocates a frame with aligned_alloc and frees it with free, which the GC pass renamed to GC_memalign/GC_free, so frames were collectable. A frame waiting to run is held only by the runtime's pool queue or a token's awaiter list, ordinary heap the collector doesn't scan. A collection freed it, and the pool resumed whatever had reused the block.

On Linux, 8 host threads awaiting at once in a gc library crashed 10/10, and now and then a call got another thread's result. GC_DONT_GC=1 made it disappear. #521's runtime crashes the same way with a timing tweak, so this predates this PR.

The GC pass now rewrites aligned_alloc(align, size) to GC_malloc_uncollectable(size). Such a block lives until GC_free and is still scanned, so whatever a frame holds across a suspension stays alive too. The alignment argument is dropped: frames ask for 8 bytes, and the collector aligns every object to two words. The JIT runtime exports GC_malloc_uncollectable, and check-x86-run.sh now expects GC_malloc_uncollectable(i32) for i686.

Tests

test-compile-foreign-thread-gc:

Before the fixes, the #523 check failed on Windows and Linux. async-pool aborted on Windows (#522) and crashed on Linux (frames); without its collection it ran fine. A regression in a load-time library costs 240 s before the test fails.

Verified locally:

  • Windows (Release): full ctest 3895/3895. A default library built from DefaultLib main against this compiler passes tests.ps1 release compile 162/162 and release JIT 162/162, plus test-jit-/test-compile-gc-defaultlib-collector. The original JIT hang repro prints [bcd]/done.
  • Linux (WSL, glibc 2.43): full ctest 3883/3883, including the -mm=gc shared library: a foreign thread's GC_allow_register_threads races a registered thread's allocation #523 check (GC_get_parallel resolves there) and async-pool with 8 host threads. The libraries that call an export during the load don't hang inside dlopen. CI runs older glibc (22.04/24.04).
  • x86: the IR check in check-x86-run.sh passes for -mm=gc (declare ptr @GC_malloc_uncollectable(i32)). Its -mm=none half was already stale on main: since Windows: modules other than gc share the process heap; attributes named once; Debug DEBUG_TYPE fix #412 that allocator is __tslang_heap_aligned_alloc. The rest of the script (it needs the x86 runtime build) wasn't run.
  • Not run: the debug (--di) suites, a Linux default library, Android (no parallel marking there).
  • Also checked: an async function returning an object under the same concurrency, collecting on 1 call in 10: no failures, so how results are stored is left alone.

Not covered

  • The x86 loader-lock offset compiles but wasn't exercised; nothing here runs a 32-bit gc DLL.
  • An unregistered host thread that calls in from inside its own DllMain stays unregistered until a later call.
  • Top-level async code in a library that starts pool work during the load runs on unhooked pool threads until the first call after the load.

🤖 Generated with Claude Code

ASDAlexander77 and others added 3 commits October 6, 2026 00:26
… after its load

Fixes #522, #523.

#521 stopped __tslang_gc_enter from enabling threads for a thread the
collector already knew, because on the thread loading the library that ran
inside DllMain and deadlocked. That left two holes:

- #522: the library's own async pool registers its threads through hooks
  that only the library's GC_enable_threads sets (each module has its own
  copy of the scheduler). When only registered threads called in, they were
  never set, and a pool thread that collected aborted ("Collecting from
  unknown thread").
- #523: for a host that is not a tslang program, the first unregistered
  thread enabled threads. GC_allow_register_threads makes the allocator take
  its lock only after it started the markers, so it raced a registered
  thread that was already allocating.

Now the first call that may enables threads, whichever thread makes it,
before it goes on. That is any call except one inside a DllMain (the
thread holds the PEB's LoaderLock) while the markers have not started. Such
a call goes on as it is, and the loading thread's next call, after the load,
enables them. No other thread can call in before the load finishes.

test-compile-foreign-thread-gc:
- The host checks that the markers run after the loading thread's calls
  whenever they run at the end. A collector with no parallel marking never
  starts them, and then there is nothing to check.
- An async-pool library works on its own pool and collects there.
Both failed before this change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the markers counted

GC_get_parallel counts the markers before they are waited for, so a call from inside another
DLL's DllMain could pass the check and wait on the thread starting them, which waits for this
thread's loader lock.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…routine finishes

The coroutine lowering allocates a frame with aligned_alloc and frees it with
free; GCPass renamed the pair to GC_memalign/GC_free, so the frame was
collectable. A frame waiting to run is held only by the runtime's pool queue
or a token's awaiter list, ordinary heap the collector does not scan, so a
collection freed it and the pool resumed whatever reused the block. On Linux,
eight host threads awaiting at once in a gc shared library crashed every time
(another call's result came back now and then); GC_DONT_GC made it go away.
It needs collections while frames wait, which several awaiting threads make
likely.

GCPass now rewrites aligned_alloc(align, size) to
GC_malloc_uncollectable(size): it lives until GC_free and is still scanned,
so what a frame holds across a suspension stays alive too. The alignment is
dropped: frames ask for 8, and the collector aligns every object to two
words. The JIT runtime exports GC_malloc_uncollectable beside GC_memalign.

check-x86-run.sh expects the i686 frame allocator as
GC_malloc_uncollectable(i32) under -mm=gc.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ASDAlexander77
ASDAlexander77 merged commit 5bf9fa7 into main Oct 6, 2026
2 checks passed
@ASDAlexander77
ASDAlexander77 deleted the gc-enable-threads-after-load branch October 6, 2026 07:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant