-mm=gc shared library: enable threads on the first call after the load; coroutine frames uncollectable - #524
Merged
Merged
Conversation
… after its load Fixes #522, #523. #521 stopped __tslang_gc_enter from enabling threads for a thread the collector already knew, because on the thread loading the library that ran inside DllMain and deadlocked. That left two holes: - #522: the library's own async pool registers its threads through hooks that only the library's GC_enable_threads sets (each module has its own copy of the scheduler). When only registered threads called in, they were never set, and a pool thread that collected aborted ("Collecting from unknown thread"). - #523: for a host that is not a tslang program, the first unregistered thread enabled threads. GC_allow_register_threads makes the allocator take its lock only after it started the markers, so it raced a registered thread that was already allocating. Now the first call that may enables threads, whichever thread makes it, before it goes on. That is any call except one inside a DllMain (the thread holds the PEB's LoaderLock) while the markers have not started. Such a call goes on as it is, and the loading thread's next call, after the load, enables them. No other thread can call in before the load finishes. test-compile-foreign-thread-gc: - The host checks that the markers run after the loading thread's calls whenever they run at the end. A collector with no parallel marking never starts them, and then there is nothing to check. - An async-pool library works on its own pool and collects there. Both failed before this change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the markers counted GC_get_parallel counts the markers before they are waited for, so a call from inside another DLL's DllMain could pass the check and wait on the thread starting them, which waits for this thread's loader lock. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…routine finishes The coroutine lowering allocates a frame with aligned_alloc and frees it with free; GCPass renamed the pair to GC_memalign/GC_free, so the frame was collectable. A frame waiting to run is held only by the runtime's pool queue or a token's awaiter list, ordinary heap the collector does not scan, so a collection freed it and the pool resumed whatever reused the block. On Linux, eight host threads awaiting at once in a gc shared library crashed every time (another call's result came back now and then); GC_DONT_GC made it go away. It needs collections while frames wait, which several awaiting threads make likely. GCPass now rewrites aligned_alloc(align, size) to GC_malloc_uncollectable(size): it lives until GC_free and is still scanned, so what a frame holds across a suspension stays alive too. The alignment is dropped: frames ask for 8, and the collector aligns every object to two words. The JIT runtime exports GC_malloc_uncollectable beside GC_memalign. check-x86-run.sh expects the i686 frame allocator as GC_malloc_uncollectable(i32) under -mm=gc. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #522. Fixes #523.
Follow-up to #521, plus a
-mm=gcasync bug that the new test found on Linux.#522 and #523: threads are enabled on the first call after the load
#521 stopped
__tslang_gc_enterfrom enabling the collector's threads for an already-registered thread, because on the loading thread that ran insideDllMainand deadlocked. That left two holes:Now the first call that may enables threads, whichever thread makes it, before it goes on. A call may do this unless the thread holds the Windows loader lock: it is inside a
DllMain, detected throughPEB.LoaderLock. The offset (0x110 on x64) was checked againstdt ntdll!_PEBin a dump of the original hang; x86 uses 0xA0. A call made during the load carries on as before, and the loading thread's next call, after the load, enables threads. No other thread can call in before the load finishes. Off Windows there is no such lock.Coroutine frames are uncollectable
The coroutine lowering allocates a frame with
aligned_allocand frees it withfree, which the GC pass renamed toGC_memalign/GC_free, so frames were collectable. A frame waiting to run is held only by the runtime's pool queue or a token's awaiter list, ordinary heap the collector doesn't scan. A collection freed it, and the pool resumed whatever had reused the block.On Linux, 8 host threads awaiting at once in a gc library crashed 10/10, and now and then a call got another thread's result.
GC_DONT_GC=1made it disappear. #521's runtime crashes the same way with a timing tweak, so this predates this PR.The GC pass now rewrites
aligned_alloc(align, size)toGC_malloc_uncollectable(size). Such a block lives untilGC_freeand is still scanned, so whatever a frame holds across a suspension stays alive too. The alignment argument is dropped: frames ask for 8 bytes, and the collector aligns every object to two words. The JIT runtime exportsGC_malloc_uncollectable, andcheck-x86-run.shnow expectsGC_malloc_uncollectable(i32)for i686.Tests
test-compile-foreign-thread-gc:GC_MARKERS=1).async-poollibrary does its work on its own pool and collects there.Before the fixes, the #523 check failed on Windows and Linux.
async-poolaborted on Windows (#522) and crashed on Linux (frames); without its collection it ran fine. A regression in a load-time library costs 240 s before the test fails.Verified locally:
tests.ps1release compile 162/162 and release JIT 162/162, plus test-jit-/test-compile-gc-defaultlib-collector. The original JIT hang repro prints[bcd]/done.GC_get_parallelresolves there) andasync-poolwith 8 host threads. The libraries that call an export during the load don't hang insidedlopen. CI runs older glibc (22.04/24.04).check-x86-run.shpasses for-mm=gc(declare ptr @GC_malloc_uncollectable(i32)). Its-mm=nonehalf was already stale on main: since Windows: modules other than gc share the process heap; attributes named once; Debug DEBUG_TYPE fix #412 that allocator is__tslang_heap_aligned_alloc. The rest of the script (it needs the x86 runtime build) wasn't run.--di) suites, a Linux default library, Android (no parallel marking there).Not covered
DllMainstays unregistered until a later call.asynccode in a library that starts pool work during the load runs on unhooked pool threads until the first call after the load.🤖 Generated with Claude Code