[Core] Fix Python 3.14 async-actor memory leak by re-anchoring stack protection to fiber stacks - #64772
Conversation
Signed-off-by: Mark Towers <mark@anyscale.com>
There was a problem hiding this comment.
Code Review
This pull request addresses a memory leak on Python 3.14+ where async actor tasks running on boost fiber stacks trigger CPython's stack overflow protection, causing GC-object deallocations to be permanently deferred. The fix introduces RayReanchorStackProtectionToCurrentFiberStack to re-anchor CPython's stack bounds to the current fiber's stack using PyUnstable_ThreadState_SetStackProtection. The feedback highlights two key improvements: documenting the critical safety constraint on used_upper_bound to prevent potential physical stack overflows, and adding a fallback definition for RTLD_DEFAULT to ensure portability and prevent compilation failures on certain Linux environments.
| char anchor; | ||
| uintptr_t top = (uintptr_t)&anchor + used_upper_bound; |
There was a problem hiding this comment.
There is a critical safety assumption in the stack limit calculation:
uintptr_t top = (uintptr_t)&anchor + used_upper_bound;
int rc = set_stack_protection(PyThreadState_Get(), (void *)(top - stack_size), stack_size);
For the calculated stack range [top - stack_size, top] to safely cover the physical fiber stack without allowing actual stack overflow, used_upper_bound must be a strict upper bound of the actual stack depth U (where U = H - &anchor, and H is the physical stack top).
If the actual stack depth U at the call site ever exceeds used_upper_bound (e.g., if the call stack is deeper than 96 KB when yielding), then:
L_est = &anchor + used_upper_bound - stack_size
L_phys = &anchor + U - stack_size
Since used_upper_bound < U, we get L_est < L_phys. This means the estimated stack limit passed to CPython is lower than the physical stack limit. CPython will believe it has more stack space than physically available, which can lead to a physical stack overflow (and a silent memory corruption or segfault) before CPython's RecursionError is triggered.
Please add a clear comment documenting this critical constraint so that future maintainers do not inadvertently deepen the call stack at these call sites without adjusting the bounds.
char anchor;
/* used_upper_bound must be a strict upper bound of the actual stack depth
* at the call site to prevent overestimating the remaining stack space,
* which could lead to physical stack overflow before RecursionError triggers. */
uintptr_t top = (uintptr_t)&anchor + used_upper_bound;
There was a problem hiding this comment.
Added a comment about this
| #if PY_VERSION_HEX >= 0x030E0000 && !defined(MS_WINDOWS) | ||
| #include <dlfcn.h> | ||
| #include "ray/core_worker/task_execution/fiber.h" |
There was a problem hiding this comment.
On some Linux environments (such as glibc-based systems), RTLD_DEFAULT is only defined in <dlfcn.h> if _GNU_SOURCE is defined. If _GNU_SOURCE is not defined globally or in the Cython compilation environment, this can lead to a compilation failure.
To ensure maximum portability and prevent compilation errors, we should provide a safe fallback definition for RTLD_DEFAULT when it is not defined. On Linux, RTLD_DEFAULT is defined as NULL (or (void *)0). On macOS, it is defined as (void *)-2, but macOS always defines it by default, so the fallback won't override it.
#if PY_VERSION_HEX >= 0x030E0000 && !defined(MS_WINDOWS)
#include <dlfcn.h>
#ifndef RTLD_DEFAULT
#define RTLD_DEFAULT ((void *)0)
#endif
#include "ray/core_worker/task_execution/fiber.h"
There was a problem hiding this comment.
Added this check
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
Reviewed by Cursor Bugbot for commit f77e4d0. Configure here.
Signed-off-by: Mark Towers <mark@anyscale.com>
Signed-off-by: Mark Towers <mark@anyscale.com>
pseudo-rnd-thoughts
left a comment
There was a problem hiding this comment.
@Kunchd I can't added a test to this yet. Does core have any general tests for memory leaks?
| char anchor; | ||
| uintptr_t top = (uintptr_t)&anchor + used_upper_bound; |
There was a problem hiding this comment.
Added a comment about this
| #if PY_VERSION_HEX >= 0x030E0000 && !defined(MS_WINDOWS) | ||
| #include <dlfcn.h> | ||
| #include "ray/core_worker/task_execution/fiber.h" |
There was a problem hiding this comment.
Added this check
…mage resolving CPython 3.14.0 (#64857) ## Description ### Release smoke test Adds a `hello_world_py314` nightly smoke release test (aws variation only), mirroring the existing `hello_world_py313` entry: - `release/ray_release/schema.json` — add `"3.14"` to the `python` enum (release-test config validation rejects `python: "3.14"` without this). - `release/ray_release/config.py` — add `"3.14"` to the cpu/cu123 BYOD python allowlist (the parallel gate the release runner walks). - `release/release_tests.yaml` — new `hello_world_py314` test, nightly, `byod: {}`, same `hello_world_compute_config.yaml` as the other hello_world tests. The py3.14 `ray-anyscale` release-test images (cpu + cuda) are already built and published on master via `.buildkite/release/build.rayci.yml`, so no image plumbing is needed here. ### Base-image fix (what made the smoke test fail) The first run of this test failed with the JobSupervisor dying at job startup: `Fatal Python error: _Py_CheckRecursiveCall: Unrecoverable stack overflow`. Root cause: the py3.14 images ship **CPython 3.14.0**, which fatally crashes any Ray async actor running on a boost fiber stack (python/cpython#141944, fixed upstream in **3.14.2**). Why the images resolve 3.14.0: `docker/base-deps/Dockerfile` installs an exact `libffi=3.4.6` pin as a *separate* conda step after installing python. That second solve downgrades python to the only 3.14 build compatible with `libffi<3.5` — which is 3.14.0. (The stale wanda layer cache compounds this, but even a fresh rebuild today re-resolves 3.14.0 because of the pin.) Fix: replace the two-step install with a **single solve using a libffi floor** (`libffi>=3.4.6`, preserving the intent of the original pin — the 3.4.2/defaults-channel libffi was buggy). No per-version special casing; every python version resolves its newest patch release with a compatible libffi. Resolved versions today: py3.10→3.10.20 (libffi 3.7.0), py3.11→3.11.15, py3.12→3.12.13, py3.13→3.13.14 (libffi 3.5.2), py3.14→**3.14.6** (libffi 3.5.2). The Dockerfile change also busts the stale wanda cache. Note: #64772 (fiber stack-protection re-anchoring) is complementary, not a fix for this crash — its `PyUnstable_ThreadState_SetStackProtection` call only exists on 3.14.2+, so it no-ops on the 3.14.0 currently in the images. Once this lands, #64772 fixes the remaining per-task async-actor memory leak. ## Verification - Reproduced the crash: async-actor repro (`ray.get(A.remote().hi.remote())` with an `async def` method) dies in `rayproject/ray:nightly-py314-cpu` (CPython 3.14.0) with the exact failure signature from release-test job `prodjob_d4dctduzm3h6eu812vrrehiuzl`. - Verified the fix: the same unpatched nightly cp314 wheel on CPython 3.14.6 (`python:3.14-slim`) runs the repro successfully. - Verified the combined solve under miniforge 24.11.3-0 (same as the Dockerfile): dry-runs for python 3.10–3.14 all resolve (versions above), plus real installs with `ctypes` smoke tests on 3.10 (libffi 3.7.0) and 3.14.6 (libffi 3.5.2). - `python -m pytest -q release/ray_release/tests/test_config.py` — 23 passed; full collection validates (319 tests) including `hello_world_py314.aws`. ## Duplicate-work note #63237 contains an earlier version of the release-test config bundled with image-build plumbing that has since landed on master through other PRs. This PR carves out the remaining release-test config plus the base-image fix; #63237 can be closed or rebased down to the raylet fix. AI assistance (Claude Code) was used for this PR; all changes reviewed by the submitter. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: elliot-barn <elliot.barnwell@anyscale.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kunchd
left a comment
There was a problem hiding this comment.
Thanks for the investigation!
@pseudo-rnd-thoughts, I'm not aware of any memory leak tests. But there are test_memory_pressure.py that does deal with memory. Perhaps we could do something similar with your repro and check for the before and after memory footprint?
| return 0; | ||
| } | ||
| #endif | ||
| """ |
There was a problem hiding this comment.
nit: Instead of embedding the logic within the cython. Could We pull this out to an actual file?
| if (<int>task_type == <int>TASK_TYPE_ACTOR_TASK | ||
| and CCoreWorkerProcess.GetCoreWorker().GetWorkerContext() | ||
| .CurrentActorIsAsync()): | ||
| RayReanchorStackProtectionToCurrentFiberStack(32 * 1024) |
There was a problem hiding this comment.
Why 32 * 1024 specifically?
Ideally we at least document how this number was determined, and how this number should be adjusted in the future if something were to go wrong.
There was a problem hiding this comment.
Instead of using these numbers, how about we keep track of each fiber stack address and query them when we do the reanchor? I made a draft at #65119
| # Re-anchor this fiber's stack before running the rest of the task on it. | ||
| # The bound is larger than at task entry because this call site is | ||
| # several C frames deeper. | ||
| RayReanchorStackProtectionToCurrentFiberStack(96 * 1024) |
There was a problem hiding this comment.
How did we arrive at this number?
|
I spent some time reproducing this independently on my devbox and can confirm both the leak and the fix. Sharing the some detailed numbers here:
0 of 300 At the same time, I'd like to help with the followup changes. I have a small follow-up prepared on top of your branch:
For the estimates themselves I'd suggest a separate PR rather than expanding this one and this is what rueian's PR is trying to achieve. |
…code comment update Signed-off-by: myan <myan@anyscale.com>
There was a problem hiding this comment.
I looked into the cython code a bit more, and I wanted to double check something. PyUnstable_ThreadState_SetStackProtection takes stack_start_addr and stack_size as arguments. Here, we passed in the current variable address as stack_start_addr and a magic number for the stack_size.
From there, PyUnstable_ThreadState_SetStackProtection invokes tstate_set_stack. This function sets soft and hard stack limits based on the following sketch (excuse my chicken scratchings):

So there's one particular thing to call out here. If the stack is growing down within the python execution of the sync function, the soft and hard limits doesn't seem to be doing anything since we're growing down from base, away from the limits.
Do we know if the stack is growing up or down here?
| # CurrentActorIsAsync() half of the entry-site condition is implied here: | ||
| # every caller of this method is already inside an is-asyncio branch. | ||
| if is_actor_task: | ||
| RayReanchorStackProtectionToCurrentFiberStack(96 * 1024) |
There was a problem hiding this comment.
The _PyOS_STACK_MARGIN_BYTES could reach 192KiB when built with Tsan (see https://github.com/python/cpython/blob/a646c99ee12fbf421394a74955d42fd3fe850b7e/Include/internal/pycore_pythonrun.h#L66). Should we increase this value?
Otherwise, we might fail this assertion: https://github.com/python/cpython/pull/141661/changes#diff-c22186367cbe20233e843261998dc027ae5f1f8c0d2e778abfa454ae74cc59deR542-R546
There was a problem hiding this comment.
Thanks for the comment! My understanding is that, the 96 * 1024 here is not the stack size, but the estimation of how much stack space we've already used. The stack size we pass in will always be 256 KiB. In this sense, it will should always be larger.
| return await coroutine | ||
| finally: | ||
| event.Notify() | ||
|
|
There was a problem hiding this comment.
Should we anchor the stack before invoking run_coroutine_threadsafe here?
There was a problem hiding this comment.
I don't think it is needed. Fibers are cooperatively scheduled on one thread, so the bounds can only be overwritten by another fiber, and a fiber can only run if the current on yields. The only fiber-suspension point from my understand is YieldCurrentFiber, which the code re-anchors immediately on resume. So when run_coroutine_threadsafe is called, there is no fiber switch so we don't need to re-anchor there.
I have a bit different understanding the above. My understanding is the following:
I might miss something from your drawing but I think the confusion here is that:
Let me know if anything doesn't makes sense or I missed anything. |
Kunchd
left a comment
There was a problem hiding this comment.
Thanks for the explanation!
Keep the tracked-fiber stack protection implementation and drop the heuristic 32 KiB / 96 KiB used-stack estimates merged via ray-project#64772. Retain master's asyncio finalizer regression test and missing-API warning. Signed-off-by: Rueian Huang <rueiancsie@gmail.com>
#65177) …protection to fiber stacks (#64772) # Description On Python 3.14 + Linux, every async-actor task permanently leaks ~518 KiB of live malloc (the per-task `asyncio.Task`, `concurrent.futures.Future`, Cython coroutine + scopes, and two msgpack `Packer`s with 256 KiB internal buffers). Closes #63290 ### Root cause **1. CPython 3.14 changed how it avoids stack overflow when freeing objects.** Freeing one object can recursively free many others (a dict frees its values, which free their contents, …), and each level is a nested C call. To keep that from overflowing the C stack, CPython has long had a safety mechanism (the "trashcan"): when it decides it's too deep, it doesn't free the object right away. Instead it parks the object on a per-thread *delete-later* list and drains the list once there's stack headroom again. Up to 3.13, "too deep" was a simple recursion counter. In 3.14 it's decided by comparing the actual machine **stack pointer** against the stack bounds CPython recorded for the thread when it attached (from pthreads, on Linux). **2. Ray async actors don't run task code on the thread's normal stack.** Each task executes on a small 256 KiB boost fiber stack allocated elsewhere in memory. The problem is that CPython still thinks the thread runs on its original pthread stack. So while a task runs on a fiber, every "am I near the stack limit?" check compares the fiber's stack pointer against the *pthread* stack's bounds. On Linux, fiber stacks happen to be allocated at lower addresses than the pthread stack, so CPython concludes the stack is hopelessly overflowed and parks **every** object freed during the task (including return-value serialization and end-of-task cleanup) on the delete-later list. That list is only ever drained by a later free on the same thread state at a healthy stack margin, which never happens here as the Ray thread only runs on fibers and Ray creates a fresh Python thread state per task and destroys it at task end. This means that CPython destroys a thread state **without draining its delete-later list** and the parked objects are orphaned permanently. That's the leak. Why the confusing symptoms: - `boost::make_fcontext` in the issue's flamegraphs just marks *where* the leaked allocations were made (on a fiber stack); the fiber stacks themselves are freed correctly. - macOS is unaffected only by luck: fiber stacks there land at *higher* addresses than the pthread stack, so the check passes. - 3.13 and earlier are unaffected because their trashcan uses the counter, not the stack pointer. ### Fix CPython 3.14.2 added an official API for exactly this situation: `PyUnstable_ThreadState_SetStackProtection` (python/cpython#141661) lets an embedder tell CPython "this thread is currently executing on *this* stack." We call it with the fiber's stack bounds: - at async-actor task entry in `task_execution_handler`, and - whenever a fiber resumes after `YieldCurrentFiber` (concurrent fibers share the thread state, so each must re-register its own stack). With the bounds correct, the near-limit check returns to normal behavior: objects are freed immediately, and the rare genuinely-deep free is parked and then properly drained. Implementation notes: the symbol is looked up via `dlsym`, so `_raylet` still imports on 3.14.0/3.14.1 (fix skipped there; those releases have a more severe, since-fixed stack-check bug anyway, python/cpython#141944). No-op below 3.14 (preprocessor-gated) and on Windows. Stack bounds are derived from the current stack pointer minus a conservative allowance for stack already used, so the protection errs toward triggering slightly early rather than missing an overflow. Side benefit: fibers gain real C-stack overflow protection (RecursionError) on 3.14, which they currently lack entirely (`boost::fibers::fixedsize_stack` has no guard pages). Also makes `FiberState::kStackSize` public so the anchoring uses the real fiber stack size. ## Related issue number Closes #63290. Supersedes #63284 (same diagnosis direction, but hand-rolled `_PyThreadStateImpl` offsets, a deliberate `gilstate_counter` leak that freezes non-main threads, and a crash premise that CPython 3.14.2 already fixed upstream). ## Checks - Verified with a locally built cp314 Linux (aarch64, python:3.14.6 docker) wheel: - refcount probe: **+4.00 refs/task → 0.00/task** (100 tasks) - `__del__` deferral probe: dealloc during return serialization on the fiber **deferred → immediate** - live-malloc probe (`mallinfo2`, 300 tasks/shape): **~518 KiB/task → ~3 KiB/task** across async call → dict/bytes, async generator, sync generator on async actor - reporter-shaped streaming workload (400 tasks, 10 concurrent sessions): live-malloc delta **0.2 MB total**, fiber-sized mapped regions 0 → 0 - async-actor smoke: correctness (echo, state, async generators, recursion), concurrency (20 overlapping 0.5 s sleeps in 0.51 s) - throughput A/B (500 sequential echo tasks, 3 runs fixed / 2 runs baseline, same container image): fixed 5624–6076 tasks/s vs unpatched 4551–4825 tasks/s meaning no regression (the unpatched build is slower while leaking) - baseline (unpatched) wheel from the same tree reproduces the bug: +4.00 refs/task, fiber dealloc deferred=True note: fable did a majority of the heavy lifting in this investigation with prompting on what to check next and validate the solution --------- (cherry picked from commit 35591ba) Signed-off-by: Mark Towers <mark@anyscale.com> Signed-off-by: myan <myan@anyscale.com> Signed-off-by: elliot-barn <elliot.barnwell@anyscale.com> Co-authored-by: Mark Towers <mark.m.towers@gmail.com> Co-authored-by: Mark Towers <mark@anyscale.com> Co-authored-by: myan <myan@anyscale.com> Co-authored-by: Mengjin Yan <mengjinyan3@gmail.com>
…protection to fiber stacks (ray-project#64772) # Description On Python 3.14 + Linux, every async-actor task permanently leaks ~518 KiB of live malloc (the per-task `asyncio.Task`, `concurrent.futures.Future`, Cython coroutine + scopes, and two msgpack `Packer`s with 256 KiB internal buffers). Closes ray-project#63290 ### Root cause **1. CPython 3.14 changed how it avoids stack overflow when freeing objects.** Freeing one object can recursively free many others (a dict frees its values, which free their contents, …), and each level is a nested C call. To keep that from overflowing the C stack, CPython has long had a safety mechanism (the "trashcan"): when it decides it's too deep, it doesn't free the object right away. Instead it parks the object on a per-thread *delete-later* list and drains the list once there's stack headroom again. Up to 3.13, "too deep" was a simple recursion counter. In 3.14 it's decided by comparing the actual machine **stack pointer** against the stack bounds CPython recorded for the thread when it attached (from pthreads, on Linux). **2. Ray async actors don't run task code on the thread's normal stack.** Each task executes on a small 256 KiB boost fiber stack allocated elsewhere in memory. The problem is that CPython still thinks the thread runs on its original pthread stack. So while a task runs on a fiber, every "am I near the stack limit?" check compares the fiber's stack pointer against the *pthread* stack's bounds. On Linux, fiber stacks happen to be allocated at lower addresses than the pthread stack, so CPython concludes the stack is hopelessly overflowed and parks **every** object freed during the task (including return-value serialization and end-of-task cleanup) on the delete-later list. That list is only ever drained by a later free on the same thread state at a healthy stack margin, which never happens here as the Ray thread only runs on fibers and Ray creates a fresh Python thread state per task and destroys it at task end. This means that CPython destroys a thread state **without draining its delete-later list** and the parked objects are orphaned permanently. That's the leak. Why the confusing symptoms: - `boost::make_fcontext` in the issue's flamegraphs just marks *where* the leaked allocations were made (on a fiber stack); the fiber stacks themselves are freed correctly. - macOS is unaffected only by luck: fiber stacks there land at *higher* addresses than the pthread stack, so the check passes. - 3.13 and earlier are unaffected because their trashcan uses the counter, not the stack pointer. ### Fix CPython 3.14.2 added an official API for exactly this situation: `PyUnstable_ThreadState_SetStackProtection` (python/cpython#141661) lets an embedder tell CPython "this thread is currently executing on *this* stack." We call it with the fiber's stack bounds: - at async-actor task entry in `task_execution_handler`, and - whenever a fiber resumes after `YieldCurrentFiber` (concurrent fibers share the thread state, so each must re-register its own stack). With the bounds correct, the near-limit check returns to normal behavior: objects are freed immediately, and the rare genuinely-deep free is parked and then properly drained. Implementation notes: the symbol is looked up via `dlsym`, so `_raylet` still imports on 3.14.0/3.14.1 (fix skipped there; those releases have a more severe, since-fixed stack-check bug anyway, python/cpython#141944). No-op below 3.14 (preprocessor-gated) and on Windows. Stack bounds are derived from the current stack pointer minus a conservative allowance for stack already used, so the protection errs toward triggering slightly early rather than missing an overflow. Side benefit: fibers gain real C-stack overflow protection (RecursionError) on 3.14, which they currently lack entirely (`boost::fibers::fixedsize_stack` has no guard pages). Also makes `FiberState::kStackSize` public so the anchoring uses the real fiber stack size. ## Related issue number Closes ray-project#63290. Supersedes ray-project#63284 (same diagnosis direction, but hand-rolled `_PyThreadStateImpl` offsets, a deliberate `gilstate_counter` leak that freezes non-main threads, and a crash premise that CPython 3.14.2 already fixed upstream). ## Checks - Verified with a locally built cp314 Linux (aarch64, python:3.14.6 docker) wheel: - refcount probe: **+4.00 refs/task → 0.00/task** (100 tasks) - `__del__` deferral probe: dealloc during return serialization on the fiber **deferred → immediate** - live-malloc probe (`mallinfo2`, 300 tasks/shape): **~518 KiB/task → ~3 KiB/task** across async call → dict/bytes, async generator, sync generator on async actor - reporter-shaped streaming workload (400 tasks, 10 concurrent sessions): live-malloc delta **0.2 MB total**, fiber-sized mapped regions 0 → 0 - async-actor smoke: correctness (echo, state, async generators, recursion), concurrency (20 overlapping 0.5 s sleeps in 0.51 s) - throughput A/B (500 sequential echo tasks, 3 runs fixed / 2 runs baseline, same container image): fixed 5624–6076 tasks/s vs unpatched 4551–4825 tasks/s meaning no regression (the unpatched build is slower while leaking) - baseline (unpatched) wheel from the same tree reproduces the bug: +4.00 refs/task, fiber dealloc deferred=True note: fable did a majority of the heavy lifting in this investigation with prompting on what to check next and validate the solution --------- Signed-off-by: Mark Towers <mark@anyscale.com> Signed-off-by: myan <myan@anyscale.com> Co-authored-by: Mark Towers <mark@anyscale.com> Co-authored-by: myan <myan@anyscale.com> Co-authored-by: Mengjin Yan <mengjinyan3@gmail.com>
…mage resolving CPython 3.14.0 (ray-project#64857) ## Description ### Release smoke test Adds a `hello_world_py314` nightly smoke release test (aws variation only), mirroring the existing `hello_world_py313` entry: - `release/ray_release/schema.json` — add `"3.14"` to the `python` enum (release-test config validation rejects `python: "3.14"` without this). - `release/ray_release/config.py` — add `"3.14"` to the cpu/cu123 BYOD python allowlist (the parallel gate the release runner walks). - `release/release_tests.yaml` — new `hello_world_py314` test, nightly, `byod: {}`, same `hello_world_compute_config.yaml` as the other hello_world tests. The py3.14 `ray-anyscale` release-test images (cpu + cuda) are already built and published on master via `.buildkite/release/build.rayci.yml`, so no image plumbing is needed here. ### Base-image fix (what made the smoke test fail) The first run of this test failed with the JobSupervisor dying at job startup: `Fatal Python error: _Py_CheckRecursiveCall: Unrecoverable stack overflow`. Root cause: the py3.14 images ship **CPython 3.14.0**, which fatally crashes any Ray async actor running on a boost fiber stack (python/cpython#141944, fixed upstream in **3.14.2**). Why the images resolve 3.14.0: `docker/base-deps/Dockerfile` installs an exact `libffi=3.4.6` pin as a *separate* conda step after installing python. That second solve downgrades python to the only 3.14 build compatible with `libffi<3.5` — which is 3.14.0. (The stale wanda layer cache compounds this, but even a fresh rebuild today re-resolves 3.14.0 because of the pin.) Fix: replace the two-step install with a **single solve using a libffi floor** (`libffi>=3.4.6`, preserving the intent of the original pin — the 3.4.2/defaults-channel libffi was buggy). No per-version special casing; every python version resolves its newest patch release with a compatible libffi. Resolved versions today: py3.10→3.10.20 (libffi 3.7.0), py3.11→3.11.15, py3.12→3.12.13, py3.13→3.13.14 (libffi 3.5.2), py3.14→**3.14.6** (libffi 3.5.2). The Dockerfile change also busts the stale wanda cache. Note: ray-project#64772 (fiber stack-protection re-anchoring) is complementary, not a fix for this crash — its `PyUnstable_ThreadState_SetStackProtection` call only exists on 3.14.2+, so it no-ops on the 3.14.0 currently in the images. Once this lands, ray-project#64772 fixes the remaining per-task async-actor memory leak. ## Verification - Reproduced the crash: async-actor repro (`ray.get(A.remote().hi.remote())` with an `async def` method) dies in `rayproject/ray:nightly-py314-cpu` (CPython 3.14.0) with the exact failure signature from release-test job `prodjob_d4dctduzm3h6eu812vrrehiuzl`. - Verified the fix: the same unpatched nightly cp314 wheel on CPython 3.14.6 (`python:3.14-slim`) runs the repro successfully. - Verified the combined solve under miniforge 24.11.3-0 (same as the Dockerfile): dry-runs for python 3.10–3.14 all resolve (versions above), plus real installs with `ctypes` smoke tests on 3.10 (libffi 3.7.0) and 3.14.6 (libffi 3.5.2). - `python -m pytest -q release/ray_release/tests/test_config.py` — 23 passed; full collection validates (319 tests) including `hello_world_py314.aws`. ## Duplicate-work note ray-project#63237 contains an earlier version of the release-test config bundled with image-build plumbing that has since landed on master through other PRs. This PR carves out the remaining release-test config plus the base-image fix; ray-project#63237 can be closed or rebased down to the raylet fix. AI assistance (Claude Code) was used for this PR; all changes reviewed by the submitter. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: elliot-barn <elliot.barnwell@anyscale.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: 400Ping <jiekaichang@apache.org>
…protection to fiber stacks (ray-project#64772) # Description On Python 3.14 + Linux, every async-actor task permanently leaks ~518 KiB of live malloc (the per-task `asyncio.Task`, `concurrent.futures.Future`, Cython coroutine + scopes, and two msgpack `Packer`s with 256 KiB internal buffers). Closes ray-project#63290 ### Root cause **1. CPython 3.14 changed how it avoids stack overflow when freeing objects.** Freeing one object can recursively free many others (a dict frees its values, which free their contents, …), and each level is a nested C call. To keep that from overflowing the C stack, CPython has long had a safety mechanism (the "trashcan"): when it decides it's too deep, it doesn't free the object right away. Instead it parks the object on a per-thread *delete-later* list and drains the list once there's stack headroom again. Up to 3.13, "too deep" was a simple recursion counter. In 3.14 it's decided by comparing the actual machine **stack pointer** against the stack bounds CPython recorded for the thread when it attached (from pthreads, on Linux). **2. Ray async actors don't run task code on the thread's normal stack.** Each task executes on a small 256 KiB boost fiber stack allocated elsewhere in memory. The problem is that CPython still thinks the thread runs on its original pthread stack. So while a task runs on a fiber, every "am I near the stack limit?" check compares the fiber's stack pointer against the *pthread* stack's bounds. On Linux, fiber stacks happen to be allocated at lower addresses than the pthread stack, so CPython concludes the stack is hopelessly overflowed and parks **every** object freed during the task (including return-value serialization and end-of-task cleanup) on the delete-later list. That list is only ever drained by a later free on the same thread state at a healthy stack margin, which never happens here as the Ray thread only runs on fibers and Ray creates a fresh Python thread state per task and destroys it at task end. This means that CPython destroys a thread state **without draining its delete-later list** and the parked objects are orphaned permanently. That's the leak. Why the confusing symptoms: - `boost::make_fcontext` in the issue's flamegraphs just marks *where* the leaked allocations were made (on a fiber stack); the fiber stacks themselves are freed correctly. - macOS is unaffected only by luck: fiber stacks there land at *higher* addresses than the pthread stack, so the check passes. - 3.13 and earlier are unaffected because their trashcan uses the counter, not the stack pointer. ### Fix CPython 3.14.2 added an official API for exactly this situation: `PyUnstable_ThreadState_SetStackProtection` (python/cpython#141661) lets an embedder tell CPython "this thread is currently executing on *this* stack." We call it with the fiber's stack bounds: - at async-actor task entry in `task_execution_handler`, and - whenever a fiber resumes after `YieldCurrentFiber` (concurrent fibers share the thread state, so each must re-register its own stack). With the bounds correct, the near-limit check returns to normal behavior: objects are freed immediately, and the rare genuinely-deep free is parked and then properly drained. Implementation notes: the symbol is looked up via `dlsym`, so `_raylet` still imports on 3.14.0/3.14.1 (fix skipped there; those releases have a more severe, since-fixed stack-check bug anyway, python/cpython#141944). No-op below 3.14 (preprocessor-gated) and on Windows. Stack bounds are derived from the current stack pointer minus a conservative allowance for stack already used, so the protection errs toward triggering slightly early rather than missing an overflow. Side benefit: fibers gain real C-stack overflow protection (RecursionError) on 3.14, which they currently lack entirely (`boost::fibers::fixedsize_stack` has no guard pages). Also makes `FiberState::kStackSize` public so the anchoring uses the real fiber stack size. ## Related issue number Closes ray-project#63290. Supersedes ray-project#63284 (same diagnosis direction, but hand-rolled `_PyThreadStateImpl` offsets, a deliberate `gilstate_counter` leak that freezes non-main threads, and a crash premise that CPython 3.14.2 already fixed upstream). ## Checks - Verified with a locally built cp314 Linux (aarch64, python:3.14.6 docker) wheel: - refcount probe: **+4.00 refs/task → 0.00/task** (100 tasks) - `__del__` deferral probe: dealloc during return serialization on the fiber **deferred → immediate** - live-malloc probe (`mallinfo2`, 300 tasks/shape): **~518 KiB/task → ~3 KiB/task** across async call → dict/bytes, async generator, sync generator on async actor - reporter-shaped streaming workload (400 tasks, 10 concurrent sessions): live-malloc delta **0.2 MB total**, fiber-sized mapped regions 0 → 0 - async-actor smoke: correctness (echo, state, async generators, recursion), concurrency (20 overlapping 0.5 s sleeps in 0.51 s) - throughput A/B (500 sequential echo tasks, 3 runs fixed / 2 runs baseline, same container image): fixed 5624–6076 tasks/s vs unpatched 4551–4825 tasks/s meaning no regression (the unpatched build is slower while leaking) - baseline (unpatched) wheel from the same tree reproduces the bug: +4.00 refs/task, fiber dealloc deferred=True note: fable did a majority of the heavy lifting in this investigation with prompting on what to check next and validate the solution --------- Signed-off-by: Mark Towers <mark@anyscale.com> Signed-off-by: myan <myan@anyscale.com> Co-authored-by: Mark Towers <mark@anyscale.com> Co-authored-by: myan <myan@anyscale.com> Co-authored-by: Mengjin Yan <mengjinyan3@gmail.com> Signed-off-by: 400Ping <jiekaichang@apache.org>

Description
On Python 3.14 + Linux, every async-actor task permanently leaks ~518 KiB of live malloc (the per-task
asyncio.Task,concurrent.futures.Future, Cython coroutine + scopes, and two msgpackPackers with 256 KiB internal buffers).Closes #63290
Root cause
1. CPython 3.14 changed how it avoids stack overflow when freeing objects. Freeing one object can recursively free many others (a dict frees its values, which free their contents, …), and each level is a nested C call. To keep that from overflowing the C stack, CPython has long had a safety mechanism (the "trashcan"): when it decides it's too deep, it doesn't free the object right away. Instead it parks the object on a per-thread delete-later list and drains the list once there's stack headroom again. Up to 3.13, "too deep" was a simple recursion counter. In 3.14 it's decided by comparing the actual machine stack pointer against the stack bounds CPython recorded for the thread when it attached (from pthreads, on Linux).
2. Ray async actors don't run task code on the thread's normal stack. Each task executes on a small 256 KiB boost fiber stack allocated elsewhere in memory. The problem is that CPython still thinks the thread runs on its original pthread stack.
So while a task runs on a fiber, every "am I near the stack limit?" check compares the fiber's stack pointer against the pthread stack's bounds. On Linux, fiber stacks happen to be allocated at lower addresses than the pthread stack, so CPython concludes the stack is hopelessly overflowed and parks every object freed during the task (including return-value serialization and end-of-task cleanup) on the delete-later list.
That list is only ever drained by a later free on the same thread state at a healthy stack margin, which never happens here as the Ray thread only runs on fibers and Ray creates a fresh Python thread state per task and destroys it at task end. This means that CPython destroys a thread state without draining its delete-later list and the parked objects are orphaned permanently. That's the leak.
Why the confusing symptoms:
boost::make_fcontextin the issue's flamegraphs just marks where the leaked allocations were made (on a fiber stack); the fiber stacks themselves are freed correctly.Fix
CPython 3.14.2 added an official API for exactly this situation:
PyUnstable_ThreadState_SetStackProtection(python/cpython#141661) lets an embedder tell CPython "this thread is currently executing on this stack." We call it with the fiber's stack bounds:task_execution_handler, andYieldCurrentFiber(concurrent fibers share the thread state, so each must re-register its own stack).With the bounds correct, the near-limit check returns to normal behavior: objects are freed immediately, and the rare genuinely-deep free is parked and then properly drained.
Implementation notes: the symbol is looked up via
dlsym, so_rayletstill imports on 3.14.0/3.14.1 (fix skipped there; those releases have a more severe, since-fixed stack-check bug anyway, python/cpython#141944). No-op below 3.14 (preprocessor-gated) and on Windows. Stack bounds are derived from the current stack pointer minus a conservative allowance for stack already used, so the protection errs toward triggering slightly early rather than missing an overflow. Side benefit: fibers gain real C-stack overflow protection (RecursionError) on 3.14, which they currently lack entirely (boost::fibers::fixedsize_stackhas no guard pages). Also makesFiberState::kStackSizepublic so the anchoring uses the real fiber stack size.Related issue number
Closes #63290. Supersedes #63284 (same diagnosis direction, but hand-rolled
_PyThreadStateImploffsets, a deliberategilstate_counterleak that freezes non-main threads, and a crash premise that CPython 3.14.2 already fixed upstream).Checks
__del__deferral probe: dealloc during return serialization on the fiber deferred → immediatemallinfo2, 300 tasks/shape): ~518 KiB/task → ~3 KiB/task across async call → dict/bytes, async generator, sync generator on async actornote: fable did a majority of the heavy lifting in this investigation with prompting on what to check next and validate the solution