Re: [RFC PATCH 1/1] sched/proxy_exec: detect cycles in proxy walks

From: K Prateek Nayak

Date: Tue Sep 22 2026 - 01:48:27 EST


On 9/22/2026 10:56 AM, John Stultz wrote:
> On Mon, Sep 14, 2026 at 9:55 AM Hui Su <sh_def@xxxxxxx> wrote:
>>
>> Proxy execution follows blocked_on relationships to find a runnable lock
>> owner. A cycle in that chain can make find_proxy_task() loop indefinitely
>> while holding rq->lock.
>>
>> Use Brent checkpoint state directly in the real owner walk. Cycle detection
>> reuses the owner resolution already performed by that walk and requires no
>> separate preflight traversal. The checkpoint, power, and span state are all
>> invocation-local.
>>
>> Keep the existing owner == p wakeup-race handling ahead of cycle detection.
>> Unlike a sequence-marker approach, this adds no task_struct or runqueue
>> state and requires no activation-time reset.
>>
>> The online walk can temporarily install a blocked_donor cycle before the
>> delayed Brent detection point. In the tested recovery path, the selected
>> task's blocked_donor was cleared before it resumed. A forced-stale control
>> also confirmed that mutex handoff revalidates the donor's blocked_on
>> relationship before consuming a backlink. Validation of this trade-off and
>> comparative measurements against the sequence-marker approach are included
>> in the cover letter.
>>
>> Signed-off-by: Hui Su <sh_def@xxxxxxx>
>
> Hey! Thanks for sending this out and apologies for my slow response to
> your patches.
>
> This looks interesting, and with Suleiman's futex work, it becomes
> more critical to include.
>
> I think similar to what Peter already mentioned, having a max depth
> counter is probably still a good idea, but it does seem like this
> would catch smaller cycles earlier without much overhead.
>
> I've included an simplified version of Zhidao Su's earlier work (using
> a fixed MAX_PROXY_CHAIN_DEPTH) in my tree, but will include this as
> well as an potential optimization.

So if we are eventually planning on returning -EDEADLK to userspace
(context: https://lore.kernel.org/lkml/20260917153649.GK4121339@xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx/)
do we care where the chain starts or can we return -EDEADLK anywhere
in the chain?

I think Brent checkpoint first converges to some point in the chain
and then requires additional traversal to find the beginning of the
chain. Is that fine?

--
Thanks and Regards,
Prateek