Re: [RFC PATCH 09/16] sched/core: Track CPU where the task was blocked on

From: K Prateek Nayak

Date: Wed Sep 16 2026 - 02:30:37 EST


Hello John,

On 9/16/2026 11:22 AM, John Stultz wrote:
> On Tue, Aug 25, 2026 at 11:32 PM K Prateek Nayak <kprateek.nayak@xxxxxxx> wrote:
>>
>> Track the CPU where task was blocked on when running with
>> sched_proxy_exec().
>>
>> This is used to re-direct the activation of blocked donors queued on the
>> blocked_head via the said CPU in the activation slowpath that will be
>> added in the subsequent patches.
>
> I'm a little confused on this, as the cpu the task was blocked on
> doesn't intuitively have much bearing on where we'd want to activate
> it, when a sleeping owner wakes up.
>
> When a lock owner sleeps, all the chains of tasks waiting on that lock
> (directly or not), will be enqueued on that sleeping owner.
>
> In my patch, when the sleeping owner wakes up, it may or may not wake
> up on the cpu it was dequeued from. It seems we would want to enqueue
> all of those blocked tasks onto the same runqueue where we activated
> the owner, so they are all present on the same rq to be selected as a
> donor to potentially run the owner and release the lock.
>
> Since its like there is a likely chance the sleeping owner will wake
> elsewhere (It looks like proxy_activate_blocked_task() in the later
> patch effectively overrides the activation on the given rq and instead
> wakes the tasks on the blocked_cpu), won't this end up activating the
> chain of blocked waiters on the wrong cpu (forcing them all to be
> immediately proxy migrated over)?

The base concept is this: When the task is blocked and is queued on a
sleeping owner (p->is_linked = 1), the whole chain is linked to one
CPU (p->blocked_cpu).

This is why, later, in Patch 13, we block with (SLEEP | MIGRATING),
and do __set_task_cpu() to owner->blocked_cpu before we transition
p->on_rq to 0.

Entire chain is linked to the p->blocked_cpu of the top level sleeping
owner.

p->wake_cpu can change (more on that below) when task is fully blocked
(p->on_rq = 0) but we *need* to have one unified CPU for the entire
chain to resume wakeup from which is why we need a second variable.

For ->is_linked tasks, any state transition (p->on_rq transitions) are
guarded by p->blocked_cpu's rq_lock() and that *cannot* change until
the owner's wakeup in proxy_activate_task() is finished.

Any external wakeup in ttwu_runnable() for p->on_rq = 0 &&
p->is_linked = 1 should go via "p->blocked_cpu". Same for all the
guard(sched_change) stuff.

Also since we do a:

set_task_cpu(owner, cpu);
ttwu_do_activate(rq, owner, flags);

we cannot simply look at task_cpu(owner) in ttwu_do_activate() to
know where the donor-chain resides. We need the p->blocked_cpu.
>> p->wake_cpu or task_cpu() is not sufficient for this purpose since
>> p->wake_cpu is not stable when !task_on_rq_queued() and there is a
>> window between set_task_cpu() and activet_blocked_task() in the wakeup
>> path that needs to be plugged in.
>
> Sorry if I'm being dim here.
>
> Do you have more details on this? Was my patch prone to the same issue
> you're avoiding in the second item here?
I think NUMA balancing and workqueue have cases where only p->wake_cpu
can be manipulated to direct wakeup of a blocked task but it is only
done when p->on_rq is 0 and p->wake_cpu is always changed to a CPU
within the task's affinity boundary so your series does not have a
problem.

Only because this RFC's approach needs to track which CPU has the
ownership of the entire chain, I need a second variable.

--
Thanks and Regards,
Prateek