Re: [PATCHSET sched_ext/for-7.4] sched_ext: Add NUMA balancing support
From: Andrea Righi
Date: Fri Oct 09 2026 - 17:03:58 EST
Hi Tejun,
On Thu, Oct 08, 2026 at 08:23:16AM -1000, Tejun Heo wrote:
> Hello, Andrea.
>
> On Sun, Oct 04, 2026 at 09:27:06AM +0200, Andrea Righi wrote:
> > This series lets a BPF scheduler opt into the NUMA hinting-fault scan with
> > SCX_OPS_NUMA_BALANCING. The existing scan is driven from the sched_ext tick for
> > the tasks of such a scheduler, and the hinting faults then work for them as they
> > do for fair tasks: the kernel maintains their NUMA fault statistics and
> > preferred node, and migrates the memory a task accesses toward the node it runs
> > on. The tasks themselves are not migrated: their placement is left to the BPF
> > scheduler.
> >
> > The resulting preferred node is exposed to BPF schedulers with
> > scx_bpf_task_numa_nid(), which they can use to keep a task close to its memory
> > when placing or balancing it. The value is advisory; nothing in the kernel acts
> > on it for these tasks.
>
> Is there a reason to gate this on a separate ops flag? An alternative would be
> an op which is called when a task's preferred node changes, with the initial
> node reported on enable like ops.set_weight(), and enabling the hinting-fault
> scan iff the op is implemented. The scheduler then gets an event it can act on
> rather than a value to poll. The node has a single writer, sched_setnuma(),
> which already dequeues and re-enqueues the task under the rq lock, so the op
> can be called from there the same way ops.set_weight() is called from the
> reweight path.
Yeah, there's no particular reason to use a separate ops flag. I think I like
the ops callback approach that reports the preferred node and enables the
hinting-fault scan when the callback is implemented. Maybe something like:
void (*set_numa_node)(struct task_struct *p, s32 nid);
Calling it from sched_setnuma() under the rq lock should also simplify the code
and lets the BPF scheduler update its placement state before the task is
re-enqueued.
I'll do some experiments with this approach and send a new version.
Thanks,
-Andrea