Re: [RFC PATCH v2 02/23] sched/topology: Introduce a NUMA distance matrix with unique distance values
From: Chen Yu
Date: Thu Oct 08 2026 - 12:48:45 EST
Hi Jianyong,
On Thu, Oct 08, 2026 at 07:35:10AM +0000, Jianyong Wu wrote:
> Hi Tim,
>
> > -----Original Message-----
> > From: Tim Chen <tim.c.chen@xxxxxxxxxxxxxxx>
> > Sent: Friday, October 2, 2026 6:04 AM
> > To: Jianyong Wu <wujianyong@xxxxxxxx>; Peter Zijlstra
> > <peterz@xxxxxxxxxxxxx>
> > Cc: Ingo Molnar <mingo@xxxxxxxxxx>; Juri Lelli <juri.lelli@xxxxxxxxxx>;
> > Vincent Guittot <vincent.guittot@xxxxxxxxxx>; Chen Yu
> > <yu.c.chen@xxxxxxxxx>; Dietmar Eggemann
> > <dietmar.eggemann@xxxxxxx>; Steven Rostedt <rostedt@xxxxxxxxxxx>;
> > Ben Segall <bsegall@xxxxxxxxxx>; Mel Gorman <mgorman@xxxxxxx>;
> > Valentin Schneider <vschneid@xxxxxxxxxx>; K Prateek Nayak
> > <kprateek.nayak@xxxxxxx>; Shrikanth Hegde <sshegde@xxxxxxxxxxxxx>;
> > Phil Auld <pauld@xxxxxxxxxx>; Andrew Morton
> > <akpm@xxxxxxxxxxxxxxxxxxxx>; David Hildenbrand <david@xxxxxxxxxx>;
> > linux-kernel@xxxxxxxxxxxxxxx; linux-mm@xxxxxxxxx;
> > jianyong.wu@xxxxxxxxxxx; Yuan Zhong <zhongyuan@xxxxxxxx>; Huangsj
> > <huangsj@xxxxxxxx>; Fengyu Wang <wangfengyu@xxxxxxxx>; Zhiwei Ying
> > <yingzhiwei@xxxxxxxx>; justin.he@xxxxxxx
> > Subject: Re: [RFC PATCH v2 02/23] sched/topology: Introduce a NUMA
> > distance matrix with unique distance values
> >
> > On Mon, 2026-09-28 at 09:39 +0000, Jianyong Wu wrote:
> > > Hi Tim,
> > >
> > >
> > > Makes sense. So, what about the following solution?
> > > Given a node affinity sequence, a move from src to dst improves the
> > > affinity of every task whose preferred node i ranks dst better than src.
> > The score
> > > is then
> > >
> > > Di = raw_dist(src, i) - raw_dist(dst, i)
> > > affinity_bias_i = position of src minus position of dst in node i's affinity
> > > sequence, counted among the nodes at the same
> > distance from
> > > i (zero when Di is not zero)
> >
> > Do we really need an affinity bias? I think your intention is to use it for
> > breaking a tie.
> > If there is a tie in affinity score (without injecting bias),
> > just use the position diff to break the tie. Having a bias distorts
> > the affinity score.
> >
> >
> I think there is a misunderstanding about the purpose of the bias.
> It is not intended to break ties between affinity scores. I want the
> score itself to reflect opportunities for aggregation, including moves
> that leave the raw NUMA distance unchanged but follow the fixed node
> affinity order.
>
> That said, I agree with your concern about using artificial node
> distances in the score calculation. The magnitude of an artificial
> distance or rank difference should not determine the weight of those
> aggregation opportunities.
>
Thanks a lot for working on this.
I think there are two level of aggregation:
LLCs aggregation, and Node aggregation.
Your current proposal uses a best-effort strategy for multi-LLC/node
aggregation. It tries to migrate towards a preferred LLC mask, and the
size of the LLC/node mask is the "range" you calculated in task_cache_work().
Take the LLC case, for example. Suppose there are 6 LLCs, and in
task_cache_work(), the range is calculated as 3 for process P, which means
that at most 3 LLCs can hold all the threads of P.
The strategy allows load balancing to move into the preferred LLC mask/range
in random order:
suppose the util of each LLCs are:
LLC0: 45% LLC1: 40% LLC2: 30%
migrate task from LLC4 to LLC1? OK
migrate task from LLC4 to LLC0? OK
migrate task from LLC2 to LLC0? OK
migrate task from LLC0 to LLC1? no
...
My understanding is that the introduction of distance comparison is
because the NUMA nodes are not symmetric to each other, so you need a
realistic sequence to saturate the nodes. However, this might not be
true for LLCs - they are symmetric, and you can start with a pref CPU
and get the next LLC via the LLC id sequence. That is to say,
dist(LLC1, LLC2) = (rank1 + rank2) % k + 1 is not mandatory. Why not
just generate the preferred_mask by adding the "range" siblings of the
preferred LLC, and use preferred_mask directly during load balance?
That seems to be straightforward. So we can get rid of the distance
comparison (at least for LLCs).
thanks,
Chenyu