Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy

From: David Hildenbrand (Arm)

Date: Thu Oct 01 2026 - 07:04:07 EST


On 9/30/26 17:10, Zi Yan wrote:
> On 30 Sep 2026, at 10:50, Gregory Price wrote:
>
>> On Wed, Sep 30, 2026 at 10:02:23PM +0800, Li Zhe wrote:
>>> The use case is closer to the second one, but the main motivation is not
>>> only startup-time performance.  The more important point is that each
>>> workload has its own DDR/CXL budget assigned by the workload manager.
>>>
>>
>> mempolicy is the wrong interface to do budget policy, you'd be better
>> off looking at Joshua's memcg tiered node solutions for that.
>>
>>> MPOL_WEIGHTED_INTERLEAVE is useful because it lets new allocations
>>> follow that per-workload budget from the beginning, instead of placing
>>> everything on one tier first and correcting the placement later. This is
>>> important for workload orchestration because the workload starts from a
>>> placement close to its assigned DDR/CXL ratio.
>>>
>>
>> This however is reasonable to me - you'd prefer to spread out the cost
>> of initial faulting placement explicitly, rather than simply take
>> fallbacks when the top-tier budget becomes pressured.
>>
>> i.e. w/o interleave:
>>
>> [node 0 ]
>> ^^^^^^^^^^^^^^ alloc until full
>> vvvvvvvvv fallback
>> [node 1 ]
>>
>> in this scenario you end up with considerable hot memory
>> on the remote node consolidated in time-space (everything
>> allocated after node0 becomes full skews heavily toward
>> node1)
>>
>> Tiering then likely takes many faults to rebalance after
>> you've already reached node0 limits.
>>
>> w/ interleave
>>
>> [node 0 ]
>> ^^^vv^^^vv^^^vv^^^vv^^^vv^^^vv....
>> [node 1 ]
>>
>> In this scenario you do an initial fill distributed by weight
>> and then let tiering figure it out without consolidating all
>> of the pressure to the point where node0 has no space left.
>>
>> That said - this seems like mostly an initial-fill problem, after
>> you initially fill your memory, a new allocation largely implies
>> the memory is hot - and you probably prefer that to be local.
>
> If this is a initial-fill problem, can userspace set weighted interleave
> initially? The program or harness can observe memory usage of the program
> or related NUMA nodes and switch the policy to numa balancing via
> set_mempolicy() or mbind() without MPOL_MF_MOVE after certain threshold
> is met?

You mean: use the weighted policy initially and then switch to a NUMA-balancing
one which doesn't involve the weights anymore?

That makes more sense to me. Although I struggle to see why an effectively
"let's put random memory on slow and others at hot" is a good starting point to
later let if be fixed up by actual balancing/tiering.

It all sounds a bit hackish. :)

--
Cheers,

David