[PATCH v2 0/6] sched/cache: Fixes for cache aware scheduling
From: Tim Chen
Date: Mon Sep 21 2026 - 20:33:18 EST
Hi all,
This is an update of the patches to fix cache aware scheduling issues
found in v7.2.
We collect the fixes in this series so it is easier to track. We have
added two new fixes for issues found since v1 of this series.
Patch 1: alb_break_llc() compares nr_pref_llc_running with
cfs.h_nr_runnable, but those count different sets - one follows queued
tasks, the other drops delay-dequeued ones. With DELAY_DEQUEUE the
equality stops holding and active balance pulls a task off its preferred
LLC. So fix the counter. Reported by Zhan Xusheng:
https://lore.kernel.org/lkml/20260827135000.735138-1-zhanxusheng@xxxxxxxxxx/
v1->v2: Minor code rearrangements in set_delayed().
Patch 2 (Lu Wang): the stopper doing active load balance builds a fresh
lb_env that doesn't inherit migration_type, so can_migrate_task() can
move a task *out* of its preferred LLC. A new LBF_ACTIVE_LB_LLC flag
and picking the stopper callback at kick time keep the intent; passing
migration_type through the stopper would muddy delayed dequeue. v4:
https://lore.kernel.org/lkml/20260903020656.3793626-1-wanglu.priv@xxxxxxxxx/
v1->v2: No change.
Patches 3-4 are for the use after free Hyunwoo Kim caught with KASAN:
https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
account_mm_sched() reaches the stats via p->mm->sc_stat, but a task can
be switching mm on one CPU while another is inside account_mm_sched(),
so the mm and the stats inside it can go away underneath. Locking the
rq in the mm free path felt like the wrong trade, so patch 4 pulls
sched_cache_stat out of mm_struct into a refcounted, RCU freed
sched_cache_group - just moving code - and patch 4 does the real fix:
each task takes its own reference (copy_mm(), exec_mmap(), dropped in
exit_mm()), and leverage call_rcu() to to protect
against UAF in account_mm_sched() so the group outlives any mm switch.
Zehnghui Yu also independentaly found this issue with memory poison.
https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@xxxxxxxxx/
@Hyunwoo and @Zhenhui, will appreciated you can test these patches
and add your Tested-by
Nice side effect: the group no longer follows the address space,
so a user defined group, or cgroup or numa_group could own it later.
These are also the grouping by prctl RFC's first two patches, sent here
so the fix isn't held up by that discussion.
v1->v2: Put cache aware related code in process exit/fork/copy in
its own functions
Patch 5 This patch makes sure kernel threads are excluded from cache
aware scheduling consideration.
v1->v2: Split from previous patch 3 as suggested by Peter Z.
Patch 6 is a new patch to fix an issue of undercomputing the LLC size
during CPU hot plug events. The LLC size was used for estimating
if a process's memory footprint will fit a LLC.
https://lore.kernel.org/all/20260916134432.11767-1-davichazbh@xxxxxxxxx/
BTW, there are two other issues in discussion currently and need
a bit more work:
1. Incorrect donor context being passed to task_tick_cache().
https://lore.kernel.org/lkml/20260909092901.2989564-1-sh_def@xxxxxxx/
It is currently under discussion and is not included in this series.
2. Cache aware scheduling interfering with ITMT.
https://lore.kernel.org/lkml/20260810033742.1688718-1-yu.c.chen@xxxxxxxxx/
https://lore.kernel.org/lkml/2fe2c681-b748-41fa-8b56-1169c86cefbc@xxxxxxxxx/
Applies on sched/urgent branch.
Tim Chen and Chen Yu
Chen Yu (1):
sched/cache: Skip kernel thread for cache aware scheduling
Davi Chaves Azevedo (1):
sched/cache: Refresh LLC capacity across CPU hotplug
Lu Wang (1):
sched/cache: Honor migrate_llc_task semantics in active load balance
Tim Chen (3):
sched/cache: Keep nr_pref_llc_running in the runnable domain
sched/cache: Decouple sched_cache_group from mm
sched/cache: Introduce task_struct->sched_cache_grp
drivers/base/cacheinfo.c | 11 +-
fs/exec.c | 1 +
include/linux/mm_types.h | 15 +-
include/linux/sched.h | 21 ++-
include/linux/sched/topology.h | 4 +-
kernel/exit.c | 28 +--
kernel/fork.c | 2 +
kernel/sched/build_utility.c | 4 +
kernel/sched/cache_sched.c | 106 +++++++++++
kernel/sched/fair.c | 310 ++++++++++++++++++++++++---------
kernel/sched/sched.h | 3 +
kernel/sched/topology.c | 22 ++-
12 files changed, 394 insertions(+), 133 deletions(-)
create mode 100644 kernel/sched/cache_sched.c