Re: [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection

From: Breno Leitao

Date: Mon Sep 21 2026 - 10:45:42 EST


On Thu, Sep 17, 2026 at 04:05:20PM +0200, Andrea Righi wrote:
> POWER7 uses SD_ASYM_PACKING at the shared-capacity SMT level to order
> hardware threads, and NVIDIA Olympus benefits from the same policy. Idle
> CPU selection does not consult that order, so a task can wake on an
> arbitrary sibling and remain there until load balancing corrects the
> placement. On these systems, that initial choice can prevent the core
> from entering its preferred lower-thread resource mode and cause a large
> and persistent performance loss.
>
> When idle selection finds an available CPU in an SMT core, choose the
> highest-priority available sibling. On SMT2 Olympus this only changes
> selection on fully idle cores. A partially idle core has only one
> available CPU. On wider SMT systems such as POWER7, it also fills
> available siblings in priority order while the core is partially busy.
>
> Apply the preference to idle-core and idle-CPU scans,
> asymmetric-capacity scans, target, previous, recently-used CPU fast
> paths and the slow path. Inspect the lowest scheduling domain directly,
> but require both CPUs to share its span because isolcpus can split
> hardware siblings across scheduling domains.
>
> Keep physical-core capacity selection independent from SMT sibling
> ordering. SD_ASYM_CPUCAPACITY first selects among cores with different
> maximum capacities, then SD_ASYM_PACKING selects the preferred available
> sibling inside the chosen core, whose siblings continue to share equal
> capacity.
>
> Reviewed-by: Srikar Dronamraju <srikar@xxxxxxxxxxxxx>
> Reviewed-by: K Prateek Nayak <kprateek.nayak@xxxxxxx>
> Tested-by: K Prateek Nayak <kprateek.nayak@xxxxxxx>
> Signed-off-by: Andrea Righi <arighi@xxxxxxxxxx>

Tested-by: Breno Leitao <leitao@xxxxxxxxxx>

I tested v6 on top of linux-next 20260918 with a KASAN + PROVE_LOCKING
arm64 kernel and etst on my own 2-socket NVIDIA Olympus system (88 cores
per socket, SMT2, 352 CPUs, 36 NUMA nodes).

With sched_smt_asym_packing=on, SD_ASYM_PACKING appears on the SMT
domain only; MC and NUMA are untouched, and the override is reapplied
on every domain rebuild (~400 hotplug operations, including offlining
and re-onlining all 176 second siblings). A task woken on the
higher-numbered sibling is redirected to the lower-numbered one on
800/800 wakeups, against 13/800 without the series. auto, off and an
invalid value behave as documented. No WARN/BUG/KASAN/lockdep under
stress-ng and perf bench sched.

I did not measure throughput.