[PATCH v2] sched/stats: Fix run_delay over-count for migrated sched_delayed tasks

From: albin_yang

Date: Sun Sep 20 2026 - 00:32:51 EST


From: Wei Yang <albinwyang@xxxxxxxxxxx>

With DELAY_DEQUEUE, a blocked task stays on the runqueue with
se.sched_delayed set and its sched_info.last_queued is cleared, so the sleep
is not counted into run_delay.

When such a delayed (sleeping) task is migrated across CPUs via the plain
migration paths (move_queued_task / move_queued_task_locked, the latter used
by __migrate_swap_task), activate_task(dst, 0) calls enqueue_task() without
ENQUEUE_RESTORE, re-arming last_queued to the migration timestamp while the
task is still sleeping. The later real wakeup (ENQUEUE_DELAYED) tries to
re-arm last_queued at wakeup time but is suppressed because last_queued is
already non-zero, so sched_info_arrive() folds the whole sleep duration
between migration and wakeup into run_delay.

Load balance is affected too: the sched_delayed check in can_migrate_task()
bails out only when env->migration_type != migrate_load, so it does not
block migration when the type is migrate_load - which active load balance
always uses (its lb_env leaves migration_type at 0 == migrate_load), and
which regular load balance can also use via calculate_imbalance(). Either
way the re-attach goes through attach_task() -> activate_task(rq, p,
ENQUEUE_NOCLOCK), without ENQUEUE_RESTORE.

Fix by not re-arming last_queued for a sched_delayed task in
sched_info_enqueue(). The wakeup path clears sched_delayed before reaching
sched_info_enqueue(), so it still re-arms at the real wakeup time. Plain
runnable tasks are unaffected.

Fixes: 152e11f6df29 ("sched/fair: Implement delayed dequeue")
Reported-by: MingTao Huang <mintaohuang@xxxxxxxxxxx>
Signed-off-by: Wei Yang <albinwyang@xxxxxxxxxxx>
Reviewed-by: Chen Yu <yu.c.chen@xxxxxxxxx>
---
v1 -> v2:
- Correct the claim that load-balance migrations are unaffected:
can_migrate_task() only skips a delayed task when migration_type !=
migrate_load, so active load balance (whose lb_env leaves
migration_type at 0 == migrate_load) can migrate a sched_delayed task,
as can regular load balance via the migrate_load branches of
calculate_imbalance(). The fix covers those paths too; only the
description was wrong. Pointed out by Kayra Cizmeci.
- Add Chen Yu's Reviewed-by.

v1: https://lore.kernel.org/all/20260909133345.1572954-1-albin_yang@xxxxxxx/

kernel/sched/stats.h | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/kernel/sched/stats.h b/kernel/sched/stats.h
index ebe0a7765f98..dc626f99ffd9 100644
--- a/kernel/sched/stats.h
+++ b/kernel/sched/stats.h
@@ -290,7 +290,7 @@ static void sched_info_arrive(struct rq *rq, struct task_struct *t)
*/
static inline void sched_info_enqueue(struct rq *rq, struct task_struct *t)
{
- if (!t->sched_info.last_queued)
+ if (!t->sched_info.last_queued && !t->se.sched_delayed)
t->sched_info.last_queued = rq_clock(rq);
}

--
2.43.7