Re: [PATCH 02/18] sched/eevdf: Reset lag when waking up on idle cpu

From: Kayra Cizmeci

Date: Sun Oct 04 2026 - 13:33:37 EST


> When several tasks wake up simultaneously on an idle CPU, their final vlag
> will depend of the ordering as the first one will lose its lag but not
> the next ones.
> Reset the lag when the enqueue happens while no fair task has already been
> picked et set as the running task.

There is a typo. ('et', also 'depend' and 'loose' Not sure that's all :>)

> As a typical example:
> CPU0 is idle
> TA with vlag 0ms and TB with vlag 5ms wake up on CPU0 simultaneously.
> Depending which grab the lock 1st the behavior will be different:
> If TA is enqueued 1st, TB will be enqueued with a positive lag and will
> be picked 1st.
> But if TB is enqueued 1st, it will loose its positive vlag and both TA and
> TB will have 0 vlag when fair will pick a task.

> diff --git a/include/linux/sched.h b/include/linux/sched.h
> index d7cc77181ef9..32d7077ef148 100644
> --- a/include/linux/sched.h
> +++ b/include/linux/sched.h
> @@ -592,6 +592,7 @@ struct sched_entity {
> u64 vruntime;
> /* Approximated virtual lag: */
> s64 vlag;
> + u32 vlag_seq;
> /* 'Protected' deadline, to give out minimum quantums: */
> u64 vprot;
> u64 slice;
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 8cda1d39b037..4c8f12fc8869 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -893,6 +893,7 @@ bool update_entity_lag(struct cfs_rq *cfs_rq, struct sched_entity *se)
> vlag = min(vlag, 0);
> }
> se->vlag = vlag;
> + se->vlag_seq = cfs_rq->idle_seq;
>
> return avruntime - vlag != se->vruntime;
> }
> @@ -914,9 +915,22 @@ void decay_entity_lag(struct cfs_rq *cfs_rq, struct sched_entity *se, int flags)
>
> rq = rq_of(cfs_rq);
>
> + /* You can't claim any lag when waking on idle CPU */
> + if (rq->curr == rq->idle) {
> + se->vlag = 0;
> + return;
> + }
> +
> + /* Accessing remote rq task clock is a cost */
> if (flags & ENQUEUE_MIGRATED)
> return;
>
> + /* CPU has been idle in between so the lag has been removed */
> + if (se->vlag_seq != cfs_rq->idle_seq) {
> + se->vlag = 0;
> + return;
> + }
> +
> /* Compute sleep time */
> delta_exec = rq_clock_task(rq) - se->exec_start;
> if (unlikely(delta_exec <= 0))
> @@ -8190,6 +8204,9 @@ static bool __dequeue_task(struct rq *rq, struct task_struct *p, int flags)
>
> dequeue_hierarchy(p, flags);
>
> + if (!cfs_rq->h_nr_queued)
> + cfs_rq->idle_seq++;

Shouldn't this skip DEQUEUE_SAVE?

> +
> if (sched_feat(PLACE_REL_DEADLINE) && !task_sleep) {
> se->deadline -= se->vruntime;
> se->rel_deadline = 1;
> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
> index b98084e1f5b0..69a2a749e188 100644
> --- a/kernel/sched/sched.h
> +++ b/kernel/sched/sched.h
> @@ -689,6 +689,7 @@ struct cfs_rq {
> u64 sum_weight;
> u64 zero_vruntime;
> unsigned int sum_shift;
> + u32 idle_seq;
>
> #ifdef CONFIG_SCHED_CORE
> unsigned int forceidle_seq;

Yeah, well AFAICT other things seems OK.

Also, while trying to get this patch series on my tip
branch I had a hard time. I was at the caves
of git and b4 figthing with... Everything.
Like there were no v2 tags on some of the patches
so b4 didn't tracked them. I had to do some
things by hand and finally it worked. (Not too greatly tho.)

But who cares? :>.


Thanks,
Kayra