Re: [PATCH] cpuidle: Add the shallow governors

From: Christian Loehle

Date: Fri Oct 09 2026 - 08:46:10 EST


On 10/9/26 13:12, Roman Kagan wrote:
> On Thu, Oct 08, 2026 at 07:29:43PM +0100, Christian Loehle wrote:
>> On 10/8/26 19:17, Christian Loehle wrote:
>>> On 10/8/26 19:06, Roman Kagan wrote:
>>>> The idle states offered by a platform trade wakeup latency for energy
>>>> savings, and there are situations where the trade is not worth making:
>>>> while a latency-sensitive workload is running, or during a live update
>>>> via kexec, where everything from the outgoing kernel stopping the
>>>> workload to the incoming kernel resuming it is downtime, and deep idle
>>>> states may lengthen it.
>>>>
>>>> The mechanisms currently available for that are all one-way.
>>>> cpuidle.off=1, idle=poll and idle=halt can only be requested in the
>>>> kernel command line and cannot be undone, and the PM QoS interfaces
>>>> (/dev/cpu_dma_latency and the per-CPU pm_qos_resume_latency_us
>>>> attribute) can only be used once user space is up, so they cannot cover
>>>> the boot of the incoming kernel.
>>>
>>> You can also disable all but the shallowest idle state in sysfs:
>>> echo 1 > /sys/devices/system/cpu/cpuX/cpuidle/stateX/disable
>>>
>>
>> And I'd probably prefer having that exposed via the cmdline rather than
>> two separate governors...
>
> Doing this cmdline configuration per-cpu per-state is non-realistic. I
> guess you mean a single option that would express a policy, like "for
> all cpus in the system, disable all but the shallowest state" or "...
> all but the shallowest non-polling". But policy is exactly what
> governors are for.

Yes, I had something like
cpuidle.max_exit_latency_us=<N>
in mind that then sets the disable attribute for the applicable states.

>
> Why exactly does having two more separate simple, narrow-purpose
> governors sound wrong to you?

Because it's a lot of duplicate code we have to maintain (and the
documentation).