Re: [PATCH v2] sched/debug, sys_info: Introduce SYS_INFO_CPU_RUNQUEUES

From: Aaron Tomlin

Date: Tue Sep 22 2026 - 14:37:17 EST


On Tue, Sep 22, 2026 at 04:31:42PM +0200, Petr Mladek wrote:
> On Fri 2026-09-11 21:32:40, Aaron Tomlin wrote:
> > When investigating kernel panics, inspectability of per-CPU runqueues
> > and runnable task states is valuable for diagnosing CPU starvation
> > priority inversion, etc.
> >
> > While debugfs (/sys/kernel/debug/sched/debug) exposes runqueue metrics
> > to userspace, these details are not captured during an automated kernel
> > panic or crash dump. Capturing per-CPU runqueue state directly into
> > log_buf fills this diagnostic gap for post-mortem crash analysis.
> >
> > Introduce SYS_INFO_CPU_RUNQUEUES and its corresponding string token
> > "cpu_runqueues" to panic_sys_info. Add sched_show_runqueues(), modelled
> > on print_rq(), to emit per-CPU scheduler diagnostics to the kernel log.
> >
> > Unlike /sys/kernel/debug/sched/debug which dumps all threads assigned to
> > a CPU, sched_show_runqueues() only emits threads that are actively
> > running or queued on the runqueue (via task_on_rq_queued() and
> > task_current()). This keeps the panic log concise, reflects the true
> > runqueue depth, and prevents overflowing the printk ring buffer on
> > systems with high thread counts.
> >
> > Additionally, to guarantee deadlock and memory safety in panic context:
> > - Acquire the runqueue lock using raw_spin_rq_trylock() with
> > READ_ONCE() and rcu_dereference() fallback, marking contended
> > queues with " (contended)"
> >
> > - Wrap the per-CPU inspection in rcu_read_lock() to protect the
> > sampled current task (comm and PID) against premature release
> > during pr_info() across other callers
> >
> > - Omit cgroup group-path printing in print_rq() to avoid acquiring
> > cgroup_mutex and traversing kernfs dentries
> >
> > --- a/Documentation/admin-guide/sysctl/kernel.rst
> > +++ b/Documentation/admin-guide/sysctl/kernel.rst
> > @@ -939,6 +939,7 @@ locks print locks info if CONFIG_LOCKDEP is on
> > ftrace print ftrace buffer
> > all_bt print all CPUs backtrace (if available in the arch)
> > blocked_tasks print only tasks in uninterruptible (blocked) state
> > +cpu_runqueues print per-CPU runqueue depth and runnable tasks
>
> I would keep is short and call it "rq".

Hi Petr,

Thank you for your review.

Acknowledged, fair enough.

> > ============= ===================================================
> >
> > --- a/include/linux/sys_info.h
> > +++ b/include/linux/sys_info.h
> > @@ -16,6 +16,7 @@
> > #define SYS_INFO_PANIC_CONSOLE_REPLAY 0x00000020
> > #define SYS_INFO_ALL_BT 0x00000040
> > #define SYS_INFO_BLOCKED_TASKS 0x00000080
> > +#define SYS_INFO_CPU_RUNQUEUES 0x00000100
>
> Similar here: SYS_INFO_RQ

Acknowledged.

> > void sys_info(unsigned long si_mask);
> > unsigned long sys_info_parse_param(char *str);
> > diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c
> > index 72236db67983..79f6b00974bb 100644
> > --- a/kernel/sched/debug.c
> > +++ b/kernel/sched/debug.c
> > @@ -1322,6 +1330,48 @@ void sysrq_sched_debug_show(void)
> > }
> > }
> >
> > +void sched_show_runqueues(void)
> > +{
> > + int cpu;
> > +
> > + pr_info("CPU Runqueues:\n");

I will drop this. In print_rq(), there is often no overarching banner line
emitted beforehand.

> > + for_each_online_cpu(cpu) {
> > + struct rq *rq = cpu_rq(cpu);
> > + struct task_struct *curr;
> > + unsigned int nr_running;
> > + u64 nr_switches;
> > + unsigned long flags;
> > + bool locked;
> > +
> > + touch_nmi_watchdog();
> > + touch_all_softlockup_watchdogs();
> > +
> > + rcu_read_lock();
> > + local_irq_save(flags);
> > + locked = raw_spin_rq_trylock(rq);
>
> Is the trylock needed for all sys_info() callers or just in panic()?
> If it is just panic() then I would use it only when oops_in_progress
> is set and use raw_spin_rq_lock() otherwise.

Yes, the trylock in sched_show_runqueues() is needed for all callers to
__sys_info(), not just panic().

Because __sys_info() is also called from NMI context (hardlockup), hardirq
context (softlockup), and khungtaskd, oops_in_progress is 0 in those paths.
Using unconditional raw_spin_rq_lock() there risks fatal self-deadlocks if
rq->lock is already held on the local CPU. Furthermore, locking arbitrary
runqueues in a loop risks AB-BA deadlocks with concurrent scheduler
load-balancing.

Using raw_spin_rq_trylock() ensures sched_show_runqueues() remains strictly
non-blocking across all __sys_info() contexts, and tagging contended queues
with " (contended)" adds useful diagnostic value.

As such, I believe the current implementation is sufficient.

>
> > + if (locked) {
> > + nr_running = rq->nr_running;
> > + nr_switches = rq->nr_switches;
> > + curr = rcu_dereference(rq->curr);
> > + raw_spin_rq_unlock(rq);
> > + } else {
> > + nr_running = READ_ONCE(rq->nr_running);
> > + nr_switches = READ_ONCE(rq->nr_switches);
> > + curr = rcu_dereference(rq->curr);
> > + }
> > + local_irq_restore(flags);
> > +
> > + pr_info("cpu#%d: nr_running:%u switches:%llu curr:%s[%d]%s\n",
> > + cpu, nr_running, nr_switches,
> > + curr ? curr->comm : "<none>",
> > + curr ? task_pid_nr(curr) : -1,
> > + locked ? "" : " (contended)");
> > +
> > + print_rq(NULL, rq, cpu, false, true);
> > + rcu_read_unlock();
> > + }
> > +}
>
> IMHO, it might be a useful feature.
>
> The main question is whether it is acceptable to scheduler
> maintainers. It adds some churn. Also they would need to keep in mind
> that it can be called in panic().

Thank you for your feedback.

Indeed. Regarding the scheduler maintainers (Cc'd Ingo Molnar, Peter
Zijlstra, Juri Lelli, and Vincent Guittot):

The churn was kept strictly confined to kernel/sched/debug.c:
1. It reuses existing print_rq()/print_task() helpers with two flags;
queued_only, to keep output concise and show_cgroup_path, to avoid
cgroup_mutex.

2. sched_show_runqueues() is completely read-only and non-blocking
(raw_spin_rq_trylock() with READ_ONCE() and rcu_dereference()
fallback), so it does not impose locking constraints or latency
risks on core scheduler code, even in panic() or NMI/hardirq
contexts.


Kind regards,
--
Aaron Tomlin