Re: hunting memory corruption bug in 6.18.x

From: Borislav Petkov

Date: Thu Oct 08 2026 - 14:24:38 EST


On Thu, Oct 08, 2026 at 02:29:36PM +0200, Nikola Ciprich wrote:
> [11402.940943] BUG: unable to handle page fault for address: ffffffff0c93001c
> [11402.942273] #PF: supervisor read access in kernel mode
> [11402.943629] #PF: error_code(0x0000) - not-present page
> [11402.945025] PGD 6d6f83a067 P4D 6d6f83b067 PUD 0
> [11402.946469] Oops: Oops: 0000 [#1] SMP NOPTI
> [11402.947950] CPU: 23 UID: 0 PID: 704950 Comm: servercare-moni Kdump: loaded Tainted: G E 6.18.55lb9.01 #1 PREEMPT(voluntary)
> [11402.951254] Tainted: [E]=UNSIGNED_MODULE
> [11402.952913] Hardware name: ASUSTeK COMPUTER INC. RS720A-E12-RS12/K14PP-D24 Series, BIOS 1201 08/25/2023
> [11402.956569] RIP: 0010:__d_lookup_rcu+0x4d/0xe0
> [11402.958452] Code: 48 8d 04 c2 f6 07 02 0f 85 a0 00 00 00 48 8b 10 48 89 d0 48 83 e0 fe 48 83 fa 01 77 0d e9 80 00 00 00 48 8b 00 48 85 c0 74 78 <44> 8b 58 fc 48 39 78 10 75 ee 48 83 78 08

What is that kernel?

6.18.55lb9.01

The other machine has a 6.18.20lb9.03-something one.

How can I look at the vmlinux you're running and the sources?

rIP points to:

[11402.958452] Code: 48 8d 04 c2 f6 07 02 0f 85 a0 00 00 00 48 8b 10 48 89 d0 48 83 e0 fe 48 83 fa 01 77 0d e9 80 00 00 00 48 8b 00 48 85 c0 74 78 <44> 8b 58 fc 48 39 78 10 75 ee 48 83 78 08
All code
========
0: 48 8d 04 c2 lea (%rdx,%rax,8),%rax
4: f6 07 02 testb $0x2,(%rdi)
7: 0f 85 a0 00 00 00 jne 0xad
d: 48 8b 10 mov (%rax),%rdx
10: 48 89 d0 mov %rdx,%rax
13: 48 83 e0 fe and $0xfffffffffffffffe,%rax
17: 48 83 fa 01 cmp $0x1,%rdx
1b: 77 0d ja 0x2a
1d: e9 80 00 00 00 jmp 0xa2
22: 48 8b 00 mov (%rax),%rax
25: 48 85 c0 test %rax,%rax
28: 74 78 je 0xa2
2a:* 44 8b 58 fc mov -0x4(%rax),%r11d <-- trapping instruction
2e: 48 39 78 10 cmp %rdi,0x10(%rax)
32: 75 ee jne 0x22
34: 48 rex.W
35: 83 .byte 0x83
36: 78 08 js 0x40

I need to be able to pinpoint it back to the source.

I asked the last time:

"Just to rule out any other issues which got fixed in the meantime, can you try
mainline Linux and see if you can reproduce your observation with it?

If so, you could share your crash core along with debug kernels yadda yadda so
that I can poke at it.

And before you do, make sure you have the latest BIOS and microcode installed on
that machine.

Also, where can I find full dmesg and /proc/cpuinfo from those machines which
trigger this?"

But still nothing.

Imagine this issue has been fixed upstream but you don't have the fix in your
kernels and we're basically chasing the same thing again...

Sorry, but I have lost my debugging crystal ball which can help me guess what
the machine does. :\

> so we now know this didn't fixed it. however I didn't have tlbi=ipi set, so I'll
> now try this.

That won't help either but if you wanna try it.

> any ideas on this new info?

Yes, see above.

Bottomline is: without sufficient debugging data and up-to-date hardware,
there's not a lot I can do.

Thx.

--
Regards/Gruss,
Boris.

https://people.kernel.org/tglx/notes-about-netiquette