[PATCH v2] x86/mce: Do not treat cache hierarchy errors as memory errors

From: Adrian Schlegel

Date: Sat Oct 03 2026 - 18:21:31 EST


On Intel and Zhaoxin, mce_is_memory_error() returns true not only for memory controller errors but also for cache hierarchy errors (MCACOD 000F 0001 RRRR TTLL) and generic cache hierarchy errors (000F 0000 0000 11LL). The cache hierarchy checks were added in commit fa92c5869426 ("x86, mce: Support memory error recovery for both UCNA and Deferred error in machine_check_poll") for uncorrected errors. All callers of mce_is_memory_error() expect it to report memory errors only, as its name says.

This is wrong for corrected errors. A corrected cache hierarchy error means a bit flipped in the cache and was fixed there. The reported address only names the line that happened to be cached. The CEC still counts these as DRAM errors, and with the Intel action threshold of 2 from commit d25c6948a6aa ("RAS/CEC: Reduce offline page threshold for Intel systems"), two such errors on the same page are enough to soft-offline it.

This was observed on an i9-9900K without ECC memory and with a faulty core. Bank 3 on CPU 1 reports thousands of corrected cache errors (MCACOD 0x0135, 0x0151, 0x0179, ...) with addresses spread over the whole physical address space. Within one hour the CEC made 502 soft-offline attempts on 159 distinct pages. 453 failed because the page was in kernel use, the others took healthy pages offline:

RAS: Soft-offlining pfn: 0x24f6c8
mce: [Hardware Error]: CPU 1: Machine Check: 0 Bank 3: cc5ffdc000100151
mce: [Hardware Error]: TSC 21903b589a2 ADDR 24f6c85c0 MISC 2516485

As a result, healthy RAM is lost until reboot and HardwareCorrupted keeps growing. Pages that cannot be offlined are retried repeatedly, and the log messages suggest failing DRAM while the defect is in the CPU.

AMD already restricts this to DRAM ECC errors since commit c6708d50f166 ("x86/MCE: Report only DRAM ECC as memory errors on AMD systems"). Do the same for Intel and Zhaoxin.

Fixes: fa92c5869426 ("x86, mce: Support memory error recovery for both UCNA and Deferred error in machine_check_poll")
Signed-off-by: Adrian Schlegel <me@xxxxxxxxxxxxxxxxxx>
---
v2:
- Fix mce_is_memory_error() itself instead of adding a CEC-local helper (Tony)
- v1: https://lore.kernel.org/r/20260927151904.30524-1-me@xxxxxxxxxxxxxxxxxx

Tested with mce-inject (sw) in a VM with an Intel CPU model. A corrected cache error (status 0xcc5ffc0000100179) is no longer counted and its page stays online. A corrected memory controller error (status 0x8c0000000000009f) is still counted and its page soft-offlined.

arch/x86/kernel/cpu/mce/core.c | 16 ++++------------
1 file changed, 4 insertions(+), 12 deletions(-)

diff --git a/arch/x86/kernel/cpu/mce/core.c b/arch/x86/kernel/cpu/mce/core.c
index ab469605f..4d2ace76a 100644
--- a/arch/x86/kernel/cpu/mce/core.c
+++ b/arch/x86/kernel/cpu/mce/core.c
@@ -542,19 +542,11 @@ bool mce_is_memory_error(struct mce *m)
/*
* Intel SDM Volume 3B - 15.9.2 Compound Error Codes
*
- * Bit 7 of the MCACOD field of IA32_MCi_STATUS is used for
- * indicating a memory error. Bit 8 is used for indicating a
- * cache hierarchy error. The combination of bit 2 and bit 3
- * is used for indicating a `generic' cache hierarchy error
- * But we can't just blindly check the above bits, because if
- * bit 11 is set, then it is a bus/interconnect error - and
- * either way the above bits just gives more detail on what
- * bus/interconnect error happened. Note that bit 12 can be
- * ignored, as it's the "filter" bit.
+ * Memory controller errors have an MCACOD of 000F 0000 1MMM CCCC:
+ * bit 7 set, bits 8-11 and 13-15 clear. Bit 12 is the "filter"
+ * bit and is ignored.
*/
- return (m->status & 0xef80) == BIT(7) ||
- (m->status & 0xef00) == BIT(8) ||
- (m->status & 0xeffc) == 0xc;
+ return (m->status & 0xef80) == BIT(7);

default:
return false;
--
2.55.0