[PATCH] riscv: Fix icache flush being skipped for a second mm mapping an exec folio

From: Nickolai Zeldovich

Date: Fri Oct 09 2026 - 18:20:09 EST


Since commit 01261e24cfab ("riscv: Only flush the mm icache when
setting an exec pte"), flush_icache_pte() flushes only the icache of
the harts that run the faulting mm (with a deferred fence.i for the
harts it migrates to later), but it still sets the folio-wide
PG_dcache_clean bit. The bit is then read as "no hart holds stale
instructions for this folio", which a per-mm flush does not establish.

So when a folio that was written through the page cache is mapped
executable first by mm A on hart X and then by a different mm B on a
hart Y outside A's cpumask, B gets no flush on Y and executes whatever
Y's icache still holds for those physical lines, e.g. the page's
previous contents. Before that commit, flush_icache_all() covered this
case.

Reproducer: a parent pinned to hart 0 and a child pinned to hart 3
share a file. The parent writes text "T1" with write(2), the child
mmap()s it PROT_EXEC and runs it (priming hart 3's icache with T1),
then unmaps it. The parent writes text "T2", maps it executable and
runs it (per-mm flush of hart 0 only, bit set). The child maps the
file executable again and runs it: no flush on hart 3, and the child
executes T1. On a StarFive JH7110 (VisionFive 2, non-coherent icache)
running v7.3-rc6, 149 of 150 iterations over three hart pairs execute
stale instructions. A control run that executes fence.i in the child
before the last mapping gets 0 of 50.

Keep the per-mm flush and make the skip decision per mm instead:
count the flushes that set the bit in a global generation, and let
every mm remember the generation of its own last flush taken in
flush_icache_pte(). An mm whose generation lags cannot trust any bit
set since, so it flushes its own harts once (local fence.i, IPIs only
to the harts currently running it, deferred fence.i for the rest) and
catches up. No global flush is issued, nothing happens while no new
executable folio is written, and the cost is bounded by one
flush_icache_mm() per mm per generation bump.

With the fix the reproducer executes 0 of 150 stale iterations on the
same board. The function-call IPI counters stay at a few hundred per
hart for the whole boot plus 200 iterations, i.e. the IPI savings of
the per-mm flush are kept.

Tested on the JH7110 with v7.3-rc6 and this patch; not tested on
32-bit. The bug does not reproduce under QEMU TCG, which invalidates
translated code on page writes.

Fixes: 01261e24cfab ("riscv: Only flush the mm icache when setting an exec pte")
Cc: stable@xxxxxxxxxxxxxxx
Assisted-by: LLM
Signed-off-by: Nickolai Zeldovich <nickolai@xxxxxxxxxxxxx>
---
arch/riscv/include/asm/mmu.h | 2 ++
arch/riscv/include/asm/mmu_context.h | 1 +
arch/riscv/mm/cacheflush.c | 20 ++++++++++++++++++++
3 files changed, 23 insertions(+)

diff --git a/arch/riscv/include/asm/mmu.h b/arch/riscv/include/asm/mmu.h
index cf8e6eac77d5..e0e7a310151c 100644
--- a/arch/riscv/include/asm/mmu.h
+++ b/arch/riscv/include/asm/mmu.h
@@ -21,6 +21,8 @@ typedef struct {
cpumask_t icache_stale_mask;
/* Force local icache flush on all migrations. */
bool force_icache_flush;
+ /* icache_folio_gen at this mm's last flush in flush_icache_pte(). */
+ u64 icache_gen;
#endif
#ifdef CONFIG_BINFMT_ELF_FDPIC
unsigned long exec_fdpic_loadmap;
diff --git a/arch/riscv/include/asm/mmu_context.h b/arch/riscv/include/asm/mmu_context.h
index dbf27a78df6c..cc0f7f65ec8b 100644
--- a/arch/riscv/include/asm/mmu_context.h
+++ b/arch/riscv/include/asm/mmu_context.h
@@ -32,6 +32,7 @@ static inline int init_new_context(struct task_struct *tsk,
{
#ifdef CONFIG_MMU
atomic_long_set(&mm->context.id, 0);
+ mm->context.icache_gen = 0;
#endif
if (IS_ENABLED(CONFIG_RISCV_ISA_SUPM))
clear_bit(MM_CONTEXT_LOCK_PMLEN, &mm->context.flags);
diff --git a/arch/riscv/mm/cacheflush.c b/arch/riscv/mm/cacheflush.c
index f8ead7cb7c7d..880c210dbfec 100644
--- a/arch/riscv/mm/cacheflush.c
+++ b/arch/riscv/mm/cacheflush.c
@@ -97,13 +97,33 @@ void flush_icache_mm(struct mm_struct *mm, bool local)
#endif /* CONFIG_SMP */

#ifdef CONFIG_MMU
+/*
+ * PG_dcache_clean is folio-wide, but flush_icache_mm() only reaches the
+ * harts of one mm. Count the flushes that set the bit; an mm whose
+ * generation lags cannot trust a bit set since its own last flush, so it
+ * flushes its harts once before relying on it.
+ */
+static atomic64_t icache_folio_gen = ATOMIC64_INIT(0);
+
void flush_icache_pte(struct mm_struct *mm, pte_t pte)
{
struct folio *folio = page_folio(pte_page(pte));
+ u64 gen;

if (!test_bit(PG_dcache_clean, &folio->flags.f)) {
+ gen = atomic64_inc_return(&icache_folio_gen);
flush_icache_mm(mm, false);
+ WRITE_ONCE(mm->context.icache_gen, gen);
set_bit(PG_dcache_clean, &folio->flags.f);
+ return;
+ }
+
+ /* Pairs with the fully ordered atomic64_inc_return() above. */
+ smp_rmb();
+ gen = atomic64_read(&icache_folio_gen);
+ if (unlikely(READ_ONCE(mm->context.icache_gen) != gen)) {
+ flush_icache_mm(mm, false);
+ WRITE_ONCE(mm->context.icache_gen, gen);
}
}
#endif /* CONFIG_MMU */
--
2.56.0