[PATCH v8 11/30] mm: split PMD swap entries into PTE swap entries

From: Usama Arif

Date: Fri Oct 02 2026 - 06:02:38 EST


Once a PMD can hold a swap entry, everything that splits a PMD - mprotect()
or munmap() over part of the range, MADV_FREE, a pagewalk with no PMD
handler - has to be able to split that entry too. A swap PMD already passes
the pmd_is_valid_softleaf() gate, so without this it reaches
__split_huge_pmd_locked() and falls through to the present-PMD path, which
pmdp_invalidate()s it and calls pmd_page() on a non-present entry.

No reference counting is needed: a swap entry pins no folio, and swap_map
is already one per slot, so the PTEs simply take over what the PMD held.

The migration-only entry point cannot reach the new branch:
page_vma_mapped_walk() never hands back a swap PMD, and
__split_huge_pmd_locked() already asserts that to_migration_entries
implies a present or device-private PMD.

Test the pre-split old_pmd rather than re-reading *pmd in the trailing
folio_remove_rmap_pmd() gate, so every entry-type test in the function
interrogates the same snapshot. That part is cosmetic: pmdp_invalidate()
leaves the PMD present as far as software is concerned.

Signed-off-by: Usama Arif <usama.arif@xxxxxxxxx>
---
mm/huge_memory.c | 22 +++++++++++++++++++++-
1 file changed, 21 insertions(+), 1 deletion(-)

diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 1df4f619620b4..24d116ae1fc30 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -3306,6 +3306,11 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
folio_add_anon_rmap_ptes(folio, page, HPAGE_PMD_NR,
vma, haddr, rmap_flags);
}
+ } else if (pmd_is_swap_entry(*pmd)) {
+ old_pmd = *pmd;
+ soft_dirty = pmd_swp_soft_dirty(old_pmd);
+ uffd_wp = pmd_swp_uffd(old_pmd);
+ anon_exclusive = pmd_swp_exclusive(old_pmd);
} else {
/*
* Up to this point the pmd is present and huge and userland has
@@ -3443,6 +3448,21 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
VM_WARN_ON(!pte_none(ptep_get(pte + i)));
set_pte_at(mm, addr, pte + i, entry);
}
+ } else if (pmd_is_swap_entry(old_pmd)) {
+ pte_t entry = softleaf_to_pte(softleaf_from_pmd(old_pmd));
+
+ if (soft_dirty)
+ entry = pte_swp_mksoft_dirty(entry);
+ if (uffd_wp)
+ entry = pte_swp_mkuffd(entry);
+ if (anon_exclusive)
+ entry = pte_swp_mkexclusive(entry);
+
+ for (i = 0, addr = haddr; i < HPAGE_PMD_NR; i++, addr += PAGE_SIZE) {
+ VM_WARN_ON(!pte_none(ptep_get(pte + i)));
+ set_pte_at(mm, addr, pte + i, entry);
+ entry = pte_next_swp_offset(entry);
+ }
} else {
pte_t entry;

@@ -3470,7 +3490,7 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
}
pte_unmap(pte);

- if (!pmd_is_migration_entry(*pmd))
+ if (!pmd_is_migration_entry(old_pmd) && !pmd_is_swap_entry(old_pmd))
folio_remove_rmap_pmd(folio, page, vma);
if (to_migration_entries)
put_page(page);
--
2.53.0-Meta