Re: [PATCH v2 01/10] KVM: Reject user accesses to guest memory if current->mm != kvm->mm
From: Sean Christopherson
Date: Thu Oct 01 2026 - 17:20:37 EST
On Thu, Oct 01, 2026, James Houghton wrote:
> On Thu, Oct 1, 2026 at 1:24 PM Sean Christopherson <seanjc@xxxxxxxxxx> wrote:
> >
> > Reject user accesses to guest memory, which are supposed to be done only
> > in the context of KVM_RUN or similar operations, if the current address
> > space is not the VM's (host userspace) address space. If KVM writes to
> > guest memory after the owning host process has exited, or if the VM is
> > being destroyed in the context of a different process, then writing using
> > the wrong address space will corrupt a different process' memory.
> >
> > Reject the access but don't WARN() or KVM_BUG_ON() event though attempting
> > to access guest memory with a mismatched address space is a blatant KVM
> > bug, because unfortunately KVM is buggy. On KVM VMX, when a vCPU is
> > destroyed while L2 is active, KVM synthesizes a nested VM-Exit to force the
> > vCPU out of L2 in order to free the nested VMX assets, and a side effect of
> > a nested VM-Exit is that it flushes the cached shadow VMCS12 back to guest
> > memory:
> >
> > vmx_vcpu_free()
> > |-> nested_vmx_free_vcpu()
> > |-> vmx_leave_nested()
> > |-> nested_vmx_vmexit(vcpu, -1, 0, 0)
> > |-> nested_flush_cached_shadow_vmcs12()
> > |-> kvm_write_guest_cached()
> > |-> __copy_to_user(ghc->hva, ...)
> >
> > Fix the bug broadly even though the "real" bug is that KVM abuses the
> > nested VM-Exit flow for non-architectural purposes, as there may be other
> > such violations lurking. For now, punt on fixing individual bugs and
> > hardening the common flows, e.g. with WARNs.
> >
> > Opportunistically provide wrappers in anticipation of adding more checks
> > and hardening, i.e. growing the logic beyond checking current->mm.
> >
> > Fixes: 61ada7488ffd ("KVM: nVMX: Cache shadow vmcs12 on VMEntry and flush to memory on VMExit")
> > Cc: stable@xxxxxxxxxxxxxxx
> > Reported-by: Jim Mattson <jmattson@xxxxxxxxxx>
> > Closes: https://lore.kernel.org/all/20260908132838.2116068-1-jmattson@xxxxxxxxxx
> > Signed-off-by: Sean Christopherson <seanjc@xxxxxxxxxx>
>
> Thanks, Sean. Feel free to add:
>
> Reviewed-by: James Houghton <jthoughton@xxxxxxxxxx>
>
> I wonder if it makes sense to add similar hardening to kvm_faultin_pfn().
> What do you think?
I'm not opposed to explicitly hardening kvm_faultin_pfn(), but I don't think it
would add much value in practice. Far more arch code uses __kvm_faultin_pfn()
directly, and that doesn't have a @vcpu or @vm pointer to do the check. We could
obviously "fix" that, but nuking the memslots (patches 3-5) will prevent all but
the most ridiculous bugs. Getting anywhere near __kvm_faultin_pfn() with the
wrong mm would either mean KVM is doing something amazingly stupid during VM
teardown, or I guess maybe the scheduler or preempt notifiers went off the rails?