Re: [PATCH v2 0/3] KVM: TDX: Syntehsize SHUTDOWN on unhandled EPT Violation
From: Sean Christopherson
Date: Fri Oct 09 2026 - 16:58:48 EST
On Thu, Oct 08, 2026, Yan Zhao wrote:
> On Tue, Sep 29, 2026 at 05:55:11PM -0700, Sean Christopherson wrote:
> > On Wed, Sep 30, 2026, Rick P Edgecombe wrote:
> > > On Tue, 2026-09-29 at 17:11 -0700, Sean Christopherson wrote:
> > > > Signal SHUTDOWN instead of -EIO if the guest accesses an unaccepted page and
> > > > has disabled EPT Violation #VEs on such accesses. Returning -EIO is all but
> > > > guaranteed to mislead the VMM into thinking KVM (or the VMM) messed up, and
> > > > will likely result in the VM being terminated instead of rebooted.
> Hmm, it's by design in my previous implementation to terminate a VM when the
> guest accesses an unaccepted page and has disabled EPT Violation #VEs on such
> accesses, according to the TDX module spec:
>
> "This happens if the TD is configured to TD-exit (instead of a #VE) on an EPT
> violation due to accessing a PENDING page. It normally indicates an error
> condition; the host VMM may decide to tear the TD down."
>
> So, for the initial implementation, we chose to invoke kvm_vm_dead() and print
> out the exact reason as a hint to the system admin.
kvm_vm_dead() should only be used when KVM *needs* to prevent userspace from
running the VM.
> (Previously, reboot was also not supported for a TDX guest).
But QEMU isn't KVM. While QEMU is the de facto reference VMM, it's important to
keep in mind that it's not the only VMM (and that's not just about Google's VMM,
there are plenty of other non-QEMU VMMs that are used in production environments).
> Returning -EIO was based on the following considerations:
> - vcpu_enter_guest() returns -EIO when kvm_test_request(KVM_REQ_VM_DEAD, vcpu)
> is true.
> - A request to handle a specific EPT violation is rejected by the firmware (the
> TDX module).
>
> > > I'm pretty sure Yan had a test for this path, and I'd love to see it actually
> > > exercised. What is the urgency on getting this fix upstream? Can we wait a week?
> I triggered the EPT violation in the kexec path, and successfully had the TD
> - killed with error msg: "qemu-system-x86_64: cpus are not resettable,
> terminating" on an old QEMU, or
> - rebooted on a new QEMU with the msg printed:
> "qemu-system-x86_64: info: virtual machine state has been rebuilt with new
> guest file handle".
>
> The TDX selftests we are using do not handle KVM_EXIT_SHUTDOWN, so if such
> an error occurs, it's silently ignored by the TDX selftests even with msg
> "kvm_intel: Guest access before accepting 0x8000c000 on vCPU 1" printed in dmesg.
If TDX selftests don't fail due to an unexpected KVM_EXIT_SHUTDOWN, then that needs
to be fixed irrespective of this change.
> > Absolutely. It can probably wait a month and no one would care. IIRC, this got
> > hit by someone (internal to Google) deliberately crashing a guest kernel and doing
> > funky things with kexec. It showed up on my radar purely because our automated
> > madness alerted on the resulting assertion (on -EIO) in the VMM.
> Though I didn't realize that returning -EIO would mislead the VMM into thinking
> KVM (or the VMM) messed up, such an error could also be introduced by a VMM bug?
Yes, but that holds true for literally every guest failure. It's like saying that
all observed errors could be due to a CPU bug, e.g. because the CPU corrupted state
or because a bit flipped in DRAM. It's technically true, but in practice the vast
majority of failures (outside of pre-production hardware) are software errors, and
so absent evidence that a hardware/VMM/KVM bug is at play, KVM shouldn't assume
anything. I.e. as above, unless KVM needs to protect itself, KVM should aways try
to provide semantics that are rooted in the CPU architecture.
E.g. an analogous failure would for non-TDX guests would be if the host zeroed a
page that happened to be necessary to handle exceptions in the guest, and as a
result the guest hit a triple fault shutdown. KVM doesn't return -EIO instead of
KVM_EXIT_SHUTDOWN just because the failure could have been introduced by a kernel
or VMM bug.
> e.g., VMM removes an accepted page without any notification to the TD configured
> with TDX_TD_ATTRIBUTES_SEPT_VE_DISABLE bit.