106 lines
5.6 KiB
Plaintext
106 lines
5.6 KiB
Plaintext
Subject: [PATCH 0/5] KVM: x86: match the architecture in four guest-observable exception cases
|
|
|
|
To: Sean Christopherson <seanjc@google.com>, Paolo Bonzini <pbonzini@redhat.com>
|
|
Cc: Thomas Gleixner <tglx@kernel.org>, Ingo Molnar <mingo@redhat.com>,
|
|
Borislav Petkov <bp@alien8.de>, Dave Hansen <dave.hansen@linux.intel.com>,
|
|
"H. Peter Anvin" <hpa@zytor.com>, x86@kernel.org,
|
|
kvm@vger.kernel.org, linux-kernel@vger.kernel.org
|
|
|
|
Four cases where a guest observes a different exception, or a different
|
|
RIP in the exception frame, than the same instruction produces on bare
|
|
metal. All four are reachable from unprivileged guest code, none needs
|
|
a debugger attached, and each is a one-hunk fix on a path KVM already
|
|
owns. The series does not add any new behaviour; it makes the existing
|
|
behaviour match the AMD64 architecture (and, for 2/5 and 3/5, what the
|
|
VMX side of KVM already does).
|
|
|
|
1/5 KVM: SVM: intercept #GP when guest EFER.SVME is clear
|
|
|
|
svm_set_efer() forces EFER_SVME=1 in the VMCB, so a CPL>0
|
|
VMLOAD/VMSAVE/STGI/CLGI passes the hardware SVME check and #GPs
|
|
before the instruction intercept can fire. Bare metal, where
|
|
SVME really is clear, #UDs. svm_recalc_instruction_intercepts()
|
|
already carries a FIXME for this. Intercept #GP while the
|
|
guest's SVME is clear so gp_interception() can decode the
|
|
instruction and route it to nested_svm_check_permissions(),
|
|
which injects the #UD. SEV guests are excluded as KVM cannot
|
|
decode their instructions, mirroring the erratum workaround.
|
|
|
|
2/5 KVM: x86: emulator: #UD, not #GP, for VMCALL at CPL > 0
|
|
|
|
group7_rm0[] declares VMCALL as Priv without PrivUD, so the CPL
|
|
check in x86_emulate_insn() injects #GP(0). 0F 01 C1 is not an
|
|
AMD instruction and is invalid outside VMX operation on Intel;
|
|
hardware #UDs at CPL > 0 on both vendors. This path is reached
|
|
on AMD through the #UD intercept (a guest built for Intel running
|
|
on AMD) and on Intel only under forced emulation.
|
|
|
|
3/5 KVM: x86: #UD for KVM hypercalls issued at CPL > 0
|
|
|
|
kvm_emulate_hypercall() returns -KVM_EPERM and skips the
|
|
instruction for a CPL > 0 caller, so the guest observes no
|
|
exception at all. VMMCALL/VMCALL is only a legal instruction
|
|
because the hypervisor intercepts it; bare metal #UDs. Inject
|
|
#UD instead. This is a guest-observable ABI change, argued
|
|
below.
|
|
|
|
4/5 KVM: SVM: intercept ICEBP and skip it before injecting its #DB
|
|
|
|
ICEBP's #DB is trap-like: the exception frame RIP is the next
|
|
instruction. SVM does not route ICEBP's #DB through the #DB
|
|
exception intercept, so the guest normally takes it correctly,
|
|
but when the delivery is interrupted by a VM-exit (an NPT fault
|
|
on a cold IDT/handler page) EXITINTINFO reports the pending #DB
|
|
with RIP saved on the ICEBP and svm_complete_interrupts()
|
|
re-injects it as-is, so the guest handler sees the ICEBP's own
|
|
address. VMX tags ICEBP (is_icebp()) and skips the instruction;
|
|
SVM has the dedicated ICEBP intercept, which fires before the #DB
|
|
exists. Enable it and skip.
|
|
|
|
5/5 KVM: selftests: verify ICEBP #DB reports RIP past the ICEBP
|
|
|
|
On the 3/5 ABI change: every KVM_HC_* already fails for a CPL > 0
|
|
caller, so only the shape of the failure changes, from a silent skip to
|
|
#UD. The Xen and Hyper-V dispatch sit before the check and keep their
|
|
own CPL semantics; kvm_hv_hypercall() already injects #UD for CPL > 0
|
|
(hyperv.c), so this makes the KVM PV path consistent with it rather
|
|
than inventing new behaviour. No in-tree guest issues a KVM hypercall
|
|
from userspace: Linux paravirt patches VMMCALL/VMCALL at CPL0 only, and
|
|
svm_patch_hypercall() / vmx_patch_hypercall() likewise. Windows guests
|
|
issue neither.
|
|
|
|
Vendor coverage: 1/5 and 4/5 are SVM-only. 2/5 and 3/5 live in
|
|
emulate.c and x86.c and affect Intel as well. On VMX a CPL > 0 VMCALL
|
|
unconditionally VM-exits to handle_vmcall() -> kvm_emulate_hypercall(),
|
|
so 3/5 changes its outcome from "skipped silently" to #UD there too,
|
|
which is what VMCALL outside VMX operation does on bare metal. The
|
|
series has been tested on AMD only (Zen 4, 7950X); the Intel reasoning
|
|
is from the code and has not been run on Intel hardware.
|
|
|
|
Nested: for 1/5, gp_interception() already routes decoded SVM
|
|
instructions from L2 through nested_svm_check_permissions(); when L1
|
|
intercepts #GP itself the exit is reflected to L1 as before. For 4/5
|
|
the #DB is reflected to L1 with L2's RIP past the ICEBP, matching what
|
|
the VMX path does today.
|
|
|
|
SEV: 1/5 does not arm the #GP intercept for SEV guests, and SEV-ES
|
|
guests additionally clear it in sev_es_init_vmcb(). 4/5 uses
|
|
svm_skip_emulated_instruction(), which already special-cases SEV-ES.
|
|
|
|
Testing: a Windows 11 guest, a Linux guest and a Hyper-V-enlightened
|
|
Windows guest boot and run normally on the patched modules. The
|
|
userspace probes that first exposed the divergences (CPL3 VMLOAD, CPL3
|
|
VMMCALL, ICEBP under SEH) now observe #UD / #UD / RIP+1, as on bare
|
|
metal, and the ICEBP one no longer regresses when the delivery hits a
|
|
cold page. 5/5 fails on unpatched SVM (Zen 4, kvm_exit trace shows the
|
|
NPT fault with the #DB pending in EXITINTINFO) and passes with 4/5. A
|
|
selftest for the CPL > 0 cases would need user-mode guest code, which
|
|
the x86 selftest library does not currently provide.
|
|
|
|
arch/x86/kvm/emulate.c | 7 ++-
|
|
arch/x86/kvm/svm/svm.c | 44 ++++++++++++---
|
|
arch/x86/kvm/x86.c | 17 +++++-
|
|
tools/testing/selftests/kvm/Makefile.kvm | 1 +
|
|
tools/testing/selftests/kvm/x86/icebp_test.c | 58 ++++++++++++++++++++
|
|
5 files changed, 117 insertions(+), 10 deletions(-)
|