16 KiB
vfio-native
Make a KVM guest present a hardware profile that is self-consistent with real silicon and real firmware, and measure the result against an independent library.
A lot of software behaves differently, or refuses to run at all, once it can tell it is inside a hypervisor. Malware samples self-disable or take a benign branch. Commercial software takes a licensing branch. Drivers and firmware take a different code path. Any of those makes a virtual machine useless for observing what the software would actually do on a physical machine - you end up measuring the sandbox instead of the subject.
Two independent problems produce that gap, and this addresses both.
KVM diverges from the AMD64 architecture in ways a guest can read directly. The wrong
exception vector for an SVM instruction at CPL>0. #GP where the architecture specifies #UD. A
hypercall at CPL>0 that raises no exception at all. Those are spec-conformance defects, and one of
them closes a FIXME the kernel already carries against itself. The three KVM patches make KVM
match the architecture and nothing more, which is why they are written to be acceptable upstream.
QEMU's emulated platform is not internally consistent with any board that was ever manufactured. ACPI link-device names no real firmware uses, PCI vendor and subsystem IDs that belong to the emulator rather than to a device, firmware table fields no shipping board sets, device identity strings that name the emulator. Correcting them makes the emulated platform look like a platform.
Alongside both there is a VFIO tuning layer - vCPU pinning onto real SMT pairs, sizing the guest to one cache domain, keeping the emulator thread off the vCPU cores, host CPU isolation while the guest runs. It is independent of the fidelity work and measured separately.
Scored against VMAware, which runs 85 detection techniques. A stock KVM guest starts at 10/85. This gets it to 1/85 on both the v2.8.1 release and current HEAD, and the one survivor needs a real GPU passed through. VMAware's own verdict:
VM brand: Unknown
VM likeliness: 20%
VM confirmation: false
===== CONCLUSION: Running on bare metal =====
The corrections cost no measurable performance. Measured against the same guest with the fidelity work abandoned entirely and every Hyper-V enlightenment enabled, and again with the patched modules against stock, the difference is smaller than the spread between runs. See docs/RESULTS.md.
Run this only on hardware you control and with software you are licensed to run.
How to run it
1. Install
Three packages, all built to AUR rules and proven in a clean extra-x86_64-build chroot. They
fetch this repo at tag v1.1.0, so from a checkout the same three makepkg -si work:
git clone https://git.archworks.co/sandwich/vfio-native
cd vfio-native/packaging
(cd vfio-native && makepkg -si)
(cd vfio-native-qemu && makepkg -si)
(cd vfio-native-kvm-dkms && makepkg -si)
| Package | What it is |
|---|---|
vfio-native |
the three commands, the patches, the ACPI tables, the benchmark, the libvirt hook |
vfio-native-qemu |
QEMU 11.1.1 with the platform-identity patches, in /opt/qemu-native |
vfio-native-kvm-dkms |
the patched KVM modules, rebuilt by DKMS for every installed 7.2.x kernel |
The modules land in updates/dkms/, which modprobe prefers, but nothing reloads them for you.
With every VM shut down:
sudo modprobe -r kvm_amd kvm && sudo modprobe kvm_amd
That gives you three commands. The post-install prints the whole flow, so you do not need to read any further to use it:
| Command | |
|---|---|
vm-native-setup |
configure a libvirt domain |
vm-native-verify |
measure whether it actually worked |
vm-native-gpu |
set up GPU passthrough |
Only full needs the QEMU and KVM packages. native is domain XML only.
Not an Arch user? The scripts in scripts/ are plain shell and work standalone, and the module
sources build as arch/x86/kvm against any 7.2.x tree; scripts/install-modules.sh is the manual
path. The packaging is convenience, not a dependency.
2. Fix the guest, once
docs/GUEST-SETUP.md is the full, ordered walkthrough from a plain Windows VM, including the two steps that have to happen before the disk moves to NVMe. The short version:
A Windows guest that runs its own hypervisor reports that truthfully, and no host-side setting changes it - the guest really is running Hyper-V. Turning it off also makes the VM faster, because VBS costs 5-15% on CPU-bound workloads.
bcdedit /set hypervisorlaunchtype off
Disable-WindowsOptionalFeature -Online -FeatureName Microsoft-Hyper-V-All -NoRestart
reg add "HKLM\SYSTEM\CurrentControlSet\Control\DeviceGuard" /v EnableVirtualizationBasedSecurity /t REG_DWORD /d 0 /f
Reboot twice. Uninstall the QEMU and SPICE guest agents. Move the disk to emulated NVMe
(bus='nvme' with a <serial>) - the patched QEMU will not boot a virtio disk.
3. Fix the host, once
sudo cpupower frequency-set -g performance
4. Configure the domain
virsh -c qemu:///system shutdown win11
vm-native-setup
It is an interview with a default for every answer: level, cores, RAM, identity, USB, disk, Secure Boot. It reads your CPU layout itself - AMD CCDs or
Intel P/E cores - sizes the guest to one cache domain, pins each vCPU onto a real SMT pair, keeps
the emulator off the vCPU cores, moves the disks to emulated NVMe, replaces the virtio device set,
wires Secure Boot with a key store it generates, writes the SMBIOS and ACPI identity, and gives
the guest a CPU identity from the host's own generation whose thread count matches what it
actually has. -r gives the deployment its own serials and MAC; -u auto passes keyboard and
mouse through.
Every flag answers one question in advance and -y takes every default. Scripted:
vm-native-setup -d win11 -p full -c 8 -m 16 -r -u auto -y
| Flag | |
|---|---|
-d |
domain |
-p |
tuned, native or full (default full) |
-c |
guest cores; SMT doubles this into vCPUs |
-m |
guest RAM in GiB |
-s |
Secure Boot with enrolled keys, on (default) or off |
-u |
USB passthrough: none, auto, or vid:pid,...,0000:bb:dd.f |
-r |
randomise the hardware identity: serials, MAC, memory module |
-y |
no prompts |
It backs the domain up first and prints the revert command. Re-running it is safe.
5. Check it worked
vm-native-verify
OK QPC cost (ns) 14.6
OK rdtsc cost (ns) 6.6
OK 1-thread (Mops) 5254.2
OK L3 latency (ns) 10.94
OK jitter p99.99 (us) 1.600
OK stalls >100us 0
vm-native-verify also prints whether CPUID passthrough is on, which is the last step.
6. Switch on CPUID passthrough, after every guest boot
The TIMER check times the world switch on an intercepted CPUID, and the only way to stop
paying it is to not exit. vm-native-setup prints the two lines for your declared SKU; on this
host they are:
echo 'AMD Ryzen 7 7700X 8-Core Processor' | sudo tee /sys/module/kvm_amd/parameters/brand_string
echo Y | sudo tee /sys/module/kvm_amd/parameters/cpuid_passthrough
Run them once the guest is up, never before: a booting Windows enumerates CPUID bits KVM
synthesises, and hangs if they vanish half way through. Switch it off again (echo N) before the
next boot. The module applies it only to vCPU threads pinned to exactly one host CPU, so on an
unpinned guest it does nothing rather than something wrong.
It is a module parameter, so it is one switch for the whole host: every pinned guest gets it, and a cold boot of any of them while it is on hits the race. With more than one such guest, switch it off before any of them boots and on again once they are all up.
For the detection score, run VMAware in the guest from the console session, not over SSH - OpenSSH lands you in session 0, which is not where an interactive desktop session runs. See docs/TESTING.md.
The three levels
All get identical performance tuning. The level only changes how much of the platform is corrected.
| Level | Score | Needs | Upkeep |
|---|---|---|---|
tuned |
13/85 | nothing | none |
native |
7/85 | nothing | none |
full |
1/85 | vfio-native-qemu + vfio-native-kvm-dkms, passthrough switched on after boot |
DKMS rebuilds on kernel updates inside 7.2.x |
tuned buys the pinning and the cache-domain sizing and nothing else.
native is domain XML only, so it survives any host update untouched. Take it if you would rather
not maintain anything.
The one thing to get right
Leave the host at least 4 physical cores.
Windows calibrates the TSC at boot, and that calibration is a race. If the host cannot schedule
the guest's vCPU threads cleanly through it, Windows abandons the TSC and
QueryPerformanceCounter costs about 1300 ns instead of 15 for the rest of that boot.
Software that polls the clock in a tight loop calls QPC thousands of times a second.
Measured over four cold boots at each size, same XML, on a 16-core host:
| Guest | Host keeps | Boots with a fast clock |
|---|---|---|
| 24 vCPU | 4 cores | 4 / 4 |
| 28 vCPU | 2 cores | 3 / 4 |
| 32 vCPU | 0 cores | 2 / 4 |
It is headroom, not a vCPU ceiling. 24 of 32 threads is reliable, which matters when the guest
does CPU-heavy work rather than latency-sensitive work alone. vm-native-setup warns if you go
below the margin.
Because it is a race, one measurement proves nothing. If vm-native-verify reports QPC over
1000 ns, reboot and measure again before changing anything.
Sizing the guest inside one cache domain is a separate trade: it takes L3 latency from about 17 ns to 10, which is worth it for latency-sensitive work and not for throughput.
None of these change the outcome, measured: useplatformclock false, useplatformtick no,
disabledynamictick yes, adding invtsc, pinning without leaving headroom, or which cache
domain the vCPUs sit on.
Do not put every vCPU on a realtime scheduler.
<vcpusched scheduler='fifo'>across a guest sized to the whole machine booted once, failed twice, ran away to 800% CPU and leftsystemctlunresponsive until QEMU was killed by hand. Realtime priority on as many threads as the host has cores starves everything else, including the emulator thread the guest needs.
The patches, by detection
Every patch is grouped by the check it clears, so any of them can be taken or left. patches/README.md is the map.
KVM - five patches for upstream and one that is not, four detections. They correct KVM
against the architecture rather than adding a layer on top of it, pass checkpatch.pl --strict,
and ship with a cover letter and a selftest in patches/kvm/.
| Patch | Clears |
|---|---|
0001-KVM-SVM-intercept-GP-when-guest-EFER.SVME-is-clear |
SVM_EXCEPTIONS |
0002-KVM-x86-emulator-UD-not-GP-for-VMCALL-at-CPL-0 |
KVM_INTERCEPTION |
0003-KVM-x86-UD-for-KVM-hypercalls-issued-at-CPL-0 |
KVM_INTERCEPTION |
0004-KVM-SVM-intercept-ICEBP-and-skip-it-before-injecting-its-DB |
DBVM |
0005-KVM-selftests-verify-ICEBP-DB-reports-RIP-past-the-ICEBP |
the test for 0004 |
EXPERIMENTAL-0006-runtime-cpuid-passthrough |
TIMER |
0002 and 0003 are a pair - the check tries two stubs and reports whichever misbehaves first, so
either alone leaves it standing. EXPERIMENTAL-0006 is built into the DKMS package but off by
default; it is not upstream material, because it hands the guest raw host CPUID.
QEMU - eight patchsets against v11.1.1: 01-firmware, 02-disk-identity, 03-pci-ids,
04-fw-cfg, 05-usb-hid, 06-audio, and 07-edid and 08-cpu-misc, which clear no check by
themselves and are carried for other detectors.
GPU passthrough
vm-native-gpu
Lists your GPUs with IOMMU groups, says which drives a display, and picks the path.
One GPU - the normal case. The host gives the card up while the guest runs and takes it back after:
sudo vm-native-gpu --single win11 0000:03:00.0
Installs a libvirt hook scoped to that domain only. The host has no display while the guest runs, so get SSH working first.
Two GPUs, one spare - easier and safer. Binds it to vfio-pci at boot so the host never claims
it:
sudo vm-native-gpu --dual 0000:0f:00.0
Undo either with sudo vm-native-gpu --revert.
When the guest is off, the host is whole
No configuration here reserves host resources while the guest is not running. Static hugepages are deliberately not used, because they would take memory permanently.
The CPU-isolation hook confines the host to the cores the guest is not using, and restores on
every exit path - clean shutdown, hard destroy and crash, all verified. The confinement is
--runtime only, so a reboot clears it regardless. It lives in /etc/libvirt/hooks/qemu.d/,
which libvirt runs after any qemu hook you already have, so it does not replace one.
Maintenance
| Event | tuned / native |
full |
|---|---|---|
| Kernel update inside 7.2.x | nothing | DKMS rebuilds the modules; reload or reboot |
| Kernel update past 7.2 | nothing | the modules stop building - bump _kver in the DKMS package |
| QEMU package update | nothing | nothing, /opt/qemu-native is its own package |
| Newer QEMU base | nothing | re-apply patches/qemu/ |
The DKMS tree carries 7.2.3's KVM sources. They build against any 7.2.x headers, and DKMS refuses anything else, so after a kernel upgrade past 7.2 the stock modules load silently, the guest boots fine and scores worse, with nothing to tell you but a pacman hook message:
cat /sys/module/kvm/srcversion /sys/module/kvm_amd/srcversion
modinfo -F srcversion kvm kvm_amd
Differ? Check dkms status. Check both modules - the SVM fixes land in kvm-amd.ko and the
hypercall fixes in kvm.ko, so verifying one reports success on a stale build of the other.
The last two checks
GPU_CAPABILITIES. Wants a display reporting a gamma ramp, which means a real GPU passed
through. Not a GPU enumeration despite the name - it is one GetDeviceCaps call. Out of scope
for an emulated display; passing a real GPU through with vm-native-gpu clears it.
TIMER - cleared, but with a workaround. Two detectors that OR together. The exception-latency
one was first measured as a hardware floor at ratio 6.2; that was a broken probe - rebuilt to
VMAware's own mechanism it reads 1.4 against a threshold of 2.5 on stock KVM, so it was never real.
The instruction-latency one is the AMD world switch on an intercepted CPUID. EXPERIMENTAL-0006
clears it, but as a workaround rather than a spec fix: it stops intercepting CPUID on a 1:1-pinned
vCPU and reprograms the brand-string MSRs per core so raw CPUID still names the declared SKU, with
RDPRU passed through because raw CPUID advertises it. Opt-in, runtime-only, and constrained to
pinned vCPUs for that reason. Window 2080 -> 199 ticks, ratio 7.9 -> 0.76. docs/ANALYSIS.md has
the whole trail, including the two wrong turns.
Layout
patches/kvm/ five upstream-style kernel patches, a cover letter, plus one experimental
patches/qemu/ eight patchsets against QEMU 11.1.1
packaging/ three AUR-shaped packages: tools, patched QEMU, DKMS modules
scripts/ setup, verify, GPU, libvirt hook, module install
scripts/generate-tables.py raw SMBIOS type 7/26/27/28/29 blobs for -smbios file=
bench/ the benchmark and the TIMER probe, C, mingw-w64
acpi/ SSDT sources for the battery, platform devices and sensor probes
docs/ANALYSIS.md why each check fires, with source references
docs/RESULTS.md every measurement
docs/TESTING.md how to measure it yourself
Longer write-up: archworks.co/docs/vfio-native
Credits
Written against VMAware by NotRequiem, which is the honest way to score this - it is the independent library, not a checklist that grades itself.
GPL-2.0, matching the kernel and QEMU patches.