Files
vfio-native/README.md

16 KiB

vfio-native

Make a KVM guest present a hardware profile that is self-consistent with real silicon and real firmware, and measure the result against an independent library.

A lot of software behaves differently, or refuses to run at all, once it can tell it is inside a hypervisor. Malware samples self-disable or take a benign branch. Commercial software takes a licensing branch. Drivers and firmware take a different code path. Any of those makes a virtual machine useless for observing what the software would actually do on a physical machine - you end up measuring the sandbox instead of the subject.

Two independent problems produce that gap, and this addresses both.

KVM diverges from the AMD64 architecture in ways a guest can read directly. The wrong exception vector for an SVM instruction at CPL>0. #GP where the architecture specifies #UD. A hypercall at CPL>0 that raises no exception at all. Those are spec-conformance defects, and one of them closes a FIXME the kernel already carries against itself. The three KVM patches make KVM match the architecture and nothing more, which is why they are written to be acceptable upstream.

QEMU's emulated platform is not internally consistent with any board that was ever manufactured. ACPI link-device names no real firmware uses, PCI vendor and subsystem IDs that belong to the emulator rather than to a device, firmware table fields no shipping board sets, device identity strings that name the emulator. Correcting them makes the emulated platform look like a platform.

Alongside both there is a VFIO tuning layer - vCPU pinning onto real SMT pairs, sizing the guest to one cache domain, keeping the emulator thread off the vCPU cores, host CPU isolation while the guest runs. It is independent of the fidelity work and measured separately.

Scored against VMAware, which runs 85 detection techniques. A stock KVM guest starts at 10/85. This gets it to 1/85 on both the v2.8.1 release and current HEAD, and the one survivor needs a real GPU passed through. VMAware's own verdict:

VM brand: Unknown
VM likeliness: 20%
VM confirmation: false
===== CONCLUSION: Running on bare metal =====

The corrections cost no measurable performance. Measured against the same guest with the fidelity work abandoned entirely and every Hyper-V enlightenment enabled, and again with the patched modules against stock, the difference is smaller than the spread between runs. See docs/RESULTS.md.

Run this only on hardware you control and with software you are licensed to run.


How to run it

1. Install

Three packages, all built to AUR rules and proven in a clean extra-x86_64-build chroot. They fetch this repo at tag v1.1.0, so from a checkout the same three makepkg -si work:

git clone https://git.archworks.co/sandwich/vfio-native
cd vfio-native/packaging
(cd vfio-native && makepkg -si)
(cd vfio-native-qemu && makepkg -si)
(cd vfio-native-kvm-dkms && makepkg -si)
Package What it is
vfio-native the three commands, the patches, the ACPI tables, the benchmark, the libvirt hook
vfio-native-qemu QEMU 11.1.1 with the platform-identity patches, in /opt/qemu-native
vfio-native-kvm-dkms the patched KVM modules, rebuilt by DKMS for every installed 7.2.x kernel

The modules land in updates/dkms/, which modprobe prefers, but nothing reloads them for you. With every VM shut down:

sudo modprobe -r kvm_amd kvm && sudo modprobe kvm_amd

That gives you three commands. The post-install prints the whole flow, so you do not need to read any further to use it:

Command
vm-native-setup configure a libvirt domain
vm-native-verify measure whether it actually worked
vm-native-gpu set up GPU passthrough

Only full needs the QEMU and KVM packages. native is domain XML only.

Not an Arch user? The scripts in scripts/ are plain shell and work standalone, and the module sources build as arch/x86/kvm against any 7.2.x tree; scripts/install-modules.sh is the manual path. The packaging is convenience, not a dependency.

2. Fix the guest, once

docs/GUEST-SETUP.md is the full, ordered walkthrough from a plain Windows VM, including the two steps that have to happen before the disk moves to NVMe. The short version:

A Windows guest that runs its own hypervisor reports that truthfully, and no host-side setting changes it - the guest really is running Hyper-V. Turning it off also makes the VM faster, because VBS costs 5-15% on CPU-bound workloads.

bcdedit /set hypervisorlaunchtype off
Disable-WindowsOptionalFeature -Online -FeatureName Microsoft-Hyper-V-All -NoRestart
reg add "HKLM\SYSTEM\CurrentControlSet\Control\DeviceGuard" /v EnableVirtualizationBasedSecurity /t REG_DWORD /d 0 /f

Reboot twice. Uninstall the QEMU and SPICE guest agents. Move the disk to emulated NVMe (bus='nvme' with a <serial>) - the patched QEMU will not boot a virtio disk.

3. Fix the host, once

sudo cpupower frequency-set -g performance

4. Configure the domain

virsh -c qemu:///system shutdown win11
vm-native-setup

It is an interview with a default for every answer: level, cores, RAM, identity, USB, disk, Secure Boot. It reads your CPU layout itself - AMD CCDs or Intel P/E cores - sizes the guest to one cache domain, pins each vCPU onto a real SMT pair, keeps the emulator off the vCPU cores, moves the disks to emulated NVMe, replaces the virtio device set, wires Secure Boot with a key store it generates, writes the SMBIOS and ACPI identity, and gives the guest a CPU identity from the host's own generation whose thread count matches what it actually has. -r gives the deployment its own serials and MAC; -u auto passes keyboard and mouse through.

Every flag answers one question in advance and -y takes every default. Scripted:

vm-native-setup -d win11 -p full -c 8 -m 16 -r -u auto -y
Flag
-d domain
-p tuned, native or full (default full)
-c guest cores; SMT doubles this into vCPUs
-m guest RAM in GiB
-s Secure Boot with enrolled keys, on (default) or off
-u USB passthrough: none, auto, or vid:pid,...,0000:bb:dd.f
-r randomise the hardware identity: serials, MAC, memory module
-y no prompts

It backs the domain up first and prints the revert command. Re-running it is safe.

5. Check it worked

vm-native-verify
  OK  QPC cost (ns)          14.6
  OK  rdtsc cost (ns)        6.6
  OK  1-thread (Mops)        5254.2
  OK  L3 latency (ns)        10.94
  OK  jitter p99.99 (us)     1.600
  OK  stalls >100us          0

vm-native-verify also prints whether CPUID passthrough is on, which is the last step.

6. Switch on CPUID passthrough, after every guest boot

The TIMER check times the world switch on an intercepted CPUID, and the only way to stop paying it is to not exit. vm-native-setup prints the two lines for your declared SKU; on this host they are:

echo 'AMD Ryzen 7 7700X 8-Core Processor' | sudo tee /sys/module/kvm_amd/parameters/brand_string
echo Y | sudo tee /sys/module/kvm_amd/parameters/cpuid_passthrough

Run them once the guest is up, never before: a booting Windows enumerates CPUID bits KVM synthesises, and hangs if they vanish half way through. Switch it off again (echo N) before the next boot. The module applies it only to vCPU threads pinned to exactly one host CPU, so on an unpinned guest it does nothing rather than something wrong.

It is a module parameter, so it is one switch for the whole host: every pinned guest gets it, and a cold boot of any of them while it is on hits the race. With more than one such guest, switch it off before any of them boots and on again once they are all up.

For the detection score, run VMAware in the guest from the console session, not over SSH - OpenSSH lands you in session 0, which is not where an interactive desktop session runs. See docs/TESTING.md.


The three levels

All get identical performance tuning. The level only changes how much of the platform is corrected.

Level Score Needs Upkeep
tuned 13/85 nothing none
native 7/85 nothing none
full 1/85 vfio-native-qemu + vfio-native-kvm-dkms, passthrough switched on after boot DKMS rebuilds on kernel updates inside 7.2.x

tuned buys the pinning and the cache-domain sizing and nothing else.

native is domain XML only, so it survives any host update untouched. Take it if you would rather not maintain anything.


The one thing to get right

Leave the host at least 4 physical cores.

Windows calibrates the TSC at boot, and that calibration is a race. If the host cannot schedule the guest's vCPU threads cleanly through it, Windows abandons the TSC and QueryPerformanceCounter costs about 1300 ns instead of 15 for the rest of that boot. Software that polls the clock in a tight loop calls QPC thousands of times a second.

Measured over four cold boots at each size, same XML, on a 16-core host:

Guest Host keeps Boots with a fast clock
24 vCPU 4 cores 4 / 4
28 vCPU 2 cores 3 / 4
32 vCPU 0 cores 2 / 4

It is headroom, not a vCPU ceiling. 24 of 32 threads is reliable, which matters when the guest does CPU-heavy work rather than latency-sensitive work alone. vm-native-setup warns if you go below the margin.

Because it is a race, one measurement proves nothing. If vm-native-verify reports QPC over 1000 ns, reboot and measure again before changing anything.

Sizing the guest inside one cache domain is a separate trade: it takes L3 latency from about 17 ns to 10, which is worth it for latency-sensitive work and not for throughput.

None of these change the outcome, measured: useplatformclock false, useplatformtick no, disabledynamictick yes, adding invtsc, pinning without leaving headroom, or which cache domain the vCPUs sit on.

Do not put every vCPU on a realtime scheduler. <vcpusched scheduler='fifo'> across a guest sized to the whole machine booted once, failed twice, ran away to 800% CPU and left systemctl unresponsive until QEMU was killed by hand. Realtime priority on as many threads as the host has cores starves everything else, including the emulator thread the guest needs.

The patches, by detection

Every patch is grouped by the check it clears, so any of them can be taken or left. patches/README.md is the map.

KVM - five patches for upstream and one that is not, four detections. They correct KVM against the architecture rather than adding a layer on top of it, pass checkpatch.pl --strict, and ship with a cover letter and a selftest in patches/kvm/.

Patch Clears
0001-KVM-SVM-intercept-GP-when-guest-EFER.SVME-is-clear SVM_EXCEPTIONS
0002-KVM-x86-emulator-UD-not-GP-for-VMCALL-at-CPL-0 KVM_INTERCEPTION
0003-KVM-x86-UD-for-KVM-hypercalls-issued-at-CPL-0 KVM_INTERCEPTION
0004-KVM-SVM-intercept-ICEBP-and-skip-it-before-injecting-its-DB DBVM
0005-KVM-selftests-verify-ICEBP-DB-reports-RIP-past-the-ICEBP the test for 0004
EXPERIMENTAL-0006-runtime-cpuid-passthrough TIMER

0002 and 0003 are a pair - the check tries two stubs and reports whichever misbehaves first, so either alone leaves it standing. EXPERIMENTAL-0006 is built into the DKMS package but off by default; it is not upstream material, because it hands the guest raw host CPUID.

QEMU - eight patchsets against v11.1.1: 01-firmware, 02-disk-identity, 03-pci-ids, 04-fw-cfg, 05-usb-hid, 06-audio, and 07-edid and 08-cpu-misc, which clear no check by themselves and are carried for other detectors.


GPU passthrough

vm-native-gpu

Lists your GPUs with IOMMU groups, says which drives a display, and picks the path.

One GPU - the normal case. The host gives the card up while the guest runs and takes it back after:

sudo vm-native-gpu --single win11 0000:03:00.0

Installs a libvirt hook scoped to that domain only. The host has no display while the guest runs, so get SSH working first.

Two GPUs, one spare - easier and safer. Binds it to vfio-pci at boot so the host never claims it:

sudo vm-native-gpu --dual 0000:0f:00.0

Undo either with sudo vm-native-gpu --revert.


When the guest is off, the host is whole

No configuration here reserves host resources while the guest is not running. Static hugepages are deliberately not used, because they would take memory permanently.

The CPU-isolation hook confines the host to the cores the guest is not using, and restores on every exit path - clean shutdown, hard destroy and crash, all verified. The confinement is --runtime only, so a reboot clears it regardless. It lives in /etc/libvirt/hooks/qemu.d/, which libvirt runs after any qemu hook you already have, so it does not replace one.


Maintenance

Event tuned / native full
Kernel update inside 7.2.x nothing DKMS rebuilds the modules; reload or reboot
Kernel update past 7.2 nothing the modules stop building - bump _kver in the DKMS package
QEMU package update nothing nothing, /opt/qemu-native is its own package
Newer QEMU base nothing re-apply patches/qemu/

The DKMS tree carries 7.2.3's KVM sources. They build against any 7.2.x headers, and DKMS refuses anything else, so after a kernel upgrade past 7.2 the stock modules load silently, the guest boots fine and scores worse, with nothing to tell you but a pacman hook message:

cat /sys/module/kvm/srcversion /sys/module/kvm_amd/srcversion
modinfo -F srcversion kvm kvm_amd

Differ? Check dkms status. Check both modules - the SVM fixes land in kvm-amd.ko and the hypercall fixes in kvm.ko, so verifying one reports success on a stale build of the other.


The last two checks

GPU_CAPABILITIES. Wants a display reporting a gamma ramp, which means a real GPU passed through. Not a GPU enumeration despite the name - it is one GetDeviceCaps call. Out of scope for an emulated display; passing a real GPU through with vm-native-gpu clears it.

TIMER - cleared, but with a workaround. Two detectors that OR together. The exception-latency one was first measured as a hardware floor at ratio 6.2; that was a broken probe - rebuilt to VMAware's own mechanism it reads 1.4 against a threshold of 2.5 on stock KVM, so it was never real. The instruction-latency one is the AMD world switch on an intercepted CPUID. EXPERIMENTAL-0006 clears it, but as a workaround rather than a spec fix: it stops intercepting CPUID on a 1:1-pinned vCPU and reprograms the brand-string MSRs per core so raw CPUID still names the declared SKU, with RDPRU passed through because raw CPUID advertises it. Opt-in, runtime-only, and constrained to pinned vCPUs for that reason. Window 2080 -> 199 ticks, ratio 7.9 -> 0.76. docs/ANALYSIS.md has the whole trail, including the two wrong turns.


Layout

patches/kvm/          five upstream-style kernel patches, a cover letter, plus one experimental
patches/qemu/         eight patchsets against QEMU 11.1.1
packaging/            three AUR-shaped packages: tools, patched QEMU, DKMS modules
scripts/              setup, verify, GPU, libvirt hook, module install
scripts/generate-tables.py   raw SMBIOS type 7/26/27/28/29 blobs for -smbios file=
bench/                the benchmark and the TIMER probe, C, mingw-w64
acpi/                 SSDT sources for the battery, platform devices and sensor probes
docs/ANALYSIS.md      why each check fires, with source references
docs/RESULTS.md       every measurement
docs/TESTING.md       how to measure it yourself

Longer write-up: archworks.co/docs/vfio-native


Credits

Written against VMAware by NotRequiem, which is the honest way to score this - it is the independent library, not a checklist that grades itself.

GPL-2.0, matching the kernel and QEMU patches.