docs: readme, analysis, results, testing and guest setup
This commit is contained in:
382
README.md
Normal file
382
README.md
Normal file
@@ -0,0 +1,382 @@
|
||||
# vfio-native
|
||||
|
||||
Make a KVM guest present a hardware profile that is self-consistent with real silicon and real
|
||||
firmware, and measure the result against an independent library.
|
||||
|
||||
A lot of software behaves differently, or refuses to run at all, once it can tell it is inside a
|
||||
hypervisor. Malware samples self-disable or take a benign branch. Commercial software takes a
|
||||
licensing branch. Drivers and firmware take a different code path. Any of those makes a virtual
|
||||
machine useless for observing what the software would actually do on a physical machine - you end
|
||||
up measuring the sandbox instead of the subject.
|
||||
|
||||
Two independent problems produce that gap, and this addresses both.
|
||||
|
||||
**KVM diverges from the AMD64 architecture** in ways a guest can read directly. The wrong
|
||||
exception vector for an SVM instruction at CPL>0. `#GP` where the architecture specifies `#UD`. A
|
||||
hypercall at CPL>0 that raises no exception at all. Those are spec-conformance defects, and one of
|
||||
them closes a `FIXME` the kernel already carries against itself. The three KVM patches make KVM
|
||||
match the architecture and nothing more, which is why they are written to be acceptable upstream.
|
||||
|
||||
**QEMU's emulated platform is not internally consistent** with any board that was ever
|
||||
manufactured. ACPI link-device names no real firmware uses, PCI vendor and subsystem IDs that
|
||||
belong to the emulator rather than to a device, firmware table fields no shipping board sets,
|
||||
device identity strings that name the emulator. Correcting them makes the emulated platform look
|
||||
like a platform.
|
||||
|
||||
Alongside both there is a VFIO tuning layer - vCPU pinning onto real SMT pairs, sizing the guest to
|
||||
one cache domain, keeping the emulator thread off the vCPU cores, host CPU isolation while the
|
||||
guest runs. It is independent of the fidelity work and measured separately.
|
||||
|
||||
Scored against [VMAware](https://github.com/NotRequiem/VMAware), which runs 85 detection
|
||||
techniques. A stock KVM guest starts at 10/85. This gets it to **1/85** on both the v2.8.1 release
|
||||
and current HEAD, and the one survivor needs a real GPU passed through. VMAware's own verdict:
|
||||
|
||||
```
|
||||
VM brand: Unknown
|
||||
VM likeliness: 20%
|
||||
VM confirmation: false
|
||||
===== CONCLUSION: Running on bare metal =====
|
||||
```
|
||||
|
||||
**The corrections cost no measurable performance.** Measured against the same guest with the
|
||||
fidelity work abandoned entirely and every Hyper-V enlightenment enabled, and again with the
|
||||
patched modules against stock, the difference is smaller than the spread between runs. See
|
||||
[docs/RESULTS.md](docs/RESULTS.md).
|
||||
|
||||
Run this only on hardware you control and with software you are licensed to run.
|
||||
|
||||
---
|
||||
|
||||
## How to run it
|
||||
|
||||
### 1. Install
|
||||
|
||||
Three packages, all built to AUR rules and proven in a clean `extra-x86_64-build` chroot. They
|
||||
fetch this repo at tag `v1.1.0`, so from a checkout the same three `makepkg -si` work:
|
||||
|
||||
```sh
|
||||
git clone https://git.archworks.co/sandwich/vfio-native
|
||||
cd vfio-native/packaging
|
||||
(cd vfio-native && makepkg -si)
|
||||
(cd vfio-native-qemu && makepkg -si)
|
||||
(cd vfio-native-kvm-dkms && makepkg -si)
|
||||
```
|
||||
|
||||
| Package | What it is |
|
||||
| --- | --- |
|
||||
| `vfio-native` | the three commands, the patches, the ACPI tables, the benchmark, the libvirt hook |
|
||||
| `vfio-native-qemu` | QEMU 11.1.1 with the platform-identity patches, in `/opt/qemu-native` |
|
||||
| `vfio-native-kvm-dkms` | the patched KVM modules, rebuilt by DKMS for every installed 7.2.x kernel |
|
||||
|
||||
The modules land in `updates/dkms/`, which modprobe prefers, but nothing reloads them for you.
|
||||
With every VM shut down:
|
||||
|
||||
```sh
|
||||
sudo modprobe -r kvm_amd kvm && sudo modprobe kvm_amd
|
||||
```
|
||||
|
||||
That gives you three commands. The post-install prints the whole flow, so you do not need to read
|
||||
any further to use it:
|
||||
|
||||
| Command | |
|
||||
| --- | --- |
|
||||
| `vm-native-setup` | configure a libvirt domain |
|
||||
| `vm-native-verify` | measure whether it actually worked |
|
||||
| `vm-native-gpu` | set up GPU passthrough |
|
||||
|
||||
Only `full` needs the QEMU and KVM packages. `native` is domain XML only.
|
||||
|
||||
Not an Arch user? The scripts in `scripts/` are plain shell and work standalone, and the module
|
||||
sources build as `arch/x86/kvm` against any 7.2.x tree; `scripts/install-modules.sh` is the manual
|
||||
path. The packaging is convenience, not a dependency.
|
||||
|
||||
### 2. Fix the guest, once
|
||||
|
||||
[docs/GUEST-SETUP.md](docs/GUEST-SETUP.md) is the full, ordered walkthrough from a plain Windows
|
||||
VM, including the two steps that have to happen before the disk moves to NVMe. The short version:
|
||||
|
||||
A Windows guest that runs its own hypervisor reports that truthfully, and no host-side setting
|
||||
changes it - the guest really is running Hyper-V. Turning it off also makes the VM **faster**,
|
||||
because VBS costs 5-15% on CPU-bound workloads.
|
||||
|
||||
```
|
||||
bcdedit /set hypervisorlaunchtype off
|
||||
Disable-WindowsOptionalFeature -Online -FeatureName Microsoft-Hyper-V-All -NoRestart
|
||||
reg add "HKLM\SYSTEM\CurrentControlSet\Control\DeviceGuard" /v EnableVirtualizationBasedSecurity /t REG_DWORD /d 0 /f
|
||||
```
|
||||
|
||||
Reboot twice. Uninstall the QEMU and SPICE guest agents. Move the disk to emulated NVMe
|
||||
(`bus='nvme'` with a `<serial>`) - the patched QEMU will not boot a virtio disk.
|
||||
|
||||
### 3. Fix the host, once
|
||||
|
||||
```sh
|
||||
sudo cpupower frequency-set -g performance
|
||||
```
|
||||
|
||||
### 4. Configure the domain
|
||||
|
||||
```sh
|
||||
virsh -c qemu:///system shutdown win11
|
||||
vm-native-setup
|
||||
```
|
||||
|
||||
It is an interview with a default for every answer: level, cores, RAM, identity, USB, disk, Secure Boot. It reads your CPU layout itself - AMD CCDs or
|
||||
Intel P/E cores - sizes the guest to one cache domain, pins each vCPU onto a real SMT pair, keeps
|
||||
the emulator off the vCPU cores, moves the disks to emulated NVMe, replaces the virtio device set,
|
||||
wires Secure Boot with a key store it generates, writes the SMBIOS and ACPI identity, and gives
|
||||
the guest a CPU identity from the host's own generation whose thread count matches what it
|
||||
actually has. `-r` gives the deployment its own serials and MAC; `-u auto` passes keyboard and
|
||||
mouse through.
|
||||
|
||||
Every flag answers one question in advance and `-y` takes every default. Scripted:
|
||||
|
||||
```sh
|
||||
vm-native-setup -d win11 -p full -c 8 -m 16 -r -u auto -y
|
||||
```
|
||||
|
||||
| Flag | |
|
||||
| --- | --- |
|
||||
| `-d` | domain |
|
||||
| `-p` | `tuned`, `native` or `full` (default `full`) |
|
||||
| `-c` | guest cores; SMT doubles this into vCPUs |
|
||||
| `-m` | guest RAM in GiB |
|
||||
| `-s` | Secure Boot with enrolled keys, `on` (default) or `off` |
|
||||
| `-u` | USB passthrough: `none`, `auto`, or `vid:pid,...,0000:bb:dd.f` |
|
||||
| `-r` | randomise the hardware identity: serials, MAC, memory module |
|
||||
| `-y` | no prompts |
|
||||
|
||||
It backs the domain up first and prints the revert command. Re-running it is safe.
|
||||
|
||||
### 5. Check it worked
|
||||
|
||||
```sh
|
||||
vm-native-verify
|
||||
```
|
||||
|
||||
```
|
||||
OK QPC cost (ns) 14.6
|
||||
OK rdtsc cost (ns) 6.6
|
||||
OK 1-thread (Mops) 5254.2
|
||||
OK L3 latency (ns) 10.94
|
||||
OK jitter p99.99 (us) 1.600
|
||||
OK stalls >100us 0
|
||||
```
|
||||
|
||||
`vm-native-verify` also prints whether CPUID passthrough is on, which is the last step.
|
||||
|
||||
### 6. Switch on CPUID passthrough, after every guest boot
|
||||
|
||||
The `TIMER` check times the world switch on an intercepted `CPUID`, and the only way to stop
|
||||
paying it is to not exit. `vm-native-setup` prints the two lines for your declared SKU; on this
|
||||
host they are:
|
||||
|
||||
```sh
|
||||
echo 'AMD Ryzen 7 7700X 8-Core Processor' | sudo tee /sys/module/kvm_amd/parameters/brand_string
|
||||
echo Y | sudo tee /sys/module/kvm_amd/parameters/cpuid_passthrough
|
||||
```
|
||||
|
||||
Run them once the guest is up, never before: a booting Windows enumerates CPUID bits KVM
|
||||
synthesises, and hangs if they vanish half way through. Switch it off again (`echo N`) before the
|
||||
next boot. The module applies it only to vCPU threads pinned to exactly one host CPU, so on an
|
||||
unpinned guest it does nothing rather than something wrong.
|
||||
|
||||
It is a module parameter, so it is one switch for the whole host: every pinned guest gets it, and
|
||||
a cold boot of any of them while it is on hits the race. With more than one such guest, switch it
|
||||
off before any of them boots and on again once they are all up.
|
||||
|
||||
For the detection score, run VMAware in the guest **from the console session, not over SSH** -
|
||||
OpenSSH lands you in session 0, which is not where an interactive desktop session runs. See
|
||||
[docs/TESTING.md](docs/TESTING.md).
|
||||
|
||||
---
|
||||
|
||||
## The three levels
|
||||
|
||||
All get identical performance tuning. The level only changes how much of the platform is
|
||||
corrected.
|
||||
|
||||
| Level | Score | Needs | Upkeep |
|
||||
| --- | --- | --- | --- |
|
||||
| `tuned` | 13/85 | nothing | none |
|
||||
| `native` | 7/85 | nothing | none |
|
||||
| `full` | **1/85** | `vfio-native-qemu` + `vfio-native-kvm-dkms`, passthrough switched on after boot | DKMS rebuilds on kernel updates inside 7.2.x |
|
||||
|
||||
`tuned` buys the pinning and the cache-domain sizing and nothing else.
|
||||
|
||||
`native` is domain XML only, so it survives any host update untouched. Take it if you would rather
|
||||
not maintain anything.
|
||||
|
||||
---
|
||||
|
||||
## The one thing to get right
|
||||
|
||||
**Leave the host at least 4 physical cores.**
|
||||
|
||||
Windows calibrates the TSC at boot, and that calibration is a race. If the host cannot schedule
|
||||
the guest's vCPU threads cleanly through it, Windows abandons the TSC and
|
||||
`QueryPerformanceCounter` costs about **1300 ns instead of 15** for the rest of that boot.
|
||||
Software that polls the clock in a tight loop calls QPC thousands of times a second.
|
||||
|
||||
Measured over four cold boots at each size, same XML, on a 16-core host:
|
||||
|
||||
| Guest | Host keeps | Boots with a fast clock |
|
||||
| --- | --- | --- |
|
||||
| 24 vCPU | 4 cores | **4 / 4** |
|
||||
| 28 vCPU | 2 cores | 3 / 4 |
|
||||
| 32 vCPU | 0 cores | 2 / 4 |
|
||||
|
||||
It is headroom, not a vCPU ceiling. 24 of 32 threads is reliable, which matters when the guest
|
||||
does CPU-heavy work rather than latency-sensitive work alone. `vm-native-setup` warns if you go
|
||||
below the margin.
|
||||
|
||||
Because it is a race, **one measurement proves nothing**. If `vm-native-verify` reports QPC over
|
||||
1000 ns, reboot and measure again before changing anything.
|
||||
|
||||
Sizing the guest inside one cache domain is a separate trade: it takes L3 latency from about
|
||||
17 ns to 10, which is worth it for latency-sensitive work and not for throughput.
|
||||
|
||||
None of these change the outcome, measured: `useplatformclock false`, `useplatformtick no`,
|
||||
`disabledynamictick yes`, adding `invtsc`, pinning without leaving headroom, or which cache
|
||||
domain the vCPUs sit on.
|
||||
|
||||
> **Do not put every vCPU on a realtime scheduler.** `<vcpusched scheduler='fifo'>` across a
|
||||
> guest sized to the whole machine booted once, failed twice, ran away to 800% CPU and left
|
||||
> `systemctl` unresponsive until QEMU was killed by hand. Realtime priority on as many threads as
|
||||
> the host has cores starves everything else, including the emulator thread the guest needs.
|
||||
|
||||
## The patches, by detection
|
||||
|
||||
Every patch is grouped by the check it clears, so any of them can be taken or left.
|
||||
[patches/README.md](patches/README.md) is the map.
|
||||
|
||||
**KVM** - five patches for upstream and one that is not, four detections. They correct KVM
|
||||
against the architecture rather than adding a layer on top of it, pass `checkpatch.pl --strict`,
|
||||
and ship with a cover letter and a selftest in `patches/kvm/`.
|
||||
|
||||
| Patch | Clears |
|
||||
| --- | --- |
|
||||
| `0001-KVM-SVM-intercept-GP-when-guest-EFER.SVME-is-clear` | `SVM_EXCEPTIONS` |
|
||||
| `0002-KVM-x86-emulator-UD-not-GP-for-VMCALL-at-CPL-0` | `KVM_INTERCEPTION` |
|
||||
| `0003-KVM-x86-UD-for-KVM-hypercalls-issued-at-CPL-0` | `KVM_INTERCEPTION` |
|
||||
| `0004-KVM-SVM-intercept-ICEBP-and-skip-it-before-injecting-its-DB` | `DBVM` |
|
||||
| `0005-KVM-selftests-verify-ICEBP-DB-reports-RIP-past-the-ICEBP` | the test for 0004 |
|
||||
| `EXPERIMENTAL-0006-runtime-cpuid-passthrough` | `TIMER` |
|
||||
|
||||
0002 and 0003 are a pair - the check tries two stubs and reports whichever misbehaves first, so
|
||||
either alone leaves it standing. `EXPERIMENTAL-0006` is built into the DKMS package but off by
|
||||
default; it is not upstream material, because it hands the guest raw host CPUID.
|
||||
|
||||
**QEMU** - eight patchsets against v11.1.1: `01-firmware`, `02-disk-identity`, `03-pci-ids`,
|
||||
`04-fw-cfg`, `05-usb-hid`, `06-audio`, and `07-edid` and `08-cpu-misc`, which clear no check by
|
||||
themselves and are carried for other detectors.
|
||||
|
||||
---
|
||||
|
||||
## GPU passthrough
|
||||
|
||||
```sh
|
||||
vm-native-gpu
|
||||
```
|
||||
|
||||
Lists your GPUs with IOMMU groups, says which drives a display, and picks the path.
|
||||
|
||||
**One GPU** - the normal case. The host gives the card up while the guest runs and takes it back
|
||||
after:
|
||||
|
||||
```sh
|
||||
sudo vm-native-gpu --single win11 0000:03:00.0
|
||||
```
|
||||
|
||||
Installs a libvirt hook scoped to that domain only. The host has no display while the guest runs,
|
||||
so get SSH working first.
|
||||
|
||||
**Two GPUs, one spare** - easier and safer. Binds it to `vfio-pci` at boot so the host never claims
|
||||
it:
|
||||
|
||||
```sh
|
||||
sudo vm-native-gpu --dual 0000:0f:00.0
|
||||
```
|
||||
|
||||
Undo either with `sudo vm-native-gpu --revert`.
|
||||
|
||||
---
|
||||
|
||||
## When the guest is off, the host is whole
|
||||
|
||||
No configuration here reserves host resources while the guest is not running. Static hugepages are
|
||||
deliberately not used, because they would take memory permanently.
|
||||
|
||||
The CPU-isolation hook confines the host to the cores the guest is not using, and restores on
|
||||
every exit path - clean shutdown, hard destroy and crash, all verified. The confinement is
|
||||
`--runtime` only, so a reboot clears it regardless. It lives in `/etc/libvirt/hooks/qemu.d/`,
|
||||
which libvirt runs after any `qemu` hook you already have, so it does not replace one.
|
||||
|
||||
---
|
||||
|
||||
## Maintenance
|
||||
|
||||
| Event | `tuned` / `native` | `full` |
|
||||
| --- | --- | --- |
|
||||
| Kernel update inside 7.2.x | nothing | DKMS rebuilds the modules; reload or reboot |
|
||||
| Kernel update past 7.2 | nothing | **the modules stop building** - bump `_kver` in the DKMS package |
|
||||
| QEMU package update | nothing | nothing, `/opt/qemu-native` is its own package |
|
||||
| Newer QEMU base | nothing | re-apply `patches/qemu/` |
|
||||
|
||||
The DKMS tree carries 7.2.3's KVM sources. They build against any 7.2.x headers, and DKMS refuses
|
||||
anything else, so after a kernel upgrade past 7.2 the stock modules load silently, the guest boots
|
||||
fine and scores worse, with nothing to tell you but a pacman hook message:
|
||||
|
||||
```sh
|
||||
cat /sys/module/kvm/srcversion /sys/module/kvm_amd/srcversion
|
||||
modinfo -F srcversion kvm kvm_amd
|
||||
```
|
||||
|
||||
Differ? Check `dkms status`. Check **both** modules - the SVM fixes land in `kvm-amd.ko` and the
|
||||
hypercall fixes in `kvm.ko`, so verifying one reports success on a stale build of the other.
|
||||
|
||||
---
|
||||
|
||||
## The last two checks
|
||||
|
||||
**`GPU_CAPABILITIES`.** Wants a display reporting a gamma ramp, which means a real GPU passed
|
||||
through. Not a GPU enumeration despite the name - it is one `GetDeviceCaps` call. Out of scope
|
||||
for an emulated display; passing a real GPU through with `vm-native-gpu` clears it.
|
||||
|
||||
**`TIMER` - cleared, but with a workaround.** Two detectors that OR together. The exception-latency
|
||||
one was first measured as a hardware floor at ratio 6.2; that was a broken probe - rebuilt to
|
||||
VMAware's own mechanism it reads 1.4 against a threshold of 2.5 on stock KVM, so it was never real.
|
||||
The instruction-latency one is the AMD world switch on an intercepted `CPUID`. `EXPERIMENTAL-0006`
|
||||
clears it, but as a workaround rather than a spec fix: it stops intercepting `CPUID` on a 1:1-pinned
|
||||
vCPU and reprograms the brand-string MSRs per core so raw `CPUID` still names the declared SKU, with
|
||||
RDPRU passed through because raw CPUID advertises it. Opt-in, runtime-only, and constrained to
|
||||
pinned vCPUs for that reason. Window 2080 -> 199 ticks, ratio 7.9 -> 0.76. `docs/ANALYSIS.md` has
|
||||
the whole trail, including the two wrong turns.
|
||||
|
||||
---
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
patches/kvm/ five upstream-style kernel patches, a cover letter, plus one experimental
|
||||
patches/qemu/ eight patchsets against QEMU 11.1.1
|
||||
packaging/ three AUR-shaped packages: tools, patched QEMU, DKMS modules
|
||||
scripts/ setup, verify, GPU, libvirt hook, module install
|
||||
scripts/generate-tables.py raw SMBIOS type 7/26/27/28/29 blobs for -smbios file=
|
||||
bench/ the benchmark and the TIMER probe, C, mingw-w64
|
||||
acpi/ SSDT sources for the battery, platform devices and sensor probes
|
||||
docs/ANALYSIS.md why each check fires, with source references
|
||||
docs/RESULTS.md every measurement
|
||||
docs/TESTING.md how to measure it yourself
|
||||
```
|
||||
|
||||
Longer write-up: [archworks.co/docs/vfio-native](https://archworks.co/docs/vfio-native/)
|
||||
|
||||
---
|
||||
|
||||
## Credits
|
||||
|
||||
Written against [VMAware](https://github.com/NotRequiem/VMAware) by NotRequiem, which is the
|
||||
honest way to score this - it is the independent library, not a checklist that grades itself.
|
||||
|
||||
GPL-2.0, matching the kernel and QEMU patches.
|
||||
Reference in New Issue
Block a user