12 KiB
Testing methodology
Two separate questions, measured separately: does the platform read as native, and is it fast. A configuration can pass one and fail the other, and most published tuning advice is never measured at all.
Measuring the detection score
The scanner
VMAware runs 85 techniques and reports which fired. Use the debug build, which prints why each one fired - that string is the only efficient way to work through them.
vmaware_debug.exe --all --no-ansi
Grading with this project's own checklist instead is generous: it tests what the project already fixed. VMAware is independent of it, which is the point.
Score against two builds: the latest release (vmaware_debug.exe from the GitHub release page,
v2.8.1 at the time of writing) and a debug build of HEAD. HEAD grows checks between releases -
DBVM arrived after v2.8.1 and fired on this guest while the release build said 2/85. HEAD builds
for Windows with clang (--target=x86_64-w64-mingw32 against the mingw sysroot) with -DVMAWARE_DEBUG;
one file with a frameless SEH leaf needs the gcc assembler, so build that object with
x86_64-w64-mingw32-gcc and link the rest with clang.
Score from the console session, not over SSH
OpenSSH on Windows drops you in session 0, the services session. That is not where an interactive application runs, and it has no real display. Run it as a scheduled task instead:
schtasks /create /tn VMAware /tr "C:\path\run.cmd" /sc once /st 00:00 /it /rl highest /f
schtasks /run /tn VMAware
schtasks /run reports success regardless of where the task actually ran, so confirm from inside
it:
(Get-Process -Id $PID).SessionId # must be 1, not 0
Run as Administrator either way - several checks need it.
This one cost real time here. A plausible theory said GPU_CAPABILITIES was a session-0 artefact,
because a non-interactive window station has no gamma LUT. Re-running it confirmed in session 1
gave the identical score with the check still firing. It was genuine, and the theory was
comfortable rather than correct.
Read the debug line, not the count
[ DETECTED ] with no debug line above it does not mean cleared - some checks return
silently. Conversely, a check firing with an unexpected debug string usually means your model of
it is wrong.
SVM_EXCEPTIONS was diagnosed entirely from the absence of a debug line: both noisy exits in
that function log something, so silence forced the conclusion that it fell through to the bare
return true, which in turn pinned down exactly what the exception must have been.
The guest ignores ACPI shutdown when it feels like it
virsh shutdown is a request, and a Windows guest with a dialog open or an update pending sits
there. Every script here waits, then asks from inside:
ssh User@guest 'shutdown /s /t 0 /f'
and only then touches modules. modprobe -r kvm with the domain still up fails, and a module
swap racing a guest that is still shutting down is how a scan ends up measuring the wrong build.
One change per boot
The ACPI failure mode is a guest that powers itself off after about 70 seconds having read nothing from disk. No error, no log line, nothing in the journal. The only way to attribute it is to have changed exactly one thing since the last known-good boot.
Snapshot before any firmware change:
qemu-img snapshot -c pre-acpi-$(date +%Y%m%d) /path/to/guest.qcow2
Chains report only their first hit
FIRMWARE walks about a dozen fingerprints and returns on the first match, so fixing one link
costs a full rebuild and boot and only reveals the next. Bundle every edit into one rebuild, or you
will spend an evening discovering links one at a time.
Measuring performance
Why not an application benchmark
The benchmark guest has no GPU, so there is no graphics figure to measure. Everything here is a CPU, memory, scheduling or clock proxy. That is a real limitation and worth stating plainly rather than implying the numbers cover a whole application.
They are chosen because each one stands in for something an interactive or CPU-bound workload actually does:
| Measurement | What it stands in for |
|---|---|
| QPC cost | software that polls the clock in a tight loop calls QPC thousands of times a second |
| single-thread throughput | the main thread of a latency-sensitive process |
| memory latency | the dominant cost in most pointer-chasing work |
| core-to-core latency | work-queue handoffs between threads |
| jitter tail | hitching in an interactive session |
| storage IO | bulk load and streaming of data off disk |
The harness
bench/vmbench.c, cross-compiled with mingw-w64 and copied in. Nothing is installed in the
guest - no Cinebench, no AIDA64, nothing that would itself be a detectable artefact.
x86_64-w64-mingw32-gcc -O2 -o vmbench.exe vmbench.c
vm-native-verify does the build, copy, run and grading for you.
How the harness avoids lying
- A volatile sink on every loop result, so the optimiser cannot delete the work.
- Pointer-chase over a random single cycle for memory latency, one node per 64-byte line. A strided walk would be prefetched and would measure bandwidth instead.
QueryPerformanceCounterfor timing, neverrdtsc. TSC behaviour is one of the things this project changes, so timing with it would measure the instrument.- Percentiles, not means, for jitter. The tail is the whole point.
- Warm-up passes before every timed section.
- Minimum across repetitions where the metric is a floor, interquartile mean where it is a distribution.
The mistakes that produce fake results
Measure a race more than once. Windows' boot-time TSC calibration either succeeds or does not, and the result holds for that whole boot. The same XML gives 15 ns on one boot and 1250 ns on the next when host headroom is tight. A single sample per configuration produced a confident, published, wrong conclusion here - a hard "never exceed 16 vCPUs" rule drawn from two data points with nothing tested between them. Four boots per configuration is the minimum for anything boot-dependent.
Never run benchmarks concurrently. This produced a phantom "3% cost of the corrections" here that survived into a written conclusion before it was caught. The repeat loop was:
vmbench clock & vmbench cpu & vmbench jitter 0 10
Three benchmarks competing for the same cores. Both configurations were measured the same way, so the comparison looked controlled, and the artefact still leaked through because the two schedule contention differently. Sequential runs removed the gap entirely.
Benchmark a settled guest. A machine that just booted is indexing, patching and starting
services, and all of it lands in the stall counts. vm-native-verify waits 60 seconds by default.
Set the host governor first. powersave costs 2-3% and makes everything noisier.
Watch for host confounders. A background compute workload on the host is the most likely source of the occasional multi-millisecond outlier in the jitter tail. It is not controlled for, and that is stated rather than left out.
Expected numbers
On a 7950X with the guest correctly configured:
| Reading | Expected | If wrong |
|---|---|---|
| QPC cost | ~15 ns | over 1000 ns means the boot TSC calibration lost its race; reboot and re-measure |
| rdtsc cost | ~7 ns | - |
| 1 thread | ~5300 Mops | host governor is not performance |
| L1 / L2 / L3 | ~0.8 / ~3.4 / ~10 ns | high L3 means vCPUs across two cache domains |
| DRAM | ~60-90 ns | check transparent hugepages are actually applying |
| memory read | ~50 GB/s | - |
| jitter p99.99 | ~2 us | pinning missing, or emulator on a vCPU core |
| stalls >100us | 0-2 per 10 s | same causes; this is the tail an interactive session feels as hitching |
Verifying hugepages actually apply
A guest can silently run entirely on 4 KiB pages. memfd with shared memory backing blocks
transparent hugepages, because shared mappings are not anonymous:
pid=$(pgrep -f "[q]emu-system-x86_64" | head -1)
sudo awk '/AnonHugePages/ {s+=$2} END {print s/1024 " MiB"}' /proc/$pid/smaps
Should be close to the guest's RAM size. It read 0 MiB here until the leftover backing was removed, at which point it read 8110 MiB of an 8 GiB guest.
That figure needs free, unfragmented host memory at the moment the guest touches its RAM. On the
same host with 6 GiB free, 14 GiB of page cache and defrag at defer+madvise, an 8 GiB guest
booted with 1154 MiB of its 7196 MiB resident on hugepages, and /proc/vmstat showed
thp_fault_fallback at 46% of thp_fault_alloc. Dropping the page cache and compacting before
the boot took it to 4064 MiB of 8261:
sudo sync; echo 3 | sudo tee /proc/sys/vm/drop_caches; echo 1 | sudo tee /proc/sys/vm/compact_memory
vm-native-verify prints the figure; without root it falls back to the host-wide
AnonHugePages, which is the guest's own number when only one guest runs. vm-native-setup says
when free memory is below the guest's RAM. Static hugepages would guarantee it and are deliberately
not used, because they take the memory from the host permanently.
Note the [q] in that pattern - an unbracketed pgrep -f matches your own shell and will hand you
the wrong PID.
The selftest
patches/kvm/0005-KVM-selftests-verify-ICEBP-DB-reports-RIP-past-the-ICEBP.patch adds
tools/testing/selftests/kvm/x86/icebp_test.c. Build it in a kernel tree with the series applied:
make -C tools/testing/selftests/kvm "$PWD/tools/testing/selftests/kvm/x86/icebp_test"
sudo tools/testing/selftests/kvm/x86/icebp_test
On unpatched SVM it fails with 0x402ec4 != 0x402ec5 (db_rip != next_rip); with 0004 it passes.
It runs against whatever kvm_amd is loaded, so it is also the quickest way to tell which build
is live without booting the Windows guest.
Reproducing the TIMER analysis
bench/timerprobe.c reproduces VMAware's instruction-latency detector so its ratio can be measured
directly, without needing a debug build of VMAware itself. It agrees with VMAware's own figure to
within a few percent, which is what makes it usable for iterating.
x86_64-w64-mingw32-gcc -O2 -o timerprobe.exe timerprobe.c
timerprobe.exe # sweep every counter placement, then detector 2
timerprobe.exe 0 2 # both detectors, measuring on cpu 0, counter on cpu 2
detector 1, instruction latency - counter placement swept:
counter=2 ratio= 7.896 (cpuid 2069 / lfence 262) DETECTED
detector 2, exception latency:
exception db= 8703.3 api= 5941.8 ratio= 1.465 pass (db traps seen: 4000)
With cpuid_passthrough on, detector 1 reads cpuid 199 / lfence 262, ratio 0.760, pass.
The sweep matters: it shows the best and worst case across placements, rather than only the one VMAware happened to choose. That is how the cross-CCD idea was tested and ruled out.
Detector 2 is where the probe itself was wrong for two days. The first version used
RaiseException under a vectored handler for the software side, which never enters the kernel,
and reported a ratio of 6.169. VMAware's software side is RtlCaptureContext plus a direct
ZwRaiseException inside an SEH frame that executes its handler, so it pays a syscall, a kernel
dispatch and an RtlUnwindEx. The probe now does the same, with a __C_specific_handler scope
table because mingw has no __try, and its number agrees with VMAware's own debug line to within
a few percent. Reproduce the mechanism, not the API name.
Host-side checks worth taking
# is the patched module actually the one loaded
cat /sys/module/kvm/srcversion
cat /sys/module/kvm_amd/srcversion
# did the CPU isolation hook fire, and did it restore
journalctl -t libvirt-cpu-isolation -n 10
# is the host whole again with the guest off
systemctl show --property=AllowedCPUs system.slice
nproc
free -g
The last group is not optional. A hook that confines the host and fails to restore is worse than no hook, and it has happened here - a hung domain start left the host on twelve of thirty-two CPUs with nothing running. Both the clean-shutdown and hard-destroy paths are now verified explicitly rather than assumed.
A rule learned the hard way
Never call virsh from inside a libvirt hook. libvirt blocks waiting for the hook to return
while the hook waits on libvirt. The deadlock wedges domain start, which wedges virsh list, which
is what virt-manager reads - so the symptom is virt-manager showing no VMs at all, which looks
nothing like the cause.