Compare commits
7 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 0c3425e5af | |||
| a956ececfd | |||
| 966a2a6390 | |||
| d3062f5e24 | |||
| 1e1078e16a | |||
| f90745e68e | |||
| 91fd1c45e1 |
@@ -75,12 +75,6 @@ Enable-NetFirewallRule -DisplayGroup "Remote Desktop"
|
||||
Put your public key in `C:\ProgramData\ssh\administrators_authorized_keys` for an Administrator
|
||||
account. `vm-native-verify` uses that key.
|
||||
|
||||
SSH into the guest fails with `Corrupted MAC on input` until the host has the e1000e offload rule
|
||||
from the `vfio-native` package. The emulated NIC's TX offloads corrupt integrity-checked traffic on
|
||||
the host side of the tap; SMB tolerates it, SSH does not. The package installs a udev rule that
|
||||
turns the offloads off on every libvirt tap as it appears, and `vm-native-setup` says so if it is
|
||||
missing.
|
||||
|
||||
## 3. Make the NVMe driver boot-critical
|
||||
|
||||
The disk is about to move from virtio to emulated NVMe, and Windows only loads boot-start drivers
|
||||
@@ -127,9 +121,14 @@ Two things it asks or warns about:
|
||||
|
||||
Before the first boot, if the host has less free memory than the guest's RAM, free and compact
|
||||
it so the guest lands on transparent hugepages; `vm-native-setup` prints the two commands when it
|
||||
applies. The NIC stays `e1000e`, so the network survives the driver removal in the next step. Do not use
|
||||
applies. The NIC stays `igb`, so the network survives the driver removal in the next step. Do not use
|
||||
virtiofs for host files: it is a virtio device the scanner names, and its shared memory backing
|
||||
blocks transparent hugepages for the whole guest. Share over SMB on the e1000e link instead.
|
||||
blocks transparent hugepages for the whole guest. Share over SMB on the `igb` link instead.
|
||||
|
||||
The interface asks for MTU 9000, and the package's hook turns GSO and GRO on for the tap. Together
|
||||
they take the inbound link from 2.9 to 14 Gbit/s; neither does anything alone. The libvirt network
|
||||
needs `<mtu size='9000'/>` too, and the guest needs *Jumbo Packet* 9014 with its interface MTU at
|
||||
9000 - setting the adapter property alone leaves the IP MTU at 1500 and gains nothing.
|
||||
|
||||
## 5. Remove the virtio drivers and the agents
|
||||
|
||||
|
||||
@@ -8,7 +8,7 @@
|
||||
# vfio-native-qemu QEMU 11.1.1 with the platform-identity patches, in /opt
|
||||
|
||||
pkgname=vfio-native
|
||||
pkgver=1.1.1
|
||||
pkgver=1.3.2
|
||||
pkgrel=1
|
||||
pkgdesc="Present a libvirt guest as a self-consistent physical machine, and tune it"
|
||||
arch=('any')
|
||||
@@ -48,7 +48,6 @@ package() {
|
||||
# this coexists with whatever hook the host already has.
|
||||
install -Dm755 scripts/libvirt-hook-cpu-isolation.sh \
|
||||
"${pkgdir}/etc/libvirt/hooks/qemu.d/10-cpu-isolation.sh"
|
||||
# e1000e offloads corrupt integrity-checked traffic on libvirt taps; host-wide by nature
|
||||
install -Dm644 scripts/99-vfio-native-vnet-offload.rules \
|
||||
"${pkgdir}/usr/lib/udev/rules.d/99-vfio-native-vnet-offload.rules"
|
||||
install -Dm755 scripts/libvirt-hook-vnet-offload.sh \
|
||||
"${pkgdir}/etc/libvirt/hooks/qemu.d/20-vnet-offload.sh"
|
||||
}
|
||||
|
||||
@@ -49,11 +49,11 @@ post_install() {
|
||||
|
||||
'native' is a good default if you would rather not maintain anything.
|
||||
|
||||
A udev rule (99-vfio-native-vnet-offload.rules) turns TX offloads off on every
|
||||
libvirt tap as it appears. The emulated e1000e NIC corrupts integrity-checked
|
||||
traffic with them on; SSH to the guest fails with "Corrupted MAC on input".
|
||||
It applies to every VM on the host; the throughput cost on a host<->guest
|
||||
link is not measurable.
|
||||
A hook at /etc/libvirt/hooks/qemu.d/20-vnet-offload.sh turns GSO and GRO on
|
||||
for the guest's tap. With the interface's MTU 9000 that takes the inbound link
|
||||
from 2.9 to 14 Gbit/s; neither does anything alone. The libvirt network needs
|
||||
<mtu size='9000'/> too, and the guest needs Jumbo Packet 9014 with its
|
||||
interface MTU at 9000.
|
||||
|
||||
A libvirt hook is installed at /etc/libvirt/hooks/qemu.d/10-cpu-isolation.sh.
|
||||
It keeps host processes off the cores the guest is pinned to, automatically,
|
||||
|
||||
@@ -1,8 +0,0 @@
|
||||
# vfio-native: the emulated e1000e NIC's TX checksum and segmentation offloads
|
||||
# corrupt packets on the host side of a libvirt tap. SMB tolerates it; SSH fails
|
||||
# with "Corrupted MAC on input" and any integrity-checked protocol breaks.
|
||||
# Measured on a Zen 4 host with QEMU 11.1.1. Disabling the offloads on every
|
||||
# libvirt tap as it appears fixes it; on a host<->guest link the throughput
|
||||
# cost is not measurable. libvirt's <driver><host .../> attributes are ignored
|
||||
# for e1000e, and a libvirt hook must not call virsh, hence udev.
|
||||
ACTION=="add", SUBSYSTEM=="net", KERNEL=="vnet*", RUN+="/usr/bin/ethtool -K %k tx off gso off gro off tso off"
|
||||
@@ -1,12 +1,18 @@
|
||||
#!/bin/bash
|
||||
# Waits for a just-started guest's network to actually come up, then turns
|
||||
# kvm_amd cpuid_passthrough on.
|
||||
# Keeps the kvm_amd cpuid_passthrough switch correct across a guest's whole life,
|
||||
# in-guest reboots included.
|
||||
#
|
||||
# The readiness signal is the guest's own tap RX counter: it is fresh every boot,
|
||||
# it needs nothing enabled inside the guest (no SSH, no RDP, no agent), and it only
|
||||
# moves once the guest's NIC driver has really loaded - which is well past the CPU
|
||||
# enumeration that the switch must not change under. A DHCP lease left over from a
|
||||
# previous boot cannot trip it early.
|
||||
# The switch must be OFF (N) while the guest enumerates CPUID at boot or Windows
|
||||
# hangs, and ON (Y) once it is up, where it clears the timer detection. The libvirt
|
||||
# hook only fires at VM start and stop, so a guest-initiated reboot would otherwise
|
||||
# re-enumerate with the switch still Y and hang. This watcher drops it to N on every
|
||||
# QEMU RESET and raises it again once the guest's NIC is back up.
|
||||
#
|
||||
# Readiness signal: the guest tap's rx_packets counter growing past a baseline. It
|
||||
# only moves once the guest NIC driver has loaded, well past CPU enumeration, and it
|
||||
# needs nothing enabled inside the guest. The counter is cumulative and does NOT
|
||||
# reset on an in-guest reboot, so readiness is growth past the value captured at the
|
||||
# reset, not an absolute threshold.
|
||||
#
|
||||
# Launched as a transient systemd unit by the cpuid-passthrough hook, so it is free
|
||||
# to call virsh (the hook itself must not - that deadlocks libvirtd).
|
||||
@@ -18,24 +24,43 @@ V="virsh -c qemu:///system"
|
||||
|
||||
running() { [ "$($V domstate "$DOMAIN" 2>/dev/null)" = running ]; }
|
||||
|
||||
tap=""
|
||||
i=0
|
||||
for _ in $(seq 1 65); do # ~195 s cap, then flip anyway if still up
|
||||
tap=$($V domiflist "$DOMAIN" 2>/dev/null | awk '$1 ~ /^(vnet|tap|macvtap)/ {print $1; exit}')
|
||||
RX="/sys/class/net/$tap/statistics/rx_packets"
|
||||
rx() { cat "$RX" 2>/dev/null || echo 0; }
|
||||
|
||||
set_N() { echo N > "$PARAM/cpuid_passthrough"; }
|
||||
set_Y() { printf '%s' "$BRAND" > "$PARAM/brand_string"; echo Y > "$PARAM/cpuid_passthrough"; }
|
||||
|
||||
# Wait until the tap rx counter grows at least 4 past $1 (guest NIC driver up again).
|
||||
# The iteration floor keeps a stray pre-OS packet (a UEFI netboot attempt) from
|
||||
# tripping the flip before the guest is even past its interrupt and timer setup.
|
||||
wait_net_up() {
|
||||
local base=$1 i=0
|
||||
for _ in $(seq 1 100); do # ~300 s cap, then raise anyway if still up
|
||||
i=$((i + 1))
|
||||
running || { sleep 3; continue; } # not "running" yet at prepare time - wait
|
||||
[ -z "$tap" ] && tap=$($V domiflist "$DOMAIN" 2>/dev/null |
|
||||
awk '$1 ~ /^(vnet|tap|macvtap)/ {print $1; exit}')
|
||||
rx="/sys/class/net/$tap/statistics/rx_packets"
|
||||
# the iteration floor keeps a stray pre-OS packet (a UEFI netboot attempt) from
|
||||
# tripping the flip before the guest is even past its interrupt and timer setup
|
||||
if [ "$i" -ge 4 ] && [ -n "$tap" ] && [ -r "$rx" ] &&
|
||||
[ "$(cat "$rx" 2>/dev/null || echo 0)" -ge 4 ]; then
|
||||
break
|
||||
fi
|
||||
running || return 1
|
||||
[ "$i" -ge 4 ] && [ -n "$tap" ] && [ -r "$RX" ] &&
|
||||
[ "$(rx)" -ge "$((base + 4))" ] && return 0
|
||||
sleep 3
|
||||
done
|
||||
return 0
|
||||
}
|
||||
|
||||
running || exit 0 # guest went away before it came up
|
||||
printf '%s' "$BRAND" > "$PARAM/brand_string"
|
||||
echo Y > "$PARAM/cpuid_passthrough"
|
||||
# initial cold boot: wait for the network, then harden
|
||||
wait_net_up 0
|
||||
running || exit 0
|
||||
set_Y
|
||||
logger -t vfio-cpuid "$DOMAIN network up: cpuid_passthrough=Y"
|
||||
|
||||
# every in-guest reboot fires a QEMU RESET: drop to N for the re-enumeration, then
|
||||
# raise it again once the guest's NIC is back. --loop streams one line per reset.
|
||||
$V qemu-monitor-event --domain "$DOMAIN" --event RESET --loop 2>/dev/null | while read -r _; do
|
||||
running || continue
|
||||
base=$(rx)
|
||||
set_N
|
||||
logger -t vfio-cpuid "$DOMAIN reset: cpuid_passthrough=N for re-enumeration"
|
||||
wait_net_up "$base" || continue
|
||||
running || continue
|
||||
set_Y
|
||||
logger -t vfio-cpuid "$DOMAIN back up: cpuid_passthrough=Y"
|
||||
done
|
||||
|
||||
@@ -57,6 +57,24 @@ group_members() {
|
||||
done
|
||||
}
|
||||
|
||||
# The functions to hand to the guest with GPU $1: every device in its IOMMU group
|
||||
# (mandatory for vfio), plus any sibling function of the same PCI device that is
|
||||
# itself a display or HDMI/DP audio controller. An APU parks its PSP and USB
|
||||
# controllers on the same PCI device in separate IOMMU groups - those are the
|
||||
# host's, so siblings are filtered by class and never pulled in blindly.
|
||||
passthrough_devs() {
|
||||
local d="$1" m cls
|
||||
{
|
||||
group_members "$d"
|
||||
for m in "/sys/bus/pci/devices/${d%.*}".*; do
|
||||
[ -e "$m" ] || continue
|
||||
m=$(basename "$m")
|
||||
cls=$(cat "/sys/bus/pci/devices/$m/class" 2>/dev/null)
|
||||
case "$cls" in 0x03*|0x0403*) echo "$m";; esac
|
||||
done
|
||||
} | sort -u
|
||||
}
|
||||
|
||||
driver_of() {
|
||||
local l
|
||||
l=$(readlink -f "/sys/bus/pci/devices/$1/driver" 2>/dev/null)
|
||||
@@ -79,7 +97,7 @@ drives_display() {
|
||||
|
||||
ids_of() { # vendor:device for vfio-pci binding
|
||||
local d
|
||||
for d in $(group_members "$1"); do
|
||||
for d in $(passthrough_devs "$1"); do
|
||||
local cls; cls=$(cat "/sys/bus/pci/devices/$d/class" 2>/dev/null)
|
||||
# only real functions of the card: skip bridges (class 0x0604xx)
|
||||
case "$cls" in 0x0604*) continue;; esac
|
||||
@@ -219,7 +237,7 @@ prepare)
|
||||
|
||||
# 1. stop whatever is holding the DRM device
|
||||
dm=$(systemctl list-units --type=service --state=running --no-legend 2>/dev/null |
|
||||
awk '{print $1}' | grep -xE '(gdm|sddm|lightdm|lxdm|greetd|display-manager)\.service' | head -1)
|
||||
awk '{print $1}' | grep -xE '((gdm|sddm|lightdm|lxdm|greetd|ly|emptty|lemurs)(@[a-z0-9-]+)?|display-manager)\.service' | head -1)
|
||||
if [ -n "$dm" ]; then
|
||||
echo "$dm" > "$STATE/dm"
|
||||
log "stopping $dm"
|
||||
@@ -303,7 +321,7 @@ apply() {
|
||||
"$v" -c qemu:///system dumpxml --inactive "$dom" > "$backup/$dom.gpu-current.xml"
|
||||
grep -q 'ua-vfionative-gpu' "$backup/$dom.gpu-current.xml" || cp "$backup/$dom.gpu-current.xml" "$backup/$dom.before-gpu.xml"
|
||||
local devs="" m
|
||||
for m in $( { group_members "$d"; ls -d "/sys/bus/pci/devices/${d%.*}".* | xargs -n1 basename; } | sort -u); do
|
||||
for m in $(passthrough_devs "$d"); do
|
||||
case "$(cat "/sys/bus/pci/devices/$m/class" 2>/dev/null)" in 0x0604*) continue;; esac
|
||||
IFS=':.' read -r dm bs sl fn <<< "$m"
|
||||
devs+=" <hostdev mode='subsystem' type='pci' managed='yes'>\n <source>\n"
|
||||
@@ -330,7 +348,7 @@ xml() {
|
||||
echo "Add this to the domain, inside <devices>. Every device in the GPU's"
|
||||
echo "IOMMU group has to go together:"
|
||||
echo
|
||||
for m in $(group_members "$d"); do
|
||||
for m in $(passthrough_devs "$d"); do
|
||||
local cls; cls=$(cat "/sys/bus/pci/devices/$m/class" 2>/dev/null)
|
||||
case "$cls" in 0x0604*) continue;; esac
|
||||
IFS=':. ' read -r dom bus slot fn <<< "$(echo "$m" | tr ':.' ' ')"
|
||||
|
||||
15
scripts/libvirt-hook-vnet-offload.sh
Executable file
15
scripts/libvirt-hook-vnet-offload.sh
Executable file
@@ -0,0 +1,15 @@
|
||||
#!/bin/bash
|
||||
# libvirt qemu hook: turn GSO and GRO on for the guest's tap.
|
||||
#
|
||||
# QEMU leaves them off and sets the tap up after udev has run, so it has to
|
||||
# happen here. Paired with MTU 9000 they are the difference between 2.5 and
|
||||
# 14 Gbit/s into the guest; neither helps alone. Tap names come from the domain
|
||||
# XML on stdin, so the hook never calls virsh, which would deadlock libvirtd.
|
||||
|
||||
[ "$2" = started ] || exit 0
|
||||
|
||||
for tap in $(grep -oE "<target dev='(vnet|tap|macvtap)[^']*'" | sed "s/.*dev='//; s/'$//"); do
|
||||
ethtool -K "$tap" gso on gro on 2>/dev/null
|
||||
done
|
||||
|
||||
exit 0
|
||||
@@ -637,8 +637,11 @@ if conformant and E["CONVERT"] == "1":
|
||||
if "device='disk'" not in d or ("bus='nvme'" in d and "<serial>" in d):
|
||||
return d
|
||||
d = re.sub(r"<target dev='([^']*)' bus='(virtio|sata|scsi)'/>", r"<target dev='\1' bus='nvme'/>", d)
|
||||
# cache='none' is O_DIRECT: on btrfs the guest can change a page while the
|
||||
# write is in flight, so the stored checksum never matches and later reads
|
||||
# fail with EIO. Buffered writes hand the filesystem a stable page.
|
||||
d = re.sub(r"<driver name='qemu' type='([^']*)'[^/]*/>",
|
||||
r"<driver name='qemu' type='\1' cache='none' io='native' discard='unmap'/>", d)
|
||||
r"<driver name='qemu' type='\1' cache='writeback' io='threads' discard='unmap'/>", d)
|
||||
d = re.sub(r"\s*<address type='(pci|drive)'[^/]*/>", "", d)
|
||||
if "<serial>" not in d:
|
||||
serial = E["NVME_SERIAL"] if n[0] == 0 else E["NVME_SERIAL"][:-1] + "0123456789ABCDEF"[n[0] % 16]
|
||||
@@ -658,7 +661,11 @@ if conformant and E["CONVERT"] == "1":
|
||||
s = re.sub(r"\s*<input type='[^']*' bus='virtio'/>", "", s)
|
||||
s = re.sub(r"<memballoon model='virtio'>.*?</memballoon>", "<memballoon model='none'/>", s, flags=re.S)
|
||||
s = re.sub(r"<memballoon model='virtio'/>", "<memballoon model='none'/>", s)
|
||||
s = re.sub(r"<model type='virtio'/>(\s*<driver [^/]*/>)?", "<model type='e1000e'/>", s)
|
||||
# jumbo frames are the single biggest win on the host<->guest link: the emulated
|
||||
# NIC is packet-rate bound, so 9000-byte frames cut the per-packet cost the guest
|
||||
# pays on receive. Measured 2922 -> 14232 Mbit/s inbound on igb, byte-exact clean.
|
||||
s = re.sub(r"<model type='virtio'/>(\s*<driver [^/]*/>)?",
|
||||
"<model type='igb'/>\n <mtu size='9000'/>", s)
|
||||
if prof == "full":
|
||||
s = re.sub(r"<video>.*?</video>", "<video>\n <model type='none'/>\n </video>", s, flags=re.S)
|
||||
|
||||
@@ -824,10 +831,19 @@ if [ "$PROFILE" = full ]; then
|
||||
fi
|
||||
|
||||
if [ ! -e /usr/lib/udev/rules.d/99-vfio-native-vnet-offload.rules ] && [ ! -e /etc/udev/rules.d/99-vfio-native-vnet-offload.rules ]; then
|
||||
echo "NOTE: the e1000e offload udev rule is not installed. SSH into the guest will fail with"
|
||||
echo "NOTE: the NIC offload udev rule is not installed. SSH into the guest will fail with"
|
||||
echo " 'Corrupted MAC on input' until it is:"
|
||||
echo " sudo install -Dm644 $SELF/scripts/99-vfio-native-vnet-offload.rules /etc/udev/rules.d/ && sudo udevadm control --reload-rules"
|
||||
fi
|
||||
|
||||
NET=$("${C[@]}" dumpxml "$DOM" 2>/dev/null | sed -n "s/.*<source network='\([^']*\)'.*/\1/p" | head -1)
|
||||
NETMTU=$([ -n "$NET" ] && "${C[@]}" net-dumpxml --inactive "$NET" 2>/dev/null | sed -n "s/.*<mtu size='\([0-9]*\)'.*/\1/p")
|
||||
if [ -n "$NET" ] && [ "${NETMTU:-1500}" -lt 9000 ]; then
|
||||
echo "NOTE: the guest interface asks for MTU 9000 but libvirt network '$NET' is at ${NETMTU:-1500}."
|
||||
echo " Jumbo needs both ends; inbound throughput is ~5x with it. Add <mtu size='9000'/> to"
|
||||
echo " the network and restart it: virsh net-edit $NET && virsh net-destroy $NET && virsh net-start $NET"
|
||||
echo " Then in the guest: set the NIC's Jumbo Packet to 9014 and the interface MTU to 9000."
|
||||
fi
|
||||
gov=$(cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor 2>/dev/null || echo unknown)
|
||||
[ "$gov" = performance ] || echo "host governor is '$gov' - run: sudo cpupower frequency-set -g performance"
|
||||
|
||||
|
||||
Reference in New Issue
Block a user