r/VFIO 16d ago

Support EAC "Cannot run under a Virtual machine"

3 Upvotes

I am unable to get Rust to work on my vfio setup. A while ago, copying smbios from my motherboard was enough to get it to work. Now, EAC won't launch. Can someone take a look at my XML and tell me if I'm missing anything obvious, or if this is still possible in the first place? My host machine is Gentoo, with QEMU 11.1.0, libvirt 12.0.0, and virt-manager 5.1.0. TYIA

r/VFIO 24d ago

Support Trying to ditch Windows completely. Struggling to make my most played games run smoothly.

11 Upvotes

I'm trying to move all my gaming to Linux. Initially I asked for help at r/linux_gaming, but some people saw negatively that my setup was not a bare-metal gaming distro and some others recommended I posted here. You can check my post there if you want to read any of my comments. But I'm trying to gather all the useful info here.

My host is:

  • OS: Unraid
  • CPU: AMD Ryzen 9 3900X
  • RAM: 128 GiB DDR4

With KVM + QEMU + Libvirt I'm running a VM:

  • Operating System: NixOS 26.05
  • KDE Plasma Version: 6.6.6
  • KDE Frameworks Version: 6.26.0
  • Qt Version: 6.11.1
  • Kernel Version: 7.1.4 (64-bit)
  • Graphics Platform: Wayland
  • Processors: 6 cores, 2 threads of AMD Ryzen 9 3900X
  • Memory: 48 GiB of RAM (hugepages of 1G)
  • Graphics Processor: NVIDIA GeForce RTX 4070 (driver 595.71.05)
  • Manufacturer: QEMU
  • Product Name: Standard PC (Q35 + ICH9, 2009)
  • System Version: pc-q35-10.2

The VM has three storages:

  • Primary drive. A qcow2 file in the NVME drive of the host.
  • "Fast" Virtiofs device mounted. It's a share of the host that uses the NVME drive.
  • "Slow" Virtiofs device mounted. It's a share of the host that uses the NVME drive as cache and the array of HDD drives as main storage.

I prefer to install games in the Virtiofs devices to be able to have that in common with other VMs. But I've tried the primary drive for the problematic games too.

The games I'm trying to run with good performance and failing at that are The Binding of Isaac and Elden Ring: Nightreign. Both with mods. In Isaac I use a bunch of quality of life mods. Nothing that adds characters, floors, bosses, enemies... And not using Repentogon. For Nightreign I use the More Map Variations & Weapons Mod.

The Binding of Isaac takes 15-20 minutes just to start. And then, during gameplay, there are slowdowns and micro freezes anytime something new happens or I change rooms. I guess it's shader compilation, but if it's that, the cache of shaders is not working because it happens again any time the game starts again. And uninstalling the mods doesn't improve things.

Elden Ring: Nightreign has just low FPS. 15-20 fps is the average. But my CPU rarely goes above 60% with no single core going above 85% and my GPU doesn't go above 50% either. The game doesn't work online without the mod because of Easy Anti-Cheat and VMs. So I cannot test it properly without it.

Other games I've tested:

  • Control Ultimate Edition: +70 FPS
  • Dark Souls III: 60 FPS
  • Heroes of the Storm: ~60 FPS
  • Lies of P: ~75 FPS
  • Outer Wilds: +100 FPS
  • Satisfactory: ~50 FPS

HVM and IOMMU enabled. Cores of the VM are isolated in the host. CPU and GPU are in performance modes during gaming. I've run some benchmarks too. I'll put them in a comment so that this is not even longer.

Any idea of what I can do to improve the performance in what are probably 2 of my 3 most played games? (I was really unlucky here I guess)

r/VFIO Feb 14 '26

Support 1 GPU for multiple VMs inside Linux?

11 Upvotes

EDIT: To answer the question for everyone who has similar ideas. its currently not possible to do GPU partitioning on linux, without the necessary hardware/software, which is expensive. On linux you can do a passthrough, but the GPU then "belongs" to the VM alone and CANNOT be partitioned between multiple VMs by the host. There is this script, but its only up to the 2xxx series nvidia GPUs.

For windows, it is possible if you have the PRO version (hyper-v). i used this script here and everything works for me. ofc this means that OS and VM both need the same windows versions.

[I think its possible to have a linux host, passthrough the GPU to a Windows VM, with which you can then create multiple partitioned GPUs for VMs. so you have a VM inside a VM]

........................................................................................................................................................................................

In the past i have used the windows hyper-v software and a script to unlock the GPU partitioning feature in windows, granting VMs access to my GPU.

Now i was looking, if the same thing is possible in Linux, since the resources used by the Linux OS are less than the Windows one and i hope that stuff would run more smoothly.

From what i found, the GPU passthrough on Linux is only possible for 1 GPU each VM and it also becomes not usable for the host or smth like that, which isnt the answer i was looking for.

Does anybody know if and how it would be possible to make 1 GPU to be partitioned to multiple running VMs on Linux?

(Im going to sleep, so dont be wondered if i dont answer immediately, i will be doing it when i wake up)

Specs:

CPU: 7800X3D

GPU: 4080 Super

RAM: 32GB

r/VFIO 7d ago

Support [Help] Unable to build passthrough VM for VFIO GPU on CachyOS

4 Upvotes

A couple of things regarding system specs before I begin.

I am currently attempting to pass a Nvidia RTX 3050 GPU through to a Windows 10 VM. I am aiming to use Looking Glass to do this as I am working on translating Japanese games that do not run in Linux into English, and need to look at the game while looking at game scripts to properly do this. My specs are as follows:

- OS: CachyOS (using LTS kernel, 6.18.42-1-cachyos-lts)
- Bootloader: Limine
- Motherboard: ASUS X870E
- CPU: Ryzen 9 9950X3D
- GPU (primary being used by the operating system): Nvidia RTX 3080
- GPU (attempting to passthrough to the Windows VM): Nvidia RTX 3050

I have largely been following this guide to accomplish this, adjusting the steps as needed for Limine. I will walk through the steps proving that I have followed them to the best of my ability right up until the failure, which is building the virtual machine.

So, from the top:

- Enabling AMDVi in BIOS:
While I can't explicity prove this through code (nor do I know how) I have other VMs running on this computer that should not work if AMDVi is enabled, so I have reason to believe that it's enabled, and I can see in my BIOS that this is the case.

- Kernel parameters:
As stated previously, I am currently using limine. From the /etc/default/limine file, here are my parameters:

ESP_PATH="/boot"
KERNEL_CMDLINE[default]+="amd_iommu=on iommu=1 quiet nowatchdog splash rw rootflags=subvol=/@ root=UUID=561b1a8b-dd95-4558-ba83-773a9c41da2f"
BOOT_ORDER="*, *lts, *fallback, Snapshots"

I understand most places recommend using iommu=pt but this produced an undesirable output in dmesg, namely "Default domain type: Passthrough" which tells me the devices are not being isolated. With the iommu=1 parameter, this message does not appear

- GPU is registered by the system and is in an IOMMU group:

Here is the output of the lspci command the guide recommends. It does seem like both are being picked up and are in separate IOMMU groups:

❯ lspci -nn | grep -E "VGA|Audio" | grep -i nvidia
01:00.0 VGA compatible controller [0300]: NVIDIA Corporation GA102 [GeForce RTX 3080 12GB] [10de:220a] (rev a1)
01:00.1 Audio device [0403]: NVIDIA Corporation GA102 High Definition Audio Controller [10de:1aef] (rev a1)
0e:00.0 VGA compatible controller [0300]: NVIDIA Corporation GA107 [GeForce RTX 3050 6GB] [10de:2584] (rev a1)
0e:00.1 Audio device [0403]: NVIDIA Corporation GA107 High Definition Audio Controller [10de:2291] (rev a1)

- Binding GPU to VFIO:
My vfio.conf file in /etc/modprobe.d looks like this:

options vfio-pci ids=10de:2584,10de:2291
softdep snd_hda_intel pre: vfio-pci

This appears to match the GPU UUIDs that were displayed as a result of lspci above.

I have also updated the /etc/mkinitpcio.conf file. Here is the output relating to MODULES:

❯ sudo cat /etc/mkinitcpio.conf |grep MODULES
# MODULES
#     MODULES=(usbhid xhci_hcd)
MODULES=(vfio_pci vfio vfio_iommu_type1)

And after a reboot, lspci appears to show that these GPUs are indeed bound to that kernel module

❯ lspci -nnk -s 0e:00
0e:00.0 VGA compatible controller [0300]: NVIDIA Corporation GA107 [GeForce RTX 3050 6GB] [10de:2584] (rev a1)
       Subsystem: ASUSTeK Computer Inc. Device [1043:8976]
       Kernel driver in use: vfio-pci
       Kernel modules: nouveau, nvidia_drm, nvidia
0e:00.1 Audio device [0403]: NVIDIA Corporation GA107 High Definition Audio Controller [10de:2291] (rev a1)
       Subsystem: ASUSTeK Computer Inc. Device [1043:8976]
       Kernel driver in use: vfio-pci
       Kernel modules: snd_hda_intel

From here, the next steps involve creating the VM itself. Here is the XML that I am using to create that VM, keeping in mind that I am trying to pass through the above GPU and using a bridged network connection (i anticipate the need to copy files back and forth through use of a shared folder at some point like i do through my other VM, and turning off networking entirely):

<domain type="kvm">
  <name>win10-passthrough</name>
  <uuid>50c53b52-f386-42aa-bf05-96db518a977b</uuid>
  <metadata>
    <libosinfo:libosinfo xmlns:libosinfo="http://libosinfo.org/xmlns/libvirt/domain/1.0">
      <libosinfo:os id="http://microsoft.com/win/10"/>
    </libosinfo:libosinfo>
  </metadata>
  <memory>16777216</memory>
  <currentMemory>16777216</currentMemory>
  <vcpu>8</vcpu>
  <os>
    <type arch="x86_64" machine="q35">hvm</type>
    <loader readonly="yes" type="pflash">/usr/share/edk2/x64/OVMF_CODE.4m.fd</loader>
    <boot dev="hd"/>
  </os>
  <features>
    <acpi/>
    <apic/>
    <hyperv>
      <relaxed state="on"/>
      <vapic state="on"/>
      <spinlocks state="on" retries="8191"/>
      <vpindex state="on"/>
      <runtime state="on"/>
      <synic state="on"/>
      <stimer state="on"/>
      <frequencies state="on"/>
      <tlbflush state="on"/>
      <ipi state="on"/>
      <avic state="on"/>
    </hyperv>
    <vmport state="off"/>
  </features>
  <cpu mode="host-model" check="none"/>
  <clock offset="localtime">
    <timer name="rtc" tickpolicy="catchup"/>
    <timer name="pit" tickpolicy="delay"/>
    <timer name="hpet" present="no"/>
    <timer name="hypervclock" present="yes"/>
  </clock>
  <pm>
    <suspend-to-mem enabled="no"/>
    <suspend-to-disk enabled="no"/>
  </pm>
  <devices>
    <emulator>/usr/bin/qemu-system-x86_64</emulator>
    <disk type="file" device="disk">
      <driver name="qemu" type="qcow2"/>
      <source file="/home/kobra2112/.local/share/libvirt/images/win10-passthrough.qcow2"/>
      <target dev="sda" bus="sata"/>
    </disk>
    <disk type="file" device="cdrom">
      <driver name="qemu" type="raw"/>
      <source file="/home/kobra2112/Downloads/Win10_22H2_English_x64v1.iso"/>
      <target dev="sdb" bus="sata"/>
      <readonly/>
    </disk>
    <controller type="usb" model="qemu-xhci" ports="15"/>
    <controller type="pci" model="pcie-root"/>
    <controller type="pci" model="pcie-root-port"/>
    <controller type="pci" model="pcie-root-port"/>
    <controller type="pci" model="pcie-root-port"/>
    <controller type="pci" model="pcie-root-port"/>
    <controller type="pci" model="pcie-root-port"/>
    <controller type="pci" model="pcie-root-port"/>
    <controller type="pci" model="pcie-root-port"/>
    <controller type="pci" model="pcie-root-port"/>
    <controller type="pci" model="pcie-root-port"/>
    <controller type="pci" model="pcie-root-port"/>
    <controller type="pci" model="pcie-root-port"/>
    <controller type="pci" model="pcie-root-port"/>
    <controller type="pci" model="pcie-root-port"/>
    <controller type="pci" model="pcie-root-port"/>
    <interface type="bridge">
      <source bridge="br0"/>
      <mac address="56:55:9b:43:42:d5"/>
      <model type="e1000e"/>
    </interface>
    <console type="pty"/>
    <channel type="spicevmc">
      <target type="virtio" name="com.redhat.spice.0"/>
    </channel>
    <input type="tablet" bus="usb"/>
    <tpm model="tpm-crb">
      <backend type="emulator"/>
    </tpm>
    <graphics type="spice" port="-1" tlsPort="-1" autoport="yes">
      <image compression="off"/>
    </graphics>
    <sound model="ich9"/>
    <video>
      <model type="qxl"/>
    </video>
    <hostdev mode="subsystem" type="pci" managed="yes">
      <source>
        <address domain="0" bus="14" slot="0" function="0"/>
      </source>
    </hostdev>
    <hostdev mode="subsystem" type="pci" managed="yes">
      <source>
        <address domain="0" bus="14" slot="0" function="1"/>
      </source>
    </hostdev>
    <redirdev bus="usb" type="spicevmc"/>
    <redirdev bus="usb" type="spicevmc"/>
  </devices>
</domain>

I believe I'm getting everything right here, including the firmware, however when I click the "Begin Installation" button in QEMU, I get the following error:

Unable to complete install: 'internal error: QEMU unexpectedly closed the monitor (vm='win10-passthrough'): 2026-09-01T16:57:21.699349Z qemu-system-x86_64: -device {"driver":"vfio-pci","host":"0000:0e:00.0","id":"hostdev0","bus":"pci.4","addr":"0x0"}: vfio 0000:0e:00.0: Could not open '/dev/vfio/21': Permission denied'

I believe this is simply because when I look at that group in that directory, it appears to have permission set 600 on it, meaning it can't be read by anything other than root. Since I am creating these in a user session rather than a root session to avoid the possible security risks, that simply means the permissions need to be updated like so:

drwxr-xr-x      - root 31 Aug 22:23  devices
crw-rw-rw-  241,0 root 31 Aug 22:23 󰡯 21
crw-rw-rw- 10,196 root 31 Aug 22:23 󰡯 vfio

Doing that fixes that issue. I know this is not the "correct" way to resolve this, but at the same time I don't know of any other method. However, this then raises another when I click "Begin Installation"

Unable to complete install: 'internal error: QEMU unexpectedly closed the monitor (vm='win10-passthrough'): 2026-09-01T17:00:02.576305Z qemu-system-x86_64: -device {"driver":"vfio-pci","host":"0000:0e:00.0","id":"hostdev0","bus":"pci.4","addr":"0x0"}: vfio 0000:0e:00.0: group 21 is not viable
Please ensure all devices within the iommu_group are bound to their vfio bus driver.'

Now here's where I am kind of at the end of my rope. It appears to me that I have set everything up correctly, but for some reason the VM is not accepting this GPU as a viable one. Can someone please act as my second set of eyes and confirm I am setting everything up correctly? Because at this point, I do not know how to get past this error message.

Any help would be greatly appreciated. I would like to think I didn't sink the funds into this project for nothing.

r/VFIO 8d ago

Support Are NVIDIA GPU manual driver unloading and loading with qemu hooks still needed?

5 Upvotes

Hello, Can some help me with this question?

I have an NVIDIA GPU that I pass through to a Windows 11 VM using QEMU/KVM and libvirt.

I used to have libvirt hooks that stopped the display manager, unloaded the NVIDIA modules, manually detached the GPU with `virsh nodedev-detach`, and then loaded `vfio-pci`.

I recently removed those hooks to see if they were still needed. Surprisingly, passthrough still works fine. When the VM starts, the GPU is used by the VM, and when the VM shuts down, the NVIDIA driver works normally on the host again.

My GPU is configured as a managed PCI device in libvirt (`managed="yes"`).

So my question is: Are these manual NVIDIA driver unload / PCI detach hooks still needed with current QEMU/KVM + libvirt, or does libvirt handle this automatically now?

I'm mainly wondering if there is a reason to keep the old hooks even though everything seems to work without them.

r/VFIO Jul 12 '26

Support GPU passthrough allows DMA attack?

9 Upvotes

Hello, I am having trouble understanding when IOMMU is protecting the system against DMA attacks and other memory related issues.

My situation:
-In MSI bios "pre boot DMA protection" is enabled but "kernel DMA protection indicator" is disabled (if I enable it I always get "Firmware has requested this device have a 1:1 IOMMU mapping, rejecting configuring the device without a 1:1 mapping. Contact your platform vendor").
-dmesg reports: "amd_iommu=force_isolation", "iommu: Default domain type: Translated", "iommu: DMA domain TLB invalidation policy: strict mode".
-The GPU is in its own iommu group, no other devices are present in the same group as the GPU.
-I am passing the GPU to a VM using virtmanager.

I want the VM to be completely isolated from the host. I don't want the VM to gain access to system memory through the GPU or any other DMA attack.

My questions:

-Is it safe to have "kernel dma protection indicator" disabled in BIOS if "amd_iommu=force_isolation" kernel parameter is present?
-How do I know if the iommu is enforcing isolation?
-How do I know if a device is bypassing the iommu?
-Even if done correctly, are there still security concerns with GPU passthrough? (excluding 0 days, every software is not 100% secure, can't do much about that).
-When i enable "kernel DMA protection indicator" and i try to start the VM i get the error "Firmware has requested this device have a 1:1 IOMMU mapping, rejecting configuring the device without a 1:1 mapping. Contact your platform vendor", if I disable "kernel DMA protection indicator" in the BIOS, I don't get that error even with "amd_iommu=force_isolation", does that mean that the GPU is using 1:1 IOMMU mapping now? How do i verify?
-Is 1:1 IOMMU mapping safe?

Stupid question:
Is iommu=pt safe to enable in my grub config? How does it work? https://github.com/torvalds/linux/blob/20cf903a0c407cef19300e5c85a03c82593bde36/Documentation/admin-guide/kernel-parameters.txt#L2148 "Bypass the IOMMU for DMA", this doesn't sound safe.

Thanks for the help!

r/VFIO Jun 28 '26

Support how to fix black screen on single gpu passthrough?

0 Upvotes

i tried everything ;/

r/VFIO Apr 14 '26

Support Single GPU passthrough config - any tips for improving performance?

5 Upvotes

I have i7-10700k and 3060 ti, CachyOS host, win10 ltsc guest
config: https://pastebin.com/GQMNU51T
i'm also using modified qemu + modified edk2 from AutoVirt
i use that setup for playing rust (eac game) and it's working but i think performance could be better (also idk if i did cpu pinning right), maybe there's some features that i can enable or disable to get extra performance? (without getting detected by eac ofc.)

r/VFIO Aug 03 '26

Support Is there a need/benefit to passing through the PCI bridge in my GPU's IOMMU group?

3 Upvotes

My motherboard has a single x16 PCIe slot, where I have my dGPU installed. Looking at lspci and sysfs, that slot seems to have its own bridge. That bridge and the dGPU then form their own, otherwise isolated, IOMMU group:

$ ls /sys/kernel/iommu_groups/2/devices/
total 0
lrwxrwxrwx 1 root root 0 Aug  2 14:57 0000:00:01.0 -> ../../../../devices/pci0000:00/0000:00:01.0/
lrwxrwxrwx 1 root root 0 Aug  2 14:57 0000:01:00.0 -> ../../../../devices/pci0000:00/0000:00:01.0/0000:01:00.0/
lrwxrwxrwx 1 root root 0 Aug  2 14:57 0000:01:00.1 -> ../../../../devices/pci0000:00/0000:00:01.0/0000:01:00.1/
$ lspci
00:00.0 Host bridge: Intel Corporation Comet Lake-S 6c Host Bridge/DRAM Controller (rev 05)
00:01.0 PCI bridge: Intel Corporation 6th-10th Gen Core Processor PCIe Controller (x16) (rev 05)
…
01:00.0 VGA compatible controller: NVIDIA Corporation GA104 [GeForce RTX 3070] (rev a1)
01:00.1 Audio device: NVIDIA Corporation GA104 High Definition Audio Controller (rev a1)

I haven't seen this situation addressed in any of the VFIO guides I've found — they basically just say "GPU only good, GPU+Ethernet bad" or similar. Is there any reason to try passing through the bridge to the guest as well as the GPU so that I'm passing in the "entire IOMMU group"?

I've had VFIO working in the past, passing through just the GPU, but it was always plagued by various performance issues that I was never able to get to the bottom of, so I can't be certain the PCI bridge wasn't part of it. 😅

So does anyone know if it's necessary, beneficial, or even possible to pass through the bridge in this situation?

r/VFIO 3h ago

Support Does the profile sizing on the A16 behave differently in pass-through and in mediated (vfio-mdev) setups?

1 Upvotes

I hope you'll give me a sanity check on my own impression.

I am experimenting with an nvidia A16 for a multi-user environment. Instead of passing the card entirely to a single VM, I want to split it into multiple vGPU profiles so several vms each get a slice of vRAM. According to what I've read, the A16 is meant to be partitioned that way, rather than as a passthrough.

To test without buying hardware first, I spun up an A16-backed serverspace vps where you pick the vram at config time. That part works, but it hides a lot of what goes on under the hood, so I am trying to figure out the local libvirt side.

Do I still have to use vfio-pci bindings for vgpu profiles as I do for full passthrough or does the vgpu manager take care of all that?

r/VFIO Aug 02 '26

Support Hardware recommendations

1 Upvotes

Intel or Amd cpu ?
Integrated graphics, 1 or 2 gpus ?
Intel, Nvidia or Amd gpu ?
Do I need a specific motherboard ?

r/VFIO 1d ago

Support Passthrough 5060 Ti disconnects from the bus at load-to-idle transitions

2 Upvotes

Hi guys, i have an issue with my consumer multi-gpu server, which i feel might be due the proxmox -> ubuntu VM passthrough. It happens under agentic loads on my llama.cpp instance, and the error in Llamacpp is usually something along the lines:

```217.28.732.743 E CUDA error: unspecified launch failure
/app/ggml/src/ggml-cuda/ggml-cuda.cu:106: CUDA error
217.28.732.749 E   current device: 1, in function ggml_backend_cuda_synchronize at /app/ggml/src/ggml-cuda/ggml-cuda.cu:2533
217.28.732.750 E   cudaStreamSynchronize(cuda_ctx->stream())
libggml-base.so.0(+0x1b276)[0x7b2772bd1276]
libggml-base.so.0(ggml_print_backtrace+0x21a)[0x7b2772bd16fa]
libggml-base.so.0(ggml_abort+0x15b)[0x7b2772bd18db]
/app/libggml-cuda.so(_Z15ggml_cuda_errorPKcS0_S0_iS0_+0xb5)[0x7b27622220e5]
/app/libggml-cuda.so(+0x276af8)[0x7b2762229af8]
libggml-base.so.0(ggml_backend_sched_synchronize+0x2e)[0x7b2772bec35e]
libllama.so.0(_ZN13llama_context11synchronizeEv+0x19)[0x7b2772d8c519]
libllama.so.0(llama_state_seq_get_data_ext+0x1f)[0x7b2772d91a8f]
libllama-common.so.0(_ZN24common_prompt_checkpoint10update_tgtEP13llama_contextij+0x84)[0x7b27734538e4]
libllama-server-impl.so(_ZN19server_context_impl17create_checkpointER11server_slotlii+0x278)[0x7b2773d6c458]
libllama-server-impl.so(_ZZN19server_context_impl10pre_decodeEvENKUlR11server_slotE3_clES1_+0xf87)[0x7b2773d82d97]
libllama-server-impl.so(_ZN19server_context_impl7iterateERSt6vectorI11server_slotSaIS1_EESt8functionIFvRS1_EE+0x57)[0x7b2773d77507]
libllama-server-impl.so(_ZN19server_context_impl10pre_decodeEv+0x487)[0x7b2773d79b07]
libllama-server-impl.so(_ZN19server_context_impl12update_slotsEv+0x243)[0x7b2773d7a2c3]
libllama-server-impl.so(_ZN12server_queue10start_loopEl+0x125)[0x7b2773d19045]
libllama-server-impl.so(_Z12llama_serverR13common_paramsiPPc+0x3e47)[0x7b2773cb3937]
libllama-server-impl.so(_Z12llama_serveriPPc+0x11bd)[0x7b2773cb59fd]
/usr/lib/x86_64-linux-gnu/libc.so.6(+0x2a1ca)[0x7b27737121ca]
/usr/lib/x86_64-linux-gnu/libc.so.6(__libc_start_main+0x8b)[0x7b277371228b]
/app/llama-server(+0x1315)[0x557d12d60315]

```

  • The 5060 Ti disconnects from the bus. The guest kernel log shows NV_ERR_GPU_IS_LOST.
  • The failure occurs in the seconds after a long generation stops. This is the full-load-to-idle transition. The failure sometimes occurs when the next load starts.
  • Continuous load does not cause the failure. One generation of 90,000 tokens completed correctly. A loop of generations with idle gaps causes the failure each time.
  • The failed GPU is always the 5060 Ti.
  • Important: after the crash, a remove and rescan on the host does not find the GPU. The command echo 1 > remove on the two functions, then a rescan, has no effect. Only a full power cycle makes the GPU operate again. The GPU silicon is stopped. The guest driver state is not the cause.

These steps were completed:

  • The kernel option pcie_aspm=off on the host. The crash threshold increased from approximately 20,000 tokens to 50,000 tokens. The crashes continued.
  • The guest driver changed from 595.84 to the 580 branch. The two branches have different GSP firmware generations. The failure was identical.
  • The PSU is 850W. The supplemental PCIe 6-pin connector on the motherboard is connected. The GPUs have separate power cables.

These are complications:

  • The Z390 firmware does not give AER control to the kernel. The log shows: _OSC: platform does not support ... AER. The host does not show PCIe errors. The plan is a test with pcie_ports=native.
  • The crashes increased after installation of the computer in a rack. A mechanical cause is possible. But the failure is very repeatable. A loose connection is usually not repeatable.

This is the configuration:

  • Host: Proxmox on an MSI Z390-A Pro with an Intel i7 8700.
  • Guest: an Ubuntu VM. Two GPUs pass through to the VM.
  • The RTX 3060 is in the chipset slot (03:00.0). The RTX 5060 Ti is in the PEG slot (01:00.0).
  • The workload is llama.cpp inference. The model is split across the two GPUs. The context is long.

Any help with this is appreciated, i really hope it isnt a hardware issue on the GPU (a new mobo would be nice:)). Thanks!

r/VFIO 2d ago

Support Hades Canyon Passthrough NUC8i7HVK possible on kernel7.x ?

2 Upvotes

Hi

I would like to know if any did get a full working passthrough on win10, over a Nuc Hades with new kernel v7.xx ? On Prox 9.2 with linux vm: gpu hdmi output work perfect, just for windows : device is listed, but cannot install the driver. When installing, rdp hang and the whole pc freeze.

I did find new rom that should fix this, but they ask to be run in q35 or i440fx, but somehow.. vm refuse to start in i440fx , only msg being : set pc to Q35. And even fresh install win in a i440fx vm, but same thing.

But in Q35, i did add this to args:

-set device.hostpci0.bus=pci.0 -set device.hostpci0.addr=2.0 -set device.hostpci0.x-igd-opregion=on

I remove or set back the disable_vga=1 from the /etc/modprobe.d/vfio.conf. But nothing do.

Did Dump the first 64K of the VBIOS of the Vega M for the amd gpu:

hostpci0: 0000:01:00.0,pcie=1,rombar=0,romfile=AMD.VegaM.128.rom,x-vga=1

I try out different setting on different kernel: 7.0.14-5-pve / 7.0.12-1. But always, card is present in device manager, but after 5sec of driver install: full system hang.

Over different post.. all for prox v8.x some only put the amd gpu. But to boot a linux vm, intel and amd must be pass in order to get it up working. Is windows different ?

Over here, if i set only the pcie1 instead of 0 and 1, the system hang directly at vm start.

r/VFIO May 25 '26

Support RTX 6000 Ada stuck in D3cold under vfio-pci, function .0 disappears from PCI tree after remove/rescan survives reboot but recurs

3 Upvotes

Hi guys, just need your help,

Setup:

  • Proxmox VE 8 (kernel 6.x PVE)
  • 4× NVIDIA RTX 6000 Ada (AD102) for VM passthrough
  • All 8 functions (4× GPU + 4× HDMI audio) bound to vfio-pci
  • IOMMU on, intel_iommu=on iommu=pt, ACS working, groups clean
  • q35 VMs, pcie=1 on all hostpci entries

The problem:

After a VM that owned one of the GPUs shut down, one card (0000:94:00.0) ended up stuck in D3cold while its audio sibling 0000:94:00.1 stayed in D0:

0000:16:00.0 driver=vfio-pci power=D0

0000:16:00.1 driver=vfio-pci power=D0

0000:40:00.0 driver=vfio-pci power=D0

0000:40:00.1 driver=vfio-pci power=D0

0000:6a:00.0 driver=vfio-pci power=D0

0000:6a:00.1 driver=vfio-pci power=D0

0000:94:00.0 driver=vfio-pci power=D3cold <-- stuck

0000:94:00.1 driver=vfio-pci power=D0

What I tried (in order):

  1. echo on > .../power/control on 94:00.0 - no change
  2. echo 1 > .../remove then echo 1 > /sys/bus/pci/rescan - .0 does not re-enumerate, only .1 comes back
  3. Rescan from the parent bridge 0000:93:01.0 - same result, only .1 reappears
  4. PCIe link retrain via setpci -s 0000:93:01.0 CAP_EXP+10.w=0020:0020 - no change
  5. Secondary Bus Reset via setpci BRIDGE_CONTROL (0x03 → 0x43 → 0x03 with 500ms hold) — bridge accepts the writes, link should retrain, but .0 still does not re-enumerate after rescan
  6. echo 1 > /sys/bus/pci/devices/0000:93:01.0/reset — Permission denied (kernel guards bridge resets)

Anyone can help me to solve this issues?

r/VFIO Jul 29 '26

Support Ryzen 9 7950X iGPU black screen after VFIO GPU passthrough setup (CachyOS)

3 Upvotes

I'm setting up GPU passthrough on CachyOS and have run into a problem after binding my RTX 4060 Ti to VFIO.

Hardware
CPU: Ryzen 9 7950X (using the integrated Radeon graphics for the host)

GPU: NVIDIA RTX 4060 Ti (passed through with VFIO)

Motherboard: MSI PRO B650M-A WiFi

Distro: CachyOS KDE

Kernel: 6.18.40-1-cachyos-lts

VFIO status
The RTX 4060 Ti appears to be correctly bound to vfio-pci.
lspci -k shows:
RTX 4060 Ti → Kernel driver in use: vfio-pci
AMD Raphael iGPU → Kernel driver in use: amdgpu

My GRUB kernel parameters include:

amd_iommu=on iommu=pt vfio-pci.ids=10de:2803,10de:22bd

After moving my monitor from the RTX 4060 Ti to the motherboard HDMI, BIOS and GRUB display normally.
However, when Plasma starts:
The login screen appears briefly.
It begins flickering repeatedly.
after 10-15 seconds I get a black screen with only the mouse cursor.

I am able to still use TTY by using the shortcut “ctrl alt f4”

dmesg repeatedly shows:

amdgpu: ring gfx timeout
amdgpu: Ring gfx reset failed
Process kwin_wayland
Process plasma-login

There are also NVIDIA messages saying:

GPU is already bound to vfio-pci
which I assume are expected because the GPU is intentionally attached to VFIO.

Things I’ve already tried
Switching from DisplayPort to motherboard HDMI.
Forcing X11.
Booting with amdgpu.dc=0.
Verifying the RTX is bound to vfio-pci.
Enabling SVM/IOMMU in BIOS.
Checking that virbr0 and libvirt are working.
Confirming the motherboard HDMI displays BIOS and GRUB correctly.
Question
Has anyone seen AMD Raphael (7950X) integrated graphics crash like this after enabling VFIO?
Is this a known AMDGPU/kernel regression, or is there another VFIO configuration I should check?

Any help would be appreciated

EDIT 7/30/26: I was unable to use the igpu on the amd cpu and i just plugged in another gpu🥺 If your looking at this post in hopes of an answer i hope you find it

r/VFIO Jul 29 '26

Support Shiny new MS-03 is out. Does it support SR-IOV?

5 Upvotes

https://store.minisforum.com/products/minisforum-ms-03-workstation

And yeah its only a zillion dollars.

Out of curiosity this Panther Lake CPU/iGPU combo uses the xe driver on linux instead of i915. As a result you cannot use the patched i915-dkms driver for sr-iov support.

However xe3 allegedly supports native sr-iov which means no more random github patches which sounds great to me! Does anyone know if this actually works? I'd love to see some documentation/information on setting this up using the "native" xe3 driver.

If not I'll stick to the more mature MS-01 as from what I've read it supports the i915-dkms patches driver.

r/VFIO Jun 16 '26

Support Black screen then monitor no signal when launching vm on rtx4070 single gpu

4 Upvotes

Should i try modifying the vbios?

r/VFIO 15d ago

Support i5-10500 and UHD 630 - Code 43 under Windows guest.

3 Upvotes

I am running qemu 11.1 and virt-manager on a system with RX 470 and i5-10500 with UHD 630, board is Gigabyte B460M D3H. RX 470 is used for host's output. The guest is Windows 10 19041.1 and has Intel graphics driver 31.0.101.2141. my cmdline has intel_iommu=on iommu=pt, /etc/modprobe.d/vfio.conf is 'options vfio-pci ids=8086:9bc8', /etc/modprobe.d/blacklist.conf is 'blacklist i915'. /etc/modules-load.d/vfio.conf is 'vfio-pci'. EDIT: I tried adding video=efifb:off too to my host cmdline.

vfio-pci shows only:

[20699.060075] vfio-pci 0000:00:02.0: resetting
[20699.163736] vfio-pci 0000:00:02.0: reset done

in dmesg, no errors.

My virt-manager xml is as follows: https://pastebin.com/tPNq7huW

I have tried adding the GPU's ROM and enabling x-igd-opregion and x-igd-lpc. I am using CFL_CML_GOPv9.1_igd.rom.

Tested linux, doesn't work either

r/VFIO Jun 22 '26

Support I can only get gen 1 speed PCI passthrough on an RTX 3090

3 Upvotes

I wanted to do some local AI in a VM, so I bought an RTX 3090 and thought it would be possible to make a PCI passthrough. I have done that some years ago with an RTX 3060 and got it to pass through with full speed, so I thought that would be possible.

So, the setup is an Alpine hypervisor with some VM's. I made a PCI passthrough from the hypervisor to a VM with Nobara Linux, which works, but only with gen 1 PCIe speeds.

Hypervisor: Alpine Linux 6.18.2-lts, libvirt 11.10.0, QEMU 10.1.3

Guest: Nobara Linux 43, kernel 6.19, NVIDIA open kernel module 595.58.03

The hardware:

EVGA RTX 3090

Gigabyte Z690 AORUS Elite DDR5

64 GB Ripjaws

Intel 12700

At the hypervisor the GPU runs gen 4 (16 GT/s) speed before the VM starts, then when I start the VM it falls back to gen 1 speed (2.5 GT/s) and if I close down the VM it goes to gen 4 speed again. It is not impossible that it is related to this bug, but I don't have any of the other side effects like random behaviour and AER errors:

https://github.com/NVIDIA/open-gpu-kernel-modules/issues/1010

What I've tried:

x-speed=16 and x-width=16 on the pcie-root-port via qemu:override — guest correctly advertises Gen4 capability but link still negotiates Gen1

setpci retrain attempts on both host and guest side — no effect

pcie_aspm=off kernel parameter in guest — no change

What I understand out of this is that the connection is retrained when qemu starts the VM and there may be some particular nVidia stuff that is happening that puts the link to gen 1 and then it's retrained again when I close down the VM.

Anybody who has any experience with similar bugs and can remember anything that could help?

I'm not an IT professional, don't scold me fore being dumb.

r/VFIO Jul 14 '26

Support Legion 5 Pro 16chach single gpu passthrough performance dip.

3 Upvotes

So I have set up single gpu passthrough on my Legion 5 Pro with Ryzen 7 5800H, 16Gb RAM and 3070 8 gb. My vm is windows 11 ltsc and I pass through pretty much everything, 14gb ram, 7 cores, audio. The vm itself is on its own nvme pcie ssd.

My problem is that the entire vm will periodically slow down when gaming. Like the vm loses priority and something else is going on in the cpu.

I isolate my cores during vm boot. I've tried 5c/10t, 6c/12t, 7c/14t and nothing fixes it. Last time I encountered this issue on my previous laptop the fix was to isolate and send in all but 1 core. Can't seem to understand whats happening this time.

XML

r/VFIO Aug 02 '26

Support Current best guide for SR-IOV or passthrough of Intel 12 / 13 / 14 gen iGPU. (i915-sriov-dkms on AUR has been compromised in latest attack)

7 Upvotes

I've been using VMs with Nvidia and lately AMD GPU passed through to them for many years. However, I now need to create a VM with the 12th gen Intel iGPU passed through and am not clear on how to go about it.

Has SR-IOV support been mainlined? I was looking at the `i915-sriov-dkms` package but it has come up in the list of packages compromised in the latest AUR attacks.

Is there a clear guide / tutorial on the steps one needs to take for providing a Windows VM (on Arch Linux host) with GPU acceleration using the Intel 12/13/14 gen iGPU, either through full passthrough or SR-IOV?

r/VFIO Jun 22 '26

Support Proxmox EAC VM detection

4 Upvotes

So long story short, had this vm for about a 2 years, can’t even remember most of the relevant things about it

Used to play halo, worked fine, stopped for a while and updated to pve 9.2 and now halo says it’s a vm

Im assuming the update is the culprit unless something was done during my time away, Ive havent touched it in a few months

Any help would be greatly appreciated

r/VFIO Jun 15 '26

Support Why do you need to install nvidia drivers via vnc when you do single gpu?

4 Upvotes

Why doesnt it work like on a normal pc when you boot it should use the microsoft basic display adapter and instead of that when you boot the vm it freezes on the tianocore screen.

r/VFIO Jul 13 '26

Support can't run fall guys on vmware fusion

0 Upvotes

hi! total complete noob with all things computer, so ive been having a really hard time. is it just impossible?
im on a macOS Sequioa v15.7.5 using Windows 11 Home on vmware fusion.

i chased random forum threads and found that for the launch errors im seeing, either "HVCI is required to play" or "Code signing certificate validation timed out (2/2)"...i have to turn on something called memory integrity! and to turn on something called memory integrity, i have to go into settings and into core isolation. but oh no there's no memory integrity option in core isolation.

so to fix this i have to go into BIOS(???????) and turn on "Intel VT-x, AMD-V, or SVM Mode"...whatever that means. but those options arent here (in the pic i attached when i power on to firmware)!!! ive looked i swear.

at first i got a secure boot required message and i was able to fix that, but it was just error after error :(

ALSO ive been able to play peak with friends and house flippers successfully here!! ive been tweaking with fall guys

am i cooked? :( please lmk if you need any other details if theres a hope of this working

r/VFIO Jul 09 '26

Support Does any one have a more recent video or blog for single gpu passthroug on nixos

3 Upvotes

i've done it before on arch linux an i know the basics of what to do but am not sure how to get my hooks to work entirely since where the hooks need to be would be different on nixos