protonscr

Sapphire RX 7900 XTX (RDNA3) Illegal opcode in command stream causing GPU reset (DX12)

vkd3dclosed
HansKristian-Work/vkd3d-proton#3004 · opened 2026-05-10 by sebadamus · updated 2026-08-30 · 28 comments · github
Ssebadamus 2026-05-10 github

### Problem description
Dying Light The Beast, Resident Evil 9 (and others) triggers a GPU hang via an "illegal opcode" in the GFX command stream when running under DX12/VKD3D-Proton. The GPU successfully resets via MODE1 (gpu_recovery=1) but the Vulkan context cannot be restored after reset (Failed to initialize parser -125 / ECANCELED), resulting in a frozen game (you can still hear the sounds of game environment and access to linux remotely ssh) but requiring restart to really recover. The hang occurs randomly during gameplay with no specific trigger identified.

What was tried had no impact on the illegal opcode / timeout, happens randomly in each kernel / parameters.

Kernel parameters (can be searched what they mean, sorry for too many, I went through several IAs testing:
amdgpu.ppfeaturemask=0xffffffff
amdgpu.gpu_recovery=1
amdgpu.gfx_off=0
amdgpu.dcdebugmask=0x410'
amdgpu.dcdebugmask=0x10
amdgpu.tmz=0
amdgpu.sg_display=0
amdgpu.runpm=0
iommu=soft'
amdgpu.ppfeaturemask=0xffffffff
amdgpu.ppfeaturemask=0xffffffff
amdgpu.ppfeaturemask=xfffd7fff
amdgpu.ppfeaturemask=0xfff7ffff
amdgpu.ppfeaturemask=0xfffd7fff
amdgpu.ppfeaturemask=0xfffdbfff
amdgpu.ppfeaturemask=0xffff7fff
amdgpu.dcdebugmask=0x410
amdgpu.dcdebugmask=0x10
amdgpu.gfx_off=0

Protons:
GE-Proton 10-34
Proton Experimental
CachyOS Proton 11.0-20260429,
CachyOS Proton 11.0-20260521

Kernels:
7.0.12-x64v3-xanmod1
6.19.14
6.18.34
6.17.12
6.17.0.35
6.12.82

Tried updating linux-firmware mainline from git repo but seems versions from default amdgpu linux-firmware from Kubuntu 24.04 LTS are the same.

None of the above prevented the illegal opcode from being generated.

### Software information
Game: Dying Light The Beast, Resident Evil 9
Proton: GE-Proton 10-34 / proton-cachyos-11.0-20260429
Settings: DX12 render mode (in-game option)
Launch options tested: VKD3D_CONFIG=no_upload_hvv,no_writeback_uploads PROTON_ENABLE_NVAPI=0 %command%
Tried with no steam options, gamemoderun %command%
amdgpu.sg_display=0 amdgpu.mcbp=0'
OS: Kubuntu 24.04 LTS (Noble)

### System information
GPU: AMD Radeon RX 7900 XTX (Navi31, RDNA3, GFX11)
Device ID: 0x744C
VRAM: 24560MB GDDR6 384-bit
Driver: amdgpu 3.64.0
SMU fw version: 78.130.0 (0x004e8200)
DMUB version: 0x07002F00

Mesa/RADV: driverVersion = 26.1.2 (109056002)

Motherboard: ASUS TUF GAMING X870 WIFI PLUS, latest firmware 1654
CPU: AMD Ryzen 9 9900X

Log files

Proton log steam-3008130.log
Dmesg log dmesg2.txt

steam-3008130.log
dmesg2.txt

TTheRealWolfick 2026-05-15 github

I have the exact same issue with the latest versions of Satisfactory and now Subnautica 2, I think it has something to do with global illumination / lumen which is often buggy?

System Information

OS: Arch Linux x86_64
Host: MS-7C56 6.0
Kernel: 7.0.7-arch1-1
CPU: AMD Ryzen 7 5800X3D (16) @ 4.552GHz
GPU: AMD ATI Radeon RX 7900 XT/7900 XTX/7900 GRE/7900M
Mesa: mesa-git 26.2.0_devel.222584
LLVM: llvm-minimal-git 23.0.0_r580641

Was on the mainline LLVM and Mesa, recently swapped to devel branches (15th May) to see if that could address the issue (it did not). Moving to the GE Proton to try and hopefully have it fixed sooner.

Crash Log

May 15 21:33:37 WufufuArch kernel: amdgpu 0000:2d:00.0: Dumping IP State

May 15 21:33:37 WufufuArch kernel: amdgpu 0000:2d:00.0: Dumping IP State Completed
May 15 21:33:37 WufufuArch kernel: amdgpu 0000:2d:00.0: [drm] AMDGPU device coredump file has been created
May 15 21:33:37 WufufuArch kernel: amdgpu 0000:2d:00.0: [drm] Check your /sys/class/drm/card1/device/devcoredump/data
May 15 21:33:37 WufufuArch kernel: amdgpu 0000:2d:00.0: ring gfx_0.0.0 timeout, signaled seq=4893878, emitted seq=4893880
May 15 21:33:37 WufufuArch kernel: amdgpu 0000:2d:00.0: Process GameThread pid 10242 thread vkd3d_queue pid 10396
May 15 21:33:37 WufufuArch kernel: amdgpu 0000:2d:00.0: Starting gfx_0.0.0 ring reset
May 15 21:33:37 WufufuArch kernel: [drm:gfx_v11_0_bad_op_irq [amdgpu]] ERROR Illegal opcode in command stream
May 15 21:33:39 WufufuArch kernel: amdgpu 0000:2d:00.0: MES failed to respond to msg=RESET
May 15 21:33:39 WufufuArch kernel: amdgpu 0000:2d:00.0: failed to reset legacy queue
May 15 21:33:39 WufufuArch kernel: amdgpu 0000:2d:00.0: reset via MES failed and try pipe reset -110
May 15 21:33:39 WufufuArch kernel: amdgpu 0000:2d:00.0: The CPFW hasn't support pipe reset yet.
May 15 21:33:39 WufufuArch kernel: amdgpu 0000:2d:00.0: Ring gfx_0.0.0 reset failed
May 15 21:33:39 WufufuArch kernel: amdgpu 0000:2d:00.0: GPU reset begin!. Source: 1
May 15 21:33:41 WufufuArch kernel: amdgpu 0000:2d:00.0: MES failed to respond to msg=REMOVE_QUEUE
May 15 21:33:41 WufufuArch kernel: amdgpu 0000:2d:00.0: failed to unmap legacy queue
May 15 21:33:41 WufufuArch kernel: [drm:gfx_v11_0_hw_fini [amdgpu]] ERROR failed to halt cp gfx
May 15 21:33:41 WufufuArch kernel: amdgpu 0000:2d:00.0: MODE1 reset
May 15 21:33:41 WufufuArch kernel: amdgpu 0000:2d:00.0: GPU mode1 reset
May 15 21:33:41 WufufuArch kernel: amdgpu 0000:2d:00.0: GPU smu mode1 reset
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: GPU reset succeeded, trying to resume
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: [drm] PCIE GART of 512M enabled (table at 0x0000008001300000).
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: VRAM is lost due to GPU reset!
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: PSP is resuming...
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: reserve 0x1300000 from 0x84fc000000 for PSP TMR
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: RAP: optional rap ta ucode is not available
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: SECUREDISPLAY: optional securedisplay ta ucode is not available
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: SMU is resuming...
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: smu driver if version = 0x0000003d, smu fw if version = 0x00000040,
smu fw program = 0, smu fw version = 0x004e8300 (78.131.0)
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: SMU driver if version not matched
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: SMU is resumed successfully!
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: [drm] DMUB hardware initialized: version=0x07002F00
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring gfx_0.0.0 uses VM inv eng 0 on hub 0
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring comp_1.0.0 uses VM inv eng 1 on hub 0
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring comp_1.1.0 uses VM inv eng 4 on hub 0
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring comp_1.2.0 uses VM inv eng 6 on hub 0
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring comp_1.3.0 uses VM inv eng 7 on hub 0
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring comp_1.0.1 uses VM inv eng 8 on hub 0
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring comp_1.1.1 uses VM inv eng 9 on hub 0
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring comp_1.2.1 uses VM inv eng 10 on hub 0
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring comp_1.3.1 uses VM inv eng 11 on hub 0
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring sdma0 uses VM inv eng 12 on hub 0
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring sdma1 uses VM inv eng 13 on hub 0
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring vcn_unified_0 uses VM inv eng 0 on hub 8
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring vcn_unified_1 uses VM inv eng 1 on hub 8
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring jpeg_dec uses VM inv eng 4 on hub 8
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: ring mes_kiq_3.1.0 uses VM inv eng 14 on hub 0
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: GPU reset(1) succeeded!
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: [drm] ERROR Failed to initialize parser -125!
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: [drm] ERROR Failed to initialize parser -125!
May 15 21:33:42 WufufuArch kernel: amdgpu 0000:2d:00.0: [drm] device wedged, but recovered through reset

Iiskela45 2026-05-19 github

I seem to have a similar issue in Warhammer 40,000: Darktide. I captured an amdgpu coredump during one of the crashes. included it in the attachments.

Ring timed out details
IP Type: 0 Ring Name: gfx_0.0.0

[gfxhub] Page fault observed
Faulty page starting at address: 0x0000000000000000
Protection fault status register: 0x0

What I tried so far:

  • VKD3D_CONFIG=no_upload_hvv
  • Proton 10.4, GE-Proton10-32
  • updating my kernel and drivers

System information

  • OS: Arch Linux x86_64
  • Kernel: 7.0.8-arch1-1
  • Mesa: 1:26.0.6-1
  • vulkan-radeon: 1:26.0.6-1
  • GPU: AMD Radeon RX 7900 XT (Navi 31, RDNA3, GFX11)
  • CPU: Intel Core i7-13700K

Log files

dmesg_darktide.txt
darktide_crash.log
amdgpu_coredump_1779125377.bin.txt

Ssebadamus 2026-05-19 github

I seem to have a similar issue in Warhammer 40,000: Darktide. I captured an amdgpu coredump during one of the crashes. included it in the attachments.

Ring timed out details
IP Type: 0 Ring Name: gfx_0.0.0

[gfxhub] Page fault observed
Faulty page starting at address: 0x0000000000000000
Protection fault status register: 0x0

What I tried so far:

* VKD3D_CONFIG=no_upload_hvv

* Proton 10.4, GE-Proton10-32

* updating my kernel and drivers

System information

* OS: Arch Linux x86_64

* Kernel: 7.0.8-arch1-1

* Mesa: 1:26.0.6-1

* vulkan-radeon: 1:26.0.6-1

* GPU: AMD Radeon RX 7900 XT (Navi 31, RDNA3, GFX11)

* CPU: Intel Core i7-13700K

Log files

dmesg_darktide.txt darktide_crash.log amdgpu_coredump_1779125377.bin.txt

Hi @iskela45,

Can you tell if you have any kind of undervolting in your GPU or GPU or any kind of "playing" tweaking thing in you GPU/CPU? do you have LACT installed?

Iiskela45 2026-05-19 github

@sebadamus No undervolting or overclocking on GPU or CPU. I do technically have LACT installed but the daemon is inactive and there's no profile being applied.

Jjoschka-w 2026-05-20 github

Experiencing the same issue with Subnautica 2 on an RX 7900 GRE (Navi31, RDNA3, GFX11).

System information

  • OS: Arch Linux x86_64
  • Kernel: 7.0.9-arch1-1
  • Mesa: 26.0.6
  • GPU: AMD Radeon RX 7900 GRE (Navi31, RDNA3, GFX11)
  • SMU fw version: 0x004e8300 (78.131.0)
  • DMUB version: 0x07002F00

What was tried (no effect)

  • GE-Proton 10-34
  • Proton 11.0 beta
  • VKD3D_CONFIG=no_upload_hvv
  • RADV_DEBUG=noopt
  • No undervolting or overclocking, no LACT installed

Relevant dmesg

[drm:gfx_v11_0_bad_op_irq [amdgpu]] *ERROR* Illegal opcode in command stream
amdgpu: Process GameThread thread vkd3d_queue
amdgpu: GPU reset succeeded, trying to resume
amdgpu: VRAM is lost due to GPU reset!
amdgpu: device wedged, but recovered through reset

Confirming @TheRealWolfick's observation that Subnautica 2 is affected. The crash occurs randomly during gameplay with no specific trigger identified.

Lluisalvarado 2026-05-20 github

I found this report because of the mentioned of subnautica 2, I have an RTX 5090 and the only way to play the game is to use these commands:

VKD3D_DISABLE_EXTENSIONS=VK_EXT_mesh_shader,VK_NV_raw_access_chains %command%

Is this related to the thread or should I start a new one? Using Kubuntu 26.04 with the 595 drivers. Note that I have played well over 25 hours already without a crash if I use this parameters.

BBlisto91 2026-05-20 github

It is unrelated and a known 5000 series issue. Think Nvidia is aware.

Edit: actually I am mixing it being 5000 series specific with another issue

Ssebadamus 2026-05-20 github

Experiencing the same issue with Subnautica 2 on an RX 7900 GRE (Navi31, RDNA3, GFX11).

System information

* OS: Arch Linux x86_64

* Kernel: 7.0.9-arch1-1

* Mesa: 26.0.6

* GPU: AMD Radeon RX 7900 GRE (Navi31, RDNA3, GFX11)

* SMU fw version: 0x004e8300 (78.131.0)

* DMUB version: 0x07002F00

What was tried (no effect)

* GE-Proton 10-34

* Proton 11.0 beta

* `VKD3D_CONFIG=no_upload_hvv`

* `RADV_DEBUG=noopt`

* No undervolting or overclocking, no LACT installed

Relevant dmesg

[drm:gfx_v11_0_bad_op_irq [amdgpu]] *ERROR* Illegal opcode in command stream
amdgpu: Process GameThread thread vkd3d_queue
amdgpu: GPU reset succeeded, trying to resume
amdgpu: VRAM is lost due to GPU reset!
amdgpu: device wedged, but recovered through reset

Confirming @TheRealWolfick's observation that Subnautica 2 is affected. The crash occurs randomly during gameplay with no specific trigger identified.

So, its just like with DL The Beast? you still hear the environment sounds, image freezed, and you can access SSH from another computer i.e.? thats whats happens to me exactly.

Now I asked @iskela45 about tweaking bios cpu/gpu undervolting and stuff changed from default because I am noticing something new. Recently I upgraded my TUF GAMING X870-PLUS WIFI to latest BIOS Version 1654 2026/04/27, so I had to reset everything to default. As you may know you can run memory in EXPO profile 6400 (UCLK = MCLK/2 for kind of stability) but you surely might have memory glitches i.e. if you run OCCT (back days in Windows) to stress the computer (it wont hang/freeze but will give errors)

So I changed my EXPO profile always to 6000 and everything else default (LACT overclocking enabled but just for leveling up the fan control)

Strangely, I am not having any freezes recently (have played some hours, not to much, will try later when I have time)

I just want to let you know and compare/test.

Ssebadamus 2026-05-21 github

Happened again unfortunately, I tried connecting through HDMI 60hz instead of Display Port 100 hz, in game its vsync off , locked 60fps but it repeats randomly.

Also tried all these kernel options combinations, following some AI guesses, the freeze repeats, too, but options tested!

GRUB_CMDLINE_LINUX_DEFAULT='quiet splash split_lock_detect=off threadirqs amdgpu.gpu_recovery=1 amdgpu.tmz=0 amdgpu.sg_display=0 amdgpu.mcbp=0 amdgpu.aspm=0 amdgpu.gfxoff=0 amdgpu.dcdebugmask=0x410 amdgpu.dcdebugmask=0x410'

# amdgpu.ppfeaturemask options:
# 0xffffffff enable overclocking full
# 0xfffd7fff 
# 0xfff7ffff
# 0xfffd7fff just for fans to work
# 0xfffdbfff only fans
# amdgpu.dcdebugmask=0x410 disables energy saving

And lastly these kernels:

6.19.14
6.18.32
6.17.12
6.17.0 hwe ubuntu 24.04
6.12.90 (this wont let you control LACT gpu fans)
7.0.6-x64v3-xanmod1

Will try updating /lib/firmware/amdgpu, and report back... (update: firmware cant be more updated, the linux-firmware for amdgpu are updated to the latest possible)

Ssebadamus 2026-06-05 github

Dont know if this really means something, but playing with different IA searching, seems my kernel (all that I tried) spects a different SMU version.

sudo dmesg | grep -E "smu.*version|amdgpu.*smu"
[ 6.955227] amdgpu 0000:03:00.0: amdgpu: detected ip block number 4
[ 7.208533] amdgpu 0000:03:00.0: amdgpu: smu driver if version = 0x0000003d, smu fw if version = 0x00000040, smu fw program = 0, smu fw version = 0x004e8200 (78.130.0)

Kernel driver interface version: 0x0000003d (61)
Firmware interface version: 0x00000040 (64)

The mismatch is that the firmware is newer (0x40) than the driver expects (0x3d)

May be you can post if this is also your case?

SSilverwolf-3D 2026-06-11 github

sGot same crash in Satisfactory 1.2

system:
OS: CachyOS
KERNEL: 7.0.11-1-cachyos
CPU: AMD Ryzen 9 7900X 12-Core
GPU: AMD Radeon RX 7900 XT (radeonsi, navi31, ACO, DRM 3.64, 7.0.11-1-cachyos)
GPU DRIVER: 4.6 Mesa 26.1.2-arch2.1
RAM: 63 GB

At the time running GE-proton10-34
Have LACT but only to set fan curve no overclock

Hope it helps

Update:
Not sure but when I turn off FSR, it looks like no more crashes.
But could only play for 2 hours, longer then I could before.

[do jun 11 10:09:25 2026] amdgpu 0000:03:00.0: Dumping IP State
[do jun 11 10:09:25 2026] amdgpu 0000:03:00.0: Dumping IP State Completed
[do jun 11 10:09:25 2026] amdgpu 0000:03:00.0: [drm] AMDGPU device coredump file has been created
[do jun 11 10:09:25 2026] amdgpu 0000:03:00.0: [drm] Check your /sys/class/drm/card1/device/devcoredump/data
[do jun 11 10:09:25 2026] amdgpu 0000:03:00.0: ring gfx_0.0.0 timeout, signaled seq=21361663, emitted seq=21361665
[do jun 11 10:09:25 2026] amdgpu 0000:03:00.0:  Process kwin_wayland pid 1263 thread kwin_wayla:cs0 pid 1273
[do jun 11 10:09:25 2026] amdgpu 0000:03:00.0: Starting gfx_0.0.0 ring reset
[do jun 11 10:09:25 2026] [drm:gfx_v11_0_bad_op_irq [amdgpu]] *ERROR* Illegal opcode in command stream
[do jun 11 10:09:27 2026] amdgpu 0000:03:00.0: MES failed to respond to msg=RESET
[do jun 11 10:09:27 2026] amdgpu 0000:03:00.0: failed to reset legacy queue
[do jun 11 10:09:27 2026] amdgpu 0000:03:00.0: reset via MES failed and try pipe reset -110
[do jun 11 10:09:27 2026] amdgpu 0000:03:00.0: The CPFW hasn't support pipe reset yet.
[do jun 11 10:09:27 2026] amdgpu 0000:03:00.0: Ring gfx_0.0.0 reset failed
[do jun 11 10:09:27 2026] amdgpu 0000:03:00.0: GPU reset begin!. Source:  1
[do jun 11 10:09:29 2026] amdgpu 0000:03:00.0: MES failed to respond to msg=REMOVE_QUEUE
[do jun 11 10:09:29 2026] amdgpu 0000:03:00.0: failed to unmap legacy queue
[do jun 11 10:09:29 2026] [drm:gfx_v11_0_hw_fini.llvm.7217976583856768137 [amdgpu]] *ERROR* failed to halt cp gfx
[do jun 11 10:09:29 2026] amdgpu 0000:03:00.0: MODE1 reset
[do jun 11 10:09:29 2026] amdgpu 0000:03:00.0: GPU mode1 reset
[do jun 11 10:09:29 2026] amdgpu 0000:03:00.0: GPU smu mode1 reset
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: GPU reset succeeded, trying to resume
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: [drm] PCIE GART of 512M enabled (table at 0x0000008001300000).
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: VRAM is lost due to GPU reset!
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: PSP is resuming...
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: reserve 0x1300000 from 0x84fc000000 for PSP TMR
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: RAP: optional rap ta ucode is not available
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: SECUREDISPLAY: optional securedisplay ta ucode is not available
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: SMU is resuming...
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: smu driver if version = 0x0000003d, smu fw if version = 0x00000040, smu fw program = 0, smu fw version = 0x004e8300 (78.131.0)
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: SMU driver if version not matched
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: SMU is resumed successfully!
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: [drm] DMUB hardware initialized: version=0x07003101
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring gfx_0.0.0 uses VM inv eng 0 on hub 0
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring comp_1.0.0 uses VM inv eng 1 on hub 0
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring comp_1.1.0 uses VM inv eng 4 on hub 0
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring comp_1.2.0 uses VM inv eng 6 on hub 0
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring comp_1.3.0 uses VM inv eng 7 on hub 0
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring comp_1.0.1 uses VM inv eng 8 on hub 0
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring comp_1.1.1 uses VM inv eng 9 on hub 0
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring comp_1.2.1 uses VM inv eng 10 on hub 0
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring comp_1.3.1 uses VM inv eng 11 on hub 0
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring sdma0 uses VM inv eng 12 on hub 0
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring sdma1 uses VM inv eng 13 on hub 0
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring vcn_unified_0 uses VM inv eng 0 on hub 8
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring vcn_unified_1 uses VM inv eng 1 on hub 8
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring jpeg_dec uses VM inv eng 4 on hub 8
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: ring mes_kiq_3.1.0 uses VM inv eng 14 on hub 0
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: GPU reset(1) succeeded!
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: [drm] *ERROR* Failed to initialize parser -125!
[do jun 11 10:09:30 2026] amdgpu 0000:03:00.0: [drm] device wedged, but recovered through reset
Ssebadamus 2026-06-26 github

@Silverwolf-3D

I also have LACT just for fan speed control only, if you can, please check if your GPU max clock mhz and max power limit shown in LACT are the real max values of your GPU. In my case card has boost clock: Up to 2680 MHz and game clock: Up to 2510 MHz shown in board specs but inside LACT shows GPU clock max 2940 MHz (I asked LACT github and says that info comes from what kernel informs) so I set that manually max to card specs.
Same thing happens with powerlimit cap (can read about that there https://github.com/ilya-zlobintsev/LACT/issues/1068)

When error occurs, can recover to desktop restarting sddm service from ssh.

What I will test now is cleaning shader caches manually on each run (mesa_shader* and radv_builtin_shader* folders inside game folder and/or ~/.cache/mesa_shader_cache) would need trying some startup variables, too like:

RADV_DEBUG=hang
RADV_DEBUG=syncshaders
RADV_DEBUG=preoptir,nocache (or is it ACO_DEBUG=noopt?)
RADV_DEBUG=checkir,nocache
RADV_DEBUG=llvm
MESA_SHADER_CACHE_DISABLE=true
RADV_DEBUG=nooutoforder,nocache
RADV_DEBUG=invariantgeom,nocache
RADV_DEBUG=nonggc,nocache
(from what I could find in in freedesktop https://gitlab.freedesktop.org/drm/amd/-/work_items interesting info there about this error)

These are my configs for now:
In my motherboard BIOS I disabled an onboard USB module (chip ASM4242) I realized I didnt use, maybe it killed 2 USB ports that shares PCIE lanes (something to do? dont know)
GPU, PCIE fixed to GEN4 (instead of auto)
I have 1 nvme disk installed in the onboard port where it does not share lanes with GPU (each motherboard have some specific configuration, in my case is the nearest nvme connector to the CPU... if you have every nvme port connected PCIE might change from 16X to 8x, share lanes or dont really know how each mother works that out)
EXPO1 6000mhz (also tried 5600 problem persist)
GRUB_CMDLINE_LINUX_DEFAULT='quiet splash split_lock_detect=off amdgpu.ppfeaturemask=0xffffffff amdgpu.gpu_recovery=1'
Changed GPU DigitalPort to the second one (used to be connected to the first near the motherboard)
GE-Proton 10-34 and 11-2

inxi -Gx
Graphics:
Device-1: AMD Navi 31 [Radeon RX 7900 XT/7900 XTX/7900M]
vendor: Sapphire NITRO+ driver: amdgpu v: kernel arch: RDNA-3
bus-ID: 03:00.0
Display: x11 server: X.Org v: 21.1.11 with: Xwayland v: 24.1.6 driver: X:
loaded: amdgpu unloaded: fbdev,modesetting,radeon,vesa dri: radeonsi
gpu: amdgpu resolution: 1920x1080
API: EGL v: 1.5 drivers: radeonsi,swrast platforms:
active: gbm,x11,surfaceless,device inactive: wayland
API: OpenGL v: 4.6 vendor: amd mesa v: PPA glx-v: 1.4 direct-render: yes
renderer: AMD Radeon RX 7900 XTX (radeonsi navi31 ACO DRM 3.64
7.1.1-070101-generic)
API: Vulkan v: 1.3.275 drivers: N/A surfaces: xcb,xlib devices: 2

None of these kernel options changed the problem:
amdgpu.ppfeaturemask=0xfffd3fff
amdgpu.dcdebugmask=0x10 amdgpu.tmz=0 amdgpu.sg_display=0 amdgpu.runpm=0 iommu=soft'
amdgpu.ppfeaturemask=0xffffffff
amdgpu.ppfeaturemask=0xffff7fff #gfxoff
amdgpu.dcdebugmask=0x410 #no power management
amdgpu.dcdebugmask=0x10 #no PSR Panel Self Refresh
amdgpu.lockup_timeout=10000
amdgpu.noretry=0
amdgpu.gfxoff=0
amdgpu.gpu_recovery=1
amdgpu.tmz=0
amdgpu.sg_display=0
amdgpu.mcbp=0
amdgpu.aspm=0
amdgpu.gfxoff=0
amdgpu.dcdebugmask=0x410

Ssebadamus 2026-07-06 github

Have you tried steam launch with: RADV_DEBUG=nocache ?
I set that, cleaned shader caches, 7.1.2-070102-generic (Kubuntu 24.04)
Display Port connection
This kernel boot params added: mem_encrypt=off amdgpu.ppfeaturemask=0xfffd7fff amdgpu.gpu_recovery=1
(might be the amdgpu mask the only one needed in my case to set max gpu/ram speed and fans curve)
This for letting my max gpu/ram to work:
sh -c 'echo "manual" > /sys/class/drm/card1/device/power_dpm_force_performance_level'
sh -c 'echo "1" > /sys/class/drm/card1/device/pp_power_profile_mode'

~~Then I set the max gpu clock to my card max spec minus 50MHz, GPU mV offset -20, VRAM default max 2500 MHz and power ~~
cap to 385w (everything less my card specs)

~~Could finished DL The Beast without any crash in 2k 100hz max setting, vsync on, raytrace, max distance, etc when I used to ~~
have random crashes all the time.

~~What I noticed doing some IA stuff is that GPU went to little over max gpu core speed settings in LACT, max spec is 2680MHz ~~
~~and it reached 2730MHz, so thats why I set it 50 MHz lower, that seemed to limit it in this kind of extreme usage that the games wont reach, so it help.

Nothing works 👎

Ttyrypyrking 2026-07-14 github

This is an issue with RX 7900 AMD series mesa driver specifically. A semi-decent fix is already live for a long time on mesa-git in AUR - newest 26.2.0 is what you need. I'm already working on a full fix that will resolve the rest of the issues introduced with the shaky driver workarounds. I have a live working driver and will test for a short while before pushing for upstream adoption.

TTheRealWolfick 2026-07-14 github

This is an issue with RX 7900 AMD series mesa driver specifically. A semi-decent fix is already live for a long time on mesa-git in AUR - newest 26.2.0 is what you need. I'm already working on a full fix that will resolve the rest of the issues introduced with the shaky driver workarounds. I have a live working driver and will test for a short while before pushing for upstream adoption.

How long typically until that gets adopted by the stable Extra repo?

Ttyrypyrking 2026-07-14 github

I'd say 1-3 month if nothing gets stale to git master, stable release is another issue and I can't say anything on that. I'll device a repo with fixes for this specific issue and all the required steps in a nice script to allow my fellow RX 7900 XT/XTX enjoyers to finally be able to play UE5 titles without artifacts/crashes/hangs caused by a faulty mitigation impl.

This is an issue with RX 7900 AMD series mesa driver specifically. A semi-decent fix is already live for a long time on mesa-git in AUR - newest 26.2.0 is what you need. I'm already working on a full fix that will resolve the rest of the issues introduced with the shaky driver workarounds. I have a live working driver and will test for a short while before pushing for upstream adoption.

How long typically until that gets adopted by the stable Extra repo?

Ttyrypyrking 2026-07-14 github

I've hit about 5 hours combined in Arc Raiders and the latest AC4:BF Resynced. And had 0 issues or crashes during that time, so luckily it's a single simple issue that is specific to only some engines that do specific pointer address truncation technique (UE5 being the single biggest one).

Ttyrypyrking 2026-08-10 github

@TheRealWolfick Update to current stable or git, it doesn't seem to be present anymore. My workaround isn't in, but there seems to be no need for it anymore.

Ggbonnema 2026-08-21 github

Follow-up observation (Borderlands 4 / RX 7600 XT / COSMIC)

Same class of hang as this issue: Illegal opcode in command stream on gfx_0.0.0, then full GPU reset (VRAM is lost). Kernel always names the game side:

  • Process information: process GameThread … thread vkd3d_queue

After the reset, the whole COSMIC Wayland session dies (Steam, browsers, mail, etc.), which is expected once VRAM/contexts are gone — but it made a pattern easier to notice.

Correlation with other GUI clients

While playing Borderlands 4 regularly I found that if I quit essentially all other GUI programs first (Firefox, Chromium, Thunderbird, my own GUI apps — leaving Steam + the game), I can play for hours without the illegal-opcode hang.

If other GUI apps are left running under COSMIC, the hang returns much more often.

One strong natural-experiment case: I had been playing for a long stretch with no other GUI apps and no hang. A scheduled local program opened a GUI window at 21:00, and the illegal-opcode / GPU reset happened shortly afterward.

I have not seen this “other desktop clients as aggravator” noted on the Borderlands FSR tickets (#3034 / #3047). Happy to gather a Proton log / amdgpu devcoredump on the next occurrence if useful.

Environment

  • OS: Pop!_OS 24.04 LTS (COSMIC Wayland session)
  • Kernel: 6.16.3-76061603-generic
  • CPU: AMD Ryzen 9 7950X
  • GPU: AMD Radeon RX 7600 XT (Navi33 / GFX11, 1002:7480), ~16 GiB VRAM
    Secondary: Raphael iGPU (1002:164e, unused for the game)
  • Mesa / RADV: mesa-vulkan-drivers 25.2.8-0ubuntu0.24.04.2
  • Compositor: cosmic-comp 0.1~1787052378~24.04~314fc67
  • Game: Borderlands 4 (Steam AppID 1285190), Unreal Engine 5.5 / DX12 via vkd3d-proton
  • Proton: Proton Experimental experimental-11.0-20260814b
  • Launch options:
    PROTON_ENABLE_WAYLAND=1 VKD3D_CONFIG=retain_descriptor_heaps,defer_resource_destruction %command%
  • In-game (relevant): Upscaling TSR (Balanced), VSync on, FPS capped (~60), HW ray tracing off, frame generation off, Textures High, Effects low

Notes

  • Already on TSR (not FSR) and the retain_descriptor_heaps / defer_resource_destruction workarounds from the BL4 FSR discussions; those helped earlier FSR-related instability, but this illegal-opcode hang still occurs.
  • Closing other GPU clients before play has been a reliable mitigation for me for over a week.
  • Speculating only lightly: extra Wayland/Vulkan/GL clients (browsers etc.) may increase concurrent GPU load / queue pressure and make a GFX11 race more likely; the dmesg culprit thread remains vkd3d_queue, so I am not claiming Firefox submits the bad packet.
Ttyrypyrking 2026-08-21 github

Environment

  • OS: Pop!_OS 24.04 LTS (COSMIC Wayland session)
  • Kernel: 6.16.3-76061603-generic
  • CPU: AMD Ryzen 9 7950X
  • GPU: AMD Radeon RX 7600 XT (Navi33 / GFX11, 1002:7480), ~16 GiB VRAM
    Secondary: Raphael iGPU (1002:164e, unused for the game)
  • Mesa / RADV: mesa-vulkan-drivers 25.2.8-0ubuntu0.24.04.2

Update to mesa 26.1.8-1, I didn't have crashes since minor ver. 6, but 6 and 7 had video decoding issues, so far minor ver. 8 is the most stable with no noticeable issues.

Ppatrickrifici 2026-08-23 github

Update to mesa 26.1.8-1, I didn't have crashes since minor ver. 6, but 6 and 7 had video decoding issues, so far minor ver. 8 is the most stable with no noticeable issues.

Hey tyrypyrkind, your fix might still be required. I'm getting this crash when playing STALKER 2 v2.02 on my 7900 XTX. Note that the crash is infrequent, I was able to play ~4-5 hours on and off today before the crash happened. At the time I had a few desktop apps open as well as a YouTube stream in Microsoft Edge.

Logs are attached, specs are:
OS: Arch Linux x86_64
Desktop Environment: Plasma 6.7.4 (Wayland)
Kernel: Linux 7.1.9-arch1-2
CPU: AMD Ryzen 7 7800X3D
GPU: AMD Radeon RX 7900 XTX
Mesa: 26.2.1
Proton Version: 11.0-2

CrashLogs.zip

Sserhii-nakon 2026-08-23 github

@tyrypyrking Hello I have exactly the same crashes instantly with The Sinking City 2 UE5 (after few minutes gameplay no diff what to do) and very rarely on STALKER 2 Cost Of Hope DLC UE5 and sometimes even in CS2

I uploaded my investigation to claude and he found this topic and told that most likely it related with Mesa
Where I can get your patches for mesa due I just can not play new game The Sinking City 2 at all - only working workaround is VKD3D_CONFIG=nodxr but game work with huge FPS drop with it

Setup
GPU RX7900XTX
CPU AMD R7 7700X
RAM 32GB
Mesa 26.3.0-devel (git-914e6ae217) - almost last one from mesa master
Linux 7.1.8+deb13-amd64
Debian 13
Proton 11.0-2 or GE 11-5
Firmware last ones from master https://gitlab.com/kernel-firmware/linux-firmware

PS: I have setup with iGPU + dGPU so when iGPU used for GUI while dGPU for specific tasks like game or ML/AI - and this setup guard me from GUI crash during this crashes because it reset whole dGPU VRAM

Also I deleted all caches in my system before login in GUI by those commands - to make sure that nothing left from older version

find ~/.local/share/Steam -type f -name 'vkd3d-proton.cache*' -delete
rm -rf .cache/AMD/ .cache/mesa_shader_cache/ .cache/radv_builtin_shaders/

Logs contain: dmesg, amdgpu dump, results of RADV_DEBUG=hang for specific games (I named it)
log_dumps.zip

Sserhii-nakon 2026-08-23 github

Just found the similar issue here https://github.com/doitsujin/dxvk/issues/5841 with DXVK - it something semi broken with Mesa????

Sserhii-nakon 2026-08-23 github

I found this one https://gitlab.freedesktop.org/mesa/mesa/-/work_items/16158 more likely it related

Sserhii-nakon 2026-08-23 github

I found out that VKD3D_CONFIG=single_queue seems works as temporal workaround

Ppatrickrifici 2026-08-28 github

I found out that VKD3D_CONFIG=single_queue seems works as temporal workaround

Can confirm this has worked around the crashing in STALKER 2, thank you very much :)

Iiskela45 2026-08-28 github

I found out that VKD3D_CONFIG=single_queue seems works as temporal workaround

Did not work on my end with Darktide

Sserhii-nakon 2026-08-30 github

Guys try downgrade amdgpu related firmware files to 20260622 version - more likely it should works well... (or 20260519)
We slightly tested different versions and seems those versions should works