protonscr

Monster Hunter Wilds: Navi31 SQC GPUVM protection fault; retain_descriptor_heaps not yet tested

vkd3dopen
HansKristian-Work/vkd3d-proton#3290 · opened 2026-09-13 by firefloc · updated 2026-09-13 · 1 comments · github
Ffirefloc 2026-09-13 github

Related to #3034.

Summary

I am seeing repeated hard AMDGPU resets while playing Monster Hunter Wilds
(Steam AppID 2246340) through Proton. This is an observation report only: I
have not yet tested VKD3D_CONFIG=retain_descriptor_heaps and make no
claim that it fixes this case.

Observed environment

  • AMD Radeon RX 7900 XTX (Navi31 / gfx1100), RADV
  • Mesa / vulkan-radeon 26.2.2-2
  • CachyOS kernel 7.2.3-1
  • Proton Experimental selected in Steam
  • The most recent affected launch had only
    PROTON_ENABLE_WAYLAND=1 game-performance %command% as user launch
    options. It had no VKD3D_CONFIG and no RADV_DEBUG=hang.

REFramework and ReShade were installed, so they remain confounders. RenoDX
was not deployed for the most recent affected run. The saved game config at
launch reported Algorithm=FSR3, AlgorithmFG=FSR3, and
FrameGenerationMode=On; REFramework also logged the FSR frame-generation
DLL being loaded. These are configuration/log observations, not an assertion
about the in-game UI state.

Kernel signature from the most recent affected run

amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:4 pasid:561)
amdgpu: Process MonsterHunterWi pid 64738 thread vkd3d_queue pid 64796
amdgpu: in page starting address 0x0000030000008000 from client 10
amdgpu: GCVM_L2_PROTECTION_FAULT_STATUS:0x00401430
amdgpu: Faulty UTCL2 client ID: SQC (data) (0xa)
amdgpu: PERMISSION_FAULTS: 0x3
amdgpu: MAPPING_ERROR: 0x0
amdgpu: ring gfx_0.0.0 timeout
amdgpu: MES -110
amdgpu: VRAM lost due GPU reset!

The display session hard-locked and was recovered with REISUB. There was no
out-of-memory report in the captured kernel output.

On an earlier occurrence, a RADV hang capture reported the same failing page
address and VA not found in its address-binding report:

Failing VM page: 0x30000008000
CLIENT_ID: (SQC (data)) 0xa
PERMISSION_FAULTS: 3
MAPPING_ERROR: 0
VA not found!

That capture used RADV_DEBUG=hang, which implicitly enabled
syncshaders and reduced performance substantially, so it is not suitable
for normal gameplay testing.

Planned single-variable follow-up

I will repeat the workload with the same baseline and add only
VKD3D_CONFIG=retain_descriptor_heaps, then update this issue with the exact
launch command, duration, performance impact, and whether this same kernel
signature recurs.

PPhotonkannon 2026-09-13 github

Same when I play Dragon's Dogma 2 (RE Engine specific??) on Fedora 43 using Proton Experimental. No launch options or mods. Vanilla install. Crashes occur randomly. Could be 10 minutes or 10 hours of game play. Usually I can get 1+ hours regularly, but about 10% of the time I get the same crash within the first hour.

System Information:

OS: Fedora Linux 43 (KDE Plasma Desktop Edition) x86_64
Kernel: Linux 7.2.4-100.fc43.x86_64
DE: KDE Plasma 6.7.5
WM: KWin (Wayland)
CPU: AMD Ryzen 5 9600X (12) @ 5.49 GHz
GPU 1: AMD Radeon RX 6750 XT


Steam Beta Branch:  Stable Client
Steam Version:  1788652215
Steam Client Build Date:  Wed, Sep 2, 2026 8:32 PM UTC -08:00
Steam Web Build Date:  Sat, Sep 5, 2026 6:13 PM UTC -08:00
Steam API Version:  SteamClient023

GPU + Driver

# lspci -vnn | grep -A3 VGA
03:00.0 VGA compatible controller [0300]: Advanced Micro Devices, Inc. [AMD/ATI] Navi 22 [Radeon RX 6700/6700 XT/6750 XT / 6800M/6850M XT] [1002:73df] (rev c0) (prog-if 00 [VGA controller])
        Subsystem: Tul Corporation / PowerColor Device [148c:2419]
        Flags: bus master, fast devsel, latency 0, IRQ 108, IOMMU group 15
        Memory at f800000000 (64-bit, prefetchable) [size=16G]

# glxinfo | grep "OpenGL version"
OpenGL version string: 4.6 (Compatibility Profile) Mesa 25.3.6


Initial logs. I'll try VKD3D_CONFIG=retain_descriptor_heaps, RADV_DEBUG=hang, and PROTON_LOG=1 launch options next time I play.

coredumpctl list --reverse
TIME                           PID  UID  GID SIG     COREFILE     EXE                                                           >
Sun 2026-09-13 00:59:02 CDT  21667 1000 1000 SIGABRT present      /usr/bin/plasma-keyboard                                      >
Sun 2026-09-13 00:44:28 CDT  21263 1000 1000 SIGABRT present      /usr/bin/Xwayland                                             >
Sun 2026-09-13 00:27:14 CDT   3718 1000 1000 SIGABRT present      /usr/bin/plasma-keyboard                                      >
Sun 2026-09-13 00:27:00 CDT   3943 1000 1000 SIGABRT present      /usr/libexec/org_kde_powerdevil                               >
Sun 2026-09-13 00:27:00 CDT   3730 1000 1000 SIGABRT present      /usr/bin/Xwayland                  
journalctl -k --since "2026-09-13 00:25" --until "2026-09-13 00:45" | grep -iE "amdgpu|gpu|reset|hang|fault|error"   
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0: [gfxhub] page fault (src_id:0 ring:24 vmid:5 pasid:606)
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:  Process DD2.exe pid 14376 thread vkd3d_queue pid 14455
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:   in page starting at address 0x0000800052347000 from client 0x1b (UTCL2)
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0: GCVM_L2_PROTECTION_FAULT_STATUS:0x00501431
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:          Faulty UTCL2 client ID: SQC (data) (0xa)
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:          MORE_FAULTS: 0x1
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:          WALKER_ERROR: 0x0
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:          PERMISSION_FAULTS: 0x3
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:          MAPPING_ERROR: 0x0
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:          RW: 0x0
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0: [gfxhub] page fault (src_id:0 ring:24 vmid:5 pasid:606)
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:  Process DD2.exe pid 14376 thread vkd3d_queue pid 14455
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:   in page starting at address 0x0000800052347000 from client 0x1b (UTCL2)
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0: [gfxhub] page fault (src_id:0 ring:24 vmid:5 pasid:606)
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:  Process DD2.exe pid 14376 thread vkd3d_queue pid 14455
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:   in page starting at address 0x0000800052347000 from client 0x1b (UTCL2)
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0: [gfxhub] page fault (src_id:0 ring:24 vmid:5 pasid:606)
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:  Process DD2.exe pid 14376 thread vkd3d_queue pid 14455
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:   in page starting at address 0x0000800052347000 from client 0x1b (UTCL2)
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0: [gfxhub] page fault (src_id:0 ring:24 vmid:5 pasid:606)
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:  Process DD2.exe pid 14376 thread vkd3d_queue pid 14455
Sep 13 00:26:51 reddevil kernel: amdgpu 0000:03:00.0:   in page starting at address 0x0000800052347000 from client 0x1b (UTCL2)
Sep 13 00:26:53 reddevil kernel: amdgpu 0000:03:00.0: Dumping IP State
Sep 13 00:26:53 reddevil kernel: amdgpu 0000:03:00.0: Dumping IP State Completed
Sep 13 00:26:53 reddevil kernel: amdgpu 0000:03:00.0: [drm] AMDGPU device coredump file has been created
Sep 13 00:26:53 reddevil kernel: amdgpu 0000:03:00.0: [drm] Check your /sys/class/drm/card1/device/devcoredump/data
Sep 13 00:26:53 reddevil kernel: amdgpu 0000:03:00.0: ring gfx_0.0.0 timeout, signaled seq=32029962, emitted seq=32029965
Sep 13 00:26:53 reddevil kernel: amdgpu 0000:03:00.0:  Process DD2.exe pid 14376 thread vkd3d_queue pid 14449
Sep 13 00:26:53 reddevil kernel: amdgpu 0000:03:00.0: Starting gfx_0.0.0 ring reset
Sep 13 00:26:53 reddevil kernel: amdgpu 0000:03:00.0: Ring gfx_0.0.0 reset failed
Sep 13 00:26:53 reddevil kernel: amdgpu 0000:03:00.0: GPU reset begin!. Source:  1
Sep 13 00:26:53 reddevil kernel: amdgpu 0000:03:00.0: MODE1 reset
Sep 13 00:26:53 reddevil kernel: amdgpu 0000:03:00.0: GPU mode1 reset
Sep 13 00:26:53 reddevil kernel: amdgpu 0000:03:00.0: GPU smu mode1 reset
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: GPU reset succeeded, trying to resume
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: [drm] PCIE GART of 512M enabled (table at 0x0000008000F00000).
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: VRAM is lost due to GPU reset!
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: PSP is resuming...
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: reserve 0xa00000 from 0x82fd000000 for PSP TMR
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: RAS: optional ras ta ucode is not available
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: SECUREDISPLAY: optional securedisplay ta ucode is not available
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: SMU is resuming...
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: use vbios provided pptable
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: SMU is resumed successfully!
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: kiq ring mec 2 pipe 1 q 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: [drm] DMUB hardware initialized: version=0x02020022
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring gfx_0.0.0 uses VM inv eng 0 on hub 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring gfx_0.1.0 uses VM inv eng 1 on hub 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring comp_1.0.0 uses VM inv eng 4 on hub 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring comp_1.1.0 uses VM inv eng 5 on hub 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring comp_1.2.0 uses VM inv eng 6 on hub 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring comp_1.3.0 uses VM inv eng 7 on hub 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring comp_1.0.1 uses VM inv eng 8 on hub 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring comp_1.1.1 uses VM inv eng 9 on hub 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring comp_1.2.1 uses VM inv eng 10 on hub 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring comp_1.3.1 uses VM inv eng 11 on hub 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring kiq_0.2.1.0 uses VM inv eng 12 on hub 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring sdma0 uses VM inv eng 13 on hub 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring sdma1 uses VM inv eng 14 on hub 0
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring vcn_dec_0 uses VM inv eng 0 on hub 8
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring vcn_enc_0.0 uses VM inv eng 1 on hub 8
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring vcn_enc_0.1 uses VM inv eng 4 on hub 8
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: ring jpeg_dec uses VM inv eng 5 on hub 8
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: GPU reset(2) succeeded!
Sep 13 00:26:54 reddevil kernel: amdgpu 0000:03:00.0: [drm] device wedged, but no recovery needed

Proton versions

Launch options