protonscr

EVE Online hangs AMD RX 6900 XT with RADV

vkd3dopen
HansKristian-Work/vkd3d-proton#2441 · opened 2025-04-16 by Venemo · updated 2025-04-23 · 18 comments · github
VVenemo 2025-04-16 github

Software information

I'm running EVE Online from Steam in D3D12 mode, with all graphics settings set to High.

  • The hang seems to happen "randomly" while playing the game. I haven't noticed any specific steps to reproduce it, but it has happened to me a few times over the past few days.
  • I haven't played the game for years so it's unclear if this is a regression or not. It seemed stable 2 years ago, but not sure if that statement is helpful or not.
  • According to a dev blog article, the game has recently switched to GPU driven rendering, which I assume means this is now one of the few games utilizing DGC.
  • It's unclear to me whether the bug is with VKD3D-Proton, or with RADV, or a game bug. I welcome any suggestions on how to diagnose this.

System information

  • GPU: AMD RX 6900 XT (Navi 21) - can test with other GPUs if that helps.
  • Driver: RADV 25.0.3
  • Proton version: latest Proton Hotfix as of 2025. april 16. Unclear how to check which version that is exactly.
  • VKD3D-Proton version: whichever version is included in the above Proton version. Unclear how to check exactly.

Log files

The issue casues a GPU hang, and I haven't found a good way to grab any Proton logs in a way that survives a hang. Let me know if there is any, or if there is any usefulness.

dmesg log

The dmesg log indicates some page faults happening for which the vkd3d_queue thread is responsible. I have no clue why the same page fault is repeated several times.

You can add an image or a code block, too.

[38174.516100] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38174.516105] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38174.516107] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800043b5f000 from client 0x1b (UTCL2)
[38174.516110] amdgpu 0000:03:00.0: amdgpu: GCVM_L2_PROTECTION_FAULT_STATUS:0x00601431
[38174.516112] amdgpu 0000:03:00.0: amdgpu: 	 Faulty UTCL2 client ID: SQC (data) (0xa)
[38174.516113] amdgpu 0000:03:00.0: amdgpu: 	 MORE_FAULTS: 0x1
[38174.516114] amdgpu 0000:03:00.0: amdgpu: 	 WALKER_ERROR: 0x0
[38174.516116] amdgpu 0000:03:00.0: amdgpu: 	 PERMISSION_FAULTS: 0x3
[38174.516117] amdgpu 0000:03:00.0: amdgpu: 	 MAPPING_ERROR: 0x0
[38174.516118] amdgpu 0000:03:00.0: amdgpu: 	 RW: 0x0
[38174.516122] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38174.516124] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38174.516126] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800043b5f000 from client 0x1b (UTCL2)
[38174.516130] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38174.516132] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38174.516134] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800043b5f000 from client 0x1b (UTCL2)
[38174.516138] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38174.516140] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38174.516141] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800043b5f000 from client 0x1b (UTCL2)
[38174.516145] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38174.516147] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38174.516149] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800043b5f000 from client 0x1b (UTCL2)
[38174.516153] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38174.516154] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38174.516156] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800043b5f000 from client 0x1b (UTCL2)
[38174.516164] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38174.516165] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38174.516167] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800043b5f000 from client 0x1b (UTCL2)
[38174.516176] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38174.516178] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38174.516179] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800043b5f000 from client 0x1b (UTCL2)
[38174.516183] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38174.516185] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38174.516187] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800043b5f000 from client 0x1b (UTCL2)
[38174.516191] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38174.516193] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38174.516195] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800065b9c000 from client 0x1b (UTCL2)
[38184.877761] amdgpu 0000:03:00.0: amdgpu: Dumping IP State
[38184.879712] amdgpu 0000:03:00.0: amdgpu: Dumping IP State Completed
[38184.879819] gmc_v10_0_process_interrupt: 7 callbacks suppressed
[38184.879822] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38184.879825] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38184.879827] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x00008000a0ffb000 from client 0x1b (UTCL2)
[38184.879829] amdgpu 0000:03:00.0: amdgpu: GCVM_L2_PROTECTION_FAULT_STATUS:0x00601431
[38184.879830] amdgpu 0000:03:00.0: amdgpu: 	 Faulty UTCL2 client ID: SQC (data) (0xa)
[38184.879832] amdgpu 0000:03:00.0: amdgpu: 	 MORE_FAULTS: 0x1
[38184.879833] amdgpu 0000:03:00.0: amdgpu: 	 WALKER_ERROR: 0x0
[38184.879834] amdgpu 0000:03:00.0: amdgpu: 	 PERMISSION_FAULTS: 0x3
[38184.879835] amdgpu 0000:03:00.0: amdgpu: 	 MAPPING_ERROR: 0x0
[38184.879836] amdgpu 0000:03:00.0: amdgpu: 	 RW: 0x0
[38184.879840] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38184.879842] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38184.879843] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800086d0d000 from client 0x1b (UTCL2)
[38184.879848] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38184.879849] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38184.879851] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x00008000a0ffb000 from client 0x1b (UTCL2)
[38184.879855] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38184.879856] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38184.879857] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800086d0d000 from client 0x1b (UTCL2)
[38184.879861] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38184.879862] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38184.879864] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x00008000a0ffb000 from client 0x1b (UTCL2)
[38184.879868] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38184.879869] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38184.879870] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800086d0d000 from client 0x1b (UTCL2)
[38184.879874] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38184.879875] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38184.879877] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x00008000a0ffb000 from client 0x1b (UTCL2)
[38184.879881] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38184.879882] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38184.879883] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800086d0d000 from client 0x1b (UTCL2)
[38184.879888] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38184.879889] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38184.879890] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x00008000a0ffb000 from client 0x1b (UTCL2)
[38184.879895] amdgpu 0000:03:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:6 pasid:32803)
[38184.879896] amdgpu 0000:03:00.0: amdgpu:  in process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38184.879897] amdgpu 0000:03:00.0: amdgpu:   in page starting at address 0x0000800086d0d000 from client 0x1b (UTCL2)
[38184.889760] amdgpu 0000:03:00.0: amdgpu: ring gfx_0.0.0 timeout, signaled seq=8134045, emitted seq=8134047
[38184.889763] amdgpu 0000:03:00.0: amdgpu: Process information: process exefile.exe pid 59901 thread vkd3d_queue pid 60100
[38184.889770] amdgpu 0000:03:00.0: amdgpu: Starting gfx_0.0.0 ring reset
[38185.100224] amdgpu 0000:03:00.0: amdgpu: Ring gfx_0.0.0 reset failure
[38185.100228] amdgpu 0000:03:00.0: amdgpu: GPU reset begin!
[38185.401731] amdgpu 0000:03:00.0: amdgpu: MODE1 reset
[38185.401735] amdgpu 0000:03:00.0: amdgpu: GPU mode1 reset
[38185.401792] amdgpu 0000:03:00.0: amdgpu: GPU smu mode1 reset
[38185.911133] amdgpu 0000:03:00.0: amdgpu: GPU reset succeeded, trying to resume
Ddoitsujin maintainer 2025-04-16 github

Do you have a full hang report or anything else that could help narrow this down?

"hangs every now and then" isn't exactly the most actionable thing in the world.

which I assume means this is now one of the few games utilizing DGC.

Doesn't have to be, trivial cases are still covered by plain DrawIndirect. It does however increase the chances of the game just having sync bugs, pretty much everything lately is broken in one way or another in that regard.

VVenemo 2025-04-17 github

Do you have a full hang report or anything else that could help narrow this down?

Running a game with RADV_DEBUG=hang for hours hoping to catch a hang isn't an option due to the perf hit. But I'm open to suggestions if you have any.

"hangs every now and then" isn't exactly the most actionable thing in the world.

I know, I'm sorry.

It does however increase the chances of the game just having sync bugs, pretty much everything lately is broken in one way or another in that regard.

Do you have some ideas to try to mitigate that?

BBillli11 2025-04-20 github

You can always try adding amdgpu.mcbp=0 to kernel parameter.

VVenemo 2025-04-20 github

You can always try adding amdgpu.mcbp=0 to kernel parameter.

What does any of this have to do with mcbp?

BBillli11 2025-04-20 github

I do not know how it help but kernel used to have random gpu page fault with mcbp enable starting with kernel 6.5. I think it is fixed though

And disabling it on my machine do help with system stability.
I'm on RDNA3 with it may not apply to you.

It does not hurt to try.

VVenemo 2025-04-22 github

According to the logs, the page fault comes from the "vkd3d_queue" thread, so I don't think this has anything to do with mcbp. But sure, I can try disabling it.

HHansKristian-Work maintainer 2025-04-22 github

0x0000800043b5f000 looks like an error with descriptor heap OOB.

VVenemo 2025-04-22 github

0x0000800043b5f000 looks like an error with descriptor heap OOB.

Do you have a suggestion as to how to diagnose this issue?

HHansKristian-Work maintainer 2025-04-22 github

Really need some actionable logs to have any chance to debug this.

  • UMR wave dumps that can prove it's faulting on loading a descriptor.
  • Breadcrumbs observing stable shader hashes
  • Running with instruction_qa_checks enabled build with expect-assume to catch OOB.
  • Descriptor_qa_checks could maybe also be used.
VVenemo 2025-04-22 github

Is there any way to get any of that without RADV_DEBUG=hang?

HHansKristian-Work maintainer 2025-04-22 github

Disable GPU recovery and run UMR wave dump on the dead machine (needs SSH).

VVenemo 2025-04-22 github

How do you get breadcrumbs without RADV_DEBUG=hang? Does vkd3d-proton have its own way to get them?

Running with instruction_qa_checks enabled build with expect-assume to catch OOB.
Descriptor_qa_checks could maybe also be used.

I'm not familiar with either of these. Is there an environment variable or do I need a custom built vkd3d-proton for these?

HHansKristian-Work maintainer 2025-04-22 github

Use bleeding-edge debug branch, then you can do VKD3D_CONFIG=breadcrumbs and it will appear in vkd3d-proton log.

For descriptor_qa_checks you need a custom build. Just try to get a wave dump for now to confirm why it's page faulting.

VVenemo 2025-04-22 github

Just try to get a wave dump for now to confirm why it's page faulting.

Do you mean without the bleeding edge debug branch or any custom build or env vars?

HHansKristian-Work maintainer 2025-04-22 github

Nothing custom, just umr after GPU is hung. Same as what RADV does automatically when RADV_DEBUG=hang is used.

VVenemo 2025-04-23 github

What umr command should I use to get the info you need?

HHansKristian-Work maintainer 2025-04-23 github

Based on https://gitlab.freedesktop.org/mesa/mesa/-/blob/main/docs/drivers/amd/hang-debugging.rst?ref_type=heads

umr -O bits,halt_waves,full_shader -go 0 -wa gfx_0.0.0 -go 1 >waves.log 2>&1

VVenemo 2025-04-23 github

Will do. The hang is very rare but I'll write here when I managed to reproduce it and got the umr info.

Proton versions

Launch options