Please do a bisect and/or try to provide an apitrace from an environment that doesn't hang, setting up random old WoW builds with private servers and whatnot isn't exactly the kind of thing we can easily do in half an hour.
Will do later.
Also I understand that not what you want, but I kept server build/installation as portable as I could.
In case you interested you could easily setup this in half an hour, just need to polish some scripts.
Ok, since I going to rebuild it to initial state now I took opportunity and doing "manual bisect":
So I'm now on 6d9e0baa2776a13b0a12f0a9cb3290996d0e2110
something happened in between 4 weeks and 3 weeks ago
I'll try to do short apitrace to fit in gh upload limit, either way it hangs almost immediately after login.
EDIT:
While on this, could you please edit wiki to use "subheader" instead '*', hard to read this.
ftom
to
btw noticed dmesg messages:
[425075.869788] amdgpu 0000:05:00.0: amdgpu: Dumping IP State
[425075.872900] amdgpu 0000:05:00.0: amdgpu: Dumping IP State Completed
[425075.883006] amdgpu 0000:05:00.0: amdgpu: ring comp_1.1.0 timeout, signaled seq=23, emitted seq=25
[425075.883014] amdgpu 0000:05:00.0: amdgpu: Process information: process WoW.exe pid 16624 thread dxvk-submit pid 16679
[425075.883016] amdgpu 0000:05:00.0: amdgpu: Starting comp_1.1.0 ring reset
[425305.246781] amdgpu 0000:05:00.0: amdgpu: Dumping IP State
[425305.249345] amdgpu 0000:05:00.0: amdgpu: Dumping IP State Completed
[425305.259375] amdgpu 0000:05:00.0: amdgpu: ring comp_1.1.0 timeout, signaled seq=38, emitted seq=39
[425305.259386] amdgpu 0000:05:00.0: amdgpu: Process information: process WoW.exe pid 17059 thread dxvk-submit pid 17115
[425305.259390] amdgpu 0000:05:00.0: amdgpu: Starting comp_1.1.0 ring reset
[431977.629820] amdgpu 0000:05:00.0: amdgpu: Dumping IP State
[431977.632673] amdgpu 0000:05:00.0: amdgpu: Dumping IP State Completed
[431977.642806] amdgpu 0000:05:00.0: amdgpu: ring comp_1.1.0 timeout, signaled seq=51, emitted seq=53
[431977.642815] amdgpu 0000:05:00.0: amdgpu: Process information: process WoW.exe pid 31889 thread dxvk-submit pid 31941
[431977.642818] amdgpu 0000:05:00.0: amdgpu: Starting comp_1.1.0 ring reset
[433806.038354] [drm:gfx_v9_0_bad_op_irq [amdgpu]] *ERROR* Illegal opcode in command stream
[433806.038708] amdgpu 0000:05:00.0: amdgpu: Dumping IP State
[433806.041293] amdgpu 0000:05:00.0: amdgpu: Dumping IP State Completed
[433806.051386] amdgpu 0000:05:00.0: amdgpu: ring comp_1.2.0 timeout, signaled seq=55, emitted seq=58
[433806.051396] amdgpu 0000:05:00.0: amdgpu: Process information: process WoW.exe pid 7066 thread dxvk-submit pid 7118
[433806.051400] amdgpu 0000:05:00.0: amdgpu: Starting comp_1.2.0 ring reset
[433806.307243] amdgpu 0000:05:00.0: amdgpu: fail to wait on hqd deactive
[433806.307250] amdgpu 0000:05:00.0: amdgpu: Ring comp_1.2.0 reset failure
Yes, you're getting a GPU hang, but given that we have next to no information to work with right now we can't debug that.
Trying to do apitrace now, for some reason it refuse to run.
Figuring out ...
I have virtually the same Error on my System with the same dxvk-git version.
The System is instantly crashing in Star Trek Online (DX11) the instant where the title screen should appear after loading the Game, even with resetted Graphic-Settings.
Even got the gfx ring timeouts...
I rolled back to commit a08579e from March 6th and everything works fine again.
Arch Linux - 6.13.6-zen1
Mesa 25.1.0-devel (git-76883e0b3c)
Linux-Firmware-git
wine-10.2-staging-tkg-amd64
AMD RX 9070 XT
I'll try to bisect as well later this day or tomorrow, but i think things must have been broken in one of the March 7th commits onwards
does apitrace work with that game at least?
Never done Apitrace before, but i'll try my best later when i'm home again.
Sorry, no luck with apitrace. Tried to copy files in different locations, game refuse to run.
But you should have a bunch of traces from other WotLK issues : e.g. https://github.com/doitsujin/dxvk/issues/3819
... I can't find if there is actual apitrace.
EDIT: or maybe WoW in general https://github.com/doitsujin/dxvk/issues?q=is%3Aissue state%3Aopen wow
I haven't reproduced the issue with either of the two d3d9 WoW installs i have. Tested rdna2 and rdna3 via Proton
Edit: They are based on 1.12 and 3.3.5 i gather
lol, now got this on same revision I marked as good ( and played for a while earlier (https://github.com/doitsujin/dxvk/commit/6d9e0baa2776a13b0a12f0a9cb3290996d0e2110))
[449942.222425] [drm:gfx_v9_0_bad_op_irq [amdgpu]] *ERROR* Illegal opcode in command stream
[449942.222819] amdgpu 0000:05:00.0: amdgpu: Dumping IP State
[449942.225793] amdgpu 0000:05:00.0: amdgpu: Dumping IP State Completed
[449942.235867] amdgpu 0000:05:00.0: amdgpu: ring comp_1.1.0 timeout, signaled seq=318, emitted seq=319
[449942.235885] amdgpu 0000:05:00.0: amdgpu: Process information: process WoW.exe pid 10573 thread dxvk-submit pid 10625
[449942.235890] amdgpu 0000:05:00.0: amdgpu: Starting comp_1.1.0 ring reset
[449942.488173] amdgpu 0000:05:00.0: amdgpu: fail to wait on hqd deactive
[449942.488185] amdgpu 0000:05:00.0: amdgpu: Ring comp_1.1.0 reset failure
[449942.488192] amdgpu 0000:05:00.0: amdgpu: GPU reset begin!
[449942.605603] amdgpu 0000:05:00.0: amdgpu: MODE2 reset
[449942.606195] amdgpu 0000:05:00.0: amdgpu: GPU reset succeeded, trying to resume
[449942.606466] [drm] PCIE GART of 1024M enabled.
[449942.606469] [drm] PTB located at 0x000000F400A00000
[449942.606491] amdgpu 0000:05:00.0: amdgpu: PSP is resuming...
[449942.626539] amdgpu 0000:05:00.0: amdgpu: reserve 0x400000 from 0xf47fc00000 for PSP TMR
[449942.688699] amdgpu 0000:05:00.0: amdgpu: RAS: optional ras ta ucode is not available
[449942.696759] amdgpu 0000:05:00.0: amdgpu: RAP: optional rap ta ucode is not available
[449942.696770] amdgpu 0000:05:00.0: amdgpu: SECUREDISPLAY: securedisplay ta ucode is not available
[449942.955570] [drm] kiq ring mec 2 pipe 1 q 0
[449942.974596] amdgpu 0000:05:00.0: amdgpu: ring gfx uses VM inv eng 0 on hub 0
[449942.974604] amdgpu 0000:05:00.0: amdgpu: ring comp_1.0.0 uses VM inv eng 1 on hub 0
[449942.974607] amdgpu 0000:05:00.0: amdgpu: ring comp_1.1.0 uses VM inv eng 4 on hub 0
[449942.974609] amdgpu 0000:05:00.0: amdgpu: ring comp_1.2.0 uses VM inv eng 5 on hub 0
[449942.974612] amdgpu 0000:05:00.0: amdgpu: ring comp_1.3.0 uses VM inv eng 6 on hub 0
[449942.974614] amdgpu 0000:05:00.0: amdgpu: ring comp_1.0.1 uses VM inv eng 7 on hub 0
[449942.974616] amdgpu 0000:05:00.0: amdgpu: ring comp_1.1.1 uses VM inv eng 8 on hub 0
[449942.974617] amdgpu 0000:05:00.0: amdgpu: ring comp_1.2.1 uses VM inv eng 9 on hub 0
[449942.974619] amdgpu 0000:05:00.0: amdgpu: ring comp_1.3.1 uses VM inv eng 10 on hub 0
[449942.974621] amdgpu 0000:05:00.0: amdgpu: ring kiq_0.2.1.0 uses VM inv eng 11 on hub 0
[449942.974624] amdgpu 0000:05:00.0: amdgpu: ring sdma0 uses VM inv eng 0 on hub 8
[449942.974625] amdgpu 0000:05:00.0: amdgpu: ring vcn_dec uses VM inv eng 1 on hub 8
[449942.974627] amdgpu 0000:05:00.0: amdgpu: ring vcn_enc0 uses VM inv eng 4 on hub 8
[449942.974629] amdgpu 0000:05:00.0: amdgpu: ring vcn_enc1 uses VM inv eng 5 on hub 8
[449942.974631] amdgpu 0000:05:00.0: amdgpu: ring jpeg_dec uses VM inv eng 6 on hub 8
[449942.984728] amdgpu 0000:05:00.0: amdgpu: [gfxhub0] no-retry page fault (src_id:0 ring:40 vmid:0 pasid:0)
[449942.984743] amdgpu 0000:05:00.0: amdgpu: in page starting at address 0x0000000000672000 from IH client 0x1b (UTCL2)
[449942.984757] amdgpu 0000:05:00.0: amdgpu: VM_L2_PROTECTION_FAULT_STATUS:0x00040A51
[449942.984765] amdgpu 0000:05:00.0: amdgpu: Faulty UTCL2 client ID: CPC (0x5)
[449942.984772] amdgpu 0000:05:00.0: amdgpu: MORE_FAULTS: 0x1
[449942.984777] amdgpu 0000:05:00.0: amdgpu: WALKER_ERROR: 0x0
[449942.984803] amdgpu 0000:05:00.0: amdgpu: PERMISSION_FAULTS: 0x5
[449942.984822] amdgpu 0000:05:00.0: amdgpu: MAPPING_ERROR: 0x0
[449942.984828] amdgpu 0000:05:00.0: amdgpu: RW: 0x1
[449942.984836] amdgpu 0000:05:00.0: amdgpu: [gfxhub0] no-retry page fault (src_id:0 ring:40 vmid:0 pasid:0)
[449942.984843] amdgpu 0000:05:00.0: amdgpu: in page starting at address 0x0000000000672000 from IH client 0x1b (UTCL2)
[449942.984851] amdgpu 0000:05:00.0: amdgpu: [gfxhub0] no-retry page fault (src_id:0 ring:40 vmid:0 pasid:0)
[449942.984857] amdgpu 0000:05:00.0: amdgpu: in page starting at address 0x0000000000672000 from IH client 0x1b (UTCL2)
[449942.984864] amdgpu 0000:05:00.0: amdgpu: [gfxhub0] no-retry page fault (src_id:0 ring:40 vmid:0 pasid:0)
[449942.984869] amdgpu 0000:05:00.0: amdgpu: in page starting at address 0x0000000000672000 from IH client 0x1b (UTCL2)
[449942.984877] amdgpu 0000:05:00.0: amdgpu: [gfxhub0] no-retry page fault (src_id:0 ring:40 vmid:0 pasid:0)
[449942.984885] amdgpu 0000:05:00.0: amdgpu: in page starting at address 0x0000000000672000 from IH client 0x1b (UTCL2)
[449942.991057] amdgpu 0000:05:00.0: amdgpu: GPU reset(10) succeeded!
[449943.010729] [drm:amdgpu_cs_ioctl [amdgpu]] *ERROR* Failed to initialize parser -125!
thinkpad /tmp #
@Gnatzelle I installed STO locally and it runs fine on RDNA2.
b6af4f04e4401e0c28ab7adb94c832755358885a fixes a validation error that the game ran it, but this should have been harmless anyway.
[#4639](https://github.com/doitsujin/dxvk/issues/4639) : https://drive.proton.me/urls/5J7XAW979M#bCiSR3m3nFWN ?
@pchome Does it hang for you when you replay that apitrace?
@doitsujin I'll try this version, or the actual latest at that point when i'm home again...
And in general...thank you and <3 for everything you all the others do with dxvk, vkd3d, etc. That had to be said ^^"
@Blisto91 Currently back to 2.5.3
I'll try to replay on 9999 (gentoo git) version later.
https://github.com/doitsujin/dxvk/issues/4746#issuecomment-2708866789 👀
Could you give it a go with a GitHub artifact instead so we know the compile is a known good? https://github.com/doitsujin/dxvk/actions/runs/13749285495
Edit: Or try with the fix in the linked comment
avx you mean?
summon @ionen (hope this is the same maintainer name on github)
or restrict this in meson, despite this is 2025
rebuilt with current git version with
"-march=native -O2 -pipe -mno-avx"
no luck
@doitsujin Narrowed down "my" problem...MSAA is causing the Lockups. Started the Game in Safe Mode, cranked up the Graphic Settings to Max, Enabled MSAA > GFX Lockup. FXAA works fine though.
Will Edit this Post (and maybe file a new issue?) when i found the latest working commit
Yeah, msaa-related things have changed a lot recently. Annoying that we're uncovering all sorts of fun driver issues with this now though. FWIW, the game defaults to 8xMSAA, which as mentioned works just fine on my RDNA2 card.
Also, yes, please file a separate issue, things get messy when there's two different things being discussed in one thread.
@Gnatzelle could check if https://github.com/doitsujin/dxvk/tree/desktop-no-secondary-cmdbufs this branch fixes it.
This essentially reverts the MSAA work, which, as we already found out, is also broken on Intel for unknown reasons.
I just found this issue and disabling the MSAA significantly added more FPS and no more random BSOD
It seems like it was indeed multisampling. But in my case it was set to 1x, setting to 8x allow game to load w/o hang.
@Crumdidlyumshis What GPU, driver and game?
I have still not been able to reproduce the issue
After updating mesa to 25.0.1 I can't reproduce the issue too. Setting multisampling back to 1x is ok now.
Maybe it was mesa bug and got fixed in https://docs.mesa3d.org/relnotes/25.0.1.html
If you running git version (e.g. Mesa 25.1.0-devel) it likely contain most fixes.
I'm testing w/ fd9b561
I'm testing w/ fd9b561
@pchome Not surprising, that you can't reproduce it then. That commit disables the code which caused it. Can you reproduce it if you start it with the environment variable DXVK_CONFIG=dxvk.tilerMode=True?
@K0bin But I was on the same commit since yesterday, before mesa update ...
Ok, something strange happening now:
dxvk.tilerMode=True + 1x msaa - hangdxvk.tilerMode=True + 8x msaa - okdxvk.tilerMode=True + 1x msaa - ok!I also noticed sclk was usually sitting on 200MHz, but now after GPU reset it on 900-1000MHz and 1x msaa work again.
I guess I could trigger the bug again after system sleep-wake process, which I suppose what was just happened.
yup, after sleep-wake first run with dxvk.tilerMode=True + 8x msaa - hang
EDIT: first run was just this:
[43459.743798] amdgpu 0000:05:00.0: amdgpu: Dumping IP State
[43459.746829] amdgpu 0000:05:00.0: amdgpu: Dumping IP State Completed
[43459.756929] amdgpu 0000:05:00.0: amdgpu: ring comp_1.1.0 timeout, signaled seq=599, emitted seq=600
[43459.756939] amdgpu 0000:05:00.0: amdgpu: Process information: process WoW.exe pid 21975 thread dxvk-submit pid 22028
[43459.756943] amdgpu 0000:05:00.0: amdgpu: Starting comp_1.1.0 ring reset
sclk stayed on 200MHz
second run with removed dxvk.tilerMode=True hang too, but with this and sclk now at ~700-900 MHz
[44246.175800] amdgpu 0000:05:00.0: amdgpu: Dumping IP State
[44246.178334] amdgpu 0000:05:00.0: amdgpu: Dumping IP State Completed
[44246.188463] amdgpu 0000:05:00.0: amdgpu: ring comp_1.1.0 timeout, signaled seq=612, emitted seq=613
[44246.188477] amdgpu 0000:05:00.0: amdgpu: Process information: process WoW.exe pid 23405 thread dxvk-submit pid 23459
[44246.188495] amdgpu 0000:05:00.0: amdgpu: Starting comp_1.1.0 ring reset
[44246.448670] amdgpu 0000:05:00.0: amdgpu: fail to wait on hqd deactive
[44246.448679] amdgpu 0000:05:00.0: amdgpu: Ring comp_1.1.0 reset failure
[44246.448683] amdgpu 0000:05:00.0: amdgpu: GPU reset begin!
[44246.565973] amdgpu 0000:05:00.0: amdgpu: MODE2 reset
[44246.566852] amdgpu 0000:05:00.0: amdgpu: GPU reset succeeded, trying to resume
[44246.567210] [drm] PCIE GART of 1024M enabled.
[44246.567234] [drm] PTB located at 0x000000F400A00000
[44246.567301] amdgpu 0000:05:00.0: amdgpu: PSP is resuming...
[44246.587365] amdgpu 0000:05:00.0: amdgpu: reserve 0x400000 from 0xf47fc00000 for PSP TMR
[44246.649167] amdgpu 0000:05:00.0: amdgpu: RAS: optional ras ta ucode is not available
[44246.658140] amdgpu 0000:05:00.0: amdgpu: RAP: optional rap ta ucode is not available
[44246.658148] amdgpu 0000:05:00.0: amdgpu: SECUREDISPLAY: securedisplay ta ucode is not available
[44246.921733] [drm] kiq ring mec 2 pipe 1 q 0
[44246.941934] amdgpu 0000:05:00.0: amdgpu: ring gfx uses VM inv eng 0 on hub 0
[44246.941943] amdgpu 0000:05:00.0: amdgpu: ring comp_1.0.0 uses VM inv eng 1 on hub 0
[44246.941946] amdgpu 0000:05:00.0: amdgpu: ring comp_1.1.0 uses VM inv eng 4 on hub 0
[44246.941948] amdgpu 0000:05:00.0: amdgpu: ring comp_1.2.0 uses VM inv eng 5 on hub 0
[44246.941950] amdgpu 0000:05:00.0: amdgpu: ring comp_1.3.0 uses VM inv eng 6 on hub 0
[44246.941952] amdgpu 0000:05:00.0: amdgpu: ring comp_1.0.1 uses VM inv eng 7 on hub 0
[44246.941953] amdgpu 0000:05:00.0: amdgpu: ring comp_1.1.1 uses VM inv eng 8 on hub 0
[44246.941955] amdgpu 0000:05:00.0: amdgpu: ring comp_1.2.1 uses VM inv eng 9 on hub 0
[44246.941956] amdgpu 0000:05:00.0: amdgpu: ring comp_1.3.1 uses VM inv eng 10 on hub 0
[44246.941958] amdgpu 0000:05:00.0: amdgpu: ring kiq_0.2.1.0 uses VM inv eng 11 on hub 0
[44246.941960] amdgpu 0000:05:00.0: amdgpu: ring sdma0 uses VM inv eng 0 on hub 8
[44246.941962] amdgpu 0000:05:00.0: amdgpu: ring vcn_dec uses VM inv eng 1 on hub 8
[44246.941964] amdgpu 0000:05:00.0: amdgpu: ring vcn_enc0 uses VM inv eng 4 on hub 8
[44246.941965] amdgpu 0000:05:00.0: amdgpu: ring vcn_enc1 uses VM inv eng 5 on hub 8
[44246.941967] amdgpu 0000:05:00.0: amdgpu: ring jpeg_dec uses VM inv eng 6 on hub 8
[44246.956912] amdgpu 0000:05:00.0: amdgpu: GPU reset(12) succeeded!
[44249.221496] [drm:amdgpu_cs_ioctl [amdgpu]] *ERROR* Failed to initialize parser -125!
and now (as expected) third run is ok
Can I help with something else?
Or we'll keep this issue open until it will eventually resolved by itself.
I can't say it dxvk bug.
OK, to be fair as last check I run DXVK: v2.5.3-293-g443e7cc9 proposed by @Blisto91 in https://github.com/doitsujin/dxvk/issues/4750#issuecomment-2709002126
runs ok from first attempt, so maybe indeed something with -march=native (except (or not) avx, as I tried disabling this before).
Too early.
err: DxvkSubmissionQueue: Command submission failed: VK_ERROR_DEVICE_LOST
EDIT: going back to stable release.
2.6.0 have regression. wow vanila 1.12 crush on start
2.5.3 work fine
WHQL-R-ID-Software-Hybrid-25.3.1-R2.5-Win10-Win11-PolarisVegaNavi-Sophronia
Is it a GPU hang? If not then please make a new issue and fill out the template.
I have been unable to reproduce this issue. On either amd rdna2 ,rdna3 or Intel Battlemage. I also switched to my old 380x today to give it a go but it seems to be unstable in general unrelated to this issue.
I guess apitrace is pointless this time, right?
But, you know, I,m kind of developer myself, so if you want more info? more debug info ... I know how to build it w/ -g
It's something in dxvk that triggering hw/mesa/drm/"we need to go deeper"/... bug, as for me.
Need more info? Just ask.
If you could do a bisect that would be great. And if you are able to get a apitrace we wouldn't say no to that either.
I have a feeling that bisecting here is just going to lead to misleading and useless results.
What we really need is a way to reliably reproduce this ourselves.
yup, but for now everything you have is 1-2 users (yet)
if this is some kind of hardware bug it's unlikely you can deal with this
If there are some points where I can set brakepoint or something else (branch w/ printf , etc...)
EDIT: not the first time, I already have tcs bug on this laptop, and have similary long conversation on rpcs3 issue tracker.
I got an apitrace of that Lego game and can reproduce the hang on my Renoir laptop. Works on everything else + no obvious validation errors, so it's most likely a very hardware-specific driver bug, but at least I can actually sort of debug that one.
https://github.com/doitsujin/dxvk/tree/vega-sparse-workaround this works around it on my end.
ok, rebuilding with this
p.s. would be good if you officially support unity builds
regular (re)builds are too long.
Seems ok, sleep-resume and regular run are ok so far.
merged workaround now.
DXVK_CONFIG=dxvk.tilerMode=True`?x1 2025-03
Decided to preview current (1b529c7) DXVK before release, got hung after login and
Software information
WoW WotLK 3.3.5a (local Azeroth Core server)
System information
Apitrace file(s)
none, by request
Log files
none, by request