I've patched and downgraded to 440.100 but with Kernel 5.8. I'll try to find a moment to leave the game running for at least 24 hours and see if I can replicate this bug. If I can't get this to happen again I'm really of the strong opinion Nvidia has introduced a bug into their drivers between the 44X and 45X series. Hopefully this information is useful for the dxvk devs.
We're aware of this but there's not an awful lot we can do about it on the DXVK side of things at this point.
So I tested 440.100 for 3 days (never restarted the client) and 0 issues 0 crashes. So if people want to continue using DXVK they should hold onto 440.100 as long as possible.
Is 450.66 the first version where this issue was introduced? The Nvidia website also has 450.57 and 450.51. Knowing the exact version might be helpful for a bug report.
450.66 was the first version to have this issue from what I tested. Thankfully I've tested the latest long lived version and this issue appears to be fixed. I ran the tests again for 3 days continuous game client usage. I experienced no errors with the following:
Final Fantasy XIV
None
I am unsure if the current short lived version (455.38 as of this post) has been fixed.
Despite all this testing I ended up with another crash on 450.80.02. Guess we wait more sadly.
Hmm that's interesting because 450.80.02 does not mention any fixes but only support for some new GPUs. Though i have no idea how consistent nvidia it with their changelog. I'm pretty sure i saw the issue pop up on 455.38 when testing something unrelated though. But that may have been triggered by something else. Going to keep an eye out for it if it pops up again.
Nvidia is aware of the problem and looking into it.
Nvidia is aware of the problem and looking into it.
Do you have a link about it?
Any updates on this issue?
That's generally not a very useful question to ask; if there was then someone would likely have posted it already. Unless Nvidia figure out whatever it is that's causing their driver to crash, this will not get fixed.
Occurs on 460.27.4 as well
455.45.01 also broken, this really makes me want to switch from nvidia to amd...
A possible fix but this is completely anecdotal. Try either clearing out your ~/.nv/GLCache/ or doing export __GL_SHADER_DISK_CACHE=0 with whatever way you're using wine. I haven't had a crash in awhile and this is the only thing I can think that I did lately that might affect DXVK.
There's now a post over at the Nvidia forums about this issue but I don't really have a high confidence in Nvidia fixing this (at least not with the replies I'm seeing on the post from them): Vk_error_device_lost in many game titles.
460.32.03 broken as well
Just to let everyone know I haven't had any issues since using export __GL_SHADER_DISK_CACHE=0. Currently on 460.32.03. Can others here test that?
Okay I had another crash so that doesn't work either, at this point I give up. People seeing this bug report need to go to the thread posted on the nvidia forums and actually post or this is never going to get fixed.
So please go add your voice to Vk_error_device_lost in many game titles.
Still broken 460.39.0.
~Just to let everyone know I haven't had any issues since using
export __GL_SHADER_DISK_CACHE=0. Currently on460.32.03. Can others here test that?~Okay I had another crash so that doesn't work either, at this point I give up. People seeing this bug report need to go to the thread posted on the nvidia forums and actually post or this is never going to get fixed.
So please go add your voice to Vk_error_device_lost in many game titles.
Hello im facing the same bug for about 6 months. I tried many blind workarounds to no avail lastly ive installed lutris added wine 6.0 and built my kernel and out of tree modules with clang. I couldnt reproduce it for 2 weeks still so early to talk but i just wonder if there are any changes in the dxvk code that may fix that or my blind tries finally gave me a solution. And i built nvidia-drivers with dynamic libraries. I hope this helps, I use Gentoo Linux where those workarounds are very easy to apply.
It's fun when you find out mpv also has the same problem with nvidia drivers: https://github.com/mpv-player/mpv/issues/8360
I still havent got any vk_error_device_lost yet with my shot but i sure ill get it at some time. Did you read the driver recommendation there? Please test one of those drivers and report here. Once i get that error again ill try one of those drivers too. Now no errors yet for 2 weeks.
January 27th, 2021 - Windows 457.88, Linux 455.50.04
Fixed a bug in a stencil-buffer optimization that could occasionally result in VK_ERROR_DEVICE_LOST
@arabcian my results are posted throughout this thread, multiple people are still having the issue and my driver version is far higher than yours sitting at 460.39.0 so I am pretty sure the fix your listed has nothing to do with the problem as was also mentioned on the mpv bug report where they don't make use of stencil-buffers either. I also clocked up 3 days of continuous play time only to have the bug randomly come back when I thought it was fixed.
@konomikitten
Driver 460.39.0 is not necessarily "newer" than 455.50.04. The 455.50.xx driver series is of the "vulkan beta" series, and contains a lot more "vulkan fixes" than the "release driver" from the 460.xx series.
To simplify: 455.50.04 can contain a vulkan fix, the 460.39 driver does NOT contain since they are from two different branches.
When it comes to iffy nVidia random crashes + DXVK - Its always best to try the latest "vulkan beta" driver branch from here: https://developer.nvidia.com/vulkan-driver
@SveSop I stand corrected, that's some confusing version numbering.
@arabcian how many hours of FFXIV have you played to test for the crash by the way?
Just to let everyone know I am currently testing 455.50.04 as @arabcian recommended, it's looking good so far but considering the difficult nature of replicating this bug I'm going to be testing for a few weeks or more to make sure it is in fact fixed. If anyone else could test too that would be helpful.
Still broken. 455.50.04
@arabcian how many hours of FFXIV have you played to test for the crash by the way?
I couldnt test it yet with FF since dont have the game but i was having same error with thief 2014 and World of Warcraft + Skyrim + Fallout 4,somehow bug is not around for a month. Still having mpd bug tho but im not sure if both are related.
Currently testing 455.50.07...
Still broken. 455.50.07
This should be fixed with 455.50.10:
Fixed a bug with the host-visible device-local memory heap, where if an allocation failed due to space constraints, it could cause the application to crash on future Vulkan function calls
@doitsujin oh thank god, I'll test it asap.
@doitsujin nope.... Still Broken 455.50.10.
@konomikitten
What is shown in syslog? (dmesg |grep -i nvrm ... or dmesg|grep -i xid)
It could be a faulty card maybe, but ofc if it is only 1 single title that have this problem, while all other (native/whatever) works fine its kinda hard to tell. New drivers sometimes push more performance from certain functions, that MAY make a borderline defective card crash more often.
What is shown in syslog? (
dmesg |grep -i nvrm... ordmesg|grep -i xid)
$ sudo dmesg | grep -Ei 'nvrm|xid'
[ 1.209379] r8169 0000:08:00.0 eth0: RTL8168g/8111g, e0:d5:5e:a6:68:bb, XID 4c0, IRQ 44
[ 4.804576] NVRM: loading NVIDIA UNIX x86_64 Kernel Module 455.50.10 Thu Mar 4 20:25:58 UTC 2021
It could be a faulty card maybe.
440.100 never crashed not even once and I can happily downgrade my kernel and nvidia drivers and never see this crash, which I have done multiple times.
What is shown in syslog? (
dmesg |grep -i nvrm... ordmesg|grep -i xid)$ sudo dmesg | grep -Ei 'nvrm|xid' [ 1.209379] r8169 0000:08:00.0 eth0: RTL8168g/8111g, e0:d5:5e:a6:68:bb, XID 4c0, IRQ 44 [ 4.804576] NVRM: loading NVIDIA UNIX x86_64 Kernel Module 455.50.10 Thu Mar 4 20:25:58 UTC 2021It could be a faulty card maybe.
440.100never crashed not even once and I can happily downgrade my kernel and nvidia drivers and never see this crash, which I have done multiple times.
I remember you giving a code to run videos with vulkan in mpv and resizing windows was reproducing the same bug we experience. What was that command? Because seems like i cant reproduce it anymore if im not mistaken with the command. Driver is 455.50.10
I remember you giving a code to run videos with vulkan in mpv and resizing windows was reproducing the same bug we experience. What was that command?
I never gave any command related to mpv I did link to https://github.com/mpv-player/mpv/issues/8360 where people using the vulkan backend with mpv were having similar crashes. As to why they get a kernel log involving Xid I imagine it's due to doing video playback, but I'm not entirely sure, the DXVK crash in this thread has never produced any error messages in the kernel log for me.
I remember you giving a code to run videos with vulkan in mpv and resizing windows was reproducing the same bug we experience. What was that command?
I never gave any command related to mpv I did link to mpv-player/mpv#8360 where people using the vulkan backend with mpv were having similar crashes. As to why they get a kernel log involving
XidI imagine it's due to doing video playback, but I'm not entirely sure, the DXVK crash in this thread has never produced any error messages in the kernel log for me.
Interesting thing is with mpv i get vk_error_device lost error with my intel igpu upon resizing and cant reproduce it with nvidia using 455.50.10
Testing 465.19.01.
Still Broken 465.19.01.
I've downgraded all the way to 440.100 again and I'll be tested that for a good month or so.
I think I have the same problem with another (so far not listed) game: Project Wingman.
So as I understand it, there is nothing we can do about it except hope that the Nvidia devs will fix the problem.
However, I didn't want to miss the "fun" of sharing my logs of some of my crashes, even if it won't really help.
These frequent crashes (after 15-60 minutes) are especially frustrating in this game, since there are only save points after finishing missions.
GOG version of "Project Wingman" installed via Lutris 0.5.8.3.
Config: set ENV "__GL_SHADER_DISK_CACHE_PATH" & "__GL_SHADER_DISK_CACHE_SKIP_CLEANUP: '1'", fsync tried on & off
[System]
OS: Manjaro Linux 21.0.1 Ornara
Arch: x86_64
Kernel: 5.11.10-1-MANJARO
Desktop: KDE
Display Server: x11
[CPU]
Vendor: AuthenticAMD
Model: AMD Ryzen 7 2700X Eight-Core Processor
Physical cores: 8
Logical cores: 16
[Memory]
RAM: 31.4 GB
Swap: 34.5 GB
[Graphics]
Vendor: NVIDIA Corporation
OpenGL Renderer: GeForce GTX 1070/PCIe/SSE2
OpenGL Version: 4.6.0 NVIDIA 460.67
OpenGL Core: 4.6.0 NVIDIA 460.67
OpenGL ES: OpenGL ES 3.2 NVIDIA 460.67
Vulkan: Supported
System:
Kernel: 5.11.10-1-MANJARO x86_64 bits: 64 compiler: gcc v: 10.2.0
parameters: BOOT_IMAGE=/boot/vmlinuz-5.11-x86_64
root=UUID=576659b2-353b-43cc-9510-aa3ac55bd386 rw quiet apparmor=1
security=apparmor resume=UUID=d4c036ba-fcfd-415b-8df9-f4bf690b2aa7
udev.log_priority=3
Desktop: KDE Plasma 5.21.3 tk: Qt 5.15.2 wm: kwin_x11 vt: 1 dm: SDDM
Distro: Manjaro Linux base: Arch Linux
Machine:
Type: Desktop Mobo: Micro-Star model: X470 GAMING PRO CARBON (MS-7B78)
v: 1.0 serial: <filter> UEFI: American Megatrends v: 2.80 date: 03/06/2019
Memory:
RAM: total: 31.37 GiB used: 10.9 GiB (34.8%)
RAM Report: permissions: Unable to run dmidecode. Root privileges required.
CPU:
Info: 8-Core model: AMD Ryzen 7 2700X bits: 64 type: MT MCP arch: Zen+
family: 17 (23) model-id: 8 stepping: 2 microcode: 800820D cache: L2: 4 MiB
bogomips: 118438
Speed: 2199 MHz min/max: 2200/3700 MHz boost: enabled Core speeds (MHz):
1: 2199 2: 2199 3: 2198 4: 2199 5: 2202 6: 2205 7: 2203 8: 2199 9: 1890
10: 2182 11: 1980 12: 2199 13: 2203 14: 2199 15: 1906 16: 1977
Flags: 3dnowprefetch abm adx aes aperfmperf apic arat avic avx avx2 bmi1
bmi2 bpext clflush clflushopt clzero cmov cmp_legacy constant_tsc cpb cpuid
cr8_legacy cx16 cx8 de decodeassists extapic extd_apicid f16c flushbyasid
fma fpu fsgsbase fxsr fxsr_opt ht hw_pstate ibpb irperf lahf_lm lbrv lm mca
mce misalignsse mmx mmxext monitor movbe msr mtrr mwaitx nonstop_tsc nopl
npt nrip_save nx osvw overflow_recov pae pat pausefilter pclmulqdq pdpe1gb
perfctr_core perfctr_llc perfctr_nb pfthreshold pge pni popcnt pse pse36
rdrand rdseed rdtscp rep_good sep sev sev_es sha_ni skinit smap smca sme
smep ssbd sse sse2 sse4_1 sse4_2 sse4a ssse3 succor svm svm_lock syscall tce
topoext tsc tsc_scale v_vmsave_vmload vgif vmcb_clean vme vmmcall wdt
xgetbv1 xsave xsavec xsaveerptr xsaveopt xsaves
Vulnerabilities: Type: itlb_multihit status: Not affected
Type: l1tf status: Not affected
Type: mds status: Not affected
Type: meltdown status: Not affected
Type: spec_store_bypass
mitigation: Speculative Store Bypass disabled via prctl and seccomp
Type: spectre_v1
mitigation: usercopy/swapgs barriers and __user pointer sanitization
Type: spectre_v2 mitigation: Full AMD retpoline, IBPB: conditional, STIBP:
disabled, RSB filling
Type: srbds status: Not affected
Type: tsx_async_abort status: Not affected
Graphics:
Device-1: NVIDIA GP104 [GeForce GTX 1070] vendor: Gigabyte driver: nvidia
v: 460.67 alternate: nouveau,nvidia_drm bus-ID: 1c:00.0 chip-ID: 10de:1b81
class-ID: 0300
Display: x11 server: X.Org 1.20.10 compositor: kwin_x11 driver:
loaded: nvidia display-ID: :0 screens: 1
Screen-1: 0 s-res: 3840x1080 s-dpi: 80 s-size: 1219x343mm (48.0x13.5")
s-diag: 1266mm (49.9")
Monitor-1: DVI-D-0 res: 1920x1080 dpi: 82 size: 597x336mm (23.5x13.2")
diag: 685mm (27")
Monitor-2: HDMI-0 res: 1920x1080 hz: 60 dpi: 96 size: 510x287mm (20.1x11.3")
diag: 585mm (23")
OpenGL: renderer: GeForce GTX 1070/PCIe/SSE2 v: 4.6.0 NVIDIA 460.67
direct render: Yes
Audio:
Device-1: Creative Labs EMU20k2 [Sound Blaster X-Fi Titanium Series]
driver: snd_ctxfi v: kernel bus-ID: 1b:00.0 chip-ID: 1102:000b
class-ID: 0403
Device-2: NVIDIA GP104 High Definition Audio vendor: Gigabyte
driver: snd_hda_intel v: kernel bus-ID: 1c:00.1 chip-ID: 10de:10f0
class-ID: 0403
Device-3: AMD Family 17h HD Audio vendor: Micro-Star MSI
driver: snd_hda_intel v: kernel bus-ID: 1e:00.3 chip-ID: 1022:1457
class-ID: 0403
Sound Server-1: ALSA v: k5.11.10-1-MANJARO running: yes
Sound Server-2: sndio v: N/A running: no
Sound Server-3: JACK v: 0.125.0 running: no
Sound Server-4: PulseAudio v: 14.2 running: yes
Sound Server-5: PipeWire v: 0.3.24 running: yes
Use of uninitialized value $args in concatenation (.) or string at /usr/bin/inxi line 2391.
Use of uninitialized value in concatenation (.) or string at /usr/bin/inxi line 2391.
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
100 13 100 13 0 0 54 0 --:--:-- --:--:-- --:--:-- 54
Network:
Device-1: Intel I211 Gigabit Network vendor: Micro-Star MSI driver: igb
v: kernel port: f000 bus-ID: 18:00.0 chip-ID: 8086:1539 class-ID: 0200
IF: enp24s0 state: up speed: 1000 Mbps duplex: full mac: <filter>
IP v4: <filter> type: dynamic noprefixroute scope: global
broadcast: <filter>
IP v6: <filter> type: noprefixroute scope: global
IP v6: <filter> type: temporary dynamic scope: global
IP v6: <filter> type: mngtmpaddr noprefixroute scope: global
IP v6: <filter> type: noprefixroute scope: link
Device-2: TP-Link UE300 10/100/1000 LAN (ethernet mode) [Realtek RTL8153]
type: USB driver: r8152 bus-ID: 1-10:2 chip-ID: 2357:0601 class-ID: 0000
serial: <filter>
IF: enp3s0f0u10 state: down mac: <filter>
WAN IP: <filter>
Bluetooth:
Message: No Bluetooth data was found.
Logical:
Message: No LVM data was found.
RAID:
Message: No RAID data was found.
Drives:
Local Storage: total: 4.55 TiB used: 24.07 TiB (529.3%)
SMART Message: Unable to run smartctl. Root privileges required.
ID-1: /dev/nvme0n1 maj-min: 259:0 vendor: Samsung model: SSD 970 EVO 500GB
size: 465.76 GiB block-size: physical: 512 B logical: 512 B speed: 31.6 Gb/s
lanes: 4 rotation: SSD serial: <filter> rev: 2B2QEXE7 temp: 39.9 C
scheme: GPT
ID-2: /dev/sda maj-min: 8:0 vendor: Samsung model: SSD 860 EVO 500GB
size: 465.76 GiB block-size: physical: 512 B logical: 512 B speed: 6.0 Gb/s
rotation: SSD serial: <filter> rev: 1B6Q scheme: MBR
ID-3: /dev/sdb maj-min: 8:16 vendor: HGST (Hitachi) model: HDN724040ALE640
size: 3.64 TiB block-size: physical: 4096 B logical: 512 B speed: 6.0 Gb/s
rotation: 7200 rpm serial: <filter> rev: A5E0 scheme: GPT
Optical-1: /dev/sr0 vendor: HL-DT-ST model: DVDRAM GH20NS15 rev: IL00
dev-links: cdrom
Features: speed: 48 multisession: yes audio: yes dvd: yes
rw: cd-r,cd-rw,dvd-r,dvd-ram state: running
Partition:
ID-1: / raw-size: 430.95 GiB size: 423.18 GiB (98.20%)
used: 356.21 GiB (84.2%) fs: ext4 dev: /dev/nvme0n1p2 maj-min: 259:2
label: Local Disk 1 uuid: 576659b2-353b-43cc-9510-aa3ac55bd386
ID-2: /boot/efi raw-size: 300 MiB size: 299.4 MiB (99.80%)
used: 280 KiB (0.1%) fs: vfat dev: /dev/nvme0n1p1 maj-min: 259:1 label: N/A
uuid: 71D1-C987
ID-3: /data/nfs/Documents raw-size: N/A size: 20.95 TiB
used: 20.63 TiB (98.5%) fs: nfs4 remote: 192.168.1.242:/volume1/Documents
label: N/A uuid: N/A
ID-4: /home/<filter>/mount/Local Disk 2 raw-size: 465.76 GiB
size: 457.45 GiB (98.22%) used: 405.69 GiB (88.7%) fs: ext4 dev: /dev/sda1
maj-min: 8:1 label: Local Disk 2 uuid: 340457fe-a771-4dfa-933b-6a4fcdd5b794
ID-5: /run/timeshift/backup raw-size: 3.64 TiB size: 3.58 TiB (98.40%)
used: 2.7 TiB (75.4%) fs: ext4 dev: /dev/sdb1 maj-min: 8:17
label: Local Disk 3 uuid: 750c820d-d1e6-4d1f-a8d7-1772fde1ec50
Swap:
Kernel: swappiness: 60 (default) cache-pressure: 100 (default)
ID-1: swap-1 type: partition size: 34.52 GiB used: 0 KiB (0.0%) priority: -2
dev: /dev/nvme0n1p3 maj-min: 259:3 label: N/A
uuid: d4c036ba-fcfd-415b-8df9-f4bf690b2aa7
Unmounted:
Message: No Unmounted partitions found.
USB:
Hub-1: 1-0:1 info: Full speed (or root) Hub ports: 14 rev: 2.0
speed: 480 Mb/s chip-ID: 1d6b:0002 class-ID: 0900
Device-1: 1-10:2
info: TP-Link UE300 10/100/1000 LAN (ethernet mode) [Realtek RTL8153]
type: Network driver: r8152 interfaces: 1 rev: 2.1 speed: 480 Mb/s
power: 180mA chip-ID: 2357:0601 class-ID: 0000 serial: <filter>
Device-2: 1-11:3 info: MCT Elektronikladen farbwerk type: HID
driver: hid-generic,usbhid interfaces: 1 rev: 2.0 speed: 12 Mb/s power: 2mA
chip-ID: 0c70:f00a class-ID: 0300 serial: <filter>
Device-3: 1-12:4 info: Logitech G Pro Gaming Mouse type: Mouse,HID
driver: hid-generic,usbhid interfaces: 2 rev: 2.0 speed: 12 Mb/s
power: 300mA chip-ID: 046d:c085 class-ID: 0300 serial: <filter>
Device-4: 1-13:5 info: Atmel WootingOne type: HID,Keyboard
driver: hid-generic,usbhid interfaces: 7 rev: 2.0 speed: 12 Mb/s
power: 400mA chip-ID: 03eb:ff01 class-ID: 0300 serial: <filter>
Hub-2: 2-0:1 info: Full speed (or root) Hub ports: 8 rev: 3.1 speed: 10 Gb/s
chip-ID: 1d6b:0003 class-ID: 0900
Hub-3: 3-0:1 info: Full speed (or root) Hub ports: 4 rev: 2.0
speed: 480 Mb/s chip-ID: 1d6b:0002 class-ID: 0900
Hub-4: 4-0:1 info: Full speed (or root) Hub ports: 4 rev: 3.0 speed: 5 Gb/s
chip-ID: 1d6b:0003 class-ID: 0900
Sensors:
System Temperatures: cpu: 41.8 C mobo: N/A gpu: nvidia temp: 50 C
Fan Speeds (RPM): N/A gpu: nvidia fan: 19%
Info:
Processes: 462 Uptime: 2h 15m wakeups: 0 Init: systemd v: 247
tool: systemctl Compilers: gcc: 10.2.0 alt: 8/9 clang: 11.1.0 Packages: 2628
pacman: 2595 lib: 585 flatpak: 12 snap: 21 Shell: Bash v: 5.1.0
running-in: konsole inxi: 3.3.03
project_wingman-crash1.log
project_wingman-crash2.log
project_wingman-crash3.log
project_wingman-crash4.log
project_wingman-crash5.log
🥲
Greetings
That's probably not the same problem. Make a separate issue.
That's probably not the same problem. Make a separate issue.
Are you sure about that? My logs from 455.50.10 look a lot like @Retardium's...
Mine:
err: DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err: DxvkDevice: waitForIdle: Operation failed
err: DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err: DxvkDevice: waitForIdle: Operation failed
err: DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err: DxvkDevice: waitForIdle: Operation failed
err: DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err: DxvkDevice: waitForIdle: Operation failed
err: DxvkSubmissionQueue: Command submission failed: VK_ERROR_DEVICE_LOST
Theirs:
warn: D3D11Device::CreateShaderModule: Class linkage not supported
warn: D3D11Device::CreateShaderModule: Class linkage not supported
warn: D3D11Device::CreateShaderModule: Class linkage not supported
warn: D3D11Device::CreateShaderModule: Class linkage not supported
warn: D3D11Device::CreateShaderModule: Class linkage not supported
warn: D3D11Device::CreateShaderModule: Class linkage not supported
warn: D3D11DXGIDevice::QueryInterface: Unknown interface query
warn: 6543dbb6-1b48-42f5-ab82-e97ec74326f6
warn: D3D11Device::CreateShaderModule: Class linkage not supported
warn: D3D11Device::CreateShaderModule: Class linkage not supported
warn: D3D11Device::CreateShaderModule: Class linkage not supported
warn: D3D11Device::CreateShaderModule: Class linkage not supported
err: DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err: DxvkDevice: waitForIdle: Operation failed
err: DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err: DxvkDevice: waitForIdle: Operation failed
err: DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err: DxvkSubmissionQueue: Command submission failed: VK_ERROR_DEVICE_LOST
err: DxvkDevice: waitForIdle: Operation failed
Edit: By the way two weeks on Kernel 5.4.109,110,111 and Nvidia 440.100 and 0 crashes, I am pretty sure I'm going to get to a month with 0 crashes as well. Really not liking Nvidia as usual.
I also get the same issue on panzer Corps 2 on my gtx770 with both 450 and 460.
err: DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err: DxvkDevice: waitForIdle: Operation failed
err: DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err: DxvkDevice: waitForIdle: Operation failed
err: DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err: DxvkDevice: waitForIdle: Operation failed
err: DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err: DxvkDevice: waitForIdle: Operation failed
unfortunately 5.4.72 and 440.100 doesn't seem to help I get the same error
440.100 doesn't seem to help I get the same error
Are you absolutely sure you downgraded to 440.100?
440.100 doesn't seem to help I get the same error
Are you absolutely sure you downgraded to 440.100?
Unfortunately yes, else I would be a happy gamer. But maybe the problem lies with the specific game (panzer corps 2) . I will try some other games next week and see if I have the same issue
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 440.100 Driver Version: 440.100 CUDA Version: 10.2 |
|-------------------------------+----------------------+----------------------+
| GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |
|===============================+======================+======================|
| 0 GeForce GTX 770 Off | 00000000:07:00.0 N/A | N/A |
| 30% 49C P8 N/A / N/A | 184MiB / 1994MiB | N/A Default |
+-------------------------------+----------------------+----------------------+
Unfortunately yes, else I would be a happy gamer.
I guess @K0bin was right then sorry about that K0bin.
Unfortunately yes, else I would be a happy gamer.
I guess @K0bin was right then sorry about that K0bin.
Well I will try to do a full purge of nvidia and everything and install 440.100 from command line so I can more properly track what's going on because I found so many different versions of nvidia packages in my installation.
give me a bit more time..
ehm do you have a guide on how to install 440.100 ? even though the drivers installs in command line pretty well now that I purged everything when I boot it loads the mesa driver. (shy)
I try to install the run file and I get to the libglvnd error. And the PPA doesn't seem to include 440 anymore. So I get now to this stage
NVRM: API mismatch: the client has the version 440.100, but
NVRM: this kernel module has the version 460.67. Please
NVRM: make sure that this kernel module and all NVIDIA driver
NVRM: components have the same version.
2 hours of clusterfuck and driver parts here and there and I managed to load a proper clean 440.100 installation (no 32bit support and no DKMS though) . I managed to actually play for more time than I ever managed on this game (like an hour) but it eventually crashed to our usual issue. I have to go to work now I will try some more with frostpunk tonight to see if it a reaccuring thing.
I've downgraded all the way to
440.100again and I'll be tested that for a good month or so.
So it's been a month now and no crashes, I'm just going to stick with this kernel and driver version. I've pretty much given up on trying to update the Nvidia Driver and not have this issue occur.
Final Fantasy XIV
I used to have a bunch of crashes with the 46x series until they "fixed" 465 and 470 works as well.
But you should be using: 455.50.19
I can highly recommend purging the 4xx drivers entirely, make sure you clean up the dkms folder as well and then do a manual install. You will probably lose video if you try this, so do it from a SSH console.
This driver gives by far the best performance on my GTX970 and zero crashes thus far.
I upgraded my laptop now i have RTX 3060 mobile GPU and i can confirm errors are no more with any driver ive tried. This error is called something like xid31 by nvidia and latest 470 drivers have a fix for it.
Just found out about this issue and reported it on ticket 2253
I have drivers 470.57.2 and this issue is NOT fixed.
That's most likely an unrelated issue.
DEVICE_LOST just means the driver encountered an error for one reason or another.
Still Broken with 470.63.01
@doitsujin considering people are still reporting this is broken, and https://github.com/doitsujin/dxvk/pull/1963 made the major version check for the nvidia work around 465, shouldn't this just default to being back on for all Nvidia versions? As far as I know this bug was the reason the work around existed in the first place. It doesn't seem like Nvidia actually fixed it.
I'm not convinced that the issues people are reporting are actually caused by the same issue. There's still a problem somewhere regarding high HVV use but I've never seen it crash since.
There's still a problem somewhere regarding high HVV use but I've never seen it crash since.
I'm currently testing 470.74 if I still manage to end up making the game crash (could take a month or more) but I've still never in at least a year of playing FFXIV had 440.100 crash, where does that ultimately put things? Is this just a rare bug that will never be fixed? Do you have any theories on this at all? I would've just escaped this by switching to an AMD card but with the whole market the way it is it hasn't been an option.
Do you have any thoughts on this at all? It's driving me nuts.
Interesting note on the 470.74 tho:
Fixed a regression which resulted in very-high system memory usage for Direct3D 12 games when run through vkd3d-proton.
Could be useful for DXVK too maybe? (Ie. high-memory related issue).
That said, nVidia hardware CAN sometimes be throwing those random XID errors when different manufacturers are pushing the limits on their "factory clocks" (Eg. Asus/MSI++ with their "OC" models and whatnot). Could be worth TRYING to underclock or powerlimit the gpu just to see.
You can check your powerlimit with nvidia-smi Then if it has like 215W (like my RTX2070 has), then you could try to use
sudo nvidia-smi -pl 200
Just mentioning this because i had a nVidia (260Ti card back when) that was really unstable - but it was one of those "OC Twin Frozr Whatever" cards, and ended up just downclocking it a tad, and issue went away :) Loads of factors playing in here - especially since XID(DEVICE LOST) errors MAY not be driver/dxvk related at all.
Could be useful for DXVK too maybe? (Ie. high-memory related issue).
No, the issue was specifically caused by the way vkd3d-proton uses Vulkan pipeline caches. DXVK is not affected by this in any way.
Honestly at this point I'll willing to concede this might just be a game bug. I've looked through the tracker and see other games that will crash with an in game bug but still have the DXVK log producing VK_ERROR_DEVICE_LOST. There are multiple numerous reports from Windows users on the FFXIV forums about DX11 crashes. I also play with my significant other who has had both the DX11 crash and sometimes a freeze while playing the game and they use Windows and AMD.
I think the biggest problem here is VK_ERROR_DEVICE_LOST is just too much of a generic error, is there anyway DXVK can give a better error code?
I think it's probably time to close this bug, I personally after all this time don't think this is a Nvidia or DXVK bug. Thoughts?
I don't think that the crashes on Windows are related to this. I have several friends using Nvidia cards (including RTX 3xxx) on Windows that afk all day in-game and never mentioned anything about the game crashing. And if this was actually a widespread issue there'd be a lot more noise about it with the current login queues. The people on windows having crashes more likely is borked installations, hardware or overclocks.
And even on Linux only newer Nvidia drivers are affected by it reproducibly, older ones and AMD cards run rock solid. Since switching to a 6900 XT I haven't had a single crash in the game. So from my point of view this still looks like a bug in Nvidias driver on Linux.
As far I know Nvidia devs also browse through these issues once in a while so keeping the issue open might make sense for the small chance of them looking into it one day, but more practically this would belong on the Nvidia forums or whatever they use for bug reports.
And if this was actually a widespread issue there'd be a lot more noise about it with the current login queues. The people on windows having crashes more likely is borked installations, hardware or overclocks.
There is an extensive thread over on the Final Fantasy Forums though: https://forum.square-enix.com/ffxiv/threads/448687-Directx11-error
Users have the following hardware in that thread:
AMD Ryzen 6600XT
AMD Ryzen 7 5800X
AMD Ryzen 7 5800X 8-Core Processor
AMD Radeon RX 6600 XT
AMD Ryzen 5 3600 6-Core Processor
AMD Radeon RX 5700 XT
AMD Ryzen 5 3600 6-Core Processor
Radeon RX 580 Series
Intel(R) Core(TM) i7-8700 CPU
NVIDIA GeForce RTX 2070
AMD Ryzen 7 5700G with Radeon Graphics
NVIDIA GeForce GTX 1080
AMD Ryzen 5950x
AMD Radeon 6900x
As one user on the thread states:
Nothing has fixed this issue. Sometimes we can play all day with no problems, the next day the DirectX 11 11000002 error comes up over and over. If you lookup that error number, you'll see MILLIONS of posts online (not just on these forums) about this FF14 issue. This is not an obscure thing, A LOT OF PEOPLE have this problem.
And another user said it's been there since Stormblood:
A large part of the community have been waiting for an answer regarding this issue since Stormblood release, so I'm not sure if we are going to get one anytime soon if we didn't get one so far.
Apart from the VK_ERROR_DEVICE_LOST I also randomly get complete nvidia driver crashes with this game. I first thought my GPU is probably defect, but Xid 16 means “Display engine hung” and “Driver Error” according to nvidias own list and I do play other games, sometimes for hours and I have never seen such a crash outside of FFXIV.
This is with a NVIDIA GeForce GTX TITAN X card.
[11303.224211] NVRM: GPU at PCI:0000:01:00: GPU-0b4cda80-07b2-13fb-9c08-f03bf0381c28
[11303.224213] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1ef6
[11311.534091] NVRM: Xid (PCI:0000:01:00): 16, pid=16, Head 00000000 Count 000b1ef7
[11319.853967] NVRM: Xid (PCI:0000:01:00): 16, pid=6226, Head 00000000 Count 000b1ef8
[11328.183850] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1ef9
[11336.503749] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1efa
[11344.813634] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1efb
[11353.133512] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1efc
[11361.463408] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1efd
[11369.783297] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1efe
[11378.093175] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1eff
[11386.413068] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1f00
[11394.732946] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1f01
[11403.062837] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1f02
[11411.372721] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1f03
[11419.692602] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1f04
[11428.022469] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1f05
[11463.772567] nvidia-modeset: WARNING: GPU:0: Timeout while waiting for idle.
[11465.772637] nvidia-modeset: ERROR: GPU:0: Idling display engine timed out: 0x0000957d:0:0:428
Made a post on their forums, sent a bug report to the "support" address, never even got an answer back.
Also should mention that the driver crash is 100% not a DXVK bug, since it also happens with pure wine (only tested once, because the performance is horrible). I'm not sure if this is related at all to the original bug here, but I just wanted to add that information, in the hope it may help. If this confuses things more, please just disregard.
I first thought my GPU is probably defect, but Xid 16 means “Display engine hung” and “Driver Error” according to nvidias own list and I do play other games, sometimes for hours and I have never seen such a crash outside of FFXIV.
And that's the problem I play other games outside of FFXIV and they never crash either, we have a bunch of windows users reporting the same thing on different GPUs and if you take a look at other bug reports like this one https://github.com/doitsujin/dxvk/issues/2349 you'll see that bugs in games can definitely cause VK_ERROR_DEVICE_LOST and not just problems with Nvidia or DXVK.
As I said I've played this game a lot and so has my significant other I've watched their PC running Windows have all the same crashes I do on DXVK and Nvidia and they're using Windows and AMD.
I believe the bug is in the game itself for this reason, I also believe that depending on hardware/driver versions the bug can occur less or more frequently. I only wish I had the money to experiment with different GPUs and end this debate for once and all.
@konomikitten How is this behaving with the 515 drivers if you are still testing this from time to time.
@Blisto91 I switched from my GTX 960 to a RX 6600 about 4 months ago and I haven't had this crash since. My significant other switched from an RX 570 to a RX 6600XT and now gets this issue on Windows 10.
I still see frequent posts from Windows users on the Final Fantasy Forums about DirectX Error 11000002. As far as I am concerned this is a game bug that Square need to fix.
@konomikitten How is this behaving with the 515 drivers if you are still testing this from time to time.
No issues with 515.48.07 (as stable as 440.100)
@konomikitten How is this behaving with the 515 drivers if you are still testing this from time to time.
No issues with
515.48.07(as stable as440.100)
I no longer use a Nvidia Graphics card as per my previous post so I cannot answer that question.
I think they are stating that the 515 drivers are stable for them.
__GL_SHADER_DISK_CACHE_PATHx1 2021-04__GL_SHADER_DISK_CACHE_SKIP_CLEANUPx1 2021-04__GL_SHADER_DISK_CACHE=0`.x2 2021-02__GL_SHADER_DISK_CACHE=0`x1 2020-12
There was a patch https://github.com/doitsujin/dxvk/commit/16a51f3c03d5bc52ba67101fb5a5cd5b8d96fa94 to help reduce issues with FFXIV and Nvidia Drivers 450.66 and possibly above but this just seems to have made the issue occur less. After about 12 hours strait of having the FFXIV client running I once again got the dreaded VK_ERROR_DEVICE_LOST. I see other games in the issues list are having this problem too with the same driver versions. I'm starting to think Nvidia just messed up version 450.66. I can confirm I never got this error with 440.100 so I am pretty sure this is a Nvidia driver issue.
Software information
Final Fantasy XIV
System information
Log files