protonscr

Final Fantasy XIV: VK_ERROR_DEVICE_LOST

dxvkclosed nvidia proprietary
doitsujin/dxvk#1791 · opened 2020-10-14 by konomikitten · updated 2022-10-16 · 70 comments · github
Kkonomikitten 2020-10-14 github

There was a patch https://github.com/doitsujin/dxvk/commit/16a51f3c03d5bc52ba67101fb5a5cd5b8d96fa94 to help reduce issues with FFXIV and Nvidia Drivers 450.66 and possibly above but this just seems to have made the issue occur less. After about 12 hours strait of having the FFXIV client running I once again got the dreaded VK_ERROR_DEVICE_LOST. I see other games in the issues list are having this problem too with the same driver versions. I'm starting to think Nvidia just messed up version 450.66. I can confirm I never got this error with 440.100 so I am pretty sure this is a Nvidia driver issue.

Software information

Final Fantasy XIV

System information

  • GPU: GeForce GTX 960
  • Driver: 450.66
  • Wine version: lutris-5.7-10-x86_64
  • DXVK version: 1.7.2

Log files

Kkonomikitten 2020-10-14 github

I've patched and downgraded to 440.100 but with Kernel 5.8. I'll try to find a moment to leave the game running for at least 24 hours and see if I can replicate this bug. If I can't get this to happen again I'm really of the strong opinion Nvidia has introduced a bug into their drivers between the 44X and 45X series. Hopefully this information is useful for the dxvk devs.

Ddoitsujin maintainer 2020-10-14 github

We're aware of this but there's not an awful lot we can do about it on the DXVK side of things at this point.

Kkonomikitten 2020-10-20 github

So I tested 440.100 for 3 days (never restarted the client) and 0 issues 0 crashes. So if people want to continue using DXVK they should hold onto 440.100 as long as possible.

Ssschroe 2020-11-02 github

Is 450.66 the first version where this issue was introduced? The Nvidia website also has 450.57 and 450.51. Knowing the exact version might be helpful for a bug report.

Kkonomikitten 2020-11-02 github

450.66 was the first version to have this issue from what I tested. Thankfully I've tested the latest long lived version and this issue appears to be fixed. I ran the tests again for 3 days continuous game client usage. I experienced no errors with the following:

Software information

Final Fantasy XIV

System information

  • GPU: GeForce GTX 960
  • Driver: 450.80.02
  • Wine version: lutris-5.7-10-x86_64
  • DXVK version: 1.7.2

Apitrace file(s)

None

Log files

I am unsure if the current short lived version (455.38 as of this post) has been fixed.

Despite all this testing I ended up with another crash on 450.80.02. Guess we wait more sadly.

Ssschroe 2020-11-03 github

Hmm that's interesting because 450.80.02 does not mention any fixes but only support for some new GPUs. Though i have no idea how consistent nvidia it with their changelog. I'm pretty sure i saw the issue pop up on 455.38 when testing something unrelated though. But that may have been triggered by something else. Going to keep an eye out for it if it pops up again.

KK0bin maintainer 2020-11-03 github

Nvidia is aware of the problem and looking into it.

Aarabcian 2020-11-09 github

Nvidia is aware of the problem and looking into it.

Do you have a link about it?

Kkhachbe 2020-12-02 github

Any updates on this issue?

Ddoitsujin maintainer 2020-12-02 github

That's generally not a very useful question to ask; if there was then someone would likely have posted it already. Unless Nvidia figure out whatever it is that's causing their driver to crash, this will not get fixed.

Aastraldawn 2020-12-20 github

Occurs on 460.27.4 as well

Kkonomikitten 2020-12-20 github

455.45.01 also broken, this really makes me want to switch from nvidia to amd...

Kkonomikitten 2020-12-29 github

A possible fix but this is completely anecdotal. Try either clearing out your ~/.nv/GLCache/ or doing export __GL_SHADER_DISK_CACHE=0 with whatever way you're using wine. I haven't had a crash in awhile and this is the only thing I can think that I did lately that might affect DXVK.

Kkonomikitten 2021-01-01 github

There's now a post over at the Nvidia forums about this issue but I don't really have a high confidence in Nvidia fixing this (at least not with the replies I'm seeing on the post from them): Vk_error_device_lost in many game titles.

Aastraldawn 2021-01-24 github

460.32.03 broken as well

Kkonomikitten 2021-01-24 github

Just to let everyone know I haven't had any issues since using export __GL_SHADER_DISK_CACHE=0. Currently on 460.32.03. Can others here test that?

Okay I had another crash so that doesn't work either, at this point I give up. People seeing this bug report need to go to the thread posted on the nvidia forums and actually post or this is never going to get fixed.

So please go add your voice to Vk_error_device_lost in many game titles.

Kkonomikitten 2021-02-08 github
Aarabcian 2021-02-16 github

~Just to let everyone know I haven't had any issues since using export __GL_SHADER_DISK_CACHE=0. Currently on 460.32.03. Can others here test that?~

Okay I had another crash so that doesn't work either, at this point I give up. People seeing this bug report need to go to the thread posted on the nvidia forums and actually post or this is never going to get fixed.

So please go add your voice to Vk_error_device_lost in many game titles.

Hello im facing the same bug for about 6 months. I tried many blind workarounds to no avail lastly ive installed lutris added wine 6.0 and built my kernel and out of tree modules with clang. I couldnt reproduce it for 2 weeks still so early to talk but i just wonder if there are any changes in the dxvk code that may fix that or my blind tries finally gave me a solution. And i built nvidia-drivers with dynamic libraries. I hope this helps, I use Gentoo Linux where those workarounds are very easy to apply.

Kkonomikitten 2021-02-20 github

It's fun when you find out mpv also has the same problem with nvidia drivers: https://github.com/mpv-player/mpv/issues/8360

Aarabcian 2021-02-20 github

I still havent got any vk_error_device_lost yet with my shot but i sure ill get it at some time. Did you read the driver recommendation there? Please test one of those drivers and report here. Once i get that error again ill try one of those drivers too. Now no errors yet for 2 weeks.

January 27th, 2021 - Windows 457.88, Linux 455.50.04
Fixed a bug in a stencil-buffer optimization that could occasionally result in VK_ERROR_DEVICE_LOST

Kkonomikitten 2021-02-20 github

@arabcian my results are posted throughout this thread, multiple people are still having the issue and my driver version is far higher than yours sitting at 460.39.0 so I am pretty sure the fix your listed has nothing to do with the problem as was also mentioned on the mpv bug report where they don't make use of stencil-buffers either. I also clocked up 3 days of continuous play time only to have the bug randomly come back when I thought it was fixed.

SSveSop 2021-02-20 github

@konomikitten
Driver 460.39.0 is not necessarily "newer" than 455.50.04. The 455.50.xx driver series is of the "vulkan beta" series, and contains a lot more "vulkan fixes" than the "release driver" from the 460.xx series.

To simplify: 455.50.04 can contain a vulkan fix, the 460.39 driver does NOT contain since they are from two different branches.

When it comes to iffy nVidia random crashes + DXVK - Its always best to try the latest "vulkan beta" driver branch from here: https://developer.nvidia.com/vulkan-driver

Kkonomikitten 2021-02-21 github

@SveSop I stand corrected, that's some confusing version numbering.

Kkonomikitten 2021-02-21 github

@arabcian how many hours of FFXIV have you played to test for the crash by the way?

Kkonomikitten 2021-02-24 github

Just to let everyone know I am currently testing 455.50.04 as @arabcian recommended, it's looking good so far but considering the difficult nature of replicating this bug I'm going to be testing for a few weeks or more to make sure it is in fact fixed. If anyone else could test too that would be helpful.

Kkonomikitten 2021-02-25 github

Still broken. 455.50.04

ffxiv_dx11_d3d11.log
ffxiv_dx11_dxgi.log

Aarabcian 2021-02-27 github

@arabcian how many hours of FFXIV have you played to test for the crash by the way?

I couldnt test it yet with FF since dont have the game but i was having same error with thief 2014 and World of Warcraft + Skyrim + Fallout 4,somehow bug is not around for a month. Still having mpd bug tho but im not sure if both are related.

Kkonomikitten 2021-03-03 github

Currently testing 455.50.07...

Kkonomikitten 2021-03-07 github

Still broken. 455.50.07

ffxiv_dx11_d3d11.log
ffxiv_dx11_dxgi.log

Ddoitsujin maintainer 2021-03-07 github

This should be fixed with 455.50.10:

Fixed a bug with the host-visible device-local memory heap, where if an allocation failed due to space constraints, it could cause the application to crash on future Vulkan function calls

Kkonomikitten 2021-03-07 github

@doitsujin oh thank god, I'll test it asap.

Kkonomikitten 2021-03-09 github

@doitsujin nope.... Still Broken 455.50.10.

ffxiv_dx11_d3d11.log
ffxiv_dx11_dxgi.log

SSveSop 2021-03-10 github

@konomikitten
What is shown in syslog? (dmesg |grep -i nvrm ... or dmesg|grep -i xid)

It could be a faulty card maybe, but ofc if it is only 1 single title that have this problem, while all other (native/whatever) works fine its kinda hard to tell. New drivers sometimes push more performance from certain functions, that MAY make a borderline defective card crash more often.

Kkonomikitten 2021-03-10 github

What is shown in syslog? (dmesg |grep -i nvrm ... or dmesg|grep -i xid)

$ sudo dmesg | grep -Ei 'nvrm|xid'
[    1.209379] r8169 0000:08:00.0 eth0: RTL8168g/8111g, e0:d5:5e:a6:68:bb, XID 4c0, IRQ 44
[    4.804576] NVRM: loading NVIDIA UNIX x86_64 Kernel Module  455.50.10  Thu Mar  4 20:25:58 UTC 2021

It could be a faulty card maybe.

440.100 never crashed not even once and I can happily downgrade my kernel and nvidia drivers and never see this crash, which I have done multiple times.

Aarabcian 2021-03-12 github

What is shown in syslog? (dmesg |grep -i nvrm ... or dmesg|grep -i xid)

$ sudo dmesg | grep -Ei 'nvrm|xid'
[    1.209379] r8169 0000:08:00.0 eth0: RTL8168g/8111g, e0:d5:5e:a6:68:bb, XID 4c0, IRQ 44
[    4.804576] NVRM: loading NVIDIA UNIX x86_64 Kernel Module  455.50.10  Thu Mar  4 20:25:58 UTC 2021

It could be a faulty card maybe.

440.100 never crashed not even once and I can happily downgrade my kernel and nvidia drivers and never see this crash, which I have done multiple times.

I remember you giving a code to run videos with vulkan in mpv and resizing windows was reproducing the same bug we experience. What was that command? Because seems like i cant reproduce it anymore if im not mistaken with the command. Driver is 455.50.10

Kkonomikitten 2021-03-13 github

I remember you giving a code to run videos with vulkan in mpv and resizing windows was reproducing the same bug we experience. What was that command?

I never gave any command related to mpv I did link to https://github.com/mpv-player/mpv/issues/8360 where people using the vulkan backend with mpv were having similar crashes. As to why they get a kernel log involving Xid I imagine it's due to doing video playback, but I'm not entirely sure, the DXVK crash in this thread has never produced any error messages in the kernel log for me.

Aarabcian 2021-03-17 github

I remember you giving a code to run videos with vulkan in mpv and resizing windows was reproducing the same bug we experience. What was that command?

I never gave any command related to mpv I did link to mpv-player/mpv#8360 where people using the vulkan backend with mpv were having similar crashes. As to why they get a kernel log involving Xid I imagine it's due to doing video playback, but I'm not entirely sure, the DXVK crash in this thread has never produced any error messages in the kernel log for me.

Interesting thing is with mpv i get vk_error_device lost error with my intel igpu upon resizing and cant reproduce it with nvidia using 455.50.10

Kkonomikitten 2021-03-31 github

Testing 465.19.01.

Kkonomikitten 2021-03-31 github

Still Broken 465.19.01.

ffxiv_dx11_d3d11.log
ffxiv_dx11_dxgi.log

Kkonomikitten 2021-04-01 github

I've downgraded all the way to 440.100 again and I'll be tested that for a good month or so.

Uupbox-org 2021-04-14 github

I think I have the same problem with another (so far not listed) game: Project Wingman.

So as I understand it, there is nothing we can do about it except hope that the Nvidia devs will fix the problem.
However, I didn't want to miss the "fun" of sharing my logs of some of my crashes, even if it won't really help.
These frequent crashes (after 15-60 minutes) are especially frustrating in this game, since there are only save points after finishing missions.

Software information

GOG version of "Project Wingman" installed via Lutris 0.5.8.3.
Config: set ENV "__GL_SHADER_DISK_CACHE_PATH" & "__GL_SHADER_DISK_CACHE_SKIP_CLEANUP: '1'", fsync tried on & off

System information

  • GPU: 1070
  • Driver: 460.67
  • Kernel: 5.11.10-1
  • OS: Manjaro Linux (KDE)
  • Wine version: tried multiple versions, all with the same result. Tested: lutris-4.21-x86_64, lutris-5.7-11-x86_64, lutris-6.4-x86_64
  • DXVK version: tried multiple versions, all with the same result. Tested: v1.5, v1.6.1+, v1.7.1, v1.7.3-4-g03f11baf, v1.8.1
More system infos
[System]
OS:              Manjaro Linux 21.0.1 Ornara
Arch:            x86_64
Kernel:          5.11.10-1-MANJARO
Desktop:         KDE
Display Server:  x11

[CPU]
Vendor:          AuthenticAMD
Model:           AMD Ryzen 7 2700X Eight-Core Processor
Physical cores:  8
Logical cores:   16

[Memory]
RAM:             31.4 GB
Swap:            34.5 GB

[Graphics]
Vendor:          NVIDIA Corporation
OpenGL Renderer: GeForce GTX 1070/PCIe/SSE2
OpenGL Version:  4.6.0 NVIDIA 460.67
OpenGL Core:     4.6.0 NVIDIA 460.67
OpenGL ES:       OpenGL ES 3.2 NVIDIA 460.67
Vulkan:          Supported
System:
  Kernel: 5.11.10-1-MANJARO x86_64 bits: 64 compiler: gcc v: 10.2.0 
  parameters: BOOT_IMAGE=/boot/vmlinuz-5.11-x86_64 
  root=UUID=576659b2-353b-43cc-9510-aa3ac55bd386 rw quiet apparmor=1 
  security=apparmor resume=UUID=d4c036ba-fcfd-415b-8df9-f4bf690b2aa7 
  udev.log_priority=3 
  Desktop: KDE Plasma 5.21.3 tk: Qt 5.15.2 wm: kwin_x11 vt: 1 dm: SDDM 
  Distro: Manjaro Linux base: Arch Linux 
Machine:
  Type: Desktop Mobo: Micro-Star model: X470 GAMING PRO CARBON (MS-7B78) 
  v: 1.0 serial: <filter> UEFI: American Megatrends v: 2.80 date: 03/06/2019 
Memory:
  RAM: total: 31.37 GiB used: 10.9 GiB (34.8%) 
  RAM Report: permissions: Unable to run dmidecode. Root privileges required. 
CPU:
  Info: 8-Core model: AMD Ryzen 7 2700X bits: 64 type: MT MCP arch: Zen+ 
  family: 17 (23) model-id: 8 stepping: 2 microcode: 800820D cache: L2: 4 MiB 
  bogomips: 118438 
  Speed: 2199 MHz min/max: 2200/3700 MHz boost: enabled Core speeds (MHz): 
  1: 2199 2: 2199 3: 2198 4: 2199 5: 2202 6: 2205 7: 2203 8: 2199 9: 1890 
  10: 2182 11: 1980 12: 2199 13: 2203 14: 2199 15: 1906 16: 1977 
  Flags: 3dnowprefetch abm adx aes aperfmperf apic arat avic avx avx2 bmi1 
  bmi2 bpext clflush clflushopt clzero cmov cmp_legacy constant_tsc cpb cpuid 
  cr8_legacy cx16 cx8 de decodeassists extapic extd_apicid f16c flushbyasid 
  fma fpu fsgsbase fxsr fxsr_opt ht hw_pstate ibpb irperf lahf_lm lbrv lm mca 
  mce misalignsse mmx mmxext monitor movbe msr mtrr mwaitx nonstop_tsc nopl 
  npt nrip_save nx osvw overflow_recov pae pat pausefilter pclmulqdq pdpe1gb 
  perfctr_core perfctr_llc perfctr_nb pfthreshold pge pni popcnt pse pse36 
  rdrand rdseed rdtscp rep_good sep sev sev_es sha_ni skinit smap smca sme 
  smep ssbd sse sse2 sse4_1 sse4_2 sse4a ssse3 succor svm svm_lock syscall tce 
  topoext tsc tsc_scale v_vmsave_vmload vgif vmcb_clean vme vmmcall wdt 
  xgetbv1 xsave xsavec xsaveerptr xsaveopt xsaves 
  Vulnerabilities: Type: itlb_multihit status: Not affected 
  Type: l1tf status: Not affected 
  Type: mds status: Not affected 
  Type: meltdown status: Not affected 
  Type: spec_store_bypass 
  mitigation: Speculative Store Bypass disabled via prctl and seccomp 
  Type: spectre_v1 
  mitigation: usercopy/swapgs barriers and __user pointer sanitization 
  Type: spectre_v2 mitigation: Full AMD retpoline, IBPB: conditional, STIBP: 
  disabled, RSB filling 
  Type: srbds status: Not affected 
  Type: tsx_async_abort status: Not affected 
Graphics:
  Device-1: NVIDIA GP104 [GeForce GTX 1070] vendor: Gigabyte driver: nvidia 
  v: 460.67 alternate: nouveau,nvidia_drm bus-ID: 1c:00.0 chip-ID: 10de:1b81 
  class-ID: 0300 
  Display: x11 server: X.Org 1.20.10 compositor: kwin_x11 driver: 
  loaded: nvidia display-ID: :0 screens: 1 
  Screen-1: 0 s-res: 3840x1080 s-dpi: 80 s-size: 1219x343mm (48.0x13.5") 
  s-diag: 1266mm (49.9") 
  Monitor-1: DVI-D-0 res: 1920x1080 dpi: 82 size: 597x336mm (23.5x13.2") 
  diag: 685mm (27") 
  Monitor-2: HDMI-0 res: 1920x1080 hz: 60 dpi: 96 size: 510x287mm (20.1x11.3") 
  diag: 585mm (23") 
  OpenGL: renderer: GeForce GTX 1070/PCIe/SSE2 v: 4.6.0 NVIDIA 460.67 
  direct render: Yes 
Audio:
  Device-1: Creative Labs EMU20k2 [Sound Blaster X-Fi Titanium Series] 
  driver: snd_ctxfi v: kernel bus-ID: 1b:00.0 chip-ID: 1102:000b 
  class-ID: 0403 
  Device-2: NVIDIA GP104 High Definition Audio vendor: Gigabyte 
  driver: snd_hda_intel v: kernel bus-ID: 1c:00.1 chip-ID: 10de:10f0 
  class-ID: 0403 
  Device-3: AMD Family 17h HD Audio vendor: Micro-Star MSI 
  driver: snd_hda_intel v: kernel bus-ID: 1e:00.3 chip-ID: 1022:1457 
  class-ID: 0403 
  Sound Server-1: ALSA v: k5.11.10-1-MANJARO running: yes 
  Sound Server-2: sndio v: N/A running: no 
  Sound Server-3: JACK v: 0.125.0 running: no 
  Sound Server-4: PulseAudio v: 14.2 running: yes 
  Sound Server-5: PipeWire v: 0.3.24 running: yes 
Use of uninitialized value $args in concatenation (.) or string at /usr/bin/inxi line 2391.
Use of uninitialized value in concatenation (.) or string at /usr/bin/inxi line 2391.
  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed
100    13  100    13    0     0     54      0 --:--:-- --:--:-- --:--:--    54
Network:
  Device-1: Intel I211 Gigabit Network vendor: Micro-Star MSI driver: igb 
  v: kernel port: f000 bus-ID: 18:00.0 chip-ID: 8086:1539 class-ID: 0200 
  IF: enp24s0 state: up speed: 1000 Mbps duplex: full mac: <filter> 
  IP v4: <filter> type: dynamic noprefixroute scope: global 
  broadcast: <filter> 
  IP v6: <filter> type: noprefixroute scope: global 
  IP v6: <filter> type: temporary dynamic scope: global 
  IP v6: <filter> type: mngtmpaddr noprefixroute scope: global 
  IP v6: <filter> type: noprefixroute scope: link 
  Device-2: TP-Link UE300 10/100/1000 LAN (ethernet mode) [Realtek RTL8153] 
  type: USB driver: r8152 bus-ID: 1-10:2 chip-ID: 2357:0601 class-ID: 0000 
  serial: <filter> 
  IF: enp3s0f0u10 state: down mac: <filter> 
  WAN IP: <filter> 
Bluetooth:
  Message: No Bluetooth data was found. 
Logical:
  Message: No LVM data was found. 
RAID:
  Message: No RAID data was found. 
Drives:
  Local Storage: total: 4.55 TiB used: 24.07 TiB (529.3%) 
  SMART Message: Unable to run smartctl. Root privileges required. 
  ID-1: /dev/nvme0n1 maj-min: 259:0 vendor: Samsung model: SSD 970 EVO 500GB 
  size: 465.76 GiB block-size: physical: 512 B logical: 512 B speed: 31.6 Gb/s 
  lanes: 4 rotation: SSD serial: <filter> rev: 2B2QEXE7 temp: 39.9 C 
  scheme: GPT 
  ID-2: /dev/sda maj-min: 8:0 vendor: Samsung model: SSD 860 EVO 500GB 
  size: 465.76 GiB block-size: physical: 512 B logical: 512 B speed: 6.0 Gb/s 
  rotation: SSD serial: <filter> rev: 1B6Q scheme: MBR 
  ID-3: /dev/sdb maj-min: 8:16 vendor: HGST (Hitachi) model: HDN724040ALE640 
  size: 3.64 TiB block-size: physical: 4096 B logical: 512 B speed: 6.0 Gb/s 
  rotation: 7200 rpm serial: <filter> rev: A5E0 scheme: GPT 
  Optical-1: /dev/sr0 vendor: HL-DT-ST model: DVDRAM GH20NS15 rev: IL00 
  dev-links: cdrom 
  Features: speed: 48 multisession: yes audio: yes dvd: yes 
  rw: cd-r,cd-rw,dvd-r,dvd-ram state: running 
Partition:
  ID-1: / raw-size: 430.95 GiB size: 423.18 GiB (98.20%) 
  used: 356.21 GiB (84.2%) fs: ext4 dev: /dev/nvme0n1p2 maj-min: 259:2 
  label: Local Disk 1 uuid: 576659b2-353b-43cc-9510-aa3ac55bd386 
  ID-2: /boot/efi raw-size: 300 MiB size: 299.4 MiB (99.80%) 
  used: 280 KiB (0.1%) fs: vfat dev: /dev/nvme0n1p1 maj-min: 259:1 label: N/A 
  uuid: 71D1-C987 
  ID-3: /data/nfs/Documents raw-size: N/A size: 20.95 TiB 
  used: 20.63 TiB (98.5%) fs: nfs4 remote: 192.168.1.242:/volume1/Documents 
  label: N/A uuid: N/A 
  ID-4: /home/<filter>/mount/Local Disk 2 raw-size: 465.76 GiB 
  size: 457.45 GiB (98.22%) used: 405.69 GiB (88.7%) fs: ext4 dev: /dev/sda1 
  maj-min: 8:1 label: Local Disk 2 uuid: 340457fe-a771-4dfa-933b-6a4fcdd5b794 
  ID-5: /run/timeshift/backup raw-size: 3.64 TiB size: 3.58 TiB (98.40%) 
  used: 2.7 TiB (75.4%) fs: ext4 dev: /dev/sdb1 maj-min: 8:17 
  label: Local Disk 3 uuid: 750c820d-d1e6-4d1f-a8d7-1772fde1ec50 
Swap:
  Kernel: swappiness: 60 (default) cache-pressure: 100 (default) 
  ID-1: swap-1 type: partition size: 34.52 GiB used: 0 KiB (0.0%) priority: -2 
  dev: /dev/nvme0n1p3 maj-min: 259:3 label: N/A 
  uuid: d4c036ba-fcfd-415b-8df9-f4bf690b2aa7 
Unmounted:
  Message: No Unmounted partitions found. 
USB:
  Hub-1: 1-0:1 info: Full speed (or root) Hub ports: 14 rev: 2.0 
  speed: 480 Mb/s chip-ID: 1d6b:0002 class-ID: 0900 
  Device-1: 1-10:2 
  info: TP-Link UE300 10/100/1000 LAN (ethernet mode) [Realtek RTL8153] 
  type: Network driver: r8152 interfaces: 1 rev: 2.1 speed: 480 Mb/s 
  power: 180mA chip-ID: 2357:0601 class-ID: 0000 serial: <filter> 
  Device-2: 1-11:3 info: MCT Elektronikladen farbwerk type: HID 
  driver: hid-generic,usbhid interfaces: 1 rev: 2.0 speed: 12 Mb/s power: 2mA 
  chip-ID: 0c70:f00a class-ID: 0300 serial: <filter> 
  Device-3: 1-12:4 info: Logitech G Pro Gaming Mouse type: Mouse,HID 
  driver: hid-generic,usbhid interfaces: 2 rev: 2.0 speed: 12 Mb/s 
  power: 300mA chip-ID: 046d:c085 class-ID: 0300 serial: <filter> 
  Device-4: 1-13:5 info: Atmel WootingOne type: HID,Keyboard 
  driver: hid-generic,usbhid interfaces: 7 rev: 2.0 speed: 12 Mb/s 
  power: 400mA chip-ID: 03eb:ff01 class-ID: 0300 serial: <filter> 
  Hub-2: 2-0:1 info: Full speed (or root) Hub ports: 8 rev: 3.1 speed: 10 Gb/s 
  chip-ID: 1d6b:0003 class-ID: 0900 
  Hub-3: 3-0:1 info: Full speed (or root) Hub ports: 4 rev: 2.0 
  speed: 480 Mb/s chip-ID: 1d6b:0002 class-ID: 0900 
  Hub-4: 4-0:1 info: Full speed (or root) Hub ports: 4 rev: 3.0 speed: 5 Gb/s 
  chip-ID: 1d6b:0003 class-ID: 0900 
Sensors:
  System Temperatures: cpu: 41.8 C mobo: N/A gpu: nvidia temp: 50 C 
  Fan Speeds (RPM): N/A gpu: nvidia fan: 19% 
Info:
  Processes: 462 Uptime: 2h 15m wakeups: 0 Init: systemd v: 247 
  tool: systemctl Compilers: gcc: 10.2.0 alt: 8/9 clang: 11.1.0 Packages: 2628 
  pacman: 2595 lib: 585 flatpak: 12 snap: 21 Shell: Bash v: 5.1.0 
  running-in: konsole inxi: 3.3.03 

Logs

project_wingman-crash1.log
project_wingman-crash2.log
project_wingman-crash3.log
project_wingman-crash4.log
project_wingman-crash5.log
🥲

Greetings

KK0bin maintainer 2021-04-14 github

That's probably not the same problem. Make a separate issue.

Kkonomikitten 2021-04-14 github

That's probably not the same problem. Make a separate issue.

Are you sure about that? My logs from 455.50.10 look a lot like @Retardium's...

Mine:

err:   DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err:   DxvkDevice: waitForIdle: Operation failed
err:   DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err:   DxvkDevice: waitForIdle: Operation failed
err:   DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err:   DxvkDevice: waitForIdle: Operation failed
err:   DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err:   DxvkDevice: waitForIdle: Operation failed
err:   DxvkSubmissionQueue: Command submission failed: VK_ERROR_DEVICE_LOST

Theirs:

warn:  D3D11Device::CreateShaderModule: Class linkage not supported
warn:  D3D11Device::CreateShaderModule: Class linkage not supported
warn:  D3D11Device::CreateShaderModule: Class linkage not supported
warn:  D3D11Device::CreateShaderModule: Class linkage not supported
warn:  D3D11Device::CreateShaderModule: Class linkage not supported
warn:  D3D11Device::CreateShaderModule: Class linkage not supported
warn:  D3D11DXGIDevice::QueryInterface: Unknown interface query
warn:  6543dbb6-1b48-42f5-ab82-e97ec74326f6
warn:  D3D11Device::CreateShaderModule: Class linkage not supported
warn:  D3D11Device::CreateShaderModule: Class linkage not supported
warn:  D3D11Device::CreateShaderModule: Class linkage not supported
warn:  D3D11Device::CreateShaderModule: Class linkage not supported
err:   DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err:   DxvkDevice: waitForIdle: Operation failed
err:   DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err:   DxvkDevice: waitForIdle: Operation failed
err:   DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err:   DxvkSubmissionQueue: Command submission failed: VK_ERROR_DEVICE_LOST
err:   DxvkDevice: waitForIdle: Operation failed

Edit: By the way two weeks on Kernel 5.4.109,110,111 and Nvidia 440.100 and 0 crashes, I am pretty sure I'm going to get to a month with 0 crashes as well. Really not liking Nvidia as usual.

Wwasteoinc 2021-04-22 github

I also get the same issue on panzer Corps 2 on my gtx770 with both 450 and 460.

err:   DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err:   DxvkDevice: waitForIdle: Operation failed
err:   DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err:   DxvkDevice: waitForIdle: Operation failed
err:   DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err:   DxvkDevice: waitForIdle: Operation failed
err:   DxvkSubmissionQueue: Failed to sync fence: VK_ERROR_DEVICE_LOST
err:   DxvkDevice: waitForIdle: Operation failed

unfortunately 5.4.72 and 440.100 doesn't seem to help I get the same error

Kkonomikitten 2021-04-23 github

440.100 doesn't seem to help I get the same error

Are you absolutely sure you downgraded to 440.100?

Wwasteoinc 2021-04-25 github

440.100 doesn't seem to help I get the same error

Are you absolutely sure you downgraded to 440.100?

Unfortunately yes, else I would be a happy gamer. But maybe the problem lies with the specific game (panzer corps 2) . I will try some other games next week and see if I have the same issue

+-----------------------------------------------------------------------------+
| NVIDIA-SMI 440.100      Driver Version: 440.100      CUDA Version: 10.2     |
|-------------------------------+----------------------+----------------------+
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|===============================+======================+======================|
|   0  GeForce GTX 770     Off  | 00000000:07:00.0 N/A |                  N/A |
| 30%   49C    P8    N/A /  N/A |    184MiB /  1994MiB |     N/A      Default |
+-------------------------------+----------------------+----------------------+
Kkonomikitten 2021-04-25 github

Unfortunately yes, else I would be a happy gamer.

I guess @K0bin was right then sorry about that K0bin.

Wwasteoinc 2021-04-26 github

Unfortunately yes, else I would be a happy gamer.

I guess @K0bin was right then sorry about that K0bin.

Well I will try to do a full purge of nvidia and everything and install 440.100 from command line so I can more properly track what's going on because I found so many different versions of nvidia packages in my installation.

give me a bit more time..

ehm do you have a guide on how to install 440.100 ? even though the drivers installs in command line pretty well now that I purged everything when I boot it loads the mesa driver. (shy)
I try to install the run file and I get to the libglvnd error. And the PPA doesn't seem to include 440 anymore. So I get now to this stage
NVRM: API mismatch: the client has the version 440.100, but
NVRM: this kernel module has the version 460.67. Please
NVRM: make sure that this kernel module and all NVIDIA driver
NVRM: components have the same version.

Wwasteoinc 2021-04-26 github

2 hours of clusterfuck and driver parts here and there and I managed to load a proper clean 440.100 installation (no 32bit support and no DKMS though) . I managed to actually play for more time than I ever managed on this game (like an hour) but it eventually crashed to our usual issue. I have to go to work now I will try some more with frostpunk tonight to see if it a reaccuring thing.

Kkonomikitten 2021-05-03 github

I've downgraded all the way to 440.100 again and I'll be tested that for a good month or so.

So it's been a month now and no crashes, I'm just going to stick with this kernel and driver version. I've pretty much given up on trying to update the Nvidia Driver and not have this issue occur.

Software information

Final Fantasy XIV

System information

  • GPU: GeForce GTX 960
  • Kernel: 5.4.115
  • Driver: 440.100
  • Wine version: lutris-6.4-x86_64
  • DXVK version: 1.8.1

Log files

Ssfjuocekr 2021-07-31 github

I used to have a bunch of crashes with the 46x series until they "fixed" 465 and 470 works as well.

But you should be using: 455.50.19

I can highly recommend purging the 4xx drivers entirely, make sure you clean up the dkms folder as well and then do a manual install. You will probably lose video if you try this, so do it from a SSH console.

This driver gives by far the best performance on my GTX970 and zero crashes thus far.

Aarabcian 2021-07-31 github

I upgraded my laptop now i have RTX 3060 mobile GPU and i can confirm errors are no more with any driver ive tried. This error is called something like xid31 by nvidia and latest 470 drivers have a fix for it.

Nnzbtuxnews 2021-08-29 github

Just found out about this issue and reported it on ticket 2253

I have drivers 470.57.2 and this issue is NOT fixed.

KK0bin maintainer 2021-08-29 github

That's most likely an unrelated issue.

DEVICE_LOST just means the driver encountered an error for one reason or another.

Nneskweek 2021-09-15 github

Still Broken with 470.63.01

Kkonomikitten 2021-09-20 github

@doitsujin considering people are still reporting this is broken, and https://github.com/doitsujin/dxvk/pull/1963 made the major version check for the nvidia work around 465, shouldn't this just default to being back on for all Nvidia versions? As far as I know this bug was the reason the work around existed in the first place. It doesn't seem like Nvidia actually fixed it.

Ddoitsujin maintainer 2021-09-20 github

I'm not convinced that the issues people are reporting are actually caused by the same issue. There's still a problem somewhere regarding high HVV use but I've never seen it crash since.

Kkonomikitten 2021-09-20 github

There's still a problem somewhere regarding high HVV use but I've never seen it crash since.

I'm currently testing 470.74 if I still manage to end up making the game crash (could take a month or more) but I've still never in at least a year of playing FFXIV had 440.100 crash, where does that ultimately put things? Is this just a rare bug that will never be fixed? Do you have any theories on this at all? I would've just escaped this by switching to an AMD card but with the whole market the way it is it hasn't been an option.

Do you have any thoughts on this at all? It's driving me nuts.

SSveSop 2021-09-20 github

Interesting note on the 470.74 tho:

Fixed a regression which resulted in very-high system memory usage for Direct3D 12 games when run through vkd3d-proton.

Could be useful for DXVK too maybe? (Ie. high-memory related issue).

That said, nVidia hardware CAN sometimes be throwing those random XID errors when different manufacturers are pushing the limits on their "factory clocks" (Eg. Asus/MSI++ with their "OC" models and whatnot). Could be worth TRYING to underclock or powerlimit the gpu just to see.

You can check your powerlimit with nvidia-smi Then if it has like 215W (like my RTX2070 has), then you could try to use
sudo nvidia-smi -pl 200
Just mentioning this because i had a nVidia (260Ti card back when) that was really unstable - but it was one of those "OC Twin Frozr Whatever" cards, and ended up just downclocking it a tad, and issue went away :) Loads of factors playing in here - especially since XID(DEVICE LOST) errors MAY not be driver/dxvk related at all.

Ddoitsujin maintainer 2021-09-21 github

Could be useful for DXVK too maybe? (Ie. high-memory related issue).

No, the issue was specifically caused by the way vkd3d-proton uses Vulkan pipeline caches. DXVK is not affected by this in any way.

Kkonomikitten 2021-12-22 github

Honestly at this point I'll willing to concede this might just be a game bug. I've looked through the tracker and see other games that will crash with an in game bug but still have the DXVK log producing VK_ERROR_DEVICE_LOST. There are multiple numerous reports from Windows users on the FFXIV forums about DX11 crashes. I also play with my significant other who has had both the DX11 crash and sometimes a freeze while playing the game and they use Windows and AMD.

I think the biggest problem here is VK_ERROR_DEVICE_LOST is just too much of a generic error, is there anyway DXVK can give a better error code?

I think it's probably time to close this bug, I personally after all this time don't think this is a Nvidia or DXVK bug. Thoughts?

Ssschroe 2021-12-22 github

I don't think that the crashes on Windows are related to this. I have several friends using Nvidia cards (including RTX 3xxx) on Windows that afk all day in-game and never mentioned anything about the game crashing. And if this was actually a widespread issue there'd be a lot more noise about it with the current login queues. The people on windows having crashes more likely is borked installations, hardware or overclocks.

And even on Linux only newer Nvidia drivers are affected by it reproducibly, older ones and AMD cards run rock solid. Since switching to a 6900 XT I haven't had a single crash in the game. So from my point of view this still looks like a bug in Nvidias driver on Linux.

As far I know Nvidia devs also browse through these issues once in a while so keeping the issue open might make sense for the small chance of them looking into it one day, but more practically this would belong on the Nvidia forums or whatever they use for bug reports.

Kkonomikitten 2021-12-22 github

And if this was actually a widespread issue there'd be a lot more noise about it with the current login queues. The people on windows having crashes more likely is borked installations, hardware or overclocks.

There is an extensive thread over on the Final Fantasy Forums though: https://forum.square-enix.com/ffxiv/threads/448687-Directx11-error

Users have the following hardware in that thread:

AMD Ryzen 6600XT
AMD Ryzen 7 5800X

AMD Ryzen 7 5800X 8-Core Processor
AMD Radeon RX 6600 XT

AMD Ryzen 5 3600 6-Core Processor
AMD Radeon RX 5700 XT

AMD Ryzen 5 3600 6-Core Processor
Radeon RX 580 Series

Intel(R) Core(TM) i7-8700 CPU
NVIDIA GeForce RTX 2070

AMD Ryzen 7 5700G with Radeon Graphics
NVIDIA GeForce GTX 1080

AMD Ryzen 5950x
AMD Radeon 6900x

As one user on the thread states:

Nothing has fixed this issue. Sometimes we can play all day with no problems, the next day the DirectX 11 11000002 error comes up over and over. If you lookup that error number, you'll see MILLIONS of posts online (not just on these forums) about this FF14 issue. This is not an obscure thing, A LOT OF PEOPLE have this problem.

And another user said it's been there since Stormblood:

A large part of the community have been waiting for an answer regarding this issue since Stormblood release, so I'm not sure if we are going to get one anytime soon if we didn't get one so far.

ZZeroPointEnergy 2021-12-23 github

Apart from the VK_ERROR_DEVICE_LOST I also randomly get complete nvidia driver crashes with this game. I first thought my GPU is probably defect, but Xid 16 means “Display engine hung” and “Driver Error” according to nvidias own list and I do play other games, sometimes for hours and I have never seen such a crash outside of FFXIV.

This is with a NVIDIA GeForce GTX TITAN X card.

[11303.224211] NVRM: GPU at PCI:0000:01:00: GPU-0b4cda80-07b2-13fb-9c08-f03bf0381c28
[11303.224213] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1ef6
[11311.534091] NVRM: Xid (PCI:0000:01:00): 16, pid=16, Head 00000000 Count 000b1ef7
[11319.853967] NVRM: Xid (PCI:0000:01:00): 16, pid=6226, Head 00000000 Count 000b1ef8
[11328.183850] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1ef9
[11336.503749] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1efa
[11344.813634] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1efb
[11353.133512] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1efc
[11361.463408] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1efd
[11369.783297] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1efe
[11378.093175] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1eff
[11386.413068] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1f00
[11394.732946] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1f01
[11403.062837] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1f02
[11411.372721] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1f03
[11419.692602] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1f04
[11428.022469] NVRM: Xid (PCI:0000:01:00): 16, pid=0, Head 00000000 Count 000b1f05
[11463.772567] nvidia-modeset: WARNING: GPU:0: Timeout while waiting for idle.
[11465.772637] nvidia-modeset: ERROR: GPU:0: Idling display engine timed out: 0x0000957d:0:0:428

Made a post on their forums, sent a bug report to the "support" address, never even got an answer back.

Also should mention that the driver crash is 100% not a DXVK bug, since it also happens with pure wine (only tested once, because the performance is horrible). I'm not sure if this is related at all to the original bug here, but I just wanted to add that information, in the hope it may help. If this confuses things more, please just disregard.

Kkonomikitten 2021-12-23 github

I first thought my GPU is probably defect, but Xid 16 means “Display engine hung” and “Driver Error” according to nvidias own list and I do play other games, sometimes for hours and I have never seen such a crash outside of FFXIV.

And that's the problem I play other games outside of FFXIV and they never crash either, we have a bunch of windows users reporting the same thing on different GPUs and if you take a look at other bug reports like this one https://github.com/doitsujin/dxvk/issues/2349 you'll see that bugs in games can definitely cause VK_ERROR_DEVICE_LOST and not just problems with Nvidia or DXVK.

As I said I've played this game a lot and so has my significant other I've watched their PC running Windows have all the same crashes I do on DXVK and Nvidia and they're using Windows and AMD.

I believe the bug is in the game itself for this reason, I also believe that depending on hardware/driver versions the bug can occur less or more frequently. I only wish I had the money to experiment with different GPUs and end this debate for once and all.

BBlisto91 2022-07-20 github

@konomikitten How is this behaving with the 515 drivers if you are still testing this from time to time.

Kkonomikitten 2022-07-21 github

@Blisto91 I switched from my GTX 960 to a RX 6600 about 4 months ago and I haven't had this crash since. My significant other switched from an RX 570 to a RX 6600XT and now gets this issue on Windows 10.

I still see frequent posts from Windows users on the Final Fantasy Forums about DirectX Error 11000002. As far as I am concerned this is a game bug that Square need to fix.

Aastraldawn 2022-10-16 github

@konomikitten How is this behaving with the 515 drivers if you are still testing this from time to time.

No issues with 515.48.07 (as stable as 440.100)

Kkonomikitten 2022-10-16 github

@konomikitten How is this behaving with the 515 drivers if you are still testing this from time to time.

No issues with 515.48.07 (as stable as 440.100)

I no longer use a Nvidia Graphics card as per my previous post so I cannot answer that question.

BBlisto91 2022-10-16 github

I think they are stating that the 515 drivers are stable for them.