protonscr

Nier: Automata crashes within a few minutes of starting a new game

protonclosed appid 524220Game compatibilityMesa driversAMD RADV
ValveSoftware/Proton#3599 · opened 2020-03-05 by nstgc · updated 2020-10-22 · 10 comments · github · game page · search this game
1 matching comments, n / p to jump
Nnstgc 2020-03-05 github

Compatibility Report

  • Name of the game with compatibility issues: Nier: Automata
  • Steam AppID of the game: 524220

System Information

  • GPU: RX 5600 XT
  • Driver/LLVM version: Mesa 19.3.4, LLVM 9.0.1
  • Kernel version: 5.5.7-arch1-1
  • System Info
    • Some additional info:
$ pacman -Q |egrep "mesa|amd|vulkan|llvm"
lib32-libva-mesa-driver 19.3.4-3
lib32-llvm-libs 9.0.1-1
lib32-mesa 19.3.4-3
lib32-mesa-vdpau 19.3.4-3
lib32-vulkan-icd-loader 1.2.132-1
lib32-vulkan-radeon 19.3.4-3
libva-mesa-driver 19.3.4-2
llvm 9.0.1-1
llvm-libs 9.0.1-1
mesa 19.3.4-2
mesa-demos 8.4.0-2
mesa-vdpau 19.3.4-2
vulkan-icd-loader 1.2.132-1
vulkan-radeon 19.3.4-2
xf86-video-amdgpu 19.1.0-1
  • Proton version: 5.0-3

I confirm:

  • [X] that I haven't found an existing compatibility report for this game.
    • I know that there are many other open reports, but I also know this is official supported. Not of them deal with crashes, so I thought I'd open a new one.
  • [X] that I have checked whether there are updates for my system available.

This game locked up my computer so hard I had to REISUB to restart it, so I will not be rerunning this game for a crash report. Instead I'll post my journalctl output:

Mar 04 23:51:51 Host kernel: [drm:amdgpu_dm_atomic_commit_tail [amdgpu]] *ERROR* Waiting for fences timed out!
Mar 04 23:51:51 Host kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ring gfx_0.0.0 timeout, signaled seq=195049, emitted seq=195051
Mar 04 23:51:51 Host kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* Process information: process NieRAutomata.ex pid 19179 thread NieRAutomata.ex pid 19179
Mar 04 23:51:51 Host kernel: [drm] GPU recovery disabled.
Mar 04 23:52:05 Host kernel: [drm:amdgpu_dm_atomic_commit_tail [amdgpu]] *ERROR* Waiting for fences timed out!
Mar 04 23:52:05 Host kernel: sysrq: HELP : loglevel(0-9) reboot(b) crash(c) terminate-all-tasks(e) memory-full-oom-kill(f) kill-all-tasks(i) thaw-filesystems(j) sak(k) show-backtrace-all-active-cpus(l) show->
Mar 04 23:52:06 Host kernel: sysrq: Keyboard mode set to system default
Mar 04 23:52:12 Host systemd-journald[441]: Journal stopped

Symptoms

I loaded up Nier: Automata, and started a new game. During the opening cut scene when you are flying and are projecting that screen thing in front of you, before the giant beam starts taking everyone out, the game locked up. Music was still playing (for a while) but the computer was unresponsive.

I didn't just mash REISUB, but rather tried it incrementally over the period of several minutes. There was over 20 minutes (closer to 30) between me pressing SysRq+R and actually hitting the final SysRq+B. During that time I couldn't even switch to another TTY.

Reproduction

Install and start Nier: Automata with Proton 5.0-3. Start a new game. Play said game. Crash.

Hhakzsam 2020-03-05 github

Does it hang with ACO?

Nnstgc 2020-03-05 github

Does it hang with ACO?

I have not tried it with ACO. I don't think this is a hang. It didn't feel like one last night, and with those journalctl entries it doesn't seem like one now.

One detail I missed is that I didn't just mash REISUB, but rather tried it incrementally over the period of several minutes. There was over 20 minutes (closer to 30) between me pressing SysRq+R and actually hitting the final SysRq+B. During that time I couldn't even switch to another TTY.

Hhakzsam 2020-03-05 github

LLVM is known to randomly hang your GPU on Navi but AFAIK ACO is more stable.

Nnstgc 2020-03-05 github

LLVM is known to randomly hang your GPU on Navi but AFAIK ACO is more stable.

To be clear, you are saying that LLVM can cause a hang that locks up the system so hard that SysRq+REI isn't enough to get me into another TTY, even after waiting 20 minutes? I'm not being argumentative, it's just I'm not sure you read my edit which provides this info.

Hhakzsam 2020-03-05 github

When it's a hard lockup like what you got, likely.

Nnstgc 2020-03-05 github

When it's a hard lockup like what you got, likely.

I'll try that this afternoon then. Thanks.

Nnstgc 2020-03-06 github

That did not solve the issue. I did get a bit further in the game, but I think that might be a random thing. Could this be hardware related?

The journal output is different this time.

Mar 05 21:16:38 Host kernel: [drm:amdgpu_dm_atomic_commit_tail [amdgpu]] *ERROR* Waiting for fences timed out!
Mar 05 21:16:38 Host kernel: [drm:amdgpu_dm_atomic_commit_tail [amdgpu]] *ERROR* Waiting for fences timed out!
Mar 05 21:16:38 Host kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ring gfx_0.0.0 timeout, signaled seq=3580233, emitted seq=3580235
Mar 05 21:16:38 Host kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* Process information: process NieRAutomata.ex pid 59497 thread NieRAutomata.ex pid 59497
Mar 05 21:16:38 Host kernel: [drm] GPU recovery disabled.
Mar 05 21:16:41 Host steam.desktop[4193]: Warning: The game hasn't rendered a frame from us in over 10 seconds
Mar 05 21:16:54 Host steam.desktop[4193]: Warning: The game hasn't rendered a frame from us in over 10 seconds
Mar 05 21:16:56 Host gsd-power[1229]: Error setting property 'PowerSaveMode' on interface org.gnome.Mutter.DisplayConfig: Timeout was reached (g-io-error-quark, 24)
Mar 05 21:16:56 Host gnome-session[969]: gnome-session-binary[969]: GnomeDesktop-WARNING: Failed to acquire idle monitor proxy: Timeout was reached
Mar 05 21:16:56 Host gnome-session-binary[969]: GnomeDesktop-WARNING: Failed to acquire idle monitor proxy: Timeout was reached
Mar 05 21:17:04 Host kernel: sysrq: HELP : loglevel(0-9) reboot(b) crash(c) terminate-all-tasks(e) memory-full-oom-kill(f) kill-all-tasks(i) thaw-filesystems(j) sak(k) show-backtrace-all-active-cpus(l) show-memory-usage(m) nice-all-RT-tasks(n) poweroff(o) show-registers(p) show-all-timers(q) unraw(r) sync(s) show-task-states(t) unmount(u) force-fb(V) show-blocked-tasks(w) dump-ftrace-buffer(z) 
Mar 05 21:17:04 Host kernel: sysrq: HELP : loglevel(0-9) reboot(b) crash(c) terminate-all-tasks(e) memory-full-oom-kill(f) kill-all-tasks(i) thaw-filesystems(j) sak(k) show-backtrace-all-active-cpus(l) show-memory-usage(m) nice-all-RT-tasks(n) poweroff(o) show-registers(p) show-all-timers(q) unraw(r) sync(s) show-task-states(t) unmount(u) force-fb(V) show-blocked-tasks(w) dump-ftrace-buffer(z) 
Mar 05 21:17:05 Host kernel: [drm:amdgpu_dm_atomic_commit_tail [amdgpu]] *ERROR* Waiting for fences timed out!
Mar 05 21:17:05 Host kernel: sysrq: HELP : loglevel(0-9) reboot(b) crash(c) terminate-all-tasks(e) memory-full-oom-kill(f) kill-all-tasks(i) thaw-filesystems(j) sak(k) show-backtrace-all-active-cpus(l) show-memory-usage(m) nice-all-RT-tasks(n) poweroff(o) show-registers(p) show-all-timers(q) unraw(r) sync(s) show-task-states(t) unmount(u) force-fb(V) show-blocked-tasks(w) dump-ftrace-buffer(z) 
Mar 05 21:17:09 Host kernel: sysrq: HELP : loglevel(0-9) reboot(b) crash(c) terminate-all-tasks(e) memory-full-oom-kill(f) kill-all-tasks(i) thaw-filesystems(j) sak(k) show-backtrace-all-active-cpus(l) show-memory-usage(m) nice-all-RT-tasks(n) poweroff(o) show-registers(p) show-all-timers(q) unraw(r) sync(s) show-task-states(t) unmount(u) force-fb(V) show-blocked-tasks(w) dump-ftrace-buffer(z) 
Mar 05 21:17:11 Host kernel: sysrq: Keyboard mode set to system default
.
.
.
Mar 05 21:17:19 Host kernel: sysrq: Keyboard mode set to system default
Mar 05 21:17:33 Host kernel: sysrq: Terminate All Tasks
Mar 05 21:17:33 Host haveged[440]: haveged: Stopping due to signal 15
Mar 05 21:17:33 Host haveged[440]: haveged starting up
Mar 05 21:17:33 Host dhcpcd[822]: received SIGTERM, stopping
Mar 05 21:17:33 Host dhcpcd[822]: enp12s0: removing interface
Mar 05 21:17:33 Host systemd-journald[441]: Journal stopped
Hhakzsam 2020-10-22 github

@nstgc Are you still able to reproduce this GPU hang with Mesa 20.2.x ?

Nnstgc 2020-10-22 github

@nstgc Are you still able to reproduce this GPU hang with Mesa 20.2.x ?

I haven't played in a while, to be honest, however when last I did play, this didn't seem to be a problem. If I recall correctly, at the time I was having this issue the Navi GPU's didn't have proper/full/whatever GPU resetting. That has since been resolved in more recent kernel versions. I don't think it has to do with Mesa.

Kkisak-valve maintainer 2020-10-22 github

Thanks for the feedback @nstgc. Closing.

Proton versions