Does it hang with ACO?
Does it hang with ACO?
I have not tried it with ACO. I don't think this is a hang. It didn't feel like one last night, and with those journalctl entries it doesn't seem like one now.
One detail I missed is that I didn't just mash REISUB, but rather tried it incrementally over the period of several minutes. There was over 20 minutes (closer to 30) between me pressing SysRq+R and actually hitting the final SysRq+B. During that time I couldn't even switch to another TTY.
LLVM is known to randomly hang your GPU on Navi but AFAIK ACO is more stable.
LLVM is known to randomly hang your GPU on Navi but AFAIK ACO is more stable.
To be clear, you are saying that LLVM can cause a hang that locks up the system so hard that SysRq+REI isn't enough to get me into another TTY, even after waiting 20 minutes? I'm not being argumentative, it's just I'm not sure you read my edit which provides this info.
When it's a hard lockup like what you got, likely.
When it's a hard lockup like what you got, likely.
I'll try that this afternoon then. Thanks.
That did not solve the issue. I did get a bit further in the game, but I think that might be a random thing. Could this be hardware related?
The journal output is different this time.
Mar 05 21:16:38 Host kernel: [drm:amdgpu_dm_atomic_commit_tail [amdgpu]] *ERROR* Waiting for fences timed out!
Mar 05 21:16:38 Host kernel: [drm:amdgpu_dm_atomic_commit_tail [amdgpu]] *ERROR* Waiting for fences timed out!
Mar 05 21:16:38 Host kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ring gfx_0.0.0 timeout, signaled seq=3580233, emitted seq=3580235
Mar 05 21:16:38 Host kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* Process information: process NieRAutomata.ex pid 59497 thread NieRAutomata.ex pid 59497
Mar 05 21:16:38 Host kernel: [drm] GPU recovery disabled.
Mar 05 21:16:41 Host steam.desktop[4193]: Warning: The game hasn't rendered a frame from us in over 10 seconds
Mar 05 21:16:54 Host steam.desktop[4193]: Warning: The game hasn't rendered a frame from us in over 10 seconds
Mar 05 21:16:56 Host gsd-power[1229]: Error setting property 'PowerSaveMode' on interface org.gnome.Mutter.DisplayConfig: Timeout was reached (g-io-error-quark, 24)
Mar 05 21:16:56 Host gnome-session[969]: gnome-session-binary[969]: GnomeDesktop-WARNING: Failed to acquire idle monitor proxy: Timeout was reached
Mar 05 21:16:56 Host gnome-session-binary[969]: GnomeDesktop-WARNING: Failed to acquire idle monitor proxy: Timeout was reached
Mar 05 21:17:04 Host kernel: sysrq: HELP : loglevel(0-9) reboot(b) crash(c) terminate-all-tasks(e) memory-full-oom-kill(f) kill-all-tasks(i) thaw-filesystems(j) sak(k) show-backtrace-all-active-cpus(l) show-memory-usage(m) nice-all-RT-tasks(n) poweroff(o) show-registers(p) show-all-timers(q) unraw(r) sync(s) show-task-states(t) unmount(u) force-fb(V) show-blocked-tasks(w) dump-ftrace-buffer(z)
Mar 05 21:17:04 Host kernel: sysrq: HELP : loglevel(0-9) reboot(b) crash(c) terminate-all-tasks(e) memory-full-oom-kill(f) kill-all-tasks(i) thaw-filesystems(j) sak(k) show-backtrace-all-active-cpus(l) show-memory-usage(m) nice-all-RT-tasks(n) poweroff(o) show-registers(p) show-all-timers(q) unraw(r) sync(s) show-task-states(t) unmount(u) force-fb(V) show-blocked-tasks(w) dump-ftrace-buffer(z)
Mar 05 21:17:05 Host kernel: [drm:amdgpu_dm_atomic_commit_tail [amdgpu]] *ERROR* Waiting for fences timed out!
Mar 05 21:17:05 Host kernel: sysrq: HELP : loglevel(0-9) reboot(b) crash(c) terminate-all-tasks(e) memory-full-oom-kill(f) kill-all-tasks(i) thaw-filesystems(j) sak(k) show-backtrace-all-active-cpus(l) show-memory-usage(m) nice-all-RT-tasks(n) poweroff(o) show-registers(p) show-all-timers(q) unraw(r) sync(s) show-task-states(t) unmount(u) force-fb(V) show-blocked-tasks(w) dump-ftrace-buffer(z)
Mar 05 21:17:09 Host kernel: sysrq: HELP : loglevel(0-9) reboot(b) crash(c) terminate-all-tasks(e) memory-full-oom-kill(f) kill-all-tasks(i) thaw-filesystems(j) sak(k) show-backtrace-all-active-cpus(l) show-memory-usage(m) nice-all-RT-tasks(n) poweroff(o) show-registers(p) show-all-timers(q) unraw(r) sync(s) show-task-states(t) unmount(u) force-fb(V) show-blocked-tasks(w) dump-ftrace-buffer(z)
Mar 05 21:17:11 Host kernel: sysrq: Keyboard mode set to system default
.
.
.
Mar 05 21:17:19 Host kernel: sysrq: Keyboard mode set to system default
Mar 05 21:17:33 Host kernel: sysrq: Terminate All Tasks
Mar 05 21:17:33 Host haveged[440]: haveged: Stopping due to signal 15
Mar 05 21:17:33 Host haveged[440]: haveged starting up
Mar 05 21:17:33 Host dhcpcd[822]: received SIGTERM, stopping
Mar 05 21:17:33 Host dhcpcd[822]: enp12s0: removing interface
Mar 05 21:17:33 Host systemd-journald[441]: Journal stopped
@nstgc Are you still able to reproduce this GPU hang with Mesa 20.2.x ?
@nstgc Are you still able to reproduce this GPU hang with Mesa 20.2.x ?
I haven't played in a while, to be honest, however when last I did play, this didn't seem to be a problem. If I recall correctly, at the time I was having this issue the Navi GPU's didn't have proper/full/whatever GPU resetting. That has since been resolved in more recent kernel versions. I don't think it has to do with Mesa.
Thanks for the feedback @nstgc. Closing.
proton 5.0-3x1 2020-03
Compatibility Report
System Information
I confirm:
This game locked up my computer so hard I had to REISUB to restart it, so I will not be rerunning this game for a crash report. Instead I'll post my journalctl output:
Symptoms
I loaded up Nier: Automata, and started a new game. During the opening cut scene when you are flying and are projecting that screen thing in front of you, before the giant beam starts taking everyone out, the game locked up. Music was still playing (for a while) but the computer was unresponsive.
I didn't just mash REISUB, but rather tried it incrementally over the period of several minutes. There was over 20 minutes (closer to 30) between me pressing SysRq+R and actually hitting the final SysRq+B. During that time I couldn't even switch to another TTY.
Reproduction
Install and start Nier: Automata with Proton 5.0-3. Start a new game. Play said game. Crash.