Debug Builds and the Debug Console · Console

Crashes, core dumps and Safe Mode

What happens when the firmware crashes: the report, the core dump, the build it is decoded against, the main-loop watchdog and Safe Mode.

Rollback (see How an update works) protects against new firmware that fails Probation. These measures cover firmware that was confirmed and then crashes: a corrupt setting, a server that sends something unexpected, a bug that takes an hour to show. They are in every build, release and Debug alike; the Debug Console is what makes them convenient. The decision is ADR 0005.

What a crash leaves behind

  • A core dump in a flash partition: ESP-IDF writes it on a panic.
  • A crash record in NVS, written first thing at the next boot: which version was running (so after a Rollback has switched slots, the firmware still knows which one crashed), and how many starts in a row followed a crash.
  • A summary on the console after the restart (task, program counter, reason, backtrace), and a Notification on screen.

The crash command prints it again whenever you like:

> crash
crash: last one in v0.9.0-1-g4ab873e-dirty+debug (panic)
crash: task loopTask, pc 0x4037e179, cause 0, address 0x00000000
crash: reason: abort() was called at PC 0x421209b3 on core 1
crash: backtrace 0x4037e179 0x4037e141 0x4038582d 0x421209b3 ...
crash: elf sha256 3c6a185e5

coredump erase forgets the dump.

Decode it: rdbg.py crash and rdbg.py coredump

An address is no use without the exact build that crashed. Every build archives its ELF in .pio/elves/, named by version and the first 16 hex digits of its SHA-256 (scripts/version.py); the core dump names the crashed firmware by the same digest. So a crash can be decoded after later builds, including a build you have since replaced:

scripts/rdbg.py crash              # the crash report, with the backtrace turned into functions and source lines
scripts/rdbg.py coredump           # fetch the whole core dump, then decode it
scripts/rdbg.py coredump my.bin    # ...to a file you name
  • crash runs the device's crash, finds the ELF by digest (or by version), and passes the backtrace to scripts/decode_backtrace.sh, which runs addr2line from the build container: one line for each frame, with function and file:line.
  • coredump fetches the raw dump over the console (it works when the main loop is stuck) and runs scripts/decode_coredump.sh: esp-coredump and GDB, giving every task's backtrace, the registers, and the crashed task's stack.
  • By hand: scripts/decode_backtrace.sh <version|digest> <address>..., and scripts/decode_coredump.sh <core.bin> <version|digest>. Both say which ELF they used.
  • If no archived ELF matches, they say so: the build was made on another machine, or .pio/ was cleaned.

The main loop is watched

Arduino-ESP32 puts only core 0's idle task on the task watchdog, and the main loop runs on core 1: a stuck loop used to hang the device for good, with a frozen screen and a console that could not run commands. Now a loop stuck for 5 seconds is a panic, with a core dump, counted towards Safe Mode. The rule that follows: nothing in the main loop may block for 5 seconds. Network and card work already run on their own tasks.

An installed update no longer depends on the main loop either: the Update Service restarts into it by itself after 90 seconds.

Safe Mode

The count of starts that follow a crash (a panic or the watchdog) is kept in NVS. After three in a row, the firmware starts Safe Mode instead of everything else:

  • only the clock, Wi-Fi, the Update Service and, in a Debug Build, the Debug Console start: no Apps, no IRC, no SD card;
  • the screen says so, with the address to push an update to;
  • only a few commands run (the command reference lists them); anything else answers not available in Safe Mode.

Any normal restart (reboot, an update) or a minute of uptime resets the count. So Safe Mode is reached by crashing three times quickly, and a crash loop costs about half a minute before the device becomes reachable. To leave it: push a fix (scripts/flash.sh --debug --ota), or reboot.

Its limit: if Wi-Fi or the Update Service is what crashes, Safe Mode cannot help, and it takes USB.

Crash on purpose

On a Debug Build:

crash abort      # abort(): a panic with a core dump
crash wdt        # hang the main loop until the task watchdog fires

Use them to see a real report and decode it, to check that the counter reaches Safe Mode, and to prove that a Debug Build you are about to rely on will survive its own crashes. The device restarts by itself after either, and the report is there to read when the console comes back.