Decisions · ADR 0005

Safe Mode, crash reports and a watched main loop, in every build

Rollback protects against new firmware that fails Probation. It does nothing for firmware that was confirmed and crashes later: a corrupt setting, a server that sends something unexpected, a bug that takes an hour to show.…

Rollback protects against new firmware that fails Probation. It does nothing for firmware that was confirmed and crashes later: a corrupt setting, a server that sends something unexpected, a bug that takes an hour to show. Without a cable, such a device would restart forever. Three measures, in release and Debug Builds alike, keep it reachable:

  • Safe Mode. The firmware counts starts that follow a crash (panic or watchdog) in NVS, first thing at boot. After 3 in a row, it starts only the clock, Wi-Fi, the Update Service and, in a Debug Build, the Debug Console: no Apps, no IRC, no SD card, and a screen that says so with the address to push an update to. Any normal restart (a reboot, an update), or a minute of uptime, resets the count.
  • Crash reports. The same boot record keeps which version was running, so after a crash the firmware knows which one crashed, even when a Rollback has switched slots since. ESP-IDF already writes a core dump to its flash partition on a panic; after the restart the firmware prints its summary (task, PC, reason, backtrace) and raises a Notification. The crash command shows it again later. In a Debug Build, scripts/rdbg.py crash decodes the backtrace and scripts/rdbg.py coredump fetches the whole dump for esp-coredump, against the ELF of that exact build (.pio/elves/, named by version and ELF digest).
  • The main loop is watched. Arduino-ESP32 subscribes only core 0's idle task to the task watchdog, and the main loop runs on core 1: a stuck loop used to hang the device for good, with the screen frozen and the Debug Console unable to run commands. enableLoopWDT() makes a loop stuck for 5 s a panic, with a core dump, counted towards Safe Mode. And an installed update no longer depends on the main loop: the Update Service restarts into it by itself after 90 s.

Consequences

  • Nothing in the main loop may block for 5 s. Network and card work already run on their own tasks.
  • Safe Mode can't help when Wi-Fi or the Update Service itself is what crashes; that still needs USB.
  • Three crashes within a minute of each restart are needed to reach Safe Mode, so a crash loop costs about half a minute before the device becomes reachable.

This page is generated from docs/adr/0005-safe-mode-crash-reports-watchdog.md in the repository. To change it, change that file and run site/tools/gen_dev_docs.py.