00
In productionProadvancedService recovery

A restart loop erases the useful log

Bound restarts and preserve the first failure evidence.

Pairs with systemd and startup order in Production ROS Operations. Read the concept first, then diagnose it here.

Unlock this lab with Pro

Operator report

A misconfigured node restarts 400 times, floods logs, and exhausts CPU.

Ubuntu 24.04, ROS 2 Jazzy, Cyclone DDS, containers, systemd, OpenTelemetry fleet sandbox

01

System boundary

Trace only the relevant path.

3 nodes
  1. 01process failure
  2. 02systemd restart policy
  3. 03host resourcesOUT
02

How this lab works

You diagnose it. No command list.

Open the repair workspace and gather your own evidence in a real terminal. No diagnostic commands are handed to you - finding the fault is the exercise. Stuck? Progressive hints unlock inside the workspace.

03

Verification

Pass more than the visible symptom.

4 checks
  • Retries are bounded
  • First failure remains queryable
  • Hidden regression behavior
  • Root cause explanation