00
In productionProadvancedIncident command
Recovery causes a second outage
Contain an overloaded discovery service without restarting the whole fleet simultaneously.
Pairs with Fleet incident response in Production ROS Operations. Read the concept first, then diagnose it here.
Unlock this lab with ProOperator report
Operators restart every robot, creating a discovery storm that overloads the network again.
Ubuntu 24.04, ROS 2 Jazzy, Cyclone DDS, containers, systemd, OpenTelemetry fleet sandbox
System boundary
Trace only the relevant path.
- 01operator recovery action→
- 02fleet restart wave→
- 03network capacityOUT
02
How this lab works
You diagnose it. No command list.
Open the repair workspace and gather your own evidence in a real terminal. No diagnostic commands are handed to you - finding the fault is the exercise. Stuck? Progressive hints unlock inside the workspace.
Verification
Pass more than the visible symptom.
- Staged recovery remains below capacity
- Abort threshold stops regression
- Hidden regression behavior
- Root cause explanation