Lesson 24 of 2455 minutes

Replanning, Timeouts, and Fallbacks

Start with the lesson question, connect the representations, and test the model with evidence.

replanningcancellationretry budgettimeoutfallbackgoal identity

Learning objectives

  • Compare reactive rules, state machines, and behavior trees.
  • Explain graph search, costs, constraints, and collision avoidance.
  • Design fallback and replanning behavior under uncertainty.
Lesson flowHook, model, explanationShow guidance

Inspect the opening phenomenon

Predict what changes, then name the evidence.

Apply in the lab

Name the evidence before reading the answer.

Read only what helps

Then use the lab and recall check.

More when needed

Transcript and resources stay available below.

Course progress

AI & Robotics Foundations · Planning and Robot Behavior · Lesson 24

Replanning, Timeouts, and Fallbacks

In progress

Decision challenge

Observe the phenomenon. Then connect the representations.

Use the opening example to make a prediction, identify evidence, and explain which model supports it.

When Should a Robot Stop Retrying? Replanning Explained

When Should a Robot Stop Retrying? Replanning Explained

When Should a Robot Stop Retrying? Replanning Explained

Reference drawerTranscript, source notes, scripts, and package status stay tucked away until you need them.7 files

Lesson reading

live

55 min

Video script

approved

Transcript fallback

available

courses/ai-robotics/modules/08-planning-and-robot-behavior/lessons/03-replanning-timeouts-and-fallbacks/video-transcript.md

Debug a Replanning Handoff

approved

25 min

Mastery check

live

7 questions / 8 min

Book section:courses/ai-robotics/modules/08-planning-and-robot-behavior/lessons/03-replanning-timeouts-and-fallbacks/book-section.md
Transcript for accessibility and fallback

Descriptive transcript — Replanning, Timeouts, and Fallbacks All visuals below are original simulation drawings. Times refer to the complete local MP4. No essential information requires color perception: labels, positions, gates and narration communicate the same distinctions. 0:00–0:09 · The doorway closes Visual: A top-down delivery robot approaches a doorway. A cart travels left in the upper room, then down into the doorway. The robot decelerates before the stop line; a red X marks its old path as invalid. Spoken: A delivery robot has a route. Then a cart blocks the doorway. Should it retry? First, stop the motion that the old plan allowed. 0:09–0:25 · Cancel is a request Visual: A stationary robot is labeled motion OFF. Requested, accepted and canceled are separate boxes with a moving control signal. The replacement remains locked during the first two states; terminal confirmation does not turn motion back on. Spoken: Inhibit motion and request cancellation. Cancel accepted is not canceled. Wait for the old action's terminal result before sending a replacement. A software acknowledgment does not prove the physical robot has stopped. 0:25–0:39 · Which clock expired? Visual: Three separate clock rings labeled Reply, Progress and Cancel. One ring expires; three possible diagnostic causes remain unresolved, represented by short labels. Spoken: A timeout means a deadline passed, not that the doorway is blocked. Server response, movement progress, and cancellation are different clocks. Preserve the error, timestamps, and action state. 0:39–0:52 · Wait, then re-sense Visual: Robot stationary behind stop line while simulated elapsed time advances 4 to 7 seconds; cart clears doorway. Sensor sweep and checked route appear only after fresh evidence. Spoken: Here is our simulation policy: wait three seconds, then inspect fresh evidence. The cart moves away. Check the route again; elapsed time alone is not permission to move. 0:52–1:07 · Verify before retrying Visual: Three checks unlock one new action. Counter reads attempt 2 of 3 / retry 1 of 2. Robot follows the newly validated path through the clear doorway to goal. Spoken: The old action has ended, the state is fresh, and the route is valid. Now dispatch retry one. Two retries means at most three attempts, including the original. It is a budget, not a target. 1:07–1:23 · Stop the retry loop Visual: Two alternatives are compared: a persistent blockage ends recovery with zero retries, while a separate ledger illustrates an initial attempt plus two justified retries. Stop, evidence and human help form the handoff. Spoken: If the doorway stays blocked, stop early. If justified retries still fail, the budget or mission deadline ends the task. Keep motion inhibited and escalate with evidence. Clearing a map cannot remove a real cart. 1:23–1:36 · Accepted. Move now? Visual: A stationary robot appears below separate cancel-accepted and terminal-result-missing boxes. WAIT or START is the retrieval question. A three-second counter gives time to decide. Spoken: The server accepted cancellation, but no terminal result arrived. The doorway is clear. Should a new motion goal start? Pause and decide. 1:36–1:53 · Wait for the evidence Visual: Three linked gates read INHIBIT, VERIFY and BOUND RETRIES. A cancellation timeout leads to escalation, not another dispatch. Narration links the principle to the failure-case lab. Spoken: No. Keep motion inhibited. A clear route does not settle the old action's status. Escalate if cancellation times out. Test these failure cases in the lesson lab, and follow Humanoid Hub for robotics you can explain. The three-second retrieval pause occurs before the answer. All time budgets are invented simulation values, not robot settings.

Reading lab

Core explanation

Connect the lesson's words, diagrams, graphs, evidence, and equations.

A delivery robot's route was valid a moment ago. Now a cart blocks its doorway. Pressing “try again” is easy; deciding what must be true before another motion goal is harder.

The useful question is what evidence permits the next attempt? Replanning changes a proposed route. It does not, by itself, stop an old action or remove an obstacle.

What you will learn

  1. Distinguish a replan trigger, a recovery action and evidence of readiness.
  2. Interpret a timeout without inventing a diagnosis.
  3. Separate cancellation requested, accepted and completed.
  4. Evaluate a trace using freshness, retry and deadline limits.
  5. Explain when a deliberate stop and human handoff is correct.

Recall 8.1: only an authorized action should control the robot. Recall 8.2: changing costs cannot make an invalid route fit. This lesson connects those ideas over time.

Before, during, and after the video

Watch the complete original Humanoid Hub tutorial. All learning is also available here, in the descriptive transcript, and in the Practice lab.

  • Before: the doorway is clear, but the old goal is still CANCELING. May a replacement start?
  • During: watch motion permission and action status separately. Identify the route check that unlocks the retry.
  • After: explain why the blocked branch uses zero retries although its budget allows two.

Five gates for a replacement goal: old action terminal, fresh valid state, valid route with relevant change, retry budget available, and mission time remaining. A failed gate keeps motion inhibited.

Three different questions

LayerQuestionExample
TriggerWhy reconsider the plan?Changed goal, invalid path, failed progress check, or policy interval
RecoveryWhat can improve the next attempt?Bounded wait and re-sensing for a moving obstruction
EvidenceWhat permits dispatch now?Old goal terminal, fresh state, and checked route

Periodic replanning can be legitimate even without an observed map change. The rule against blind retries here concerns repeated failed attempts, not every nominal planning tick. Nav2's documented tree illustrates event- and time-based replanning. Its numbers are not universal recommendations.

A timeout names a missing event

“Timed out” is incomplete evidence. Ask which event was expected, and by when?

ClockWhat was missingWhat it does not prove
Server responseAcknowledgment within an intervalThat the robot has no route
Movement progressRequired displacement within a windowWhich component caused poor progress
CancellationExpected cancellation progress/resultThat the old action or physical motion ended
Mission deadlineCompletion before the overall budgetThat another attempt is permitted

Nav2 documents distinct acknowledgment and cancellation bounds and a positional progress checker. Preserve the clock, error, goal ID, pose timestamp and last state. A timeout alone cannot diagnose a blocked path, network delay, unavailable server or stale localization.

Cancel accepted is not canceled

A client asks to cancel; the server may accept or reject the request. In ROS 2, CANCELING is active and CANCELED is terminal. Acceptance can precede cleanup and the final result. See the ROS 2 action design.

Our conservative policy inhibits motion, requests cancellation, and waits for the old goal's terminal result before dispatching a replacement. It does not infer completion from elapsed time or an accepted request. Unresolved cancellation leads to escalation, not potentially competing goals.

Even a terminal software result is not a physical stopping guarantee. The lab idealizes motion inhibition as instantaneous. Real machinery needs independently verified stopping behavior, appropriate safety controls and supervision. Nav2 Collision Monitor distinguishes velocity-limiting monitoring from detection-only reporting. Neither this lesson nor its replay establishes physical safety or certification. Never transfer this lesson's values directly to hardware.

Worked example with explicit budgets

Check whose result arrived

A terminal result must belong to the goal being tracked. ROS 2 uses a unique identifier for each goal; the lab uses readable tokens A, R1 and R2 instead of UUIDs. A delayed result tagged A must not erase R1's active-goal record. See the goal-identifier and result-service sections of the ROS 2 action design.

Predict this: A is canceled at t=4, R1 starts at t=7, and a duplicate cancellation result for A arrives at t=8. R1 remains tracked. The event is logged as unrelated; it cannot authorize another retry. The Practice lab also tests wrong IDs, missing IDs and repeated blockage reports. Deadline checks still run even when a message is unrelated.

Trace the main scenario

These are original toy event-replay values, not ROS/Nav2 parameters:

  • One original attempt plus at most two additional retries: maximum three attempts.
  • Mission deadline at simulated t=20 seconds, never reset by retry.
  • Evidence age no greater than 1 second at dispatch.
  • In this fixture, cancellation requested at t=2 must finish before t=5. At the exact deadline, the model stops and escalates.
  • After terminal cancellation at t=4, wait three seconds before observing at t=7.
TimeObservation and actionRetries usedMotion permission
0Original goal A executing on a checked route0On
2Path invalid; inhibit and request cancel A0Off
3Cancellation accepted; A still active0Off
4A reports CANCELED; begin bounded wait0Off
7Cart cleared; fresh valid state and route; dispatch R11On
9R1 reports success1Off

This is simulated time. The video slows and separates events for teaching; video seconds are not these clock values. Waiting offers a chance for conditions to change, but re-sensing must establish that they did.

When stopping is the correct outcome

Change t=7 to “doorway still blocked.” Our policy escalates with zero retries. A retry budget is an upper bound, not a quota. In a different trace, fresh evidence justifies R1 and R2 but both fail: a third replacement is denied. The overall deadline can end the task earlier.

Nav2's RecoveryNode provides bounded recovery control flow. Our readiness checks are an original teaching policy, not guarantees provided by that node. Recovery success permits another primary attempt; it is not mission success.

Clearing stale observations may help a diagnosed data problem. It cannot remove a real cart. Backing up is not a safe default when rear clearance is unknown. Escalate with last error, elapsed time, retries, goal ID and state evidence. Report “stopping unverified” when that is what is known.

Retrieval and transfer

  1. Cancellation accepted, route clear, no terminal result: what is missing?
  2. Retry limit two: how many total attempts are possible? What ends the task earlier?
  3. New timestamp, unchanged blocked route: is freshness sufficient?
  4. Transfer to docking: a station moved. Which observation must be refreshed, and what remains inhibited?

Attempt these before inspecting lab traces. A complete explanation distinguishes action status from physical stopping evidence.

Summary

Inhibit the invalid action, verify the handoff, inspect fresh conditions and bound retries. A timeout is a missed deadline, not a diagnosis. A clear route is necessary but not sufficient. The correct fallback may be to stop, preserve evidence and ask for help.

See the primary-source references in this lesson for source limits and the Practice lab for executable and paper-only alternatives.

Practice labDebug a Replanning HandoffOpen this when you are ready to apply the model, collect evidence, and check your explanation.25 min

Lab — Debug a Replanning Handoff

This is simulation only. No robot, paid API, ROS installation or cloud account is needed. Use a recent Node.js installation supporting ES modules, or the paper trace below. JavaScript here is a list of events, conditions and state updates.

Objective

Diagnose a failed handoff and justify each retry using explicit evidence.

Materials

Predict before running

Predict final state, total attempts and retries for: cart clears, still blocked, cancellation stalls, stale evidence, two failed retries, mission deadline and unchanged conditions. Count the original attempt separately from retries.

Steps

Save the downloadable model as model.mjs and run node model.mjs from its folder. Output includes seven full event traces and regression results.

Record receipt time, event, event goal ID, supervisor state, active goal, motion permission, attempts, retries and reason. The motion flag is an idealized permission gate, not measured velocity. Nothing drives hardware. A, R1 and R2 are readable stand-ins for unique goal identifiers.

The states are our supervisor's states, not a complete ROS action-server implementation. BLOCKED means a path-check result for the named goal; it is not a global collision or emergency-stop signal. This exercise does not implement those independent safety functions.

Investigate four failures

  1. Cancellation stalls: cancel requested t=2, accepted t=3. Clear-route evidence at t=4 must not dispatch another goal. At t=5 the timeout causes STOPPED_ESCALATED. Explain why the old goal can remain unresolved although motion permission is off.
  2. Persistent blockage: terminal cancellation at t=4, fresh observation t=7, route invalid. Explain why zero retries is correct.
  3. Budget: fresh evidence permits replacements at t=7 and t=13; terminal failures occur at t=10 and t=16. At t=19 another replacement is refused. Count three attempts and two retries.
  4. Deadline: clear evidence at t=20 arrives too late. The overall budget wins; do not reset its clock on retry.

Negative control

The blind retry record has old goal A and new goal B active after only cancellation acceptance. It is deliberately invalid. Identify the missing terminal result and explain why a clear doorway does not fix the action-state problem. Never implement it on hardware.

Delayed-message challenge

First predict each outcome without running the code:

  1. A was canceled at t=4; R1 starts at t=7. At t=8, a delayed FAILED message tagged A reaches the supervisor. Should it remove R1?
  2. Cancellation of A is pending. A CANCELED message names another goal. Is the handoff ready?
  3. Blockage reports for A arrive at t=2, t=3 and t=4. Does each report restart the cancellation deadline?
  4. A valid terminal failure for A arrives during cancellation. What old timer must be cleared before a later retry?

Inspect goalEvents.traces in the output after committing your predictions. For each event, compare eventGoal with active, then inspect ignored, cancelStarted and reason. Receipt time advances monotonically; an old goal's message can arrive later without moving the clock backward.

The fixed model keeps R1 tracked when an A message arrives. Wrong or missing IDs cannot release a cancellation gate. Repeated blockage cannot postpone its deadline. A cancellation rejection turns motion permission off and escalates without claiming that A terminated. These are deliberately conservative teaching rules, not a full middleware implementation.

Transfer check: after docking retry R2 starts, a duplicate success notification for R1 arrives. Explain which goal the message describes, what the supervisor should retain, and what evidence would legitimately finish R2.

Change one variable

Copy the model to a separate scratch folder first. Change maxRetries to 0, 1 and 3; predict the total-attempt cap. Test evidence age 1.000 versus 1.001 seconds. The matrix covers four budgets, six ages, state validity, route validity and changed conditions. Equality at one second is allowed in this toy policy, not a hardware recommendation.

Paper-only equivalent

Use cards EXECUTING, CANCEL_REQUESTED, CANCELING, WAITING and STOPPED_ESCALATED. Begin with goal token A and two retry tokens. Replay t=2 blocked, t=3 accepted, t=4 canceled, t=7 observation. Remove A only at its matching terminal result. Spend a retry token only after every readiness check passes and before the deadline. For the stalled-cancel case omit the terminal card. For the delayed-message case, keep R1 on the table while placing the late A card in an unrelated-event pile. Submit the same table in plain text; no color or screenshot required.

Reflection Questions

Submit predictions, traces or card records, the failed guard for each case, and one docking transfer example. Preserve last error, old goal, evidence age, elapsed time and retry count in a handoff packet.

If Node cannot find the file, check your current directory. If changed results surprise you, restore the scratch copy and change one value at a time. A missing terminal result is not fixed by increasing retries. Ask an instructor if the model and source semantics appear inconsistent.

State what this lab does not prove: physical stopping, collision freedom under unmodeled dynamics, real middleware timing or safety certification. Hardware use requires independently reviewed stopping behavior, emergency-stop access, supervision, manufacturer limits and sensor validity.

Expected Result

The cart-clears trace completes with one retry and two attempts. Persistent blockage escalates without spending retries. Cancellation acceptance alone never permits dispatch. The supplied model reports 2454 assertions, including 70 goal-identity regressions. A paper trace should justify the same decisions.

Extension Challenge

Design a docking trace with a delayed old-goal message and fresh but invalid route evidence. Predict which record stays active and explain which gate forbids a retry. Do not connect the model to hardware.