DecisionHotelFlow tutorial
7. Fallbacks and two-tier testing
Give every decision boundary an explicit outage policy, cover it with a deterministic adapter contract, and use the live replay only for real-provider semantic evidence.
DecisionHotelFlow has two sources of probabilistic behavior: the chat model and
the decision provider. Routine tests replace both with controlled adapters. A
separate opt-in suite asks whether the real providers can complete the same
journey.
Three step-local fallbacks
DecisionStep.onDecisionError(context) runs before the Flow-level fallback.
Each decision boundary uses the saved state it owns:
| Step | Fallback | Policy |
|---|---|---|
RouterStep |
show notice, ask first missing criterion, or render a safe summary | degrade without pretending to understand a new arbitrary request |
CriteriaReadinessJudgeStep |
route validation issue or search valid criteria | fail open only for this read-only search |
PresentationJudgeStep |
render saved results in code | fail closed on unreviewed generated prose |
Returning null would delegate to Flow.onDecisionError(context). If both
return null, PicoFlow propagates the original provider error. Caller
cancellation bypasses both hooks.
Configuration and input errors are not provider outages. A missing provider, invalid question map, reserved facts field, or non-JSON fact should fail loudly instead of entering business fallback policy.
Decision retry policy
maxRetries: 2 means at most three attempts. Only failures the provider adapter
marks transient are retried. The TypeSafe adapter treats timeouts, connection
errors, HTTP 408, HTTP 429, and 5xx responses as transient. Invalid answer
envelopes and handler exceptions are not transport retries.
This is separate from chat retryAttempts, where the configured number is the
total attempt budget.
Deterministic contract
Run:
cd pico-demo
npm run test:decision-hotel-flow:contract
The contract creates a real FlowEngine with an in-memory session store, a
scripted chat adapter, and a hand-written DecisionProviderAdapter. The decision
adapter reads the exact facts and question keys it receives, returns schema-valid
Choice, Score, and Noul answers, and throws at selected calls to exercise all
three fallbacks.
The twenty-two-turn scenario covers:
- invalid calendar dates, inverted budgets, and negative distances;
- a correction that crosses from amenities to dates and back;
- explicit criteria review and a readiness rejection;
- a model draft rejected for invented results;
- no matches followed by a budget revision;
- provider outages at router, readiness, and presentation boundaries;
- final grounded presentation and validated booking;
- persisted criteria, judge state, selected hotel, confirmation, usage, and completion.
Stable assertions target completed, the active step, and saved state. Text
checks are limited to the semantic content needed for that turn.
Live semantic evaluation
Run:
npm run test:decision-hotel-flow
With all required credentials, the E2E test boots the demo, uses the real
OpenAI and TypeSafe providers, sends the sixteen scenario requests through one
session, semantically grades each reply, and checks active step plus final state.
Only after every assertion succeeds does it write
test/.tmp/decision-hotel-flow/live.json.
The tutorial’s replay is a snapshot of a successful artifact. A failed or partial run must not replace it. Without credentials the live test is skipped; that is neither success nor evidence of provider behavior.
Separate usage accounting
Decision calls accumulate under sessionDoc.decisionUsage:
{
calls: 22,
inputTokens: 27778,
outputTokens: 2442,
totalTokens: 30220,
}
Chat usage remains under tokens. A provider response is counted when observed,
including one later rejected by framework validation. Failed transport attempts
and late results after cancellation are not visible in the aggregate, so
provider billing may be higher.
The values above came from the recorded live replay. They are not a benchmark: prompts, provider behavior, and conversation paths change the totals.
What each tier proves
| Tier | Proves | Does not prove |
|---|---|---|
| deterministic contract | routing policy, state ownership, validation, fallbacks, and completion on every run | that real models choose those answers |
| live replay | that the configured real providers completed one realistic scenario | general routing accuracy, threshold quality, availability, or cost |
Threshold calibration needs a separate labeled dataset. Production readiness also needs load, latency, provider-failure, security, and idempotency testing beyond this tutorial.
Common mistakes
- Calling a credential-gated skip a live pass. Report it as skipped.
- Mocking only happy paths. Force every decision fallback and every correctable validation branch.
- Returning malformed fake answers. The runner validates test adapters with the same contracts as Jev.
- Testing routes by exact prose. Assert cursor and saved state; reserve semantic wording checks for the live tier.
- Combining chat and decision usage. They are separate runtime mechanisms with separate policy and accounting.
Next
Return to the DecisionHotelFlow overview,
or read the complete DecisionStep reference.