No EZGraph key required. Free production use. No runtime fee. Bring your own model provider. Support Jev. Run in your own infrastructure. Optional support is available.

Attachments and testing

Handle temporary file attachments and verify graphs with deterministic tests and opt-in provider evaluation.

File attachments

ToolResponse supports model-visible attachment messages and cleanup. A file tool can remain in the same conversational agent loop while the model analyzes the attachment and submits its result.

return stay(JSON.stringify({ attached: true, fileName: name, fileId: upload.fileId }))
  .withCleanup(upload.cleanup)
  .withMessages([
    new HumanMessage({ content: [
      { type: "text", text: "Analyze the attached file and submit the extraction." },
      upload.contentPart,
    ] }),
  ]);

Test in two tiers

Keep deterministic graph tests separate from opt-in provider evaluation. The example below uses the router from Decision nodes with Jev.

{
  "test:quote-graph": "node --import tsx --test test/quote-graph/*.spec.ts",
  "test2:quote-graph": "USE_ENV=1 KEEP_SESSION=1 node --import tsx --test test/quote-graph/quote-graph.e2e.spec.ts",
  "test:decision-hotel-graph": "node --import tsx --test test/decision-hotel-graph/*.spec.ts",
  "test2:decision-hotel-graph": "USE_ENV=1 KEEP_SESSION=1 node --import tsx --test test/decision-hotel-graph/decision-hotel-graph.e2e.spec.ts"
}

USE_ENV=1 is the single live-provider switch. KEEP_SESSION=1 retains the session only when a replay needs inspection. Assertions about durable state belong in deterministic tests; semantic judges and provider calls remain explicitly opt-in.

The deterministic tier runs real GraphEngine turns against an in-memory session store. Import the helpers from @picoflow/ezgraph/testing. scriptedGateway() dictates chat-model output. scriptedDecisions() queues Jev answers or failures, so no TypeSafe key is needed:

import {
  createTurnHarness,
  scriptedDecisions,
  scriptedGateway,
} from "@picoflow/ezgraph/testing";

const decisions = scriptedDecisions()
  .answers({
    destination: {
      type: "choice",
      choice: "review",
      confidence: 0.97,
      probabilities: { dates: 0.01, budget: 0.01, review: 0.97, search: 0.01, exit: 0, unclear: 0 },
    },
  })
  .fail(new Error("simulated Jev outage")); // exercises onDecisionError()

// HotelGraph registers the RouterDecisionNode from the decision-node example.
const harness = createTurnHarness<HotelStateType>({
  graph: HotelGraph,
  gateway: scriptedGateway(),
  decisions,
});

const review = await harness.send("show my criteria");
assert.equal(review.currentNode, "RouterDecisionNode");

const outage = await harness.send("search");
assert.equal(outage.currentNode, "DateRangeNode"); // deterministic fallback
assert.ok(decisions.drained);

const sent = decisions.calls[0]!.request.state as { request: string };
assert.equal(sent.request, "show my criteria");

Scripted answers pass the same validation as live ones, so every declared label needs a probability. decisions.calls records each exact request, and decisions.drained confirms that every queued answer was used. A test can also pass a hand-written DecisionProviderAdapter that derives answers from request.state, which is how the demo’s 23-turn hotel contract works.

For conversational error policy, script invocation failures with scriptedGateway().fail(...), empty responses with .empty(), and rejected candidates with .text(...) or .callsTool(...). Verify node-to-graph precedence, retry counts, blocked responses, and cancellation. Also test failure after a successful tool: the durable state and tool feedback should survive handled recovery, and completed handlers should execute only once.

For temporary model recovery, assert the gateway’s recorded model sequence: primary, alternate, alternate for a tool follow-up, then primary on a new turn. Verify that alternate calls receive params.retries: 0, completed tools are not replayed, and an alternate failure can return a fixed reply. Include fallback after an empty or rejected primary candidate and cancellation during the alternate call. These are deterministic checks; they do not verify a live provider’s availability or answer quality.

Verify internal work and fan-out

Test the same LlmNode capability in conversational and internal execution. For onEnter(), inspect the gateway’s recorded history: supplied batch input should precede the first model call, replace only the current input, and appear once. Verify pass-through without duplication, empty-history seeding when no message is supplied, and that tool rounds and retries do not call entry again. A new invocation should call it again. Include a graph with no external user message and an internal worker whose batch message remains ephemeral. For nested calls, assert the known caller ID; ordinary graph entry leaves priorNode undefined. Entry-hook failures should propagate without model recovery, and concurrent invocations must keep their input isolated. For asynchronous entry, hold preprocessing pending and assert that neither prompt construction nor model calls have begun. Then release it and verify that facts saved by entry appear in the prompt. Include rejected preprocessing, cancellation during the await, and concurrent invocations that finish in a different order from their start order.

For onExit(), inspect the final discriminated outcome for replies, typed task results, go(), fanout(), directTo(), and finish(). Include both tool-selected outcomes and outcomes returned by onResponse(). Verify that exit sees staged state, follows JSON completion repair, and runs exactly once after retries and tool follow-up. Hold an async exit pending and assert that downstream nodes have not started; after release, their state must include exit writes. Check unhandled failures, exit rejection, cancellation during its await, internal state restrictions, and concurrent shared-node invocations.

Script accepted text, taskResult() output, blocked responses, and exhausted model errors. Assert that onResponse() saves the expected local state in both placements and that the same node-to-graph error-hook precedence applies. Internal results must contain only the worker’s node update and token usage; conversation history, reply, completion, and cursor remain owned by the caller.

For conditional fan-out, verify the full scheduling boundary:

  • Successful movie recording saves parent input and starts each child once.
  • Ordinary replies, invalid movie arguments, and termination start neither child.
  • Delay the workers by different amounts; assert that their calls overlap and the join runs once only after both local outputs are available.
  • Keep the parent as currentNode during worker execution, then assert that the joined conversational stage becomes the next turn owner.
  • Account for the parent, both children, and the join exactly once. Worker text must not leak into conversation history or the user reply.
  • An unrecovered child failure must not run the joined conversational stage. Also verify that internal workers cannot return routing or completion effects.

For nested execution, publish the returned child update in the caller’s result and verify that its state and usage survive. These checks exercise the actual node loop and LangGraph scheduling with scripted model output; they are not evidence of live-provider semantics or long-running transport behavior.

See the QuoteGraph walkthrough for a complete guided application, the DecisionHotelGraph tutorial for decision nodes backed by Jev, and the tutorial for a small runnable graph.