Agent Work Steps
The worker-side AI step types — ai_action for a single inference and ai_agent for an execution-bound agentic loop — that run inside a DAG workflow run.
The chat agent is the interactive face of Ductor's agent loop. This page is the headless one: two step types that run AI inside a DAG workflow run, driven by workers rather than a browser. Same underlying machinery, different host — a workflow step has no operator watching a stream, so it is execution-bound, resumable, and metered as part of the run.
| Step type | What it does |
|---|---|
ai_action | A single LLM inference through the inference proxy, with optional structured-output validation. |
ai_agent | A workflow host for the shared durable agent runtime, with preset resolution, governed tools, approvals, and execution bounds. |
ai_action — one inference
ai_action (executed by AIActionExecutor) is the simple case: it parses the step
args, builds an inference request, calls RouteInference on the AI service, and — when
the step declares an output JSON Schema — validates the response against it before
returning a structured result. Use it when you want the model to transform or classify
some data as one deterministic step in a pipeline, not to run a multi-turn agent.
ai_agent — a durable runtime host
ai_agent is executed by DurableAgentExecutor. It maps a DAG node onto the shared
durable agent runtime: the model proposes tool calls,
the governed broker authorizes and executes them, results are journaled, and the
runtime continues until completion, a durable input request, or an execution bound.
- Preset resolution. The step can reference a tenant
agent preset (by id, optionally version-pinned); the
agent_preset_resolverloads and resolves it before the loop starts. - Governed tool execution. The canonical registry resolves a content-addressed tool definition; the policy receipt and argument hash are bound to the invocation before the connector action runs.
- Resumable approval state. An approval-gated invocation becomes a workflow approval wait. The node persists the agent session, turn, request id, and expected journal sequence. Resume must match all four before the approved invocation runs.
- Execution bounds.
max_tool_calls/max_requestscap the loop, and crossing a bound is recorded distinctly. - Typed external input. Approval, question, budget extension, external-auth, and reconciliation pauses use the same actor-bound, expiring, single-use protocol.
- Bounded delegation. Inline definitions may declare exact/versioned specialist presets, tool intersections, depth/fanout/concurrency ceilings, child budgets, and JSON output schemas. A model supplies only task ids and prompts.
- Scoped memory. A pinned memory policy may recall only verified tenant/subject/entity records. Recalled text is always untrusted user-role context.
How it relates to the rest of the surface
- Versus durable chat. Shared runtime, different host adapter. Chat exposes committed journal events over resumable SSE. A workflow maps the same outcomes to node completion, failure, or an approval wait in durable run state.
- Versus agent presets. An
ai_agentstep resolves a tenant agent preset to get its model, instructions, tool set, and bounds — the same presets a chat session references. The compile-time preset templates are a separate thing: they run as ordinary connector actions, which anai_action-style step invokes directly. - Versus the inference proxy. Both step types issue every LLM call through the one provider-agnostic proxy, so provider choice and cost guards are uniform with the rest of Ductor's AI surface.
Observability
The executors emit dedicated metric families — ai_action_* (request counts, token
totals and per-request distribution, duration) and ai_agent_* (tool-call counts,
LLM-iteration counts, bound-exceeded counts, total duration). Model, provider, and tool
labels are normalized to a closed enum so a misconfigured workflow cannot inflate
metric cardinality.
Inference usage is also written to the durable usage plane with workflow run and step correlation. When a step is part of an experiment, pass the experiment id and immutable exposure id as inference metadata so request, token, cost, latency, and error events can be attributed to the exact variant. See Durable usage and attribution and Experiments.
Enable the durable workflow host
workflow_runtime:
enabled: true
ai_inference:
enabled: true
providers:
primary:
type: gemini
api_key_env: GEMINI_API_KEY
models: [gemini-2.5-flash-lite]
enabled: trueThere is no legacy workflow-agent flag or fallback executor. When workflow inference and connector actions are enabled, startup wires the canonical journal, model-policy, processor, tool-policy, memory, and delegation services; missing required dependencies fail startup rather than selecting another loop.
Specialist definitions are resolved and pinned before parent admission. Child tools are the exact intersection of parent tools, child tools, and the declared allowlist. Children receive only their task prompt and content-addressed lineage—not the parent transcript or memory—and join in stable declared order even when they finish in parallel. Undeclared children, duplicate task ids, recursive depth overflow, schema failure, or aggregate budget escape fail before child admission.
An approval is not a fresh tool proposal
Resume uses the invocation that was already authorized and parked. A stale sequence, wrong request id, different session, or different turn is rejected. The executor never asks the model to recreate an approval-gated mutation after the operator acts.
Related
Agent Presets
Two distinct things called "preset" — compile-time preset templates surfaced as connector actions, and tenant-owned versioned agent presets referenced by chat sessions and workflow steps.
Managing Resources
The Ductor management API surface — how to provision and operate tenants, keys, pools, rules, workflows, experiments, connections, and configuration over REST and Connect-RPC.