Clearing Surfaces
What each clearing-layer gate guarantees, what it depends on, what fails closed, and the blast radius of turning it on — pricing, spend caps, assurance, settlement, receipts, and agent RFQ awards.
Every surface below is off by default except one. This page states what each guarantees, what it needs to run, and what happens when that dependency is gone. Read it before enabling anything; the rollout order is in the configuration reference (Clearing surfaces).
The bias throughout is fail closed: a plane that cannot verify refuses rather than assumes. The two places that bias is deliberately inverted — and the two places a guarantee is weaker than it looks — are called out as warnings, not footnotes.
Config gates
Every gate, its default, and what you are changing.
| Gate | Default | Enabling does | Disabling restores | Blast radius if wrong |
|---|---|---|---|---|
connector_shared_executor.enabled | true | Routes notification dispatch, AI-agent tool calls, and DAG StepWorker connector calls through the shared executor, which carries execution admission | The prior local executors, which bypass admission entirely | Highest. Enforces quota/budget/entitlement on three paths that never had it. Most refusals are non-retryable |
connector_pricing.enabled | false | Builds the price resolver; connector actions become billable for tenants with rows in connector_action_pricing | No resolver is built; every action resolves UNPRICED and nothing bills | Billing starts. Also enables agent spend caps, which have no separate flag |
connector_assurance.enabled | false | Resolves the certified/community trust axis and projects it on catalog reads | Nothing is resolved; every action projects community | Low. Conservative in both states |
connector_assurance.require_certified_for_agents | false | Agent-initiated connector actions must resolve certified or are refused | Agent callers admitted exactly as today | Refuses agent actions on uncertified inventory. Requires enabled: true |
artifact_pricing.strategy.enabled | false | Resolves a strategy's price at the promotion/adoption gate | Nothing resolves; no queries issued | A resolution failure refuses promotion |
artifact_pricing.skill.enabled | false | Resolves a skill's price on every skill read | Nothing resolves; no queries issued | Adds a resolve to the skill read path |
settlement.journal_enabled | false | Every pricing charge co-writes a balanced double-entry posting | Charges post as before; no journal rows written | A posting that cannot balance is refused, not recorded |
work_receipt.enabled | false | Receipts are issued and signed | Receipts are not issued or signed — but the read/verify API and JWKS still serve configured public keys | Needs a signing key; see below |
agent_rfq.enabled | false | Task-spec publication and disclosure-scoped RFQ discovery | No task-spec store constructed; routing markets, ping/post, and recipient bidding unaffected | Low on its own |
agent_rfq.bidder_enabled | false | In-cluster AI agents spend model budget evaluating RFQs and bid autonomously | Discovery runs for external workers with no agent spending | Agents spend model budget. Requires a budget ceiling |
agent_rfq.award_execution_enabled | false | Launches awarded work, observes outcome, captures or releases money, issues the receipt | Awards clear but nothing executes or settles | This is the gate that moves money. Requires enabled |
connector_shared_executor is the only gate defaulting on, because the behavior
it enables — enforcing admission — is correct and the paths it covers should
never have bypassed it. It exists as a switch so a deployment can revert without
rolling back the binary.
Which listener serves them
The server opens one request listener, api.addr (default :50052). REST,
Connect, MCP, the RFQ front door (/api/agent-rfqs), work receipts
(/api/work-receipts), the JWKS document, and /health are all served there.
server.http_addr (default :8080) opens no listener; it is a legacy key, and
the server warns at boot when it disagrees with api.addr. The compose files
set DUCTOR_API_ADDR=:8080 so the published port and the listener coincide.
When /api/_meta/services reports a clearing service as
dependencies_unavailable, the gate for that service is off in the config of
the server you reached, not a second listener. Read /api/_meta/release to
confirm which build answered before pointing agents at a deployment.
Dependencies and what fails when they are gone
| Surface | Needs | If the dependency is unavailable |
|---|---|---|
| Action pricing | Postgres | Execution is refused, not run unpriced. Deliberate — see below |
| Agent spend caps | Redis + pricing enabled | Cap state is lost; see the durability warning |
| Assurance | Postgres | Resolves community; never certified |
| Settlement journal | Postgres | Postings refused; charges do not silently skip the journal |
| Work receipts | Postgres + an Ed25519 signing key | Without a key, receipts are not signed |
| RFQ / awards | Postgres | Discovery and award execution stop; committed money is unaffected |
A pricing outage refuses execution rather than running for free
When connector_pricing.enabled is true and the price cannot be resolved, the
action is aborted before any provider call — it does not fall through as
unpriced. A database outage must not silently make paid actions free. The
refusal is recorded on the execution log with tenant, provider, and action, so
a pricing outage is visible where you already look. If you would rather
degrade than refuse, the switch is connector_pricing.enabled: false, not a
retry.
Admission enforcement is a live behavior change
connector_shared_executor.enabled defaults on, and three paths that previously
bypassed execution admission now enforce quota, budget, and entitlement:
- notification dispatch,
- AI-agent tool calls,
- workflow-step connector actions.
Only queued, deferred, and blocked_by_capacity are retryable.
blocked_by_quota, blocked_by_budget, rejected, and
blocked_by_entitlement are not. A tenant whose entitlement was never
exercised on these paths — because these paths never checked — will see
permanent failures, not transient ones.
Verify entitlement coverage before first deploy
The failure mode is a workflow step or notification that used to succeed and
now fails and stays failed. Confirm every tenant has entitlement covering
connector actions on these three paths, or set
connector_shared_executor.enabled: false and enable it per cohort. Disabling
restores the previous behavior without a binary rollback.
Agent spend caps
A tool exposure rule declaring max_cost_cents draws its connector-action spend
against that ceiling: price resolved before any provider call, reserved against
the cap, committed at actual cost. Over-cap claims are refused before the
provider is reached. Enforcement follows connector_pricing.enabled — the price
meter is the pricing resolver, so a deployment without pricing has no spend a
cap could bound, and there a capped rule is refused rather than run uncapped.
Rounding is up, against the spender, at a single cents↔micros boundary (10,000 micros per cent), so sub-cent actions cannot draw indefinitely against a ceiling that never moves.
Redis is not a durable store for these caps
Committed spend is held in Redis, keyed to the agent session, on a 24-hour
window. A flush, a failover to an empty replica, or key eviction resets a
cap and lets an agent spend again.
Redis must not run an eviction policy that can evict these keys — under
allkeys-lru every agent is silently uncapped. Use noeviction, or a policy
that cannot reach them.
This is acceptable for a throttle: nothing is owed on these numbers, and
actual billing flows through the priced usage event and the durable settlement
journal. It would not be acceptable for an accounting record, and it is not
one.
Recorded spend can under-report
The spend lease has a 5-minute TTL as a crash-recovery backstop. There is no
default action timeout, and max_duration_millis is optional, so an action can
outlive its lease. When it does the reservation has already been swept back
into the budget, another claim may have been admitted against the same
headroom, and the commit is pinned to the ceiling — the money was already
spent by the time anything could refuse it.
The cap is never exceeded in the recorded total, but that total can
under-report true spend. It surfaces as an agent_tool_spend_lease_expired
log event with tenant, session, and tool. There is no metric for it yet; see
Observability.
A cap lowered mid-session is not retroactive. A lease admitted against the old ceiling commits against the old ceiling; new claims use the new one.
Pricing: catalog and billing can disagree
Catalog reads resolve tenant-wide price rows only — there is no environment
on the request context, so the projector passes an empty environment. Execution
resolves with the real EnvironmentID.
An environment-specific override is billed but not displayed
A tenant with an environment-scoped price sees the tenant-wide price in the catalog and is billed the environment-specific one. Billing is correct; the display is incomplete. If you use environment overrides, do not treat the catalog as a quote.
Unpriced is not free. An action, strategy, or skill with no price row
resolves unpriced and does not bill — that is an unknown commercial status,
not a zero price. The wire projection carries an explicit unpriced basis
rather than omitting the field, because an omitted price reads as zero and zero
reads as free.
Assurance never resolves upward
Undeclared, unreadable, expired, and revoked all yield community. Nothing
reaches certified without a valid, unexpired, unrevoked declaration — a store
error included. Expiry and revocation apply at read time, so a lapsed
certification downgrades the instant it lapses without waiting for a sweeper,
and the row stays readable so an auditor can still see the evidence and
attribution behind the original claim.
Assurance is orthogonal to price: a free action can be certified and a premium
one community-maintained. ActionPricing.tier (free/included/standard/premium/
metered) is a commercial tier and is not the assurance axis.
Receipts and settlement
Receipts are signed with Ed25519 as ed25519:<key_id>:<signature>, and
public keys are published at /.well-known/work-receipts/jwks.json so a
counterparty verifies offline, without calling Ductor. Keys carry
not_before / not_after windows plus disabled and revoked flags, so
rotation is additive: publish the new key, move active_key_id, and keep the
old public key served until every receipt signed under it is out of its
retention window. The signing private key is a secret — inject it via
DUCTOR_WORK_RECEIPT_SIGNING_PRIVATE_KEY_PEM, never in config.
With work_receipt.enabled: false, receipts are not issued or signed, but the
read/verify API and JWKS still serve whatever public keys are configured — so
previously-issued receipts remain verifiable after you turn issuance off.
Settlement postings are double-entry and must sum to zero per currency. That
is enforced by a DEFERRABLE INITIALLY DEFERRED database constraint checked at
COMMIT, so a whole balanced set is validated together and an unbalanced
transaction is rejected by Postgres rather than by application code. Money is
int64 micros end to end. A claw-back is a compensating posting; a posted
transaction is never rewritten.
Award execution
A committed award launches its work under the exact agent definition the winning bid pinned. If that definition has drifted, the award is voided rather than executed against a substituted agent. Success captures the held funds, failure releases the hold, and either way a signed receipt records the outcome.
Award status transitions are enforced in SQL, with legal predecessors derived
from the same transition table the in-process validator reads, so the storage
rule and the domain rule cannot drift. Without it, an illegal awarded → settled
move would capture money for work that never ran.
The tool-manifest pin is not an independent check
The award records both an agent-definition pin and a tool-manifest pin, but the manifest hash is derived from the definition hash. If the definition matches, the manifest necessarily matches — so the manifest comparison restates the definition comparison rather than adding coverage. A tool-manifest change that leaves the agent definition unchanged is not detected at award time. The real manifest is built from the tools a specific agent session requests, and no session exists when an award is made, so there is nothing authentic to reconcile against at that point. Manifest-level drift detection belongs at session admission.
A run that has been purged resolves after a grace window (default 24h) by releasing the hold, never capturing — without a run there is no evidence of success.
Failure modes
| Scenario | Behavior |
|---|---|
| Redis down, pricing on | Spend caps cannot reserve. A capped rule is refused rather than run uncapped. Uncapped rules are unaffected |
| Redis flushed | Committed spend is lost; caps reset. See the durability warning above |
| Postgres down, settlement on | Postings fail. Charges are refused rather than silently skipping the journal |
| Postgres down, pricing on | Execution is refused before any provider call, logged on the execution log |
| Signing key unavailable | Receipts are not signed. Verification of previously-signed receipts is unaffected — it needs only the published public key |
| Run purged mid-award | After the grace window the hold is released, never captured. The award reaches a terminal state rather than sweeping forever |
| Definition drifts after a bid wins | Award voided; nothing launches, no money moves |
| Unbalanced settlement posting | Rejected by the database at COMMIT |
Where to go next
Certified Promotions & Rollout
The governed release control plane for config and strategy artifacts — promotion candidates, rollout-evidence packs, and runtime-conformance gates.
Queue operations
The operator runbook for the tiered fair queue and its dead-letter queue — pause, resume, drain, admission control, and DLQ replay.