Skip to main content

Run a fleet across hosts, observe it as one

Production fleets don't run on one host. This recipe takes a working fleet-starter and splits it across three hosts talking over Kafka — then uses fleet.yaml#hosts[] to let one operator run declaragent fleet ps / events / dlq / logs once and see every host.

Under the hood this is Slice 3 cross-host fan-out (packages/cli/src/fleet-cross-host-cli.ts) + the CrossHostControlPlaneClient. One bad host is tagged in the output; survivors still return results.

Prerequisites

  • A Kafka cluster (MSK / Confluent Cloud / self-hosted; SASL_SSL supported).
  • Three daemon hosts that can each reach Kafka.
  • Network: each host exposes its control plane on :9464. Bearer tokens must not cross the internet in cleartext — terminate TLS at an ingress or use DECLARAGENT_BIND_ADDRESS=127.0.0.1 + a sidecar.
  • @declaragent/cli ≥ 0.7.4 on each host.

Step 1 — Pick the broker transport

In each agent's rpc-peers.yaml, point the peer at Kafka (peer entries are agent: + transports[]; per-peer auth: declares how that peer authenticates):

# agents/pr-reviewer/rpc-peers.yaml
version: 1
peers:
- agent: agent://concierge
transports:
- kind: kafka
brokers: ["kafka-prod.acme.internal:9093"]
topics:
requests: agents.concierge.requests
auth:
provider: hmac
keyId: acme-rpc-1
secretRef: "secret://platform/agent-rpc-shared-secret"

Broker credentials (SASL/SCRAM + TLS) belong on the agent's own transport declaration in capabilities.yaml, not on the peer entry:

# agents/pr-reviewer/capabilities.yaml (excerpt)
transports:
- kind: kafka
brokers: ["kafka-prod.acme.internal:9093"]
ssl: true
sasl:
mechanism: scram-sha-512
username: "${env:KAFKA_USER}"
passwordRef: "secret://platform/kafka-password"
topics:
requests: agents.pr-reviewer.requests

nats peers follow the same shape (servers + subjects). The sqs/amqp/mqtt/jetstream transports exist as library factories but are not auto-constructed by fleet run yet — see Reference → RPC.

Step 2 — Declare hosts in fleet.yaml

# fleet.yaml
version: 1
name: orders
agents:
- { id: concierge, path: ./agents/concierge }
- { id: pr-reviewer, path: ./agents/pr-reviewer }
- { id: triage, path: ./agents/triage }

hosts:
# host entries accept name / url / auth / timeoutMs only — encode the
# region in the host name (a `region:` key fails schema validation)
- name: prod-us-east-1
url: https://declaragent-use1.acme.internal:9464
auth: { bearer: env:DECLARA_TOKEN_USE1 }
timeoutMs: 5000
- name: prod-us-west-2
url: https://declaragent-usw2.acme.internal:9464
auth: { bearer: env:DECLARA_TOKEN_USW2 }
- name: prod-eu-west-1
url: https://declaragent-euw1.acme.internal:9464
auth: { bearer: env:DECLARA_TOKEN_EUW1 }

Schema constraints (packages/core/src/fleet/manifest-schema.ts):

  • name must be URL-safe (no spaces) and unique.
  • url must be a valid http/https URL.
  • auth.bearer accepts env:FOO — resolved at invocation time, never logged.

Step 3 — Deploy per-host

Each host runs the same artifact (built via GitOps). No cross-host state; Kafka is the source of truth for inter-agent messages.

Step 4 — Observe the whole fleet

From any operator workstation with a checkout of the fleet repo:

# Every running agent across all three hosts
declaragent fleet ps

# Merged audit event stream (JSON envelope is { events, failures })
declaragent fleet events --json | jq '.events[] | select(.kind == "tool_call")'

# Cross-host DLQ snapshot (ingress + dispatch)
declaragent fleet dlq list --kind dispatch

# Tail logs — capped at 50 watchers total, coalesced per-agent
declaragent fleet logs -f

Scope to one host with --host <name>:

declaragent fleet events --host prod-eu-west-1
declaragent fleet logs --host prod-us-east-1 -f

Machine-readable output for dashboards and CI checks:

# each row is { host: <host object>, ok, status | error }
declaragent fleet ps --json | jq '.hosts[] | {host: .host.name, ok: .ok}'

Step 5 — Cross-host DLQ mutations (0.7.5+)

Requeue or drop a stuck dispatch-DLQ row — name the host that holds it (find it in the fleet dlq list output first), or sweep every host with --all-hosts --yes:

declaragent fleet dlq requeue --kind dispatch --id <id> --host prod-us-east-1
declaragent fleet dlq drop --kind dispatch --id <id> --host prod-us-east-1
# or across the whole fleet:
declaragent fleet dlq drop --kind dispatch --id <id> --all-hosts --yes

Both mutations are audited with host = <name> so your SIEM can attribute the operator action (packages/core/src/audit/types.ts:184).

Partial failure — one host unreachable

When a host is down, cross-host verbs degrade gracefully:

$ declaragent fleet ps --json | jq '.failures'
[
{ "host": "prod-eu-west-1", "error": "ETIMEDOUT" }
]

Survivors still return. The exit code is non-zero so CI catches the regression without losing the data from healthy hosts.

Metrics to watch

Wire the Grafana dashboard — row 3 ("Rate limits + dispatch") aggregates across all hosts scraped by Prometheus. Key per-host alerts:

MetricAlert
source_messages_dlq{id="<source-id>"}Rate > 0 for 5 min
source_inflight{id="<source-id>"}Stuck > N for 10 min = consumer stall
Up probe on :9464/healthDown > 2m = host evacuated from hosts[]

Troubleshooting

SymptomCauseFix
no hosts: block in fleet.yaml — use declaragent ps for local viewhosts[] missingAdd the block; cross-host verbs are opt-in
One host times out every calltimeoutMs too tight for a cross-region hopRaise to 15000 for EU↔US
Merged events sorted wrongClock skew between hostsEnforce NTP; each event carries its own ts so ingest-order is deterministic
AUTH_REJECTED on RPC at bootKafka peer missing an auth: blockSee Zero-trust RPC migration