Telecraft

Serve configurations

Serving is the optional third rung: an OpAMP endpoint that hands rendered artefacts from git to connected collectors. It stores nothing durable, so stopping it loses delivery and never the record. GitOps stays an equal alternative, chosen per collector rather than per estate.

This page is the collectors' half of telecraft serve. The same process serves the console and the API on a second address, which is Run an Instance.

Nothing here sits in your telemetry path. If the server is down, collectors keep running the configuration they already have.

This guide assumes you have rendered an estate. The server serves what the renderer wrote, so there is nothing to serve until you have rendered.

Run it against a local checkout

The standalone shape is a single binary plus a directory:

./telecraft serve -estate ../estate-demo -listen 127.0.0.1:4320
console and API on http://127.0.0.1:4321
OpAMP on 127.0.0.1:4320
the session key was drawn at start, so sessions last as long as this process
serve: serving head 870c9b8a26458402c1982359bcdea90fdb7ef73d on 127.0.0.1:4320, fetch interval 30s

The head SHA in that line is the git head of the source the server is reading. Each artefact carries its own stamp, set by the render that produced it, and that stamp is what later joins a collector's reading back to the artefact it is running. The two SHAs differ whenever an earlier commit than head stamped the committed rendered/ tree.

The OpAMP endpoint is at /v1/opamp and speaks WebSocket. A plain HTTP GET gets a refusal:

GET / -> 404
GET /v1/opamp -> 400
serve: Cannot upgrade HTTP connection to WebSocket: websocket: the client is not using the websocket protocol: 'upgrade' token not found in 'Connection' header

Probe the process on the other address instead: /healthz and /readyz are in Run an Instance.

Stop the server with a signal. It exits 0 after a clean shutdown, 1 if it could not start or stop, and 2 on a usage error.

Exactly one flag names the source:

serve: exactly one of -estate or -repo names the source

The estate has to be an estate. The server refuses a directory with no teams/ tree at startup, rather than serving an empty config map:

serve: initial repo snapshot: /tmp/empty has no teams/ tree: the estate layout is teams/<team>/{tiers,services}/<name>.yaml

Run it against a git repository

-repo names a git URL the server fetches and polls. Use this shape when the estate lives somewhere other than the machine the server runs on:

./telecraft serve \
  -repo https://github.com/telecraft-dev/estate-demo.git \
  -cache /tmp/telecraft-estate-cache \
  -listen 127.0.0.1:4322 \
  -http 127.0.0.1:4323 \
  -fetch-interval 10s
console and API on http://127.0.0.1:4323
OpAMP on 127.0.0.1:4322
serve: serving head 870c9b8a26458402c1982359bcdea90fdb7ef73d on 127.0.0.1:4322, fetch interval 10s

-fetch-interval is the only freshness setting, and it bounds how stale a served artefact can be. -cache is a cache of git, not state: losing it costs one re-clone. Omit it and the server uses a fresh temporary directory, which it removes on exit.

A failed fetch does not take the server down. The server logs it and keeps serving the previous head, because a stale head is better than no head.

Which source to pick

-estate -repo
Where the estate lives A checkout on the same machine Any git URL, file:// included
Freshness Whatever wrote the directory, picked up on the poll Bounded by -fetch-interval
Suits Air-gapped and standalone instances Instances beside a hosted estate repository

Both are stateless. Membership, matching, and delivery are pure functions of the head and the collector's reported attributes, so replicas need no coordination and no leader. One Instance is one process, though: the reading of Served collectors is the reading of the connections one process holds, and a second replica would see half the estate.

What collectors need

A served collector runs an OpAMP Supervisor beside it. The Supervisor speaks OpAMP, writes the collector's configuration file, restarts the collector, and reports back what it is running.

For every Tier that declares a serving: block, the renderer writes the Supervisor's configuration for you, beside the collector artefact:

# Generated by the telecraft renderer. Do not edit by hand: change the source in git and render again.
# Supervisor config for Tier data-flow/gateway, commit dea1233454af2285b1c7c2f130d396a6046b9293.
server:
  endpoint: wss://opamp.telecraft.internal/v1/opamp
capabilities:
  # Off upstream by default; on here so the Supervisor accepts served config.
  accepts_remote_config: true
  reports_effective_config: true
  reports_health: true
agent:
  # Off upstream by default; on here so a failed config reverts on its own.
  automatic_config_rollback: true
storage:
  # Mount a durable volume here: an ephemeral directory gives the
  # collector a new identity every time the pod is replaced.
  directory: /var/lib/telecraft/supervisor

Four of those settings matter, and three of them are not upstream defaults:

  • accepts_remote_config lets the Supervisor take a served configuration at all.
  • reports_effective_config is what makes the Effective reading exist. Without it, Telecraft can see what it sent and never what is running.
  • automatic_config_rollback reverts a configuration the collector fails to start on. That turns a bad artefact into a report rather than an outage.
  • storage.directory must be a durable volume. An ephemeral directory mints a new identity on every pod replacement, and the population count churns.

That file is not runnable as delivered. It says where the server is and how the Supervisor behaves, and nothing about which collector this is or which binary to run. Both are yours to supply, and installing a served collector shows where they go.

A Tier with no serving: block renders no supervisor artefact. Those collectors are git-delivered, which Telecraft calls the Foreign path. It is a legitimate path, not a lesser one, and each collector shows which path it uses.

How a collector is matched

You never author a collector. It connects, reports its identifying attributes, and the server matches those attributes against the Tier selectors at head:

selector:
  telecraft.tier: gateway
  deployment.environment: production

Matching is equality over every authored pair. The most specific satisfied selector, meaning the one with the most pairs, wins. An equal-specificity tie resolves to the first Tier in id order, so replicas can't disagree.

The attributes on the other side of that comparison come from the collector's installation. Nothing in git supplies them, which is what installing a served collector covers.

The Unmatched artefact

A collector that matches no selector is not ignored, and it is not sent an empty config map. It gets the Unmatched artefact: a real, commit-stamped configuration with self-telemetry on and no data pipelines.

# Generated by the telecraft renderer. Do not edit by hand: change the source in git and render again.
# The Unmatched artefact, commit dea1233454af2285b1c7c2f130d396a6046b9293: served to a collector
# that matches no Tier selector. Self-telemetry only, no data
# pipelines. This is not the quarantine destination, which routes
# data rather than collectors.
service:
  telemetry:
    resource:
      k8s.node.name: ${env:TELECRAFT_NODE_NAME}
      telecraft.commit: dea1233454af2285b1c7c2f130d396a6046b9293
      # Marks this collector as matching no Tier, so the console can
      # show it and offer to onboard it.
      telecraft.unmatched: true
    metrics:
      level: normal
      readers:
        - periodic:
            exporter:
              otlp:
                endpoint: https://otlp.observability.internal:4318
                protocol: http/protobuf

Generated comments and the matching logs block are cut for length. The telecraft.unmatched: true label is the point. An ungoverned collector served this artefact reports itself, so it shows up in the console as governed by nobody rather than staying silent. The artefact lives at rendered/_estate/unmatched.yaml, and the root Team owns it.

Ungoverned collectors count against nobody: no compliance ratio counts them. They do appear in the estate view, which is a different thing.

Install a served collector

The rendered Supervisor artefact is not a finished configuration. Start a Supervisor on it unchanged and you get a collector the server can't place. The file carries what the estate decides: where the server is, which capabilities to turn on, and where to keep state. It carries nothing about which collector this is, and nothing about which binary to run. You supply both at install time, by merging an overlay over the rendered file.

Two keys are missing, for two different reasons:

  • agent.description.identifying_attributes is the identity the collector reports, and the only thing the server matches on. Without it, the collector reports nothing a selector can satisfy, so the server serves it the Unmatched artefact. That is the server behaving correctly, and it is not what you wanted.
  • agent.executable is the collector binary the Supervisor starts as a child. It belongs to the image or the package you installed, not to the estate, so nothing in git knows it.

The overlay, and the selector it satisfies

These two files are two halves of one sentence. The Tier authors the first half:

# teams/data-flow/tiers/gateway.yaml, authored
selector:
  telecraft.tier: gateway
  deployment.environment: production

The install supplies the second half, merged over rendered/data-flow/gateway.supervisor.yaml:

agent:
  # The collector binary the Supervisor forks: the package path on a VM,
  # the image path in a container.
  executable: /usr/bin/otelcol-contrib
  description:
    identifying_attributes:
      # Every pair the selector names, spelled as the selector spells it.
      # Matching is equality, so a typo here is an Unmatched collector.
      telecraft.tier: gateway
      deployment.environment: production
      # Unique per collector, and not part of matching. It is how you tell
      # two collectors in the same Tier apart.
      service.instance.id: gateway-1

Report more pairs than the selector names and matching still succeeds: a selector is equality over the pairs it authors, and says nothing about the rest. Report fewer and matching fails. So a collector meant for a Tier carries every pair in that Tier's selector, plus whatever else you want to see.

Merge the overlay over the rendered file as a deep merge: maps merge key by key, and everything else replaces. Keep the two halves in separate files. The base is reproducible from the estate at a commit, the overlay is yours, and a single hand-edited file loses that boundary.

The local development environment runs this pattern for real. devenv/identity/ holds one overlay per collector, each naming the Tier whose artefact it starts from, and telecraft-devenv prepare merges each over that artefact before the Supervisors start.

Where the values come from

The pairs the selector names are the same for every collector in the Tier, so one file serves the whole workload. The values that are unique per collector come from wherever the collector runs, and Kubernetes and systemd hand them over differently.

Kubernetes

Nothing upstream deploys a supervised collector. Neither the OpenTelemetry Operator nor the collector Helm chart knows the Supervisor exists, and a sidecar container doesn't work either, because the Supervisor forks the collector and signals it directly. The Supervisor and the collector go in one container, in an image you build, with the Supervisor as the entry point. So the workload is yours to write, and that is where the values come from:

# The gateway Tier. A StatefulSet rather than a Deployment because the
# Supervisor keeps its identity on disk, so each replica needs its own
# durable volume and a stable name.
apiVersion: apps/v1
kind: StatefulSet
spec:
  template:
    spec:
      containers:
        - name: collector
          image: registry.internal/telecraft/supervised-collector:0.159.0
          env:
            # The rendered collector artefact reads this as
            # ${env:TELECRAFT_NODE_NAME}. The Downward API is what makes one
            # manifest yield per-node identity.
            - name: TELECRAFT_NODE_NAME
              valueFrom:
                fieldRef:
                  fieldPath: spec.nodeName
          volumeMounts:
            - name: supervisor-config
              mountPath: /etc/telecraft
              readOnly: true
            - name: supervisor-storage
              mountPath: /var/lib/telecraft/supervisor
      volumes:
        - name: supervisor-config
          configMap:
            # The rendered artefact with your overlay merged over it.
            name: gateway-supervisor
  volumeClaimTemplates:
    - metadata:
        name: supervisor-storage
      spec:
        accessModes: [ReadWriteOnce]
        resources:
          requests:
            storage: 1Gi

The collector reads TELECRAFT_NODE_NAME, not the Supervisor. It lands in the collector artefact's self-telemetry resource as k8s.node.name, which is a reading rather than an identity, and it takes no part in matching. The renderer emits the indirection so that one manifest yields per-node identity across a DaemonSet; feeding the variable from spec.nodeName is the half you own.

The Supervisor reads one file, so a value that differs per collector has to differ in that file. Either template the file per replica, or leave service.instance.id out and let the identity the Supervisor persists be the unique one. Nothing in matching depends on either choice.

systemd

On a VM the Supervisor has real packaging. The opampsupervisor deb and rpm install a unit, a system user, an EnvironmentFile at /etc/opampsupervisor/opampsupervisor.conf, and an example configuration. The unit's StateDirectory=opampsupervisor gives you a durable /var/lib/opampsupervisor for free. Three steps turn that into a served collector:

  1. Install the collector and the Supervisor as separate packages, then disable the collector's own unit. Under the Supervisor, the collector runs as a child on a configuration file the Supervisor generates, so leaving otelcol-contrib.service enabled gets you two collectors racing for the same ports. Nothing upstream disables it, and nothing warns you.
  2. Write the merged configuration to /etc/opampsupervisor/config.yaml, with agent.executable pointing at the collector binary the package installed. The unit asserts that the path exists and doesn't start until it does, so a fresh install is enabled and inert until you write the file.
  3. Set TELECRAFT_NODE_NAME in the EnvironmentFile. There is no Downward API here, and the collector inherits the Supervisor's environment, so this is where the value the artefact reads comes from.

Give the storage directory a durable volume

The Supervisor mints a UUID on first run and keeps it under storage.directory. That UUID is the collector's identity on the wire, and it is meant to survive restarts. Point the directory at something ephemeral and every restart mints a new one: the server reads an arrival rather than a return, and the Tier's population churns while the collectors under it sit still.

The rendered artefact already sets the path. Making the path durable is your job:

  • A packaged VM install has it right by default, because systemd owns /var/lib/opampsupervisor. Either create the rendered path and give it to the Supervisor's user, or point storage.directory at systemd's directory in your overlay.
  • In Kubernetes, an emptyDir is the wrong answer, and it is the answer you get by not deciding. Use a volume claim per replica, or a hostPath where identity should last for the node's lifetime rather than the pod's.
  • Whichever volume it is, the user the Supervisor runs as has to be able to write to it. A fresh volume owned by root under an unprivileged container is the same failure with a different message.

Why the renderer does not fill this in

The renderer knows the identifying attributes: they are the Tier's selector, authored a few lines away. It does not write them into the artefact, because an artefact that carries the attributes its own selector matches is self-matching: matching would confirm the renderer's output instead of reading the collector. Identity stays reported, so a collector in the wrong place stays visible.

Deploy the live-check service

A Requirement with placement: live is judged against the findings a weaver registry live-check service emits. That service is upstream Weaver's, and it is yours to deploy, the way the collectors and the Supervisor are: Telecraft renders the pipeline that feeds it and reads the findings that come home, and ships no binary of its own into your telemetry path.

Three pieces have to line up, and each lives where you already manage its kind:

  1. The feed. A Tier opts in with a live_check block, and the renderer adds a sampled pipeline beside its lanes, exporting over gRPC to the endpoint live-check.yaml declares. Tiers covers the block and Estate layout the file. The Tier's own data pipelines are unchanged: if the service is down, you lose findings, never data.
  2. The service. Run weaver registry live-check wherever you run workloads, from the image or binary the Weaver project publishes. Point it at the same registry source your estate imports as its Schema Registry, and have it consume OTLP over gRPC on the port the live-check.yaml endpoint names. The command and its flags are upstream's, documented by weaver registry live-check --help.
  3. The way home. Start the service with --emit-otlp-logs, sending its findings to the same backend the rest of your telemetry lands in. The platform reads them back from there like any other telemetry, so a finding the backend never receives is a finding the evaluation never sees.

The service holds no state worth keeping: replacing it costs the findings it would have emitted while down, and the Requirements that read them show unknown for the window rather than a pass. See Placement for how that verdict is drawn.

Compare what was sent with what is running

telecraft delivery crosses the Intended reading (the artefact in git) with the Effective reading (what the collector reports), through the normaliser. The same computation runs for both delivery paths, so a git-delivered collector gets the same surface as a served one.

./telecraft delivery \
  -intended ../estate-demo/rendered/edge-ops/edge.yaml \
  -effective /tmp/reported.yaml \
  -path git

When they agree:

path              git
profile           exact
remote            known=false cause="a file comparison carries no RemoteConfigStatus reading: the OpAMP server reads it live"
intended_commit   dea1233454af2285b1c7c2f130d396a6046b9293
effective_commit  dea1233454af2285b1c7c2f130d396a6046b9293
comparison        in_sync

When they don't:

path              git
profile           exact
remote            known=false cause="a file comparison carries no RemoteConfigStatus reading: the OpAMP server reads it live"
intended_commit   dea1233454af2285b1c7c2f130d396a6046b9293
effective_commit  dea1233454af2285b1c7c2f130d396a6046b9293
comparison        drifted
changes:
  processors.batch/batcher.send_batch_size: 4096 -> 512

-path names the collector's delivery path and selects the mutation profile the comparison runs under. git compares exactly. served allows for the Supervisor's own injections, so a Supervisor-managed collector is not reported as drifted for doing its job.

A file comparison carries no delivery status, so the remote axis prints known=false with its cause. The server reads that axis live. Like observe, delivery prints rather than gates: it exits 0 for every computed status, including drifted, and exit 2 means the status couldn't be computed at all.

What next

  • Stage a Rollout moves one Tier's population onto a new Blueprint version in cohorts.
  • Check conformance answers the question delivery can't: whether the configuration that arrived worked.