Telecraft
Compose it. Apply it. Validate it.
Telecraft is how a platform team runs OpenTelemetry for everyone else. Teams build their own collector configuration from a library of parts you have already approved, Telecraft applies it to their collectors, and then checks what actually arrived in your backend.
See it running Get a verdict in five minutes
Early software Composing, applying and validating all work end to end, with a console over all four workspaces. Interfaces are still free to change.
-
Intended
traces, logs, metrics
-
Effective
traces, logs, metrics
-
Observed
traces, logs
broken_pipeline Metrics were configured here and none arrived.
Owner platform/ingest
Compose it
Parts are approved once, centrally. After that a team composes what it needs without asking anybody, because everything it can reach has already been approved.
A library teams can help themselves to
-
Catalogue
Every collector component that exists in the release you run, with its stability for each signal. Telecraft generates it from upstream and keeps one per release.
-
Allow-list
The part of the Catalogue a Team may use. A child Team can narrow the list it inherits, and never widen it.
-
Component
A configured, named, versioned part with an Owner: a receiver, processor, exporter, connector or extension. Shared across Teams, or local to one Blueprint.
-
Blueprint
A composition of Components, one ordered lane per signal. A Tier binds exactly one Blueprint version, so what a collector runs is always a version somebody chose.
-
Endorsement
The organisation standing behind a Blueprint at a named version, across the whole Estate.
A Blueprint names the Requirements it means to satisfy, and Telecraft checks that claim rather than trusting it. A team that composed from the library is compliant because of what it built, not because somebody reviewed it afterwards.
traces otlp · k8sattributes · batch · otlphttp
logs otlp · k8sattributes · batch · otlphttp
metrics otlp · k8sattributes · batch · otlphttp
satisfies:
traces-to-primary
resource-attributes-present
no-unbatched-export
Endorsed by platform, and bound to Tier C1.
Apply it
The renderer compiles a Blueprint to plain OpenTelemetry Collector YAML and opens it in git as a change proposal. Review it, approve it and merge it with the rules your repository already has: history, rollback, approval and audit are git’s job, and Telecraft does not replace any of them.
Two ways to reach a collector, chosen one collector at a time
-
Telecraft serves it
A stateless OpAMP server hands the merged configuration to an OpAMP Supervisor beside each collector. A Rollout moves a Cohort at a time, and each collector reports back what it is actually running.
Costs you a Supervisor beside each served collector
-
Your GitOps serves it
The rendered YAML sits in git and your existing delivery pipeline applies it, exactly as it applies everything else. Telecraft is not in the loop at all.
Costs you nothing you are not already running
Neither route puts Telecraft in the telemetry path. It is never a collector, a gateway or a hop, so if it stops, nothing stops flowing.
rendered/production/c1/collector.yaml
processors:
batch:
timeout: 5s
+ k8sattributes:
+ passthrough: false
service:
pipelines:
metrics:
- processors: [batch]
+ processors: [k8sattributes, batch]
Merged, then rolled out one Cohort at a time.
Validate it
Telecraft reads the estate three ways. Each reading has one definition and one source; they compose, and they are never blended.
-
Intended
The configuration in git, pinned to a commit rather than a branch tip.
-
Effective
The collector’s own report of what it is running. Never what an applier holds, and never what the platform believes it sent.
-
Observed
The telemetry that landed in your backend over a window, judged against the window each Requirement asked for.
Every peak is telemetry arriving and the flat line is the window with nothing on it. Where Intended and Effective have an arrival and Observed does not, nothing reached the backend.
Both of these Services report zero metrics
They are not the same problem, and they do not go to the same team. A dashboard scores both as zero. Crossing what a collector reports with what arrived is what sends the first to the platform team as a defect, and the second to the team that owns the Service as work they have not done yet.
Every reading also carries a Known flag, so a provider that cannot answer says so, with a cause, and no evidence is never quietly treated as a pass. The full set of outcomes is in the documentation.
Composing, applying and validating are three capabilities over one model. Adopt any one of them without the others, in any order; using one is never a condition of using another.
configured yes
arriving no
Somebody configured this and it is not working.
Owner platform/ingest
configured no
arriving no
Nobody configured it at all.
Owner team/basket
One binary. A command, or a platform.
There is nothing else to install. The same artefact answers a question in your pipeline and serves the console your teams sign in to, over the same estate, producing the same verdicts.
$ telecraft check \
-library ./requirements \
-estate ./rows.yaml
rows 5
failing_rows 4
counting_failures 4
waived 0
library_drift 0
$ echo $?
1
One JSON report on standard output, and an exit code taken from the result. It runs in CI, in a cron job, or on a laptop.
Costs you a connection string
-
checkout-api
traces
logs
metrics
-
basket-svc
traces
logs
metrics
The console, the platform API and the OpAMP endpoint, from one process over one estate. Sign-in, health probes, and no state that outlives a restart.
Costs you one process, and a Supervisor beside each served collector
Run it yourself, or let us run it
The same release either way. Nothing sits in the telemetry path in either shape, and your estate is a git repository you own in both.
Cloud
One Organisation is one Instance at an address of its own, your-name.cloud.telecraft.dev, with its own estate and its own people. It runs Standard Edition, which is the whole product: there is no capability here that a deployment of the same release on your own hardware does not have. What you are buying is the running of it.
Your infrastructure
The telemetry path
Telecraft is not on this line. If it stops, none of this stops.
cloud.telecraft.dev
Reads the estate, reads your backend, reads what your collectors report.
- Who runs it We do. The address, the certificate, the upgrades, the backups, and somebody whose job it is to notice.
- Signing in Google or Microsoft Entra ID to begin with, and your own provider added by a change to your estate.
- Getting one Sign-up is a request, not a form that provisions. A person reads it and merges it, so you wait for us rather than for a machine.
- Serving Not offered here yet. Hosted Organisations read, judge and deliver through git, and the OpAMP endpoint is not published on the internet.
- Leaving git clone. There is no export format, because the estate is the whole of your authored work and a clone is a complete copy of it.
- Costs you Nothing during the beta. There is no subscription, no plan and no price, and you will be told before that changes.
Self-managed
One container image, one Helm chart, or one binary you build. It needs a git repository and a connection to your telemetry backend, and it needs nothing that we operate. An air-gapped deployment is a first-class shape: nothing phones home, and no start-up path reaches the internet.
Your infrastructure
The telemetry path
Telecraft is not on this line. If it stops, none of this stops.
One process, on your hardware, inside your boundary.
$ helm install telecraft oci://ghcr.io/telecraft-dev/charts/telecraft \
--namespace telecraft --create-namespace \
--set estate.sync.repo=https://forge.example/acme/estate.git
- Who runs it You do, on your own hardware, at your own address.
- Getting it The chart above, the image at ghcr.io/telecraft-dev/telecraft, or a binary you build from source.
- Signing in Basic auth to bootstrap, then any OIDC provider you author into the estate.
- Serving All of it. The OpAMP endpoint is yours, beside the console, on a port you choose.
- Air gaps First class. There is no licence server to reach and no telemetry sent back.
- Costs you Nothing. Standard Edition is unrestricted in production and commercially.
Deploy on Kubernetes, run the container image, or run an Instance
Five things that stay true
These hold across every part of the product, in every deployment shape.
- Nothing sits in the telemetry path. If Telecraft is down, no telemetry stops flowing. It is never a collector, a gateway, or a hop.
- Git is the source of truth. History, rollback, approval and audit come from git. The console opens change proposals; it never writes to a cluster.
- Configurations, never binaries. No collector distribution, no container image of your collectors, no chart. Telecraft ships configuration to collectors you already run.
- The core is vendor-neutral. Backends and fleet managers sit behind seams as plugins, and a lint keeps vendor names out of the core, so independence stays greppable.
- Air gaps come first. Nothing depends on a hosted service or on a particular git host.
See it, then run it
The live demo
The real console over a public demo estate: read-only, no sign-in, and built from the latest release rather than from main, so it shows what you can run today.
A first verdict
One downloaded file and git. No toolchain, and nothing to compile. Take the CLI for your machine from the latest release, clone the public demo estate beside it, then:
./telecraft check \
-library estate-demo/requirements \
-estate estate-demo/demo/rows.yaml \
-exemptions estate-demo/exemptions
The quickstart walks the rest of it, and the documentation covers everything else.