Telecraft

Continuous integration

Four workflows live in .github/workflows/. One guards changes, one publishes releases, and two tell downstream repositories that something they build from has moved. Static analysis runs beside them as GitHub code scanning's default setup, configured in the repository settings rather than in a workflow file, and reports its findings on the same pull requests.

Workflow Fires on Does
ci.yml every pull request, every push to main, and a weekly schedule Builds, tests and lints: the checks review waits for
release.yml a v* tag Publishes the release: the image, the chart, the CLI for every platform, and the design artefacts
demo-dispatch.yml a successful release, or a manual run Moves the release pointer and asks the demo to rebuild
docs-dispatch.yml a push to main touching docs/** Asks telecraft.dev to rebuild the documentation

The weekly run exists for what changes when nothing here does: an advisory published against a dependency nobody has bumped, and a live credential that quietly expired. The second of those is the one ci.yml's schedule is strict about; the live suites section has the detail.

Nothing here deploys the product. Telecraft has no hosted instance to deploy to: an adopter runs it themselves, and the two public surfaces, demo.telecraft.dev and telecraft.dev, are built by the repositories that own them, from a ref this one publishes.

What decides which jobs run

ci.yml opens with a changes job that diffs the pull request against its base and sets the outputs below. Every other job carries an if: naming one of them.

Output True when the change touches
code anything that is not docs/**, README.md, or docs-dispatch.yml
console console/**, internal/console/**, or ci.yml itself, and otherwise inherits code
image Dockerfile, .dockerignore, tools/image/**, or ci.yml itself
chart charts/**, tools/chartlint/**, tools/chart/**, cmd/telecraft/serve.go, internal/instance/api.go, or ci.yml itself

The reason is cost rather than tidiness. The live suites stand up a real Elasticsearch and open real pull requests against a fixture repository; running them because someone fixed a typo in a guide spends a service container and leaves external side effects behind for no reading. When the runners are slow, that wait is the whole review latency.

Five checks are the deliberate exception. All five run on everything, always, because none of them reads only code, which makes a documentation-only change exactly the change that can break them. The vendor-word lint reads code and prose alike, and a vendor word arrives as easily through a guide as through a Go file (ADR-0001). The front-matter check reads the published pages, and it exists because the site is built in a different repository (issue #74): the front matter and docs/nav.yaml are the whole contract between here and there, so a block that does not parse fails over there, after merge, in a build nobody here is watching. That is not hypothetical: docs/reference/estate-layout.md carried an unquoted colon in its description and took the whole documentation build down. The tracked-executable check reads the index, so its subject is not a language at all: a build artefact committed beside a guide is as much a hit as one committed beside a Go file. The formatting check reads the index too, and the tree holds Go that code counts as documentation: docs/prototypes/normaliser-spike is a module of its own, so a branch touching only it sets code false, and go build ./... never compiles it either. The workflow lint reads .github/workflows/, which code also counts as code, but the file that defines a gate is the one file whose own errors the gate cannot be trusted to catch, so it runs ungated beside the other four.

Two details worth knowing before you rely on this:

  • A skipped job reports success. That is what makes the gating safe to put behind branch protection: a required check that is skipped is not a check that is missing. A workflow that never ran would leave the check pending forever, which is why the filtering is at job level and not on the workflow's own trigger.
  • An unusable diff base runs everything. A first push, or a force-push, can leave before naming a commit the runner cannot resolve. Rather than guess at a range, changes sets both outputs true. Over-running is a cost; under-running is a hole.

The checks

Job Runs Guards
What changed a diff against the base Nothing. It decides what the rest of the table does
Vendor-word lint (ADR-0001) go run ./tools/vendorlint The neutral core holds: no vendor word in cmd/, internal/, console/ or the normative docs, and provider implementations stay product-qualified
Documentation front matter go run ./tools/docslint Every published page carries front matter the site can read, the contract between this repository and the site built from it
No tracked binaries (issue #122) go run ./tools/binlint No tracked file is a compiled executable. Build artefacts are built, never committed, and never shipped in a clone or a source tarball
Go formatting (issue #146) go run ./tools/fmtlint Every tracked Go file is gofmt clean, and the ones that are not are named
Workflow lint actionlint over .github/workflows/ The workflow files parse, their expressions resolve, and their scripts clear shellcheck at warning level; a broken gate fails here rather than on the event that needed it
Helm chart (ADR-0068) helm lint, go run ./tools/chartlint, tools/chart/golden.sh, then tools/chart/kind.sh over an image the job builds from the same commit The chart still matches the command it deploys, the rendered manifests are the ones in the goldens, and both estate shapes install on a stock cluster and serve a console and an OpAMP endpoint
Build and test go build ./..., go vet ./..., go test -race -shuffle=on ./..., govulncheck The core compiles, vets and passes its unit tests under the race detector, and its dependency tree carries no known advisory the code reaches, with no Docker and no network
TelemetryProvider live the Live suite against a single-node Elasticsearch service container The telemetry queries work against a real backend, not only a test double
Forge adapter live the Live suite against the GitHub App and estate-fixture The pull-request flow works against the real forge API
Console (ADR-0045) typecheck, test, check:palette, build, check:zero-cdn, check:bundle-budget, bundle and its staging assertion, e2e The console typechecks, its unit and Playwright suites pass, the palette clears its floors, the built bundle reaches no external host, its entry chunk stays within its gzipped ceiling, and the bundle lands under internal/consoleassets/ where go build embeds it. A failed Playwright run uploads its traces and screenshots as a playwright-output artifact
Demo snapshot and bundle build:demo, check:zero-cdn, telecraft snapshot, the entry-document check The public demo's two halves keep working: the snapshot the real evaluators produce, and the console bundle that reads it
Container image (ADR-0068) tools/image/stage.sh, a two-architecture docker buildx build, then tools/image/offline.sh over the loaded image The image assembles for both architectures, and the one it built serves the console and the OpAMP endpoint with networking disabled
Workstation binaries (ADR-0081) go build for darwin/arm64, darwin/amd64, windows/amd64 and linux/arm64, output discarded Every platform a release attaches, minus the one the test job compiles natively, still cross-compiles. The image job also builds both Linux architectures, but only when the image's own files change, so on an ordinary Go change this loop is the only cross-compile there is. Without it a build constraint or a platform-specific import is discovered by a tag push, which is the one moment there is no way back

Ten of those guard rules that are otherwise unenforceable in review.

The vendor-word lint exists because ADR-0001's neutral core is a property of the whole tree, and no reviewer reads the whole tree.

The front-matter check exists because its failure lands in someone else's repository. A reviewer here sees a sentence changed in a guide and has no reason to think about YAML.

The tracked-executable check exists because a reviewer reading a diff sees a path rather than a file type. A 3.4 MB Mach-O vendorlint binary sat at the repository root from the first scaffolding commit until issue #122, reaching every clone and every generated source tarball, and nothing in review was ever going to notice it. The check reads magic bytes rather than the executable bit, which is set on every checked-in shell script and says nothing about what a file is.

The formatting check exists because nothing else in the repository reads layout. go build and go test compile and run the code, go vet reports suspicious constructs rather than formatting, and a reviewer reading a diff sees the lines that changed rather than the ones beside them. Three test files were unformatted for months and went green on every check there was (issue #146).

It is go run ./tools/fmtlint rather than a gofmt -l . step, and the reason is worth knowing before you replace it with the shorter thing: gofmt -l exits 0 whether or not it names a file. A run: gofmt -l . step prints the offending paths and passes, which is the failure mode this whole page exists to avoid. The tool exits 1 on a finding, asks git for the tracked files rather than walking the working tree, and formats with go/format from the standard library, so the verdict comes from the toolchain that builds it rather than from whichever gofmt binary is first on PATH. Its self-test runs the check over this repository, which means the next drift fails go test ./... on your own machine before it reaches a runner.

Scope is every tracked .go file, with no exemption for testdata. The only Go under a testdata directory here is the repository's own lint fixtures; the vendored upstream fixtures under internal/catalogue/testdata are go.mod and metadata.yaml files and carry no Go at all. The exemption list would be empty, and an empty exemption is a door held open for the next file that wants through.

The chart check exists because the chart is a copy of a contract: its whole surface is one binary's flag set, and a flag that moves under it fails silently, first for an adopter. chartlint reads cmd/telecraft/serve.go and internal/instance/api.go rather than repeating what they say, so a flag the chart passes that the command no longer defines, a listen port that moved, or a probe path the server stopped serving is a failure here. Its self-test runs the check over this repository's own chart, so the drift fails go test ./... on your machine before it reaches a runner.

tools/chart/golden.sh renders the three documented shapes and compares them with the manifests in charts/telecraft/testdata/golden/, so a change to a template shows up in review as the change it makes to what an operator installs. It then renders the combinations the decisions refuse and requires each to fail with the phrase an operator needs to read. It is a shell script rather than a Go test for a reason worth knowing before somebody moves it: the platform runs no toolchain from Go, Helm is one, and TestNoToolchainBinaryIsInvoked holds every tracked Go file to that. What chartlint checks needs no renderer, which is why it runs in go test and this does not.

Neither of those runs a manifest, so tools/chart/kind.sh installs the chart on a kind cluster twice, once in each estate shape, over a bare repository and a checkout mounted into the node rather than anything on a network. It requires both installs to answer /readyz, serve the console document and answer on the OpAMP port through the Service the chart created, then lands a commit in the repository and requires the sidecar to pull it and the pod not to restart. What it catches is the half a renderer cannot see: an image the pod cannot pull, a probe on the wrong port, a mount two containers disagree about, a security context the kubelet refuses.

check:zero-cdn runs over the built bundle, not the source, because the air-gap rule is about what the browser fetches and a bundler can introduce a request no import statement shows (ADR-0019, ADR-0045 §5). In the demo job it deliberately runs before the snapshot is placed beside the bundle: a snapshot carries endpoints and module paths the console prints and never fetches, and it is estate data rather than a bundled artefact.

check:palette exists because ADR-0047's accessibility floors are numbers, and a number can be regressed. It resolves both themes, checks every token against the ground it sits on, measures the severity triad under simulated deuteranopia and protanopia, and enforces the rule that every colour is defined in exactly two blocks and never inside a media query. Before it existed, docs/branding/design-system.md recorded the method and the expected values, and the documented palette turned out not to clear its own floors, which is precisely the failure a document cannot catch.

The image job builds and never pushes. A pull request cannot publish, and the thing worth catching before a tag is not whether a registry accepts an upload: it is whether the image assembles at all, and whether what it assembled runs. So the job stages the context, builds the index for both architectures, loads one of them, and starts it on no network. An image that needed one fetch to become ready fails there rather than in somebody's air gap.

The properties that need no daemon are checked without one. go run ./tools/imagelint reads the Dockerfile and reports an image that grew a build stage, lost its digest pin, started running as root, or drifted from the address flags it mirrors. It carries a self-test over this repository's own Dockerfile, so go test ./... fails on the drift before a runner does, and it needs no job of its own.

The deployment compose file is checked the same way. go run ./tools/composelint reads deploy/compose/, and reports an estate mounted writable, a secret carried as a value rather than placed as a file, an image reference a mirror cannot replace, a published port nothing listens on, and a variable the compose file reads that .env.example never names. Proving the deployment end to end needs a container runtime, a certificate and an estate; these are the properties that need none of them. It carries the same self-test, so it needs no job of its own either.

check:bundle-budget exists because a code split is easy to undo by accident. The console loads each Workspace's code on navigation (issue #125), and an ordinary-looking import in the wrong file pulls a Workspace back into the chunk every reader downloads first. The check measures the gzipped entry chunk against a stated ceiling, prints every chunk's size as it goes, and says in its failure message what to do: move the weight behind a route, or raise the ceiling in the same commit and say why. The console page carries the ceiling and the reasoning for it.

The live suites, and why they can be green without credentials

Both live jobs follow the conformance-kit discipline of ADR-0036: absent credentials never fail a suite. The Forge suite skips loudly, printing what it would have needed, when the FORGE_* secrets are not exposed, so a fork's pull request stays green rather than failing on something the contributor cannot provide.

This is a deliberate trade. It means a credential that quietly stops working reads as a pass, so the log is the only place that says the suite skipped. When you change provider code, read it rather than trusting the tick.

The weekly scheduled run is the backstop for that trade. On the schedule event, and only on it, the forge job refuses to skip: absent secrets fail the run, because on a quiet Monday over an unchanged tree a skip means a dead credential and not a considerate fork. A rotting credential therefore surfaces within a week rather than whenever someone next reads a log.

The same suites run locally; the development page has the environment variables, and internal/provider/telemetry/demo.sh stands up the backend the Elasticsearch suite wants.

Reproducing a failure locally

Every check is a command you can run yourself: there is nothing CI does that the repository cannot.

go build ./... && go vet ./... && go test -race -shuffle=on ./...
go run golang.org/x/vuln/cmd/[email protected] ./...
go run ./tools/vendorlint
go run ./tools/docslint
go run ./tools/binlint
go run ./tools/fmtlint
go run ./tools/imagelint
go run ./tools/composelint
go run github.com/rhysd/actionlint/cmd/[email protected]

cd console
npm ci
npm run typecheck && npm test && npm run check:palette
npm run build && npm run check:zero-cdn && npm run check:bundle-budget
npm run e2e

The image needs a container runtime, so it is the one check that is not simply a command:

tools/image/stage.sh
PLATFORMS=linux/amd64 LOAD=1 STAGED=1 IMAGE=telecraft VERSION=local tools/image/build.sh
tools/image/offline.sh telecraft:local

If npm run e2e fails on a missing browser, npx playwright install chromium fetches it. Do not add --with-deps. CI leaves it off on purpose: it shells out to apt-get against the Ubuntu mirrors, and measured over three consecutive runs it cost 138s, 362s, and then hung past 25 minutes, against 14s for the suite it exists to enable. The GitHub-hosted image already ships the shared libraries Chromium needs, and a missing library fails the next step immediately with its name, which is a better failure than a hang.

How a change reaches the public surfaces

Neither public site is deployed from this repository. Both build from a ref here, and both are told when that ref moves.

The documentation site. A push to main touching docs/** makes docs-dispatch.yml send a repository dispatch to telecraft-dev/telecraft.dev, which rebuilds. The navigation comes from docs/nav.yaml, so a new page is published by adding it here and nothing else (see writing documentation).

The demo. demo-dispatch.yml runs when release.yml finishes and acts only when it succeeded, so the demo builds from a release, a bug on main cannot reach it, and a tag whose release failed its checks moves nothing (ADR-0049 §4, issue #86). It repeats the release workflow's shape and ancestry guards rather than trusting the ordering, because a guard that lives in another file is a guard someone edits away without noticing what relied on it. Between releases the demo lags, deliberately: there is no staging site, because ci.yml's demo job already builds the snapshot and the demo bundle on every pull request that touches them, and a staging site would be a second deployment and a second public claim to catch what CI catches first.

The ref the demo checks out is a moving release tag rather than a version. That is not the obvious design and it is worth understanding before changing it: estate-demo builds on three events (a push to its own estate, a manual run, and the dispatch) and only the manual run can carry a ref. Whatever its fallback names is therefore the pin for all three, so a ref travelling in the dispatch payload would mean an estate content push building against something else, and which platform version the public site ran would depend on which event fired last.

So demo-dispatch.yml reads a tag and reaches one of three conclusions:

Tag Pointer Demo
the newest stable version moves to it rebuilds
an older stable version unmoved unchanged: a fix on an older line must not drag the demo backwards
a pre-release (v0.1.0-rc.1) unmoved unchanged

That third row is what makes the release path rehearsable: a pre-release exercises release.yml end to end and publishes its artefacts while the public demo stays exactly where it is (ADR-0049 §6). The version tags stay immutable; release is the only ref that moves, and git rev-parse release answers what the demo is built from.

Cutting a release is its own page.

Dependency updates

.github/dependabot.yml covers three ecosystems: the root Go module, the console's npm tree, and the actions the workflows use. It is not a workflow, and it fires on a schedule rather than on anything that happens here, which is why it is absent from the table at the top of this page.

Each ecosystem is grouped and weekly, so a week of upstream releases arrives as one pull request per ecosystem rather than one per dependency. Those are ordinary pull requests: ci.yml gates them exactly as it gates yours, so a grouped bump either goes green or names the dependency that broke it, and the reviewer's question is the one CI already answers. The file carries the full reasoning in its comments (issue #127).

Two things to know before you rely on it:

  • Security updates ignore the schedule. Dependabot raises a fix when it learns of the advisory, not on the next Monday.
  • Only the listed directories are watched. The root go.mod and console/package.json are the tree's shipping manifests. The go.mod files under docs/prototypes/ and internal/catalogue/testdata/ are a spike record and a vendored fixture set, and nothing builds them, so nothing updates them either.

Changing a workflow

A change to ci.yml sets console true, so the console and demo jobs run against their own new definition rather than the previous one.

Workflows in this repository carry their reasoning in comments, at more length than is usual. That is deliberate: a CI file is read most often by someone under time pressure trying to understand why it is failing, and the decisions in these files (the missing --with-deps, the skipped-is- success gating, the moving pointer) all look like mistakes until the reason is at hand. Keep that up in anything you add.