Continuous integration
Four workflows live in .github/workflows/. One guards changes, one
publishes releases, and two tell downstream repositories that something
they build from has moved. Static analysis runs beside them as GitHub
code scanning's default setup, configured in the repository settings
rather than in a workflow file, and reports its findings on the same
pull requests.
| Workflow | Fires on | Does |
|---|---|---|
ci.yml |
every pull request, every push to main, and a weekly schedule |
Builds, tests and lints: the checks review waits for |
release.yml |
a v* tag |
Publishes the release: the image, the chart, the CLI for every platform, and the design artefacts |
demo-dispatch.yml |
a successful release, or a manual run | Moves the release pointer and asks the demo to rebuild |
docs-dispatch.yml |
a push to main touching docs/** |
Asks telecraft.dev to rebuild the documentation |
The weekly run exists for what changes when nothing here does: an advisory
published against a dependency nobody has bumped, and a live credential
that quietly expired. The second of
those is the one ci.yml's schedule is strict about; the
live suites section
has the detail.
Nothing here deploys the product. Telecraft has no hosted instance to deploy to: an adopter runs it themselves, and the two public surfaces, demo.telecraft.dev and telecraft.dev, are built by the repositories that own them, from a ref this one publishes.
What decides which jobs run
ci.yml opens with a changes job that diffs the pull request against its
base and sets the outputs below. Every other job carries an if: naming one
of them.
| Output | True when the change touches |
|---|---|
code |
anything that is not docs/**, README.md, or docs-dispatch.yml |
console |
console/**, internal/console/**, or ci.yml itself, and otherwise inherits code |
image |
Dockerfile, .dockerignore, tools/image/**, or ci.yml itself |
chart |
charts/**, tools/chartlint/**, tools/chart/**, cmd/telecraft/serve.go, internal/instance/api.go, or ci.yml itself |
The reason is cost rather than tidiness. The live suites stand up a real Elasticsearch and open real pull requests against a fixture repository; running them because someone fixed a typo in a guide spends a service container and leaves external side effects behind for no reading. When the runners are slow, that wait is the whole review latency.
Five checks are the deliberate exception. All five run on
everything, always, because none of them reads only code, which makes a
documentation-only change exactly the change that can break them. The
vendor-word lint reads code and prose alike, and a vendor word arrives as
easily through a guide as through a Go file (ADR-0001). The front-matter
check reads the published pages, and it exists because the site is built
in a different repository (issue #74): the front matter and
docs/nav.yaml are the whole contract between here and there, so a block
that does not parse fails over there, after merge, in a build nobody here
is watching. That is not hypothetical: docs/reference/estate-layout.md
carried an unquoted colon in its description and took the whole
documentation build down. The tracked-executable check reads the index, so
its subject is not a language at all: a build artefact committed beside a
guide is as much a hit as one committed beside a Go file. The formatting
check reads the index too, and the tree holds Go that code counts as
documentation: docs/prototypes/normaliser-spike is a module of its own,
so a branch touching only it sets code false, and go build ./... never
compiles it either. The workflow lint reads .github/workflows/, which
code also counts as code, but the file that defines a gate is the one
file whose own errors the gate cannot be trusted to catch, so it runs
ungated beside the other four.
Two details worth knowing before you rely on this:
- A skipped job reports success. That is what makes the gating safe to put behind branch protection: a required check that is skipped is not a check that is missing. A workflow that never ran would leave the check pending forever, which is why the filtering is at job level and not on the workflow's own trigger.
- An unusable diff base runs everything. A first push, or a force-push,
can leave
beforenaming a commit the runner cannot resolve. Rather than guess at a range,changessets both outputs true. Over-running is a cost; under-running is a hole.
The checks
| Job | Runs | Guards |
|---|---|---|
| What changed | a diff against the base | Nothing. It decides what the rest of the table does |
| Vendor-word lint (ADR-0001) | go run ./tools/vendorlint |
The neutral core holds: no vendor word in cmd/, internal/, console/ or the normative docs, and provider implementations stay product-qualified |
| Documentation front matter | go run ./tools/docslint |
Every published page carries front matter the site can read, the contract between this repository and the site built from it |
| No tracked binaries (issue #122) | go run ./tools/binlint |
No tracked file is a compiled executable. Build artefacts are built, never committed, and never shipped in a clone or a source tarball |
| Go formatting (issue #146) | go run ./tools/fmtlint |
Every tracked Go file is gofmt clean, and the ones that are not are named |
| Workflow lint | actionlint over .github/workflows/ |
The workflow files parse, their expressions resolve, and their scripts clear shellcheck at warning level; a broken gate fails here rather than on the event that needed it |
| Helm chart (ADR-0068) | helm lint, go run ./tools/chartlint, tools/chart/golden.sh, then tools/chart/kind.sh over an image the job builds from the same commit |
The chart still matches the command it deploys, the rendered manifests are the ones in the goldens, and both estate shapes install on a stock cluster and serve a console and an OpAMP endpoint |
| Build and test | go build ./..., go vet ./..., go test -race -shuffle=on ./..., govulncheck |
The core compiles, vets and passes its unit tests under the race detector, and its dependency tree carries no known advisory the code reaches, with no Docker and no network |
| TelemetryProvider live | the Live suite against a single-node Elasticsearch service container |
The telemetry queries work against a real backend, not only a test double |
| Forge adapter live | the Live suite against the GitHub App and estate-fixture |
The pull-request flow works against the real forge API |
| Console (ADR-0045) | typecheck, test, check:palette, build, check:zero-cdn, check:bundle-budget, bundle and its staging assertion, e2e |
The console typechecks, its unit and Playwright suites pass, the palette clears its floors, the built bundle reaches no external host, its entry chunk stays within its gzipped ceiling, and the bundle lands under internal/consoleassets/ where go build embeds it. A failed Playwright run uploads its traces and screenshots as a playwright-output artifact |
| Demo snapshot and bundle | build:demo, check:zero-cdn, telecraft snapshot, the entry-document check |
The public demo's two halves keep working: the snapshot the real evaluators produce, and the console bundle that reads it |
| Container image (ADR-0068) | tools/image/stage.sh, a two-architecture docker buildx build, then tools/image/offline.sh over the loaded image |
The image assembles for both architectures, and the one it built serves the console and the OpAMP endpoint with networking disabled |
| Workstation binaries (ADR-0081) | go build for darwin/arm64, darwin/amd64, windows/amd64 and linux/arm64, output discarded |
Every platform a release attaches, minus the one the test job compiles natively, still cross-compiles. The image job also builds both Linux architectures, but only when the image's own files change, so on an ordinary Go change this loop is the only cross-compile there is. Without it a build constraint or a platform-specific import is discovered by a tag push, which is the one moment there is no way back |
Ten of those guard rules that are otherwise unenforceable in review.
The vendor-word lint exists because ADR-0001's neutral core is a property of the whole tree, and no reviewer reads the whole tree.
The front-matter check exists because its failure lands in someone else's repository. A reviewer here sees a sentence changed in a guide and has no reason to think about YAML.
The tracked-executable check exists because a reviewer reading a diff
sees a path rather than a file type. A 3.4 MB Mach-O vendorlint binary
sat at the repository root from the first scaffolding commit until issue
#122, reaching every clone and every generated source tarball, and nothing
in review was ever going to notice it. The check reads magic bytes rather
than the executable bit, which is set on every checked-in shell script and
says nothing about what a file is.
The formatting check exists because nothing else in the repository
reads layout. go build and go test compile and run the code, go vet
reports suspicious constructs rather than formatting, and a reviewer
reading a diff sees the lines that changed rather than the ones beside
them. Three test files were unformatted for months and went green on every
check there was (issue #146).
It is go run ./tools/fmtlint rather than a gofmt -l . step, and the
reason is worth knowing before you replace it with the shorter thing:
gofmt -l exits 0 whether or not it names a file. A run: gofmt -l .
step prints the offending paths and passes, which is the failure mode this
whole page exists to avoid. The tool exits 1 on a finding, asks git for
the tracked files rather than walking the working tree, and formats with
go/format from the standard library, so the verdict comes from the
toolchain that builds it rather than from whichever gofmt binary is
first on PATH. Its self-test runs the check over this repository, which
means the next drift fails go test ./... on your own machine before it
reaches a runner.
Scope is every tracked .go file, with no exemption for testdata. The
only Go under a testdata directory here is the repository's own lint
fixtures; the vendored upstream fixtures under
internal/catalogue/testdata are go.mod and metadata.yaml files and
carry no Go at all. The exemption list would be empty, and an empty
exemption is a door held open for the next file that wants through.
The chart check exists because the chart is a copy of a contract:
its whole surface is one binary's flag set, and a flag that moves under it
fails silently, first for an adopter. chartlint reads
cmd/telecraft/serve.go and internal/instance/api.go rather than
repeating what they say, so a flag the chart passes that the command no
longer defines, a listen port that moved, or a probe path the server stopped
serving is a failure here. Its self-test runs the check over this
repository's own chart, so the drift fails go test ./... on your machine
before it reaches a runner.
tools/chart/golden.sh renders the three documented shapes and compares
them with the manifests in charts/telecraft/testdata/golden/, so a change
to a template shows up in review as the change it makes to what an operator
installs. It then renders the combinations the decisions refuse and requires
each to fail with the phrase an operator needs to read. It is a shell script
rather than a Go test for a reason worth knowing before somebody moves it:
the platform runs no toolchain from Go, Helm is one, and
TestNoToolchainBinaryIsInvoked holds every tracked Go file to that. What
chartlint checks needs no renderer, which is why it runs in go test
and this does not.
Neither of those runs a manifest, so tools/chart/kind.sh installs the
chart on a kind cluster twice, once in each estate shape, over a bare
repository and a checkout mounted into the node rather than anything on a
network. It requires both installs to answer /readyz, serve the console
document and answer on the OpAMP port through the Service the chart created,
then lands a commit in the repository and requires the sidecar to pull it
and the pod not to restart. What it catches is the half a renderer cannot
see: an image the pod cannot pull, a probe on the wrong port, a mount two
containers disagree about, a security context the kubelet refuses.
check:zero-cdn runs over the built bundle, not the source, because
the air-gap rule is about what the browser fetches and a bundler can
introduce a request no import statement shows (ADR-0019, ADR-0045 §5). In
the demo job it deliberately runs before the snapshot is placed beside
the bundle: a snapshot carries endpoints and module paths the console
prints and never fetches, and it is estate data rather than a bundled
artefact.
check:palette exists because ADR-0047's accessibility floors are
numbers, and a number can be regressed. It resolves both themes, checks
every token against the ground it sits on, measures the severity triad
under simulated deuteranopia and protanopia, and enforces the rule that
every colour is defined in exactly two blocks and never inside a media
query. Before it existed, docs/branding/design-system.md recorded the
method and the expected values, and the documented palette turned out not
to clear its own floors, which is precisely the failure a document cannot
catch.
The image job builds and never pushes. A pull request cannot publish, and the thing worth catching before a tag is not whether a registry accepts an upload: it is whether the image assembles at all, and whether what it assembled runs. So the job stages the context, builds the index for both architectures, loads one of them, and starts it on no network. An image that needed one fetch to become ready fails there rather than in somebody's air gap.
The properties that need no daemon are checked without one. go run ./tools/imagelint reads the Dockerfile and reports an image that grew a
build stage, lost its digest pin, started running as root, or drifted from
the address flags it mirrors. It carries a self-test over this repository's
own Dockerfile, so go test ./... fails on the drift before a runner does,
and it needs no job of its own.
The deployment compose file is checked the same way. go run ./tools/composelint reads deploy/compose/, and reports an estate mounted
writable, a secret carried as a value rather than placed as a file, an image
reference a mirror cannot replace, a published port nothing listens on, and a
variable the compose file reads that .env.example never names. Proving the
deployment end to end needs a container runtime, a certificate and an estate;
these are the properties that need none of them. It carries the same
self-test, so it needs no job of its own either.
check:bundle-budget exists because a code split is easy to undo by
accident. The console loads each Workspace's code on navigation (issue
#125), and an ordinary-looking import in the wrong file pulls a Workspace
back into the chunk every reader downloads first. The check measures the
gzipped entry chunk against a stated ceiling, prints every chunk's size as
it goes, and says in its failure message what to do: move the weight behind
a route, or raise the ceiling in the same commit and say why. The
console page carries the ceiling and the
reasoning for it.
The live suites, and why they can be green without credentials
Both live jobs follow the conformance-kit discipline of ADR-0036: absent
credentials never fail a suite. The Forge suite skips loudly, printing
what it would have needed, when the FORGE_* secrets are not exposed, so
a fork's pull request stays green rather than failing on something the
contributor cannot provide.
This is a deliberate trade. It means a credential that quietly stops working reads as a pass, so the log is the only place that says the suite skipped. When you change provider code, read it rather than trusting the tick.
The weekly scheduled run is the backstop for that trade. On the schedule event, and only on it, the forge job refuses to skip: absent secrets fail the run, because on a quiet Monday over an unchanged tree a skip means a dead credential and not a considerate fork. A rotting credential therefore surfaces within a week rather than whenever someone next reads a log.
The same suites run locally; the development page has
the environment variables, and internal/provider/telemetry/demo.sh stands up the backend
the Elasticsearch suite wants.
Reproducing a failure locally
Every check is a command you can run yourself: there is nothing CI does that the repository cannot.
go build ./... && go vet ./... && go test -race -shuffle=on ./...
go run golang.org/x/vuln/cmd/[email protected] ./...
go run ./tools/vendorlint
go run ./tools/docslint
go run ./tools/binlint
go run ./tools/fmtlint
go run ./tools/imagelint
go run ./tools/composelint
go run github.com/rhysd/actionlint/cmd/[email protected]
cd console
npm ci
npm run typecheck && npm test && npm run check:palette
npm run build && npm run check:zero-cdn && npm run check:bundle-budget
npm run e2e
The image needs a container runtime, so it is the one check that is not simply a command:
tools/image/stage.sh
PLATFORMS=linux/amd64 LOAD=1 STAGED=1 IMAGE=telecraft VERSION=local tools/image/build.sh
tools/image/offline.sh telecraft:local
If npm run e2e fails on a missing browser, npx playwright install chromium fetches it. Do not add --with-deps. CI leaves it off on
purpose: it shells out to apt-get against the Ubuntu mirrors, and
measured over three consecutive runs it cost 138s, 362s, and then hung
past 25 minutes, against 14s for the suite it exists to enable. The
GitHub-hosted image already ships the shared libraries Chromium needs, and
a missing library fails the next step immediately with its name, which is
a better failure than a hang.
How a change reaches the public surfaces
Neither public site is deployed from this repository. Both build from a ref here, and both are told when that ref moves.
The documentation site. A push to main touching docs/** makes
docs-dispatch.yml send a repository dispatch to
telecraft-dev/telecraft.dev, which rebuilds. The navigation comes from
docs/nav.yaml, so a new page is published by adding it here and nothing
else (see writing documentation).
The demo. demo-dispatch.yml runs when release.yml finishes and
acts only when it succeeded, so the demo builds from a release, a bug on
main cannot reach it, and a tag whose release failed its checks moves
nothing (ADR-0049 §4, issue #86). It repeats the release workflow's shape
and ancestry guards rather than trusting the ordering, because a guard
that lives in another file is a guard someone edits away without noticing
what relied on it. Between releases the demo lags, deliberately:
there is no staging site, because ci.yml's demo job already builds the
snapshot and the demo bundle on every pull request that touches them, and
a staging site would be a second deployment and a second public claim to
catch what CI catches first.
The ref the demo checks out is a moving release tag rather than a
version. That is not the obvious design and it is worth understanding
before changing it: estate-demo builds on three events (a push to its
own estate, a manual run, and the dispatch) and only the manual run can
carry a ref. Whatever its fallback names is therefore the pin for all
three, so a ref travelling in the dispatch payload would mean an estate
content push building against something else, and which platform version
the public site ran would depend on which event fired last.
So demo-dispatch.yml reads a tag and reaches one of three conclusions:
| Tag | Pointer | Demo |
|---|---|---|
| the newest stable version | moves to it | rebuilds |
| an older stable version | unmoved | unchanged: a fix on an older line must not drag the demo backwards |
a pre-release (v0.1.0-rc.1) |
unmoved | unchanged |
That third row is what makes the release path rehearsable: a pre-release
exercises release.yml end to end and publishes its artefacts while the
public demo stays exactly where it is (ADR-0049 §6). The version tags stay
immutable; release is the only ref that moves, and git rev-parse release answers what the demo is built from.
Cutting a release is its own page.
Dependency updates
.github/dependabot.yml covers three ecosystems: the root Go module, the
console's npm tree, and the actions the workflows use. It is not a workflow,
and it fires on a schedule rather than on anything that happens here, which
is why it is absent from the table at the top of this page.
Each ecosystem is grouped and weekly, so a week of upstream releases arrives
as one pull request per ecosystem rather than one per dependency. Those are
ordinary pull requests: ci.yml gates them exactly as it gates yours, so a
grouped bump either goes green or names the dependency that broke it, and
the reviewer's question is the one CI already answers. The file carries the
full reasoning in its comments (issue #127).
Two things to know before you rely on it:
- Security updates ignore the schedule. Dependabot raises a fix when it learns of the advisory, not on the next Monday.
- Only the listed directories are watched. The root
go.modandconsole/package.jsonare the tree's shipping manifests. Thego.modfiles underdocs/prototypes/andinternal/catalogue/testdata/are a spike record and a vendored fixture set, and nothing builds them, so nothing updates them either.
Changing a workflow
A change to ci.yml sets console true, so the console and demo jobs run
against their own new definition rather than the previous one.
Workflows in this repository carry their reasoning in comments, at more
length than is usual. That is deliberate: a CI file is read most often by
someone under time pressure trying to understand why it is failing, and
the decisions in these files (the missing --with-deps, the skipped-is-
success gating, the moving pointer) all look like mistakes until the
reason is at hand. Keep that up in anything you add.