CI/CD
Thirteen workflows guard this repository. This page explains what each one protects, the decisions behind how they are wired, and the failure modes that shaped them.
Pipeline at a glance
| Workflow | Trigger | Guards against |
|---|---|---|
ci.yml | PR, main | Unformatted code, clippy warnings, failing tests (including the shell integration in real bash, zsh, fish and pwsh), MSRV drift, broken shell scripts, stale reference docs, malformed .po, files missing from the published crate |
integration-test.yml | PR, main | Migration breaking for pyenv / conda / virtualenvwrapper users |
coverage.yml | PR, main | Untested code paths going unnoticed |
bench.yml | PR, main | Performance regressions in parsing and validation |
mutants.yml | PR (diff), weekly (full) | Tests that execute code without asserting on it |
security.yml | PR, main, weekly | Vulnerable dependencies, disallowed licences, untrusted sources |
msrv-check.yml | Cargo.toml/Cargo.lock changes | A declared MSRV that no longer compiles |
docker-build.yml | docker/** changes, weekly | Broken published images, vulnerable image contents |
fuzz.yml | Weekly | Parser crashes on hostile input |
docs-check.yml | PR touching docs/src/**, docs/po/**, docs/book.toml, docs/theme/** | A docs edit that breaks the mdBook build or leaves ko.po stale reaching a release tag |
docs.yml | v* tags | Broken documentation site |
release-plz.yml | main | Manual release mistakes |
cache-cleanup.yml | PR closed, weekly | Closed PRs’ cache copies and superseded main rust caches eating the 10 GB allowance and forcing evictions mid-export |
Cross-cutting decisions
Cancel PR runs, never cancel main
Every workflow that runs on both uses the same concurrency block:
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
Pushing a fixup to a PR should abandon the previous run — nobody needs
results for a commit that no longer exists. A run on main is different:
it produces state that later runs depend on. bench.yml writes the
benchmark baseline to gh-pages, and a cancelled main run leaves the
next PR comparing against a stale baseline.
release-plz.yml uses the same block with cancel-in-progress: false: a
release must never be cancelled, but two releases must never run at once
either, and a group delivers both.
The MSRV is verified twice, deliberately
ci.yml has an msrv job that builds and tests on 1.89. msrv-check.yml
separately runs cargo msrv verify. These answer different questions:
ci.yml— does the code work on the version we claim?msrv-check.yml— is the version we claim still the lowest one that works?
The second catches the case where a dependency bump silently raises the
real floor while Cargo.toml still advertises the old one. It only runs
when Cargo.toml or Cargo.lock changes, because that is the only way
the answer can change.
cargo-msrv itself is pinned to ^0.18: 0.19 requires rustc 1.91, which
this project’s own toolchain pin cannot satisfy.
Gate what is stable, track what is noisy
Not every measurement can be a gate. bench.yml splits its benchmarks by
how reproducible they are on shared runners:
| Group | Benches | Observed spread | Behaviour |
|---|---|---|---|
| CPU | parsing, validation | 1.71x-2.37x | Fails the build past 250% |
| Filesystem | path_lookup | ~3.4x | Recorded, never fails |
find_executable_in calls stat. Across 38 runs of unchanged code on
main it produced anywhere from 574 to 1931 ns depending on which runner
the job landed on. No threshold both tolerates that and catches a real
regression, so the group is tracked on gh-pages for trend visibility
and fail-on-alert: false keeps it out of the gate.
This matters because a noisy gate is worse than no gate: it blocks unrelated work and trains reviewers to ignore red.
The CPU row was written as “~1.4x, fails past 150%”, and the numbers did not
hold. Measured across the 60 runs on gh-pages, the group swings 1.71x to
2.37x — its noise floor was above its own threshold, so it could not tell a
regression from a runner. The way it failed is worth remembering, because a
re-run does not clear it: the action compares against the single previous
data point, and a run that lands on a fast runner measures low across every
bench at once. Commit 9a9241da came in ~1.7x fast, and every PR afterwards
compared against that and failed until the next commit reached main and
replaced the baseline. 250% clears the measured ceiling. It is a coarse gate
on purpose — the regressions worth catching here are algorithmic, and those
land far beyond 2.5x.
Mutation testing runs at two scopes
cargo test proves a line executed. It does not prove a test would
notice if that line were wrong. cargo-mutants injects deliberate
defects — flipping > to >=, replacing a return value — and reports
any that the suite fails to catch.
- On PRs (
--in-diff): only the lines this PR changed. Fast enough to gate on. - Weekly (full): every candidate in scope. Tracked rather than gated for now — the step tolerates exit 2 (mutants missed) and nothing else, so a broken baseline or a timeout still fails loudly while the backlog in Known gaps is worked through.
.cargo/mutants.toml carries the exclusions, each with a written
rationale. Most exclude code whose mutations can only be killed by
spawning a real uv or python — a gap we accept rather than a test
hole. Add exclusions there with the reason, not silently.
exclude_re matches the whole mutant description, not the function name.
A bare "foo" therefore drops every mutant in foo, including the ones
existing tests already kill — the exclusion outlives the reason given for
it, and --in-diff never gates that code again. Exclude the specific
description instead ("delete match arm \[major\] in foo"), and confirm
what actually left the gate by listing both configs and diffing them:
cargo mutants --config <before>.toml --list --file '<glob>' | sort > before.txt
cargo mutants --list --file '<glob>' | sort > after.txt
comm -23 before.txt after.txt
Some mutants cannot be separated that way. Two delete ! in one function
produce a byte-identical description, so a whole-function exclusion is the
only option available — say so in the comment when that is the reason,
instead of letting it read as an equivalence claim.
New Check trait implementations and thin wrappers need direct dispatch
tests, or the PR gate reports them as MISSED.
One cache per toolchain, one writer per cache
Swatinem/rust-cache keys on shared-key, so jobs sharing a key share a
cache:
shared-key | Jobs |
|---|---|
stable-ci | ci.yml lint / test / shellcheck, coverage.yml, mutants.yml |
msrv | ci.yml msrv |
msrv-verify | msrv-check.yml |
bench | bench.yml |
MSRV artifacts are kept separate on purpose: a different compiler produces incompatible output, and sharing would mean both jobs permanently invalidating each other.
Within a shared key there is one writer. ci.yml marks lint and
shellcheck save-if: false so only the test job — which produces the
richest artifacts, including test binaries — populates the cache; in
mutants.yml the incremental job writes and the weekly full job reads.
Note that rust-cache folds workflow-level env into its key hash. Two
jobs can declare the same shared-key and still land on different
caches if their workflows set different environment variables.
Failure modes worth remembering
These cost real debugging time. They are recorded so the next person recognises them faster.
Criterion errors corrupt the benchmark parser output
rust-cache restores target/ with the criterion tree present but its
sample.json baselines pruned out. Criterion then fails to load a
baseline it can see should exist, and writes the error to stdout
mid-line:
test clap_parse_create ... Criterion.rs ERROR: error: Failed to access file
".../base/sample.json": No such file or directory
bench: 41,347 ns/iter (+/- 2,364)
benchmark-action’s parser needs test NAME ... bench: N ns/iter on one
line. Split in two, it reports “no benchmark result” even though every
benchmark ran. A cold cache passes precisely because there is no tree to
half-load.
Each bench step therefore starts with rm -rf target/criterion. The gate
compares against gh-pages history and never against criterion’s local
baseline, so removing it costs nothing.
Two benchmark-action invocations collide on gh-pages
The first invocation fetches gh-pages into a local branch and commits
its entry onto it — even on PRs, where it simply never pushes. A second
invocation’s identical fetch is then a non-fast-forward update and git
rejects it, failing the step before any comparison happens.
skip-fetch-gh-pages: true on the second step reuses what the first
fetched.
cargo bench | tee swallows failures
Without set -o pipefail the step exits with tee’s status, so a
genuinely failing cargo bench reports success and resurfaces one step
later as a confusing parse error.
Docs guards run on PRs only for docs paths
docs.yml — the mdBook build, the ko.po staleness round-trip and the
Pages deploy — triggers only on v* tags. docs-check.yml runs the same
build and the same round-trip on pull requests, but only when the PR
touches docs/src/**, docs/po/**, docs/book.toml or docs/theme/**.
It is a separate workflow rather than a trigger on docs.yml so a PR run
needs neither pages: write nor id-token: write and never enters the
pages concurrency group. Two checks in the ci.yml Lint job cover the
facts that live outside those paths:
scripts/check-doc-references.pycompares facts copied intoREADME.md,CONTRIBUTING.md,llms.txt,llms-full.txtand the docs against the code that owns them — version and MSRV fromCargo.toml, reserved names fromsrc/validate.rs, key count fromlocales/app.yml. Every check is doc-against-code; comparing the threellmsfiles to each other would pass with all three wrong.msgfmt --checkondocs/po/*.pocatches structural damage such as amsgstrwhose trailing newline no longer matches itsmsgid.
Editing any page under docs/src/ still requires regenerating ko.po in
the same commit. See Docs Translation.
Secrets
| Secret | Used by | Purpose |
|---|---|---|
RELEASE_PLZ_TOKEN | release-plz.yml | Push release PRs and tags |
CARGO_REGISTRY_TOKEN | release-plz.yml | Publish to crates.io |
CODECOV_TOKEN | coverage.yml | Upload coverage reports |
GITHUB_TOKEN | bench, docker | Push gh-pages, publish to ghcr.io |
A workflow triggered by Dependabot reads the Dependabot secret store, not
the Actions one. A secret both need has to be registered twice —
gh secret set CODECOV_TOKEN --app dependabot — or the job fails on every
Dependabot PR while passing everywhere else.
Default workflow permissions are read. Workflows that need more declare
it explicitly — bench.yml needs contents: write for gh-pages,
docker-build.yml needs packages: write for the registry.
Releases
release-plz reads Conventional Commits on main and prepares the
release; release-plz.toml holds the policy:
release_commits = "^(feat|fix|perf|refactor|revert)"— documentation and chore commits do not trigger a release on their own.features_always_increment_minor = true— afeatbumps the minor version even in0.x, where cargo’s default would only bump the patch.- Changelog generation goes through
git-cliff(cliff.toml).
release-plz only rewrites Cargo.toml, Cargo.lock and CHANGELOG.md. The
scuv <version> samples in README.md, CLAUDE.md, installation.md and
api.md are not its business, so every release PR used to fail the Lint job on
check-doc-references.py until someone hand-synced those four files. There is
no hook that runs while release-plz builds the PR — the hooks it does have live
on the publish path, which is too late — so release-plz.yml carries a step
that runs check-doc-references.py --fix and commits the result onto the
release branch. A release PR therefore arrives with an extra
docs: sync version samples to <version> commit; that is the automation
working, not drift. The step pushes with RELEASE_PLZ_TOKEN because the
default GITHUB_TOKEN does not re-trigger workflows, and re-runs the full
guard afterwards so a reference --fix does not cover still fails the job.
Merging the release PR is what publishes: it creates the tag, the GitHub
release, and the crates.io upload — and the v* tag is also what deploys
the documentation site.
Known gaps
Accurate as of the last revision of this page. Verify before relying on any of these being fixed.
codecov/projecthas never posted.codecov.ymlstates a target for it, and the config is live — settingcomment.require_changeschanged comment behaviour on the next PR. The status itself has never appeared, though, on any PR back through #165, which predates that config by months. The cause is on the Codecov side, in account or organisation settings this repository cannot read, so the relative check (“did this PR drop overall coverage”) is not being enforced. An absolute floor stands in for it incoverage.yml(cargo llvm-cov report --fail-under-lines 80), which catches a collapse but not a slow slide. Requiring the status before it is known to post would leave every PR waiting on a check that never arrives.- Unreviewed mutation escapes outside
migrate/. The last full-run artifact listed 55 mutants no test kills, 46 of them undersrc/core/migrate/**. That module is now closed — 111 mutants, 0 missed — so what remains is the return-value backlog elsewhere (UvClient::list_pythons -> Ok(vec![])and similar). Those may need a realuvorcondato observe, in which case they belong in.cargo/mutants.tomlwith a rationale, but each needs checking rather than assuming. The next completed weekly run supersedes that artifact; until the triage happens the job tolerates exit 2. - The coverage floor is absolute, not relative.
--fail-under-lines 80is measured against llvm-cov, which reads 80.97% where Codecov reads 78.5%; the two count different things, so a number taken from the Codecov dashboard would be wrong in the workflow. Coverage can still drift from 81% to 80.1% without tripping it — closing that needscodecov/project, see the gap above. - The cache hit its limit. 10.49 GB against the 10 GB allowance on
2026-09-26: 139 caches for
main(about 2 GB of themv0-rust-*keyed to the Cargo.lock hash from before the 0.16.0cargo update) and 118 copies left behind by merged PRs. LRU eviction then ran while a BuildKit export was writing layers and failed a green Docker Integration build witherror writing layer blob: not_found(#198). Three changes: every saving step (Swatinem/rust-cachesave-if, BuildKitcache-to) now writes only frommain— PR runs restoremain’s entries and leave no copy; BuildKit exports areignore-error=true, so a failed cache write is a warning rather than a failed build; andcache-cleanup.ymldeletes whatever a PR still leaves when it closes. The one-time deletion of the PR copies and the stale-lockfile rust caches brought usage to 6.4 GB.mainstill kept every lockfile generation of each rust cache (up to four per job, about 4.4 GB, by 2026-10-04); a weekly job incache-cleanup.ymlnow keeps only the newest entry under each restore key (the cache key minus its lockfile hash), the one a later run with that toolchain and environment falls back to.
Recently closed
Left here because the reasoning is worth keeping, not because anything is outstanding.
- Docs were only verified on release tags: the MSRV 1.89 bump edited pages
under
docs/src/, its PR went green, and thev0.15.4tag then failed at theko.poround-trip. The crate published and the tag was fine — only the Pages deploy stopped, which is the quiet half of the failure and the reason it went unnoticed until someone opened the site. Closed bydocs-check.yml, which runs the build and the round-trip on PRs that touch the docs paths. - Coverage uploads were rejected for months with
Token required because branch is protectedwhile the job reported success, becausefail_ci_if_error: falsehid it. The org allows tokenless uploads, but that path only covers fork PRs. Fixed by wiringCODECOV_TOKENand letting a rejected upload fail the job. That fix was half of it: Dependabot-triggered runs read a different secret store, so every Dependabot PR kept failing the same way until the token was registered there too. - Every release PR failed the Lint job because release-plz bumps
Cargo.tomlwithout touching the version samples in prose. Hand-fixed at v0.15.3, hit again at v0.15.4, and now handled by the sync step described under Releases — the direction the v0.15.3 fix asked for: teach the release to update them rather than loosen the check. - Dependabot does not read
rust-version, so it raisedrust-i18npast the MSRV and every job died at dependency resolution before a line compiled. The group could not land without the MSRV moving to 1.89, which it did..github/dependabot.ymlnow carries anignorefor the next known case (serial_test4.x needs rustc 1.93.1), marked as debt to drop when the MSRV catches up. - The weekly full mutation run had never once completed — 60 minutes killed it every time, and GitHub reports a timed-out job as cancelled, which reads as benign. Measured at 385/446 mutants in 60 minutes, so the timeout is now 90.
aquasecurity/trivy-action@masterandastral-sh/setup-uv@v7are now pinned to release tags. Dependabot could not follow either: it cannot bump a branch ref, and setup-uv stopped publishing moving major tags at v8.stable-cihad four writers and, becauserust-cachehashes workflow-levelenv, was never actually shared withcoverage.ymlormutants.ymlat all. Each now names the key it really uses.release-plz.ymlhad no concurrency group. It has one now, withcancel-in-progress: false— which delivers the “must always complete” intent that omitting the block only half-achieved.