Repository navigation
Releases: jdx/tak
Release list
v0.0.14: Per-benchmark gates, tak detect, tak log, local baselines and commit backfill
tak compare gating can now be configured per benchmark, with report-only benchmarks, an absolute instruction floor and named acceptances. The gate policy is now read from the base revision, and a comparison that compared nothing now fails. New commands and flags cover regressions that already landed on main (tak detect), history over time (tak log), local before/after measurement (tak run --baseline), per-function attribution (tak explain), building and measuring past commits (tak backfill --commits), and two kinds of extra metrics that are reported but never gate. tak is pre-v1: every new flag, key, report wording and file layout below may change. Read Breaking Changes before upgrading CI.
Highlights
- Gating is more precise and harder to get around. Each benchmark can have its own gate, a specific regression can be accepted by name, and a pull request can no longer loosen its own gate by editing
tak.toml. - More ways to read history.
tak detectflags steps that reached main,tak logshows each series over time, andtak backfill --commitsbuilds past commits so there is history to compare against. - More context next to instruction counts. Per-function profiles, custom metrics such as binary size, and DHAT heap-allocation counts are recorded and reported. Only instruction counts gate.
Added
Gating
-
Per-benchmark gates, report-only benchmarks and an instruction floor. A benchmark or subject can set its own
gatetable, which overrides[gate]key by key. A regression now has to exceed bothpctand the newmin_deltafloor. The default floor is 0, so existing verdicts don't change.enabled = falsemakes a benchmark report-only: a rise past its threshold is marked(not gated)and never fails the run. The floor can also be set with--gate-min-delta/TAK_GATE_MIN_DELTA. A benchmark's own gate wins over--gate-pct. When any gate differs from[gate], the report adds agatecolumn and the verdict readsN benchmark(s) above their gate. Otherwise the report is byte-for-byte unchanged. tak now rejects invalid gate values (a negative, NaN or infinitepct, a negative or fractionalmin_delta, unknown keys) when it loadstak.toml. By @jdx in #158.[gate] pct = 1.0 min_delta = 0 [bench.startup] cmd = ["./target/release/mycli", "--version"] gate = { pct = 5.0, min_delta = 20000 } [bench.help] cmd = ["./target/release/mycli", "--help"] gate = { enabled = false } # report only
-
Accept a regression in named benchmarks.
tak compare --accept BENCHwaives the gate for one exact benchmark name. It is repeatable, and every other benchmark still gates. Accepted rows are marked(accepted)and listed with their source. The CI guide shows how to drive the flag from atak-accept:NAMElabel that only maintainers can apply.Tak-Accept: BENCHcommit trailers inBASE..REVare honoured only when[gate] accept_trailers = trueorTAK_ACCEPT_TRAILERS=1is set. When the setting is off, the report lists any trailers it found and says they were ignored. Each trailer names exactly one benchmark using its whole value, commas included, so repeat the trailer to accept several. By @jdx in #153 and #164. -
allow_emptysetting. With[gate] allow_empty = true,TAK_ALLOW_EMPTY=1or--allow-empty, an empty comparison passes with a warning on stderr instead of failing. Under GitHub Actions the warning is a::warning::annotation. Regressions still fail.--allow-emptyis now a global flag, so it also works withtak run --baseline --gate.tak comparereads the setting from the base revision'stak.toml. With jdx/tak-action, also setfail-on-nothing-compared: falseto warn instead of fail. By @jdx in #173.
New commands
-
tak detectchecks for instruction-count regressions that already reached main. Run it aftertak run --recordandtak push. It walks REV's first-parent history (--window, default 20 recorded commits) and compares consecutive recorded points in each series:- It fails only on a step onto REV that exceeds that series' gate, so a main workflow fails once, on the push that introduced the step.
- A step across unrecorded commits is shown as a git range.
- Earlier steps and sustained drift are listed but don't fail.
- It supports per-benchmark gates,
--accept, trailers,--no-gateand--allow-empty. It fails when nothing was compared, including on a first recording. - It needs full history (
fetch-depth: 0).
- name: Detect regressions that landed run: | set -o pipefail tak detect | tee -a "$GITHUB_STEP_SUMMARY"
-
tak logshows each benchmark's history over REV's first-parent commits. It prints one table per bench/tool/runner series, with the instruction-count Δ from the previous measurement and the wall time.-n(default 30) counts recorded commits.--benchfilters, and an unknown name is an error.--html PATHwrites a self-contained HTML report with SVG charts and no scripts or external assets. The CI guide includes a GitHub Pages example.tak historyoutput is unchanged. By @jdx in #160 and #166. -
tak explainandtak run --profile-dirshow which functions an instruction-count change came from.--profile-dir DIRkeeps the cachegrind profile from the run that gave the reported minimum, atDIR/<bench>/<subject>.cachegrind.out.tak explain BASE HEADmatches two such directories and lists the functions whose self cost changed most (--top N, default 10). It only reports and always exits 0. Functions are matched by name, and stripped binaries show up as???. Profiles are not stored in notes, so the base profile has to come from measuring the base yourself. The new attribution guide describes two ways to set that up in CI. By @jdx in #157.tak run --bench startup --profile-dir base -- ./mycli hello # rebuild with the change tak run --bench startup --profile-dir head -- ./mycli hello tak explain base head
Measuring
-
Local baselines.
tak run --save-baseline NAMEsaves a run's records to<git-common-dir>/tak/baselines/NAME.jsonl, which is never committed or pushed and is shared across worktrees.tak run --baseline NAMEmeasures again and prints thetak comparetable, without writing notes. Saving merges by series, so--bench a --save-baseline xupdates onlya. With--gate, the run exits non-zero on a regression past the benchmark's gate. It also fails when it can't check everything it was asked to: nothing ran, a count failed, a series is missing from the baseline or was saved only on another runner class, or a check failed. When everything measured is report-only,--gatepasses and prints a note saying nothing was gated. An older tak refuses to save over a baseline that holds newer-schema records. By @jdx in #155 and #165.tak run --save-baseline before # edit, rebuild tak run --baseline before --gate -
tak backfill --commits RANGEbuilds and measures past commits. Each first-parent commit, newest first, is checked out into a temporary detached worktree. tak runs the new[build]table there, measures the benchmarks from the currenttak.toml, and appends the results to that commit's note, with the commit date asts.- Subjects already recorded for the runner class are skipped.
--forcemeasures again, adding a second record. --limit(default 20) caps builds per run, so rerunning continues where the last run stopped.- Build failures and per-benchmark measurement failures are remembered locally and passed over on later runs.
--force, a[build]change, a runner-class change or (for measurement failures) any edit totak.tomlretries them. A failed instruction count records nothing for that benchmark. --dry-runlists commits and builds nothing.- Programs and directories from the checkout must resolve inside it. tak checks them before and after each
setup,prepare, sample,checkand valgrind run. - On Unix, an interrupt kills only tak's own build and measured process groups, removes the worktrees and exits with the signal's status.
Backfilled numbers are only comparable with CI's if they were built the same way. By @jdx in #159, #171 and #172.
[build] cmd = ["cargo", "build", "--release", "--locked"]
tak backfill --commits main~20..main --dry-run tak backfill --commits main~20..main tak push
- Subjects already recorded for the runner class are skipped.
-
Custom metrics. A
metric.NAMEtable on a benchmark, subject or[defaults]records either a file's size in bytes (file) or the single non-negative number a command prints (cmd). Metrics appear in run output,--export-jsonand notes.tak compareshows them in a separate table and never gates on them. Names must match^[a-z][a-z0-9_]*$.instructions,wall_*andalloc_*are reserved. A metric that can't be taken fails the run, and nothing is recorded. By @jdx in #156.[bench.startup.metric.binary_bytes] file = "target/release/mycli"
-
Heap allocations via valgrind DHAT.
allocations = true(on a benchmark, subjec...
v0.0.13: Shared settings and templates in tak.toml, plus setup, check and ok_exit_codes
tak.toml can now declare settings and programs once and reuse them across benchmarks, and its values can be tera templates. Benchmarks also gain a one-time setup, a per-sample check, when conditions and ok_exit_codes. tak run adds --config and --dry-run, and --export-json now records how and where a run was produced. All of this is pre-v1 and the keys, output and export format may change.
Added
-
[defaults], shared subjects and templates. A[defaults]table sets values every benchmark starts from. A top-level[subject.NAME]declares a program once, and benchmarks measure it by listing it insubjects = [...]. Values incmd,prepare,dir,envandvarsare tera templates, withenv,bench,subjectandvarsavailable, so commands that need a path known only at run time no longer needsh -c. By @jdx in #135.[defaults] runs = "auto" dir = "{{ env.BENCH_DIR }}/project-{{ subject }}" [subject.mycli] cmd = ["{{ env.MYCLI_BIN }}", "install", "--lockfile", "{{ vars.lockfile }}"] vars = { lockfile = "mycli.lock" } [subject.othertool] cmd = ["othertool", "install"] [bench.warm] subjects = ["mycli", "othertool"] [bench.cold] subjects = ["mycli", "othertool"] [bench.cold.subject.mycli] # override a shared subject for one benchmark cmd = ["{{ env.MYCLI_BIN }}", "install", "--no-cache"]
- Settings stack as defaults, then benchmark, then shared subject, then the benchmark's own subject table. The most specific value wins, except
envandvars, which merge key by key. varsare template-only values and aren't passed to the command.- tak checks template syntax for the whole file when it loads. Values are rendered only for the benchmarks being run, before any sample. An undefined variable is an error that names the benchmark and subject, not an empty string.
- Settings stack as defaults, then benchmark, then shared subject, then the benchmark's own subject table. The most specific value wins, except
-
tak run --dry-runprints each selected subject exactly as it would run: layers applied, templates rendered, paths anchored and command-line overrides such as--warmuptaken into account. It measures nothing. By @jdx in #136. -
tak run --config PATHreads the named file instead of searching upward fortak.toml, including its[env],[gate]and[runner]settings. Commands resolve relative to that file. A missing file is an error. By @jdx in #136. -
JSON schema for
tak.toml. tak now publishes a schema at/schema/tak.jsonon the tak docs site (source:docs/public/schema/tak.json). Point a#:schemacomment at it and taplo-based editors (such as Even Better TOML) will complete keys and flag unknown ones, which catches typos the parser ignores. By @jdx in #136. -
whenconditions. A benchmark or subject can declare an expr condition. If it's false, that benchmark or subject is skipped, and tak says so on stderr. Conditions can useenv,os,arch,ci,benchandsubject. They're evaluated before templates render, so a skipped subject's variables don't have to be set. Passing--subjectfor a subject whosewhenis false is an error.??binds more loosely than!=, so a fallback needs parentheses, as below. By @jdx in #138.[subject.vlt] when = '(env.VLT_BIN ?? "") != ""' cmd = ["{{ env.VLT_BIN }}", "install"]
-
Warnings about suspect samples. After each subject with at least 5 samples, tak warns on stderr about outliers (modified z-score above 3.5) and about a first timed sample that took more than twice the median of the rest. Samples are never dropped or adjusted. By @jdx in #138.
-
setup: a one-time untimed step per subject. It runs once for each measured subject, in name order, before the benchmark's first warmup. It always runs from the directory containingtak.toml, never fromdir, so it can create or recreatedir. It layers likeprepare, supports templates, and doesn't count towardbudgetor the progress estimate. A failing setup drops its subject, as a failingpreparedoes. A shared subject listed by two benchmarks is set up once per benchmark. By @jdx in #141. -
check: verify every timed sample. Acheckcommand runs untimed after each timed sample, but not after warmups or cachegrind runs. The summary line shows the pass count (checks 18/20), and--export-jsonadds a per-resultchecksobject withpassed,totaland per-samplesamplesin the same order astimes. A failing check keeps the sample, warns on stderr and doesn't change the exit status. With--record, though, any failed check means nothing is written and the run exits non-zero. A check that can't be spawned at all drops the subject. By @jdx in #142.[bench.fix] dir = "fixture" prepare = ["git", "reset", "--hard", "--quiet", "dirty"] check = ["git", "diff", "--quiet", "clean"]
-
ok_exit_codeslists the exit codes that count as success for a benchmark or subject, for tools that exit non-zero by design, such as pre-commit, linters orgrep. The default is[0]. The most specific list replaces the default, so[1]rejects exit 0. It applies to warmups, timed samples and cachegrind runs.setup,prepareandcheckmust still exit 0, and a command killed by a signal always fails. On Unix, tak refuses to run a selected subject whose list has no code in 0–255. By @jdx in #143.[bench.pre-commit] cmd = ["pre-commit", "run", "--all-files"] ok_exit_codes = [0, 1]
-
Run details in
--export-json. The export gains the top-level keystak_version,seed(as a string),runner,time, andmachine.machineholds OS, OS version, kernel, arch, CPU model,cpusandmemory_bytes.cpuscounts the logical CPUs tak could actually use, so it honorstaskset, cpusets and cgroup quotas. Any value tak can't read isnull. Hyperfine's shape is unchanged, and none of this is written to git-notes records. Windows detection is best effort and untested in CI. By @jdx in #136 and #144. -
version_cmdrecords each subject's version. It runs once per subject aftersetupand before warmups, in the subject'sdirandenv, with a 10-second limit. The first non-empty line of stdout (or of stderr, if stdout is empty) is exported as that result'sversionand stored in the record'sversionfield with--record. If the command fails, prints nothing or times out, tak prints a warning and exports"version": null, but still measures the subject. By @jdx in #144.
Changed
- Values containing literal
{{,{%or{#are now treated as tera templates. Other existingtak.tomlfiles parse and behave as before. By @jdx in #135. - The
exit_codesfield in--export-jsonnow holds each sample's real exit code. Before, it was always 0. By @jdx in #143. - The release binary grows from about 2.7 MB to about 5 MB, mostly because of the new tera and expr-lang dependencies.
Fixed
tak runno longer hangs forever when aprepare,setuporcheckstep exits but leaves a background process holding its stderr open, such assh -c './server >&2 & exit 0'. tak now moves on as soon as the step's own process exits. It doesn't kill or wait for processes the step leaves behind. A leftover process's stderr goes to an anonymous temp file for as long as it runs, so redirect heavy logging elsewhere. By @jdx in #145.
Breaking Changes
- If any value in
tak.tomlcontains literal{{,{%or{#, it's now rendered as a template. Escape it with tera syntax (for example{{ "{{" }}) to keep the literal text.
Full Changelog: v0.0.12...v0.0.13
v0.0.12: Progress and time estimates in tak run, and runs = "auto"
tak run now shows progress on stderr while it measures, with an estimate of the time left. A new runs = "auto" setting picks each subject's number of runs based on how long its samples take. Multi-subject results also print more compactly. As with the rest of tak before v1, all of this may change.
Added
-
Progress while measuring. Until now, a long comparison printed nothing until it finished. Now progress goes to stderr. In a terminal, a single bar redraws in place:
install [###########-------------------] 12/33 36% 2s elapsed, ~3s left slowAnywhere else, such as a CI log, tak prints a plain line each time another tenth of the run is done, or every 30 seconds. The 30-second lines keep coming even during one long sample or
preparestep. To estimate the time left, tak multiplies each subject's remaining samples by that subject's own average sample time, includingprepare. That way a 20 s cold install and a 300 ms one in the same run are each counted at their own speed.--no-progressturns progress off. By @jdx in #132. -
runs = "auto". Programs in one comparison can differ in speed by 50x, and no single fixedrunssuits both. Withruns = "auto", tak times each subject after its warmups. It then gives the subject as many runs as its fastest sample fits intobudget, but never fewer thanmin_runsor more thanmax_runs. With the defaults, a 1.7 s install gets 17 runs and a 21 s one gets 5. By @jdx in #132.[bench.install] runs = "auto" budget = "30s" # wall time per subject, prepare included (default 30s) min_runs = 5 # default 5 max_runs = 50 # default 50
- You can set
budget,min_runsandmax_runson the benchmark or on a single subject, and a subject can still fix its ownruns.budgetacceptsms,s,mandh; a bare number means seconds. - A subject with
warmup = 0is sized from its first real sample, and that sample counts toward its runs. tak run --runs autoswitches a benchmark to auto for one run.--runsstill accepts a number.- The counts depend on measured timings, so
--seedrepeats the same sampling order only if the counts come out the same. runsstill defaults to a fixed 20.- tak rejects these when it parses
tak.toml: arunsword other thanauto, abudgetit can't parse or that isn't positive,min_runs = 0, andmin_runsgreater thanmax_runs.
- You can set
Changed
-
Multi-subject summary output. Multi-subject benchmarks used to print each subject's full command followed by one metric per line. Now each subject gets one line with min, p50, mean ± stddev, max and the sample count. If a subject has instruction counts enabled, they appear on the same line. Single-command benchmarks print the same as before. If you parse this output, update your parser, or use
--export-jsoninstead. By @jdx in #132.install: 2 subjects, interleaved (--seed 3) fast min 102.62 p50 106.99 mean 106.49 ± 1.50 max 108.41 ms n=28 slow min 803.47 p50 807.21 mean 806.24 ± 2.44 max 808.05 ms n=3
Full Changelog: v0.0.11...v0.0.12
v0.0.11: Compare several commands in one benchmark with interleaved samples
tak run can now compare several programs in one benchmark. It takes samples from each program in turn, in a new random order every round, so drift during the run is shared across all of them. This release also adds a tak completion command and ships shell completions with each release. Both features are pre-v1 and may change.
Added
-
Multi-subject benchmarks. Add
[bench.NAME.subject.SUBJECT]tables to compare programs within one benchmark. Each round takes one sample of every subject, in a newly shuffled order. When runs are sequential, as with hyperfine, host contention, thermal throttling or caches filling up get blamed on whichever program happened to be running. Interleaving spreads that drift across all of them. By @jdx in #131.[bench.install] runs = 10 warmup = 1 prepare = ["sh", "-c", "rm -rf node_modules"] # untimed, before every sample [bench.install.subject.mycli] cmd = ["./target/release/mycli", "install"] dir = "fixtures/mycli" [bench.install.subject.othertool] cmd = ["othertool", "install"] dir = "fixtures/othertool" env = { HOME = "/tmp/othertool-home" } runs = 5 # spread evenly across the run
- Each subject has its own
cmdand can override the benchmark'sprepare,dir,env,runsandwarmup. Each subject is recorded as a separate series, with the subject name as itstool. prepare,dirandenvalso work on ordinary single-command benchmarks.prepareruns untimed before every sample, including warmups, and doesn't go through a shell.envis applied after the variables inenv.denyare removed.- New
tak runflags:--seed Nrepeats a sampling order. Every multi-subject run prints the seed it used.--subject NAMEmeasures only the named subjects.--export-json PATHwrites every sample in hyperfine's--export-jsonformat, plusbenchandsubjectfields. It leaves outuserandsystem.
- Subjects are measured by wall clock only unless you set
counters = true, so an upgrade to another program can't trip the instruction-count gate. When counters are on,preparealso runs before each cachegrind run. - If one subject fails, tak drops it and keeps measuring the others. The run then exits non-zero and
--recordwrites nothing, so an incomplete set of results isn't recorded as if it were complete. The JSON export still includes the subjects that succeeded. - Records now include a
wall_stddev_msmetric, and runs print it. This applies to single-command benchmarks too. The record schema version is unchanged.
- Each subject has its own
-
Shell completions.
tak completion <SHELL>prints a self-contained completion script forbash,zsh,fish,powershell/pwsh,elvishandnu/nushell. Releases now also publishtak.usage.kdland completion scripts for Bash, Zsh, Fish and PowerShell as signed resources in the packslip, so installers can set up completions without running tak. By @jdx in #99.tak completion zsh > ~/.zfunc/_tak
Changed
tak.tomlvalidation is stricter. tak now rejects a config withruns = 0when it reads the file. It also rejects the new mistakes this release makes possible: a benchmark that sets bothcmdandsubject, or an emptyprepare. Otherwise, single-command benchmarks are measured, printed and recorded the same as before. By @jdx in #131.
Full Changelog: v0.0.10...v0.0.11
v0.0.10: Signed packslip with each release
A small release. The one user-facing change publishes a signed provenance document alongside each release's archives. Everything else is CI, docs, and dependency maintenance.
Added
- Each release now publishes
packslip.sigstore.jsonnext to the archives andSHA256SUMS. WhereSHA256SUMSonly proves the files did not change in transit, the packslip is a single keyless-signed document that additionally records, for every archive, its sha256 and sha512, the executable inside it (tak, ortak.exeon Windows), the shared objects the build loads from the host (musl builds come out static; gnu builds name what they need), a per-file build-provenance attestation linked by digest, and the source repo, tag, and commit. It is signed with the release workflow's OIDC identity, so a consumer — includingtak backfillpicking a download — can verify an asset against the identitygithub.com/jdx/takrather than a key this project would have to distribute. By @jdx in #92.
Full Changelog: v0.0.9...v0.0.10
v0.0.9: Trusted measurement handoff
A small release. The one user-facing change adds a tak artifact command pair so a read-only measurement job can hand its results to a separate job that holds repository-write credentials. Everything else is CI and dependency maintenance.
Added
-
tak artifact exportandtak artifact publishsplit measuring from publishing. A persistent or self-hosted runner that never receives a write token exports its local git-notes records into a bounded, versioned JSON file, and a separate trusted job validates and publishes it. By @jdx in #72.# in the read-only measurement job tak artifact export --output tak-measurement.json # in the trusted publishing job tak artifact publish tak-measurement.json --expect "$GITHUB_SHA"
The publisher never trusts the artifact to choose its target commit:
--expectsupplies the revision independently from the workflow, and a file that names a different commit is rejected. On import, records are re-validated (unknown fields, wrong schema versions, non-positive instruction counts, and non-canonical lines are all rejected) and then written through the same non-forced,cat_sort_uniqmerge-and-retry push path astak push, so existing records and concurrent writers are preserved. Export enforces a 16 MiB / 10,000-record limit. The artifact format is pre-v1 and may change between releases.
Full Changelog: v0.0.8...v0.0.9
v0.0.8: Reject failed benchmark subjects
A small release. The one user-facing change hardens how tak runs the commands it measures so a crashing subject can no longer be recorded as a fast run, and tightens the backfill working directory. Everything else in this release is CI and dependency maintenance.
Fixed
- Benchmark runs now fail when the measured command exits with a nonzero status, in both the wall-clock timing path and the Cachegrind/Valgrind path. Previously a subject that started and then exited unsuccessfully could still emit an instruction summary, so a crash or fast failure could be recorded as an improvement. By @jdx in #63.
tak backfillnow creates its working directory as a private, randomized temporary directory (mode0700on Unix) instead of a predictable/tmp/tak-backfill-<pid>path. The old path let another local user pre-create it and swap out a downloaded executable before tak spawned it. Downloads are still cleaned up on every exit path. By @jdx in #63.
Changed
- Documentation and the
env_denyhelp text now state plainly that scrubbing environment variables only removes them from the subject's direct environment — it is not a credential sandbox. A hostile binary can still inspect accessible same-user processes and files, sotak backfillshould run on an isolated, credential-free machine when the release binaries it downloads are not trusted. By @jdx in #63.
Full Changelog: v0.0.7...v0.0.8
v0.0.7: CLI and settings move onto usage
A small release. Two internal refactors move the CLI and settings machinery onto usage, and both carry a handful of user-visible behavior changes worth noting.
Changed
-
CLI parsing moved from clap to usage. Subcommands, flags, and field types are unchanged, but unknown flags now error instead of being accepted. The usage spec is now emitted directly from the CLI's static metadata. By @jdx in #47.
-
Settings resolution moved from
settings.tomlcodegen plusbuild.rsto a single#[derive(usage::Config)]on theSettingsstruct. Precedence is the same (CLI > env >tak.toml> default), and the same corner cases still hold: a blank runner value defers per-layer rather than recording under an empty class,TAK_ENV_DENY=is still a deliberate empty list, and a malformedtak.tomlis still an error that names the file rather than a silent fallback.tak.tomlremains tak's general config file, and only the dotted keys the registry binds (env.deny,gate.pct,report.credit, …) are read, so[bench]is never scanned. By @jdx in #48.Boolean env var parsing changed as a side effect:
TAK_CREDITnow follows usage-config's spelling —true/1/yes/y/onandfalse/0/no/n/off, lowercase and untrimmed. Values likeTRUEor1that used to be accepted now warn and fall through to the default, andTAK_CREDIT=now meansfalse.
Added
- A generated configuration reference page in the docs, and
tak usagenow emits the settingsconfigblock alongside the CLI spec. By @jdx in #48.
Breaking Changes
- Unknown flags now cause an error instead of being tolerated. If you were passing flags tak does not recognize, they will now fail.
TAK_CREDITboolean values outside usage-config's accepted spellings (for exampleTRUEor a value with surrounding whitespace) now warn and fall through to the default instead of being honored.TAK_CREDIT=now resolves tofalse.
Full Changelog: v0.0.6...v0.0.7
v0.0.6: Runner class as a setting
A small release. The main change makes the runner class a proper, discoverable setting instead of an undocumented environment variable, and a documentation site goes live.
Added
-
The runner class is now a first-class setting. Series are partitioned on it — tak refuses to compare instruction counts across classes because absolute numbers shift between machine types by more than a real regression does — but until now the only way to set it was the
TAK_RUNNERenvironment variable, which was read straight from the environment: absent from the registry, absent fromtak settings, documented nowhere. It is nowrunner_class, settable via--runner,TAK_RUNNER, orrunner.classintak.toml, and shown bytak settingsandtak doctor. An empty value still means derive it (gha-<os>-<arch>under GitHub Actions,local-<os>-<arch>otherwise); whitespace counts as empty at every layer, so--runner ""or an exported-but-blankTAK_RUNNERfalls through to the derived name rather than merging every machine into one blank class. Set it explicitly when the environment changes in a way the derived name cannot see — most commonly a toolchain or runner-image bump, which otherwise gets attributed to the code and fails a one-percent gate. Encoding the version into the class starts a fresh series at the bump instead. By @jdx in #25.TAK_RUNNER=gha-linux-x64-rust1.85 tak run --record
Fixed
tak doctornow tolerates atak.tomlit cannot parse. It reports the runner class, so it resolves settings; but the command whose job is diagnosing a broken setup must not itself be stopped by the broken file. It falls back to defaults with a warning while still honoring--runnerand the environment, so the class it prints matches what a recording would use. By @jdx in #25.
Documentation
- A documentation site is now published. By @jdx in #28, with theme-aware logo and favicon fixes in #32 and #33.
- Experimental disclaimers and "experiment" framing were replaced with plain pre-v1 warnings, and a note on project adoption was added. By @jdx in #34, #37, and #35.
Full Changelog: v0.0.5...v0.0.6
v0.0.5: Trend and outlier fixes for compare reports
A small release with two fixes to how tak compare renders its report. Both were review findings from #22 that were pushed to that branch after it had already merged, so they shipped in v0.0.4.
Fixed
-
The trend sparkline now agrees with the table beside it.
gather_trendplotted every record's value whilecomparefolds duplicates to the minimum, so a commit carrying several records for one series — a CI re-run, a retry — contributed several points and a noisy re-run put a spike in the sparkline that the table did not show. Trend assembly now takes the minimum per commit, the same rulecompareuses. It also places the compared revision in chronological order when that SHA already appears in the walked history — for exampletak compare v1.33.0 --rev v1.30.0, which compares against an ancestor — instead of appending its value at the end and drawing a line that reads left-to-right but is not in time order. By @jdx in #23. -
The added/removed lists now keep tool identity. Two series differing only by tool rendered identically, so a report could say the same benchmark both started and stopped gating while meaning two different programs. Both lists now name a series by bench, tool (omitted for
self), and runner, matching how the table already names it. By @jdx in #23.
Full Changelog: v0.0.4...v0.0.5