Skip to content

Releases: jdx/tak

v0.0.14: Per-benchmark gates, tak detect, tak log, local baselines and commit backfill

Choose a tag to compare

@jdx jdx released this 02 Oct 10:19
Immutable release. Only release title and notes can be modified.
e375937

tak compare gating can now be configured per benchmark, with report-only benchmarks, an absolute instruction floor and named acceptances. The gate policy is now read from the base revision, and a comparison that compared nothing now fails. New commands and flags cover regressions that already landed on main (tak detect), history over time (tak log), local before/after measurement (tak run --baseline), per-function attribution (tak explain), building and measuring past commits (tak backfill --commits), and two kinds of extra metrics that are reported but never gate. tak is pre-v1: every new flag, key, report wording and file layout below may change. Read Breaking Changes before upgrading CI.

Highlights

  • Gating is more precise and harder to get around. Each benchmark can have its own gate, a specific regression can be accepted by name, and a pull request can no longer loosen its own gate by editing tak.toml.
  • More ways to read history. tak detect flags steps that reached main, tak log shows each series over time, and tak backfill --commits builds past commits so there is history to compare against.
  • More context next to instruction counts. Per-function profiles, custom metrics such as binary size, and DHAT heap-allocation counts are recorded and reported. Only instruction counts gate.

Added

Gating

  • Per-benchmark gates, report-only benchmarks and an instruction floor. A benchmark or subject can set its own gate table, which overrides [gate] key by key. A regression now has to exceed both pct and the new min_delta floor. The default floor is 0, so existing verdicts don't change. enabled = false makes a benchmark report-only: a rise past its threshold is marked (not gated) and never fails the run. The floor can also be set with --gate-min-delta / TAK_GATE_MIN_DELTA. A benchmark's own gate wins over --gate-pct. When any gate differs from [gate], the report adds a gate column and the verdict reads N benchmark(s) above their gate. Otherwise the report is byte-for-byte unchanged. tak now rejects invalid gate values (a negative, NaN or infinite pct, a negative or fractional min_delta, unknown keys) when it loads tak.toml. By @jdx in #158.

    [gate]
    pct = 1.0
    min_delta = 0
    
    [bench.startup]
    cmd = ["./target/release/mycli", "--version"]
    gate = { pct = 5.0, min_delta = 20000 }
    
    [bench.help]
    cmd = ["./target/release/mycli", "--help"]
    gate = { enabled = false }   # report only
  • Accept a regression in named benchmarks. tak compare --accept BENCH waives the gate for one exact benchmark name. It is repeatable, and every other benchmark still gates. Accepted rows are marked (accepted) and listed with their source. The CI guide shows how to drive the flag from a tak-accept:NAME label that only maintainers can apply. Tak-Accept: BENCH commit trailers in BASE..REV are honoured only when [gate] accept_trailers = true or TAK_ACCEPT_TRAILERS=1 is set. When the setting is off, the report lists any trailers it found and says they were ignored. Each trailer names exactly one benchmark using its whole value, commas included, so repeat the trailer to accept several. By @jdx in #153 and #164.

  • allow_empty setting. With [gate] allow_empty = true, TAK_ALLOW_EMPTY=1 or --allow-empty, an empty comparison passes with a warning on stderr instead of failing. Under GitHub Actions the warning is a ::warning:: annotation. Regressions still fail. --allow-empty is now a global flag, so it also works with tak run --baseline --gate. tak compare reads the setting from the base revision's tak.toml. With jdx/tak-action, also set fail-on-nothing-compared: false to warn instead of fail. By @jdx in #173.

New commands

  • tak detect checks for instruction-count regressions that already reached main. Run it after tak run --record and tak push. It walks REV's first-parent history (--window, default 20 recorded commits) and compares consecutive recorded points in each series:

    • It fails only on a step onto REV that exceeds that series' gate, so a main workflow fails once, on the push that introduced the step.
    • A step across unrecorded commits is shown as a git range.
    • Earlier steps and sustained drift are listed but don't fail.
    • It supports per-benchmark gates, --accept, trailers, --no-gate and --allow-empty. It fails when nothing was compared, including on a first recording.
    • It needs full history (fetch-depth: 0).

    By @jdx in #154.

    - name: Detect regressions that landed
      run: |
        set -o pipefail
        tak detect | tee -a "$GITHUB_STEP_SUMMARY"
  • tak log shows each benchmark's history over REV's first-parent commits. It prints one table per bench/tool/runner series, with the instruction-count Δ from the previous measurement and the wall time. -n (default 30) counts recorded commits. --bench filters, and an unknown name is an error. --html PATH writes a self-contained HTML report with SVG charts and no scripts or external assets. The CI guide includes a GitHub Pages example. tak history output is unchanged. By @jdx in #160 and #166.

  • tak explain and tak run --profile-dir show which functions an instruction-count change came from. --profile-dir DIR keeps the cachegrind profile from the run that gave the reported minimum, at DIR/<bench>/<subject>.cachegrind.out. tak explain BASE HEAD matches two such directories and lists the functions whose self cost changed most (--top N, default 10). It only reports and always exits 0. Functions are matched by name, and stripped binaries show up as ???. Profiles are not stored in notes, so the base profile has to come from measuring the base yourself. The new attribution guide describes two ways to set that up in CI. By @jdx in #157.

    tak run --bench startup --profile-dir base -- ./mycli hello
    # rebuild with the change
    tak run --bench startup --profile-dir head -- ./mycli hello
    tak explain base head

Measuring

  • Local baselines. tak run --save-baseline NAME saves a run's records to <git-common-dir>/tak/baselines/NAME.jsonl, which is never committed or pushed and is shared across worktrees. tak run --baseline NAME measures again and prints the tak compare table, without writing notes. Saving merges by series, so --bench a --save-baseline x updates only a. With --gate, the run exits non-zero on a regression past the benchmark's gate. It also fails when it can't check everything it was asked to: nothing ran, a count failed, a series is missing from the baseline or was saved only on another runner class, or a check failed. When everything measured is report-only, --gate passes and prints a note saying nothing was gated. An older tak refuses to save over a baseline that holds newer-schema records. By @jdx in #155 and #165.

    tak run --save-baseline before
    # edit, rebuild
    tak run --baseline before --gate
  • tak backfill --commits RANGE builds and measures past commits. Each first-parent commit, newest first, is checked out into a temporary detached worktree. tak runs the new [build] table there, measures the benchmarks from the current tak.toml, and appends the results to that commit's note, with the commit date as ts.

    • Subjects already recorded for the runner class are skipped. --force measures again, adding a second record.
    • --limit (default 20) caps builds per run, so rerunning continues where the last run stopped.
    • Build failures and per-benchmark measurement failures are remembered locally and passed over on later runs. --force, a [build] change, a runner-class change or (for measurement failures) any edit to tak.toml retries them. A failed instruction count records nothing for that benchmark.
    • --dry-run lists commits and builds nothing.
    • Programs and directories from the checkout must resolve inside it. tak checks them before and after each setup, prepare, sample, check and valgrind run.
    • On Unix, an interrupt kills only tak's own build and measured process groups, removes the worktrees and exits with the signal's status.

    Backfilled numbers are only comparable with CI's if they were built the same way. By @jdx in #159, #171 and #172.

    [build]
    cmd = ["cargo", "build", "--release", "--locked"]
    tak backfill --commits main~20..main --dry-run
    tak backfill --commits main~20..main
    tak push
  • Custom metrics. A metric.NAME table on a benchmark, subject or [defaults] records either a file's size in bytes (file) or the single non-negative number a command prints (cmd). Metrics appear in run output, --export-json and notes. tak compare shows them in a separate table and never gates on them. Names must match ^[a-z][a-z0-9_]*$. instructions, wall_* and alloc_* are reserved. A metric that can't be taken fails the run, and nothing is recorded. By @jdx in #156.

    [bench.startup.metric.binary_bytes]
    file = "target/release/mycli"
  • Heap allocations via valgrind DHAT. allocations = true (on a benchmark, subjec...

Read more

v0.0.13: Shared settings and templates in tak.toml, plus setup, check and ok_exit_codes

Choose a tag to compare

@jdx jdx released this 24 Sep 21:20
Immutable release. Only release title and notes can be modified.
f2f2444

tak.toml can now declare settings and programs once and reuse them across benchmarks, and its values can be tera templates. Benchmarks also gain a one-time setup, a per-sample check, when conditions and ok_exit_codes. tak run adds --config and --dry-run, and --export-json now records how and where a run was produced. All of this is pre-v1 and the keys, output and export format may change.

Added

  • [defaults], shared subjects and templates. A [defaults] table sets values every benchmark starts from. A top-level [subject.NAME] declares a program once, and benchmarks measure it by listing it in subjects = [...]. Values in cmd, prepare, dir, env and vars are tera templates, with env, bench, subject and vars available, so commands that need a path known only at run time no longer need sh -c. By @jdx in #135.

    [defaults]
    runs = "auto"
    dir = "{{ env.BENCH_DIR }}/project-{{ subject }}"
    
    [subject.mycli]
    cmd = ["{{ env.MYCLI_BIN }}", "install", "--lockfile", "{{ vars.lockfile }}"]
    vars = { lockfile = "mycli.lock" }
    
    [subject.othertool]
    cmd = ["othertool", "install"]
    
    [bench.warm]
    subjects = ["mycli", "othertool"]
    
    [bench.cold]
    subjects = ["mycli", "othertool"]
    [bench.cold.subject.mycli]     # override a shared subject for one benchmark
    cmd = ["{{ env.MYCLI_BIN }}", "install", "--no-cache"]
    • Settings stack as defaults, then benchmark, then shared subject, then the benchmark's own subject table. The most specific value wins, except env and vars, which merge key by key.
    • vars are template-only values and aren't passed to the command.
    • tak checks template syntax for the whole file when it loads. Values are rendered only for the benchmarks being run, before any sample. An undefined variable is an error that names the benchmark and subject, not an empty string.
  • tak run --dry-run prints each selected subject exactly as it would run: layers applied, templates rendered, paths anchored and command-line overrides such as --warmup taken into account. It measures nothing. By @jdx in #136.

  • tak run --config PATH reads the named file instead of searching upward for tak.toml, including its [env], [gate] and [runner] settings. Commands resolve relative to that file. A missing file is an error. By @jdx in #136.

  • JSON schema for tak.toml. tak now publishes a schema at /schema/tak.json on the tak docs site (source: docs/public/schema/tak.json). Point a #:schema comment at it and taplo-based editors (such as Even Better TOML) will complete keys and flag unknown ones, which catches typos the parser ignores. By @jdx in #136.

  • when conditions. A benchmark or subject can declare an expr condition. If it's false, that benchmark or subject is skipped, and tak says so on stderr. Conditions can use env, os, arch, ci, bench and subject. They're evaluated before templates render, so a skipped subject's variables don't have to be set. Passing --subject for a subject whose when is false is an error. ?? binds more loosely than !=, so a fallback needs parentheses, as below. By @jdx in #138.

    [subject.vlt]
    when = '(env.VLT_BIN ?? "") != ""'
    cmd = ["{{ env.VLT_BIN }}", "install"]
  • Warnings about suspect samples. After each subject with at least 5 samples, tak warns on stderr about outliers (modified z-score above 3.5) and about a first timed sample that took more than twice the median of the rest. Samples are never dropped or adjusted. By @jdx in #138.

  • setup: a one-time untimed step per subject. It runs once for each measured subject, in name order, before the benchmark's first warmup. It always runs from the directory containing tak.toml, never from dir, so it can create or recreate dir. It layers like prepare, supports templates, and doesn't count toward budget or the progress estimate. A failing setup drops its subject, as a failing prepare does. A shared subject listed by two benchmarks is set up once per benchmark. By @jdx in #141.

  • check: verify every timed sample. A check command runs untimed after each timed sample, but not after warmups or cachegrind runs. The summary line shows the pass count (checks 18/20), and --export-json adds a per-result checks object with passed, total and per-sample samples in the same order as times. A failing check keeps the sample, warns on stderr and doesn't change the exit status. With --record, though, any failed check means nothing is written and the run exits non-zero. A check that can't be spawned at all drops the subject. By @jdx in #142.

    [bench.fix]
    dir = "fixture"
    prepare = ["git", "reset", "--hard", "--quiet", "dirty"]
    check = ["git", "diff", "--quiet", "clean"]
  • ok_exit_codes lists the exit codes that count as success for a benchmark or subject, for tools that exit non-zero by design, such as pre-commit, linters or grep. The default is [0]. The most specific list replaces the default, so [1] rejects exit 0. It applies to warmups, timed samples and cachegrind runs. setup, prepare and check must still exit 0, and a command killed by a signal always fails. On Unix, tak refuses to run a selected subject whose list has no code in 0–255. By @jdx in #143.

    [bench.pre-commit]
    cmd = ["pre-commit", "run", "--all-files"]
    ok_exit_codes = [0, 1]
  • Run details in --export-json. The export gains the top-level keys tak_version, seed (as a string), runner, time, and machine. machine holds OS, OS version, kernel, arch, CPU model, cpus and memory_bytes. cpus counts the logical CPUs tak could actually use, so it honors taskset, cpusets and cgroup quotas. Any value tak can't read is null. Hyperfine's shape is unchanged, and none of this is written to git-notes records. Windows detection is best effort and untested in CI. By @jdx in #136 and #144.

  • version_cmd records each subject's version. It runs once per subject after setup and before warmups, in the subject's dir and env, with a 10-second limit. The first non-empty line of stdout (or of stderr, if stdout is empty) is exported as that result's version and stored in the record's version field with --record. If the command fails, prints nothing or times out, tak prints a warning and exports "version": null, but still measures the subject. By @jdx in #144.

Changed

  • Values containing literal {{, {% or {# are now treated as tera templates. Other existing tak.toml files parse and behave as before. By @jdx in #135.
  • The exit_codes field in --export-json now holds each sample's real exit code. Before, it was always 0. By @jdx in #143.
  • The release binary grows from about 2.7 MB to about 5 MB, mostly because of the new tera and expr-lang dependencies.

Fixed

  • tak run no longer hangs forever when a prepare, setup or check step exits but leaves a background process holding its stderr open, such as sh -c './server >&2 & exit 0'. tak now moves on as soon as the step's own process exits. It doesn't kill or wait for processes the step leaves behind. A leftover process's stderr goes to an anonymous temp file for as long as it runs, so redirect heavy logging elsewhere. By @jdx in #145.

Breaking Changes

  • If any value in tak.toml contains literal {{, {% or {#, it's now rendered as a template. Escape it with tera syntax (for example {{ "{{" }}) to keep the literal text.

Full Changelog: v0.0.12...v0.0.13

v0.0.12: Progress and time estimates in tak run, and runs = "auto"

Choose a tag to compare

@jdx jdx released this 24 Sep 17:38
Immutable release. Only release title and notes can be modified.
e065738

tak run now shows progress on stderr while it measures, with an estimate of the time left. A new runs = "auto" setting picks each subject's number of runs based on how long its samples take. Multi-subject results also print more compactly. As with the rest of tak before v1, all of this may change.

Added

  • Progress while measuring. Until now, a long comparison printed nothing until it finished. Now progress goes to stderr. In a terminal, a single bar redraws in place:

      install [###########-------------------] 12/33  36%  2s elapsed, ~3s left  slow
    

    Anywhere else, such as a CI log, tak prints a plain line each time another tenth of the run is done, or every 30 seconds. The 30-second lines keep coming even during one long sample or prepare step. To estimate the time left, tak multiplies each subject's remaining samples by that subject's own average sample time, including prepare. That way a 20 s cold install and a 300 ms one in the same run are each counted at their own speed. --no-progress turns progress off. By @jdx in #132.

  • runs = "auto". Programs in one comparison can differ in speed by 50x, and no single fixed runs suits both. With runs = "auto", tak times each subject after its warmups. It then gives the subject as many runs as its fastest sample fits into budget, but never fewer than min_runs or more than max_runs. With the defaults, a 1.7 s install gets 17 runs and a 21 s one gets 5. By @jdx in #132.

    [bench.install]
    runs = "auto"
    budget = "30s"   # wall time per subject, prepare included (default 30s)
    min_runs = 5     # default 5
    max_runs = 50    # default 50
    • You can set budget, min_runs and max_runs on the benchmark or on a single subject, and a subject can still fix its own runs. budget accepts ms, s, m and h; a bare number means seconds.
    • A subject with warmup = 0 is sized from its first real sample, and that sample counts toward its runs.
    • tak run --runs auto switches a benchmark to auto for one run. --runs still accepts a number.
    • The counts depend on measured timings, so --seed repeats the same sampling order only if the counts come out the same.
    • runs still defaults to a fixed 20.
    • tak rejects these when it parses tak.toml: a runs word other than auto, a budget it can't parse or that isn't positive, min_runs = 0, and min_runs greater than max_runs.

Changed

  • Multi-subject summary output. Multi-subject benchmarks used to print each subject's full command followed by one metric per line. Now each subject gets one line with min, p50, mean ± stddev, max and the sample count. If a subject has instruction counts enabled, they appear on the same line. Single-command benchmarks print the same as before. If you parse this output, update your parser, or use --export-json instead. By @jdx in #132.

      install: 2 subjects, interleaved (--seed 3)
        fast  min    102.62  p50    106.99  mean    106.49 ± 1.50     max    108.41 ms  n=28
        slow  min    803.47  p50    807.21  mean    806.24 ± 2.44     max    808.05 ms  n=3
    

Full Changelog: v0.0.11...v0.0.12

v0.0.11: Compare several commands in one benchmark with interleaved samples

Choose a tag to compare

@jdx jdx released this 24 Sep 15:04
Immutable release. Only release title and notes can be modified.
e139140

tak run can now compare several programs in one benchmark. It takes samples from each program in turn, in a new random order every round, so drift during the run is shared across all of them. This release also adds a tak completion command and ships shell completions with each release. Both features are pre-v1 and may change.

Added

  • Multi-subject benchmarks. Add [bench.NAME.subject.SUBJECT] tables to compare programs within one benchmark. Each round takes one sample of every subject, in a newly shuffled order. When runs are sequential, as with hyperfine, host contention, thermal throttling or caches filling up get blamed on whichever program happened to be running. Interleaving spreads that drift across all of them. By @jdx in #131.

    [bench.install]
    runs = 10
    warmup = 1
    prepare = ["sh", "-c", "rm -rf node_modules"]   # untimed, before every sample
    
    [bench.install.subject.mycli]
    cmd = ["./target/release/mycli", "install"]
    dir = "fixtures/mycli"
    
    [bench.install.subject.othertool]
    cmd = ["othertool", "install"]
    dir = "fixtures/othertool"
    env = { HOME = "/tmp/othertool-home" }
    runs = 5    # spread evenly across the run
    • Each subject has its own cmd and can override the benchmark's prepare, dir, env, runs and warmup. Each subject is recorded as a separate series, with the subject name as its tool.
    • prepare, dir and env also work on ordinary single-command benchmarks. prepare runs untimed before every sample, including warmups, and doesn't go through a shell. env is applied after the variables in env.deny are removed.
    • New tak run flags:
      • --seed N repeats a sampling order. Every multi-subject run prints the seed it used.
      • --subject NAME measures only the named subjects.
      • --export-json PATH writes every sample in hyperfine's --export-json format, plus bench and subject fields. It leaves out user and system.
    • Subjects are measured by wall clock only unless you set counters = true, so an upgrade to another program can't trip the instruction-count gate. When counters are on, prepare also runs before each cachegrind run.
    • If one subject fails, tak drops it and keeps measuring the others. The run then exits non-zero and --record writes nothing, so an incomplete set of results isn't recorded as if it were complete. The JSON export still includes the subjects that succeeded.
    • Records now include a wall_stddev_ms metric, and runs print it. This applies to single-command benchmarks too. The record schema version is unchanged.
  • Shell completions. tak completion <SHELL> prints a self-contained completion script for bash, zsh, fish, powershell/pwsh, elvish and nu/nushell. Releases now also publish tak.usage.kdl and completion scripts for Bash, Zsh, Fish and PowerShell as signed resources in the packslip, so installers can set up completions without running tak. By @jdx in #99.

    tak completion zsh > ~/.zfunc/_tak

Changed

  • tak.toml validation is stricter. tak now rejects a config with runs = 0 when it reads the file. It also rejects the new mistakes this release makes possible: a benchmark that sets both cmd and subject, or an empty prepare. Otherwise, single-command benchmarks are measured, printed and recorded the same as before. By @jdx in #131.

Full Changelog: v0.0.10...v0.0.11

v0.0.10: Signed packslip with each release

Choose a tag to compare

@jdx jdx released this 05 Sep 17:50
Immutable release. Only release title and notes can be modified.
58cf5a7

A small release. The one user-facing change publishes a signed provenance document alongside each release's archives. Everything else is CI, docs, and dependency maintenance.

Added

  • Each release now publishes packslip.sigstore.json next to the archives and SHA256SUMS. Where SHA256SUMS only proves the files did not change in transit, the packslip is a single keyless-signed document that additionally records, for every archive, its sha256 and sha512, the executable inside it (tak, or tak.exe on Windows), the shared objects the build loads from the host (musl builds come out static; gnu builds name what they need), a per-file build-provenance attestation linked by digest, and the source repo, tag, and commit. It is signed with the release workflow's OIDC identity, so a consumer — including tak backfill picking a download — can verify an asset against the identity github.com/jdx/tak rather than a key this project would have to distribute. By @jdx in #92.

Full Changelog: v0.0.9...v0.0.10

v0.0.9: Trusted measurement handoff

Choose a tag to compare

@jdx jdx released this 31 Aug 13:00
Immutable release. Only release title and notes can be modified.
a3d3fc3

A small release. The one user-facing change adds a tak artifact command pair so a read-only measurement job can hand its results to a separate job that holds repository-write credentials. Everything else is CI and dependency maintenance.

Added

  • tak artifact export and tak artifact publish split measuring from publishing. A persistent or self-hosted runner that never receives a write token exports its local git-notes records into a bounded, versioned JSON file, and a separate trusted job validates and publishes it. By @jdx in #72.

    # in the read-only measurement job
    tak artifact export --output tak-measurement.json
    
    # in the trusted publishing job
    tak artifact publish tak-measurement.json --expect "$GITHUB_SHA"

    The publisher never trusts the artifact to choose its target commit: --expect supplies the revision independently from the workflow, and a file that names a different commit is rejected. On import, records are re-validated (unknown fields, wrong schema versions, non-positive instruction counts, and non-canonical lines are all rejected) and then written through the same non-forced, cat_sort_uniq merge-and-retry push path as tak push, so existing records and concurrent writers are preserved. Export enforces a 16 MiB / 10,000-record limit. The artifact format is pre-v1 and may change between releases.

Full Changelog: v0.0.8...v0.0.9

v0.0.8: Reject failed benchmark subjects

Choose a tag to compare

@jdx jdx released this 26 Aug 20:08
Immutable release. Only release title and notes can be modified.
a843bda

A small release. The one user-facing change hardens how tak runs the commands it measures so a crashing subject can no longer be recorded as a fast run, and tightens the backfill working directory. Everything else in this release is CI and dependency maintenance.

Fixed

  • Benchmark runs now fail when the measured command exits with a nonzero status, in both the wall-clock timing path and the Cachegrind/Valgrind path. Previously a subject that started and then exited unsuccessfully could still emit an instruction summary, so a crash or fast failure could be recorded as an improvement. By @jdx in #63.
  • tak backfill now creates its working directory as a private, randomized temporary directory (mode 0700 on Unix) instead of a predictable /tmp/tak-backfill-<pid> path. The old path let another local user pre-create it and swap out a downloaded executable before tak spawned it. Downloads are still cleaned up on every exit path. By @jdx in #63.

Changed

  • Documentation and the env_deny help text now state plainly that scrubbing environment variables only removes them from the subject's direct environment — it is not a credential sandbox. A hostile binary can still inspect accessible same-user processes and files, so tak backfill should run on an isolated, credential-free machine when the release binaries it downloads are not trusted. By @jdx in #63.

Full Changelog: v0.0.7...v0.0.8

v0.0.7: CLI and settings move onto usage

Choose a tag to compare

@jdx jdx released this 23 Aug 00:35
Immutable release. Only release title and notes can be modified.
280c83f

A small release. Two internal refactors move the CLI and settings machinery onto usage, and both carry a handful of user-visible behavior changes worth noting.

Changed

  • CLI parsing moved from clap to usage. Subcommands, flags, and field types are unchanged, but unknown flags now error instead of being accepted. The usage spec is now emitted directly from the CLI's static metadata. By @jdx in #47.

  • Settings resolution moved from settings.toml codegen plus build.rs to a single #[derive(usage::Config)] on the Settings struct. Precedence is the same (CLI > env > tak.toml > default), and the same corner cases still hold: a blank runner value defers per-layer rather than recording under an empty class, TAK_ENV_DENY= is still a deliberate empty list, and a malformed tak.toml is still an error that names the file rather than a silent fallback. tak.toml remains tak's general config file, and only the dotted keys the registry binds (env.deny, gate.pct, report.credit, …) are read, so [bench] is never scanned. By @jdx in #48.

    Boolean env var parsing changed as a side effect: TAK_CREDIT now follows usage-config's spelling — true/1/yes/y/on and false/0/no/n/off, lowercase and untrimmed. Values like TRUE or 1 that used to be accepted now warn and fall through to the default, and TAK_CREDIT= now means false.

Added

  • A generated configuration reference page in the docs, and tak usage now emits the settings config block alongside the CLI spec. By @jdx in #48.

Breaking Changes

  • Unknown flags now cause an error instead of being tolerated. If you were passing flags tak does not recognize, they will now fail.
  • TAK_CREDIT boolean values outside usage-config's accepted spellings (for example TRUE or a value with surrounding whitespace) now warn and fall through to the default instead of being honored. TAK_CREDIT= now resolves to false.

Full Changelog: v0.0.6...v0.0.7

v0.0.6: Runner class as a setting

Choose a tag to compare

@jdx jdx released this 19 Aug 12:23
Immutable release. Only release title and notes can be modified.
ba9b34c

A small release. The main change makes the runner class a proper, discoverable setting instead of an undocumented environment variable, and a documentation site goes live.

Added

  • The runner class is now a first-class setting. Series are partitioned on it — tak refuses to compare instruction counts across classes because absolute numbers shift between machine types by more than a real regression does — but until now the only way to set it was the TAK_RUNNER environment variable, which was read straight from the environment: absent from the registry, absent from tak settings, documented nowhere. It is now runner_class, settable via --runner, TAK_RUNNER, or runner.class in tak.toml, and shown by tak settings and tak doctor. An empty value still means derive it (gha-<os>-<arch> under GitHub Actions, local-<os>-<arch> otherwise); whitespace counts as empty at every layer, so --runner "" or an exported-but-blank TAK_RUNNER falls through to the derived name rather than merging every machine into one blank class. Set it explicitly when the environment changes in a way the derived name cannot see — most commonly a toolchain or runner-image bump, which otherwise gets attributed to the code and fails a one-percent gate. Encoding the version into the class starts a fresh series at the bump instead. By @jdx in #25.

    TAK_RUNNER=gha-linux-x64-rust1.85 tak run --record
    

Fixed

  • tak doctor now tolerates a tak.toml it cannot parse. It reports the runner class, so it resolves settings; but the command whose job is diagnosing a broken setup must not itself be stopped by the broken file. It falls back to defaults with a warning while still honoring --runner and the environment, so the class it prints matches what a recording would use. By @jdx in #25.

Documentation

  • A documentation site is now published. By @jdx in #28, with theme-aware logo and favicon fixes in #32 and #33.
  • Experimental disclaimers and "experiment" framing were replaced with plain pre-v1 warnings, and a note on project adoption was added. By @jdx in #34, #37, and #35.

Full Changelog: v0.0.5...v0.0.6

v0.0.5: Trend and outlier fixes for compare reports

Choose a tag to compare

@jdx jdx released this 27 Jul 16:50
Immutable release. Only release title and notes can be modified.
672dd21

A small release with two fixes to how tak compare renders its report. Both were review findings from #22 that were pushed to that branch after it had already merged, so they shipped in v0.0.4.

Fixed

  • The trend sparkline now agrees with the table beside it. gather_trend plotted every record's value while compare folds duplicates to the minimum, so a commit carrying several records for one series — a CI re-run, a retry — contributed several points and a noisy re-run put a spike in the sparkline that the table did not show. Trend assembly now takes the minimum per commit, the same rule compare uses. It also places the compared revision in chronological order when that SHA already appears in the walked history — for example tak compare v1.33.0 --rev v1.30.0, which compares against an ancestor — instead of appending its value at the end and drawing a line that reads left-to-right but is not in time order. By @jdx in #23.

  • The added/removed lists now keep tool identity. Two series differing only by tool rendered identically, so a report could say the same benchmark both started and stopped gating while meaning two different programs. Both lists now name a series by bench, tool (omitted for self), and runner, matching how the table already names it. By @jdx in #23.

Full Changelog: v0.0.4...v0.0.5