Benchmark configuration
tak searches upward from the working directory for tak.toml. Commands run relative to the directory containing that file, so CI and local runs use the same paths. tak run --config PATH reads a named file instead, including the settings it holds, and commands then run relative to that file. That's useful for a second set of benchmarks kept apart from the main ones.
Editors that read JSON Schema, such as Even Better TOML or taplo, can complete and check tak.toml against https://tak.jdx.dev/schema/tak.json. Add this as the file's first line:
#:schema https://tak.jdx.dev/schema/tak.jsonTo see exactly what a run would do, with defaults and shared subjects applied, templates rendered and command-line overrides taken into account, run it with --dry-run. Nothing is measured:
tak run --bench install --dry-runBenchmarks
Each [bench.NAME] table needs a command:
[bench.startup]
cmd = ["./target/release/mycli", "--version"]
runs = 10
warmup = 2cmd may be an argument list or a whitespace-split string. tak deliberately never starts a shell because shell startup would add work and variance to the subject.
A benchmark or subject name can be any text except control characters, newlines included. Names are printed in reports that CI reads line by line, and a newline in a name could make part of it read as one of tak's verdicts. tak rejects such a name when tak.toml loads, before anything is measured. The same rule applies to --bench and the runner class. It also applies to TAK_TOOL where that becomes the recorded tool name: a single-command benchmark or tak run -- CMD.
Command-line values override the file:
tak run --bench startup --runs 20 --warmup 3Resetting state before each sample
prepare runs before every sample, warmups included, and is not timed. Use it when each sample has to start from the same state, such as an install benchmark that needs an empty node_modules:
[bench.install]
cmd = ["./target/release/mycli", "install"]
prepare = ["sh", "-c", "rm -rf node_modules"]
dir = "fixtures/app"
env = { MYCLI_OFFLINE = "1" }prepare uses the same syntax as cmd, and there is still no implicit shell: write ["sh", "-c", "…"] when you need one. A shell costs nothing here because prepare is outside the measurement. dir is relative to tak.toml and only sets the working directory. A program path containing a /, such as ./target/release/mycli, is still found relative to tak.toml; a bare name like mycli is looked up on PATH. Variables in env are set after tak removes the ones in env.deny, so a variable written here reaches the command even when it is denied by default.
Setting up once before measuring
setup runs once for each subject of a benchmark, before that benchmark takes its first sample, and is not timed. Use it for work every sample needs but that only has to happen once, such as cloning a fixture and committing each tool's configuration to it, or priming a cache:
[defaults]
dir = ".work/{{ subject }}"
setup = ["./bench/setup.sh", "{{ subject }}", ".work/{{ subject }}"]
prepare = ["sh", "-c", "git reset -q --hard && git clean -qfd"]
[subject.hk]
cmd = ["hk", "check", "--all"]
[subject.lefthook]
cmd = ["lefthook", "run", "check", "--all-files"]
[bench.check-all]
subjects = ["hk", "lefthook"]Here setup.sh clones the fixture into .work/hk or .work/lefthook, commits that tool's configuration and runs the tool once to fill its cache. prepare then only has to reset the checkout before each sample.
- It runs in the directory holding
tak.toml, not indir. Setup usually creates or recreatesdir, and it can't start inside a directory that doesn't exist yet or that it's about to delete.diris written relative totak.tomltoo, so the same path works as an argument tosetup, as above. To run something insidedir,cdinto it:["sh", "-c", "cd .work/hk && hk check --all"]. Otherwisesetupis likeprepare: the same syntax, templates, the subject'senv, and no implicit shell. - It runs before the benchmark's sampling starts. Every subject's setup runs, in name order, before the benchmark's first warmup, so no sample shares the machine with a setup. It doesn't count toward
budgetorruns = "auto"'s sizing or the progress estimate. Progress shows which subject is being set up. A subject'sversion_cmdruns after all the setups, so it reports on whatsetupinstalled. - It runs once per benchmark. A shared subject listed by two benchmarks is set up again for the second one, because the first benchmark's samples may have changed its state. Make an expensive setup reuse what it built before when that is safe.
- It only runs for subjects that will be measured. A subject left out with
--subjector switched off bywhenisn't set up, and--dry-runprintssetupwithout running it. - A failing setup drops the subject, as a failing
preparedoes: tak measures the rest, exits non-zero, and--recordwrites nothing. - Its output is hidden. When it fails, tak reports the last line of its stderr. Run the command by hand to see the rest.
Settings stack the same way as prepare: a subject's own setup replaces the benchmark's, which replaces the one in [defaults]. There is no matching teardown. The next run's setup can clean up whatever the last one left, and keeping it around lets you inspect a subject's directory after a run.
Recording each program's version
A comparison has to say which version of each program it measured. version_cmd names a command that prints a subject's version:
[defaults]
runs = "auto"
[subject.mycli]
cmd = ["./target/release/mycli", "run", "pre-commit"]
version_cmd = ["./target/release/mycli", "--version"]
[subject.othertool]
cmd = ["othertool", "run", "pre-commit"]
version_cmd = "othertool version"
[bench.hooks]
subjects = ["mycli", "othertool"]Before sampling, tak runs each subject's version_cmd once, untimed, in the subject's dir and env. It runs after every subject's setup and before any warmup, so a setup that installs or builds the program is done before tak asks for its version. A subject whose setup fails isn't asked. The command runs with the same variables removed, so it sees what the measured command sees. The first non-empty line it prints becomes that subject's version in --export-json. tak reads stdout, or stderr when stdout is empty, as with java -version. --record stores the same string as the measurement's version.
version_cmd layers like prepare, and it's a template, so one line in [defaults] can cover subjects that answer the same flag:
[defaults]
version_cmd = ["{{ env.BIN_DIR }}/{{ subject }}", "--version"]A version_cmd that exits non-zero, prints nothing, or takes longer than 10 seconds doesn't drop the subject. tak stops one that runs too long, prints a warning, measures the subject as usual, and exports its version as null. On Linux and macOS, stopping it also stops any processes it started, so nothing it left behind runs during the samples. The same happens if tak is interrupted (Ctrl-C, or SIGTERM or SIGHUP) while a version_cmd runs. On Windows only the command itself is stopped. A non-zero exit always means null, even when the command printed something that looks like a version: a failing command's output is an error or a usage message. ok_exit_codes doesn't apply to version_cmd. Only the first 8 KiB of each output stream is kept; the rest is read and discarded, so a long banner doesn't stop the command from finishing. If the command leaves a background process holding its output open, tak uses what arrived before the command exited instead of waiting for that process. tak doesn't stop a background process left by a command that finished in time, because a tool may start a daemon on purpose. A subject without version_cmd has no version key. --dry-run lists each subject's version_cmd.
Accepting other exit codes
tak drops a subject when its command exits with anything but 0, because a failed run usually didn't do the work being measured. Some programs exit non-zero by design: pre-commit exits 1 whenever a hook modifies files, linters and test runners exit 1 when they find problems, and grep exits 1 when nothing matches. List the codes that count as success with ok_exit_codes:
[bench.pre-commit]
cmd = ["pre-commit", "run", "--all-files"]
prepare = ["git", "checkout", "--", "."]
ok_exit_codes = [0, 1] # 1: a hook modified files, which is the case being measured- The default is
[0]. A list replaces the default rather than adding to it, so leave 0 out to require a non-zero code, such as[1]for agrepthat must not match. - It applies to warmups, timed samples and the instruction-count run under valgrind. Any other code still drops the subject, and so does a command killed by a signal, whatever the list holds.
setupandpreparemust still exit 0. A failed setup or reset would leave every later sample starting from the wrong state. Acheckpasses only when it exits 0, too, and aversion_cmdthat exits non-zero always exports anullversion.--export-jsonrecords each sample's real exit code inexit_codes.- The list can't be empty, and duplicates are ignored. Unix only ever reports codes 0 to 255. Windows passes a program's 32-bit exit code through as a signed number, so write an NTSTATUS such as
0xC0000005as its negative decimal value,-1073741819. - A
tak.tomlshared between platforms can list both, such as[0, 1, -1073741819]. On Unix, tak warns about the codes that can never match there and runs with the rest. If none of a subject's codes can match on the current platform,tak runfails before anysetupor sample runs, naming the benchmark and subject.--dry-runreports the same warning or error. Both checks only cover the subjects being run, afterwhen,--benchand--subjectare applied, so a Windows-only subject switched off withwhen = 'os == "windows"'doesn't stop the rest of the file from running on Unix.
Checking every sample
A fast time is only worth reporting if the command did the right work. check runs after every timed sample, untimed, in the same directory and environment as the command. Exit status 0 means the sample passed; anything else means it failed, whatever ok_exit_codes allows the command. Use it when a command's output can be wrong some of the time, such as formatters that race when run concurrently on the same files:
[bench.fix]
runs = 20
warmup = 1
dir = "fixture" # a git repository with tags `dirty` (unformatted) and `clean` (the expected result)
prepare = ["git", "reset", "--hard", "--quiet", "dirty"]
check = ["git", "diff", "--quiet", "clean"]
# Two fixers one after the other.
[bench.fix.subject.serial]
cmd = ["sh", "-c", "sed -i 's/foo/bar/' a.txt; sed -i 's/ *$//' a.txt"]
# The same two fixers at once, on the same file.
[bench.fix.subject.parallel]
cmd = ["sh", "-c", "sed -i 's/foo/bar/' a.txt & sed -i 's/ *$//' a.txt & wait"]git diff --quiet clean exits 1 when the working tree differs from the clean commit. In this run the concurrent fixers lost an edit every time, and were faster for it:
fix: 2 subjects, interleaved (--seed 1)
warning: fix (parallel): check failed after 20 of 20 samples (sample 1, 2, 3, 4, 5, 6, 7, 8, …); first: check `git` exited with exit status: 1
parallel min 1.96 p50 2.13 mean 2.15 ± 0.15 max 2.47 ms n=20 checks 0/20
serial min 2.49 p50 2.72 mean 2.78 ± 0.21 max 3.24 ms n=20 checks 20/20- A failing check doesn't drop the subject, and without
--recordit doesn't fail the run. Every sample is kept, and the pass count is the result: a race that loses one sample in twenty shows up aschecks 19/20on the subject's summary line. tak warns on stderr, listing the failed sample numbers and the last line the first failing check wrote to stderr. The summary doesn't say which time came from which sample, so it can't tell you whether the minimum is from a failed one. The export can. - A check that can't be started at all, such as a mistyped program, is a mistake in
tak.tomlrather than a result. That drops the subject like a failingcmd. - Warmups aren't checked, because they aren't kept. The check isn't counted toward a
runs = "auto"budget, which is sized from prepare and the command alone. - The instruction-count runs of a subject with
counters = trueare separate from the timed samples, and the check doesn't run after them. checkuses the same syntax ascmdandprepare: no implicit shell, and a program path containing a/is found relative totak.toml. It stacks likeprepare, so a subject's owncheckreplaces the benchmark's.--export-jsonaddschecksto each result that has one:{"passed": 18, "total": 20, "samples": [true, …]}, withsamplesin the same order astimes, so each verdict can be matched to its time. The hyperfine fields are unchanged.--recordwrites nothing if any check failed. Git notes keep timings but not verdicts, so the timings of a run with a failed check would be stored as if it had passed.tak comparekeeps each metric's minimum and treats lower as better, and a pass rate fits neither rule, so it isn't stored alongside them. Instead, tak names each subject whose check failed and how many of its samples failed, writes no notes, and exits non-zero, as it does when a subject is dropped.--export-jsonis still written, with the verdicts.--save-baselinefollows the same rule, because a local baseline stores the same records.
Counting heap allocations
Instruction counts don't show a change that makes a command allocate more memory without doing much more work. allocations = true also runs the command under valgrind's DHAT and records three counts:
| metric | what DHAT reports |
|---|---|
alloc_blocks | heap blocks allocated over the whole run, freed or not |
alloc_bytes | bytes allocated over the whole run, freed or not |
alloc_peak_bytes | bytes live at the run's peak (DHAT's t-gmax) |
[bench.version]
cmd = ["git", "--version"]
runs = 5
allocations = true version git --version
alloc_blocks 33
alloc_bytes 7649
alloc_peak_bytes 7336
instructions 288808
wall_max_ms 0.75
wall_mean_ms 0.63
wall_min_ms 0.52
wall_p50_ms 0.61
wall_stddev_ms 0.09tak run --allocations turns it on for every subject in the run, including a command given after --.
- Allocations are recorded and reported, not gated.
tak compareshows them in a separate table under the instruction counts, only for series that have them on both sides, and they never fail the comparison. The totals repeated exactly in the measurements on the methodology page, but the peak moves with thread scheduling. allocationsis a setting likeruns: it can go in[defaults], a benchmark, or a subject, and a more specific layer's value replaces a less specific one. It's off unless set, because each DHAT run costs about as much as a cachegrind run.- Like instruction counting, it takes 3 runs after the timed samples, each after the subject's
prepare, and records each metric's minimum.checkdoesn't run after them.--no-countersdoesn't turn it off. - It needs valgrind 3.15 or newer. Without valgrind, tak prints a note and records timing only.
- If counting allocations fails, tak warns and records the rest of the measurement without them.
tak backfill --commitsfollows the same rule, so a DHAT failure doesn't stop a commit from being recorded, unlike a failed instruction count. The containment check that--commitsruns before each valgrind run applies to DHAT runs too, and a subject it refuses is dropped. - DHAT counts only allocations it can intercept. A statically linked binary, or one with its own allocator built in (jemalloc, mimalloc), reports zero. tak records the zero, and warns that it's probably wrong.
- tak warns when the totals vary across the 3 runs by more than 0.5%. That means the command's work changed, such as a cache it fills on the first run. A peak that varies while the totals don't gets a note instead: threads that allocate at the same time reach different peaks depending on how valgrind interleaves them.
- Valgrind measures the process it starts, not the programs that process runs. For
["sh", "-c", "…"], it counts the shell's allocations.
Recording other metrics
Some changes show up in a number that no timing captures: a binary's size, a bundle's size, a count from a tool's own output. A benchmark can record such numbers beside its timings:
[bench.startup]
cmd = ["./mycli", "--version"]
runs = 10
# The size of a file, in bytes.
[bench.startup.metric.binary_bytes]
file = "mycli"
# The one number a command prints.
[bench.startup.metric.help_lines_count]
cmd = ["sh", "-c", "./mycli --help | wc -l"]A single-command benchmark lists them with the timings, in name order, and a multi-subject benchmark adds them to the end of each subject's summary line. With GNU ls copied to mycli:
startup …/mycli --version
binary_bytes 142312
help_lines_count 138
wall_max_ms 3.06
…fileis a path relative totak.toml, not todir, like a program path. Its size in bytes is the value. It must be a regular file (a symlink to one is followed); a directory is an error.cmdruns likecheck: in the subject'sdirandenv, with the same variables removed, and no implicit shell. It must exit 0 and print exactly one non-negative number on stdout, such as1234,12.5or1.2e6. Surrounding whitespace is ignored. A sign,inf, a unit, a thousands separator, a second number or any other text is an error rather than something tak tries to pick a number out of. A command that prints more than 1 KiB is stopped as soon as it does. So is anything it leaves running in the background with its stdout still open, and the metric is an error, since more output could still arrive. stderr is ignored unless the command fails, when tak reports its last line.- Each metric is taken once, after the subject's samples and its instruction and allocation counts, and isn't part of any of them.
setupor the samples can create what it measures.--dry-runlists each metric without taking it. - A metric that can't be taken fails the run. A missing file or a command that exits non-zero or prints something other than one number is reported on stderr. Every benchmark is still measured and
--export-jsonis still written, with that metric asnull. Then tak exits non-zero, recording or not, and neither--recordnor--save-baselinewrites anything: a stored run with a declared metric missing would look like one where the metric had been removed. - Names use lowercase letters, digits and
_, start with a letter, and are at most 64 characters.instructionsand anything starting withwall_oralloc_are reserved for what tak measures itself. End the name with its unit,_bytes,_kb,_count,_ms, so a report reader doesn't have to guess. tak doesn't enforce the suffix. - Lower is taken as better. Several records of one series on a commit reduce to their minimum, as for every metric. Record a size or a count, not a score where higher is better.
- Custom metrics never gate.
tak comparereports them in a table of their own, after the main one, and only when there are some. A file's size is deterministic for a given build, but a command's output is not known to be.
metric tables layer by name, like env: each layer can add metrics, and a more specific layer's table replaces a same-named one entirely. file and cmd are templates, so a single table in [defaults] can cover every subject:
[defaults.metric.binary_bytes]
file = "bin/{{ subject }}"--record stores each metric in the same record as the subject's timings, keyed by its name. --export-json adds a metrics object to the subject's result. This is tak compare between a commit where mycli was GNU true and one where it was ls:
| benchmark | metric | value | Δ |
|---|---|---:|---:|
| startup | binary_bytes | 26,936 → 142,312 | +428.33% |
| startup | help_lines_count | 14 → 138 | +885.71% |
<sub>Metrics declared in tak.toml are reported, never gated. As for every metric, lower is taken as better.</sub>Comparing several programs
To compare programs against each other, declare them as subjects of one benchmark instead of giving the benchmark a cmd:
[bench.install]
runs = 10
warmup = 1
prepare = ["sh", "-c", "rm -rf node_modules"]
[bench.install.subject.mycli]
cmd = ["./target/release/mycli", "install"]
dir = "fixtures/mycli"
[bench.install.subject.othertool]
cmd = ["othertool", "install"]
dir = "fixtures/othertool"
env = { HOME = "/tmp/othertool-home" }
runs = 5tak interleaves the samples: every round takes one sample of each subject in a freshly shuffled order, rather than every sample of one subject and then the next. See methodology for why.
- Subjects inherit the benchmark's
runs,warmup,setup,prepare,check,dir,env,ok_exit_codesandallocations. A subject's ownsetup,prepareorcheckreplaces the benchmark's, and itsenventries override matching keys. - A subject with fewer
runsthan the others is spread evenly across the run. - Each subject is recorded as its own series, with the subject name as the tool. Instruction counts are off for subjects unless they set
counters = true, so another program's upgrade cannot trip the gate. - If a subject fails, tak drops it and keeps measuring the others. The run then exits non-zero and
--recordwrites nothing, because a partial set of measurements would look complete.
Choosing the number of runs
Programs in one comparison can differ in speed by a factor of 50: a 300 ms install next to a 20 s one. A fixed runs either spends minutes on the slow program or leaves the fast one under-sampled. runs = "auto" lets tak decide per subject:
[bench.install]
runs = "auto"
budget = "30s" # wall time to spend per subject, prepare included, setup not (default 30s)
min_runs = 5 # never fewer (default 5)
max_runs = 50 # never more (default 50)After the warmups, tak times each subject and gives it as many runs as fit in budget, within min_runs and max_runs. With the defaults, a 1.7 s install gets 17 runs and a 21 s one gets the minimum of 5. A subject with warmup = 0 is sized from its first real sample, which still counts toward its runs. Each sample is measured from the start of its prepare step, because that's how long it actually takes.
budget, min_runs and max_runs can be set on the benchmark or per subject, and a subject can still fix its own runs. tak run --runs auto switches a benchmark to auto for one run. Because the counts depend on measured timings, --seed only repeats the same order when the counts come out the same.
Output
While a benchmark runs, tak shows progress on stderr: a bar in a terminal, or a plain line every tenth of the way (or every 30 seconds) anywhere else, such as CI logs. The time remaining is estimated separately for each subject from its own samples so far, so a slow subject's remaining samples are counted at its own speed. Pass --no-progress to turn it off.
tak also warns on stderr when a subject's samples look suspect:
- Slow outliers (a modified z-score above 3.5; when over half the samples are identical, anything more than 25% away from them): something else ran, or the command's work varies. These don't change the minimum.
- Fast outliers: these can be the minimum, so check the command did the same work every time before trusting the headline number.
- A slow first sample: the first timed sample took over twice the median of the rest, so the warmup didn't fill some cache.
Nothing is dropped or adjusted. A warning means the comparison may be worth running again. A multi-subject benchmark prints one summary line per subject. This is the output of a real run comparing sleep 0.1 (fast) with sleep 0.8 (slow), using runs = "auto", budget = "3s" and min_runs = 3:
install: 2 subjects, interleaved (--seed 3)
fast min 102.62 p50 106.99 mean 106.49 ± 1.50 max 108.41 ms n=28
slow min 803.47 p50 807.21 mean 806.24 ± 2.44 max 808.05 ms n=3Every multi-subject run prints its seed. Pass it back with --seed to repeat an order. --subject NAME limits a run to the named subjects.
Exported results
--export-json PATH writes every sample in hyperfine's --export-json shape, so scripts that read hyperfine's file can read tak's. tak only adds keys; none of hyperfine's change meaning.
tak run --bench install --seed 1234 --export-json results.json{
"tak_version": "0.0.12",
"seed": "1234",
"runner": "local-linux-x86_64",
"time": "2026-09-24T19:40:38Z",
"machine": {
"os": "linux",
"os_version": "Ubuntu 24.04.4 LTS",
"kernel": "6.8.0-45-generic",
"arch": "x86_64",
"cpu": "AMD Ryzen 9 7950X3D 16-Core Processor",
"cpus": 2,
"memory_bytes": 100294041600
},
"results": [
{
"command": "mycli",
"bench": "install",
"subject": "mycli",
"version": "mycli 2.4.0",
"mean": 0.412,
"stddev": 0.006,
"median": 0.411,
"min": 0.404,
"max": 0.425,
"times": [0.411, 0.404, 0.425],
"exit_codes": [0, 0, 0]
}
]
}The top-level keys record how the run was made. seed is a string so large seeds survive JavaScript and jq, and runner is the class --record would store the run under.
machine describes what the run was measured on, so a results page can state it:
| key | value |
|---|---|
os, arch | as Rust names them: linux, macos, windows; x86_64, aarch64 |
os_version | PRETTY_NAME from /etc/os-release on Linux, the product version on macOS; null on Windows |
kernel | the kernel release on Linux, the Darwin release on macOS, 10.0.<build> on Windows |
cpu | the CPU model name; null where the OS does not report one, as on most arm64 Linux kernels |
cpus | logical CPUs tak could run on, which is what the subjects could use. On Linux it honours the affinity mask (taskset, a container cpuset) and cgroup CPU quotas, so it can be fewer than the machine has |
memory_bytes | total physical memory |
A value tak can't read is null rather than an error. None of this is stored by --record: series are partitioned by runner, and a kernel update shouldn't split one.
Each result has bench and subject. These keys appear only when the subject asks for them:
version, for a subject with aversion_cmd: the first line it printed, ornullif it failed.checks, for a subject with acheck: how many samples passed, out of how many, and each sample's verdict in the same order astimes.allocations, for a subject withallocations = truewhen valgrind was there to count them:{"blocks": 33, "bytes": 7649, "peak_bytes": 7336}, the same minimums--recordstores.metrics, for a subject with custom metrics: each one's value by name, such as{"binary_bytes": 142312.0}, ornullfor one that could not be taken.
user and system are absent because tak doesn't measure CPU time.
Sharing settings between benchmarks
A comparison usually runs the same programs in several scenarios, such as a warm install and a cold one. Declare each program once as a top-level subject and list it from each benchmark:
[defaults]
runs = "auto"
min_runs = 5
[subject.mycli]
cmd = ["./target/release/mycli", "install"]
[subject.othertool]
cmd = ["othertool", "install"]
env = { OTHERTOOL_CACHE = "/tmp/othertool" }
[bench.warm]
subjects = ["mycli", "othertool"]
prepare = ["sh", "-c", "rm -rf node_modules"]
[bench.cold]
subjects = ["mycli", "othertool"]
prepare = ["sh", "-c", "rm -rf node_modules ~/.cache/mycli /tmp/othertool"]
# A benchmark can override a shared subject, or add one of its own.
[bench.cold.subject.mycli]
cmd = ["./target/release/mycli", "install", "--no-cache"]Settings stack from least to most specific: [defaults], then the benchmark, then the shared [subject.NAME], then the benchmark's own [bench.B.subject.NAME]. Each layer's setting replaces the one before, except env, vars and metric, which merge key by key. [defaults] takes every benchmark setting (runs, warmup, budget, min_runs, max_runs, ok_exit_codes, setup, prepare, check, version_cmd, dir, env, vars, metric, allocations). It is a separate table because [env] already holds env.deny and env.allow.
Templates
Values in cmd, setup, prepare, check, version_cmd, dir, env, vars and a metric's file or cmd are tera templates, the same syntax mise uses. tak renders them itself before anything runs, so a command can use a path that only exists at run time and still be a plain argument list, without a shell:
[defaults]
dir = "{{ env.BENCH_DIR }}/project-{{ subject }}"
env = { HOME = "{{ env.BENCH_DIR }}/home-{{ subject }}" }
[subject.mycli]
cmd = ["{{ env.MYCLI_BIN }}", "install", "--lockfile", "{{ vars.lockfile }}"]
vars = { lockfile = "mycli.lock" }A template can use:
| name | value |
|---|---|
env | tak's own environment, such as {{ env.HOME }} |
bench | the benchmark's name |
subject | the subject's name (self for a single-command benchmark) |
vars | the subject's vars tables, merged like env; not passed to the command |
vars values are rendered first, so they can build on env. Filters work as in mise, for example {{ env.MYCLI_BIN | default(value="mycli") }}. Using a variable that isn't set is an error when the benchmark is loaded, before any sample runs, rather than an empty string that sends a command to the wrong path. Template syntax is checked for the whole file up front; values are rendered only for the benchmarks being run, so a variable needed by one benchmark doesn't have to be set to run another.
Conditions
when limits a benchmark or subject to the times an expr condition holds, the expression language mise uses for its own conditions. A comparison can list every tool and leave out whichever isn't installed:
[subject.vlt]
when = '(env.VLT_BIN ?? "") != ""'
cmd = ["{{ env.VLT_BIN }}", "install"]
[bench.linux-only]
when = 'os == "linux"'
cmd = ["./target/release/mycli", "--version"]A condition can use env (tak's environment), os and arch (as Rust names them: linux, macos, x86_64, aarch64), ci (whether CI is set to anything but empty or false), bench, and subject. It must evaluate to true or false. Conditions are decided before templates are rendered, so a skipped subject's variables don't need to be set. tak prints each skipped benchmark or subject on stderr. A subject asked for with --subject that no selected benchmark will run is an error. If when switches off everything selected, --record and --export-json fail rather than succeed with nothing written. A when in a benchmark's own subject table replaces the shared subject's; a table without one keeps the shared condition, like every other setting. Conditions are only evaluated for the benchmarks and subjects being run, so an unrelated subject's condition can't stop a --subject run. Conditions are syntax-checked when tak.toml is loaded.
Building past commits
[build] tells tak backfill --commits how to build the project in a fresh checkout of an old commit. tak run never runs it and measures whatever is already built.
[build]
cmd = ["cargo", "build", "--release", "--locked"]
# Optional. Relative to tak.toml's directory in the checkout, and must stay inside it.
dir = "."
# Optional. Added to tak's own environment for the build.
env = { CARGO_INCREMENTAL = "0" }cmd follows the same rules as a benchmark's: a list or a whitespace-split string, no shell, and a program path containing / is relative to the directory holding tak.toml. A build that needs several steps can name an interpreter explicitly, such as ["sh", "-c", "git submodule update --init && make"]. The build is not measured, so the shell adds no noise there. [build] values are not templates, and template syntax in one is an error. A project with nothing to build can declare cmd = ["true"].
env_deny does not apply to the build. It exists to keep a token from changing what a measured command does, and a build that fetches a private dependency may need exactly that token.
A commit whose build failed, or a benchmark whose measurement failed at a commit, is remembered, and later backfills pass over it; see how failures are remembered. Changing [build] retries every failure remembered under the old one. Any other edit to tak.toml retries the failed measurements, but not the failed builds.
Environment and runner settings
tak removes known sources of non-determinism from measured commands. Inspect every resolved setting and its source with:
tak settings --docsThe most important setting is the runner class. Measurements from different runner classes must never share a series:
[runner]
class = "gha-linux-x64-rust-1.90"Use an explicit class when a compiler, base image, or other invisible input changes the measurement without changing the machine name.
Regression gate
tak compare fails when an instruction count rises by more than the gate. The default is 1%, which leaves room above the observed instruction-count noise without making timing part of the decision:
[gate]
pct = 0.5The command line and environment can override the project value:
tak compare origin/main --gate-pct 2
TAK_GATE_PCT=2 tak compare origin/mainOnly instruction counts are gated. Wall-clock changes and custom metrics are displayed but never fail the comparison. Use tak compare --no-gate when a report must not fail on the comparison; errors, such as an invalid tak.toml, still fail it. To let one benchmark regress on purpose while the others still gate, pass --accept BENCH; see accepting an intentional regression.
Tak-Accept: commit trailers do the same, but only when the project opts in. Trailers are written by the change being gated, so they are ignored by default:
[gate]
accept_trailers = truetak compare reads this key from the base revision's tak.toml, as it reads the rest of the gate (see where gates are read from), so a pull request cannot turn trailers on for itself. TAK_ACCEPT_TRAILERS=0 in a workflow overrides the file. Use it in a pull-request workflow if tak.toml turns trailers on for the main branch.
Comparing nothing
When no series was measured on both sides, tak compare, tak detect and tak run --baseline --gate fail by default: a gate that checked nothing would otherwise look like one that passed. allow_empty makes that case a warning instead:
[gate]
allow_empty = trueThe command exits 0, the report keeps its **Nothing was compared first line and adds a line naming allow_empty, and a warning goes to stderr, as a ::warning:: annotation under GitHub Actions. A regression still fails.
Turn it on when comparing nothing is routine rather than a symptom: a runner class that encodes a CI image or compiler version, so each update starts a new series that the main branch has not been measured on yet, or a project still adopting tak. The cost is that a broken setup, such as a base that was never recorded or notes that were never fetched, also only warns. The CI guide has the details.
--allow-empty turns it on for one run and TAK_ALLOW_EMPTY=0 turns it off for one. Like the rest of [gate], tak compare reads it from the base revision's tak.toml, so a pull request that turns it on still fails on an empty comparison until it is merged.
An absolute floor
A percentage alone serves small benchmarks badly. On a 450k-instruction --version, 1% is 4,500 instructions, and most of such a benchmark can be the dynamic loader relocating the binary before main runs, which grows with every dependency. min_delta sets a floor: a count has to rise by more than pct percent and by more than min_delta instructions before tak compare fails.
[gate]
pct = 1.0
min_delta = 20000The default is 0, meaning no floor. It can also be set with --gate-min-delta or TAK_GATE_MIN_DELTA.
Per-benchmark gates
A benchmark or subject can have its own gate. The gate table takes the same pct and min_delta keys as [gate], plus enabled:
[bench.startup]
cmd = ["./target/release/mycli", "--version"]
gate = { pct = 5.0, min_delta = 20000 }
[bench.install]
cmd = ["./target/release/mycli", "install"]
# no gate table: held to [gate]
[bench.help]
cmd = ["./target/release/mycli", "--help"]
gate = { enabled = false } # report only- Report only.
enabled = falsekeeps the benchmark in the report and holds it to its threshold, but never fails the command on it. A rise beyond the threshold is marked(not gated)in the table and listed under the verdict as a report-only benchmark. It's meant for a benchmark you want to watch but can't gate yet, not for turning off a gate that fails, which you should investigate first. - Key by key. A
gatetable only replaces the keys it sets. The rest come from[gate], sogate = { enabled = false }keeps the project's percentage. - Subjects stack. A subject's
gatestacks on its benchmark's, in the same order as other settings: the benchmark, then the shared[subject.NAME], then[bench.B.subject.NAME]. Only subjects withcounters = truehave an instruction count to gate. - Not in
[defaults]. A gate for every benchmark is[gate]. - The file wins over the flag.
--gate-pct,--gate-min-deltaand their environment variables set the[gate]values, which are what a benchmark without its own gate uses. They don't override a benchmark that declares its own. - Validated when the file loads.
tak runandtak compareboth reject a negative, NaN or infinitepct, a negative or fractionalmin_delta, and unknown keys in agatetable. The same check covers[gate]itself, whether it comes fromtak.toml, a flag or an environment variable, so a bad value failstak runbefore anything is measured.
When any row in a comparison has a gate other than [gate], the report adds a gate column that shows every row's gate, and the verdict lists each regression with its gate:
| benchmark | instructions | Δ | gate | wall (min) | Δ |
|---|---:|---:|---|---:|---:|
| install | 104,882,113 → 105,001,234 | **+0.11%** | 1% | 88.40 → 87.95ms | -0.51% |
| startup | 452,310 → 463,102 | **+2.39%** (not gated) | report only (1%) | 1.21 → 1.19ms | -1.65% |
No gated benchmark rose beyond its gate.
**1 report-only benchmark(s) above their gate, not failing:** `startup` +2.39% (gate 1%)If every row uses [gate], the report looks the same as it did before per-benchmark gates.
Where gates are read from
Gates are not recorded. They're policy rather than measurement, so tak compare BASE reads them from the tak.toml in BASE's tree, not from the working tree. That covers [gate] (pct, min_delta, accept_trailers and allow_empty) and every gate table. In CI the working tree is the pull request, and a pull request that could edit its own gate could waive it. Flags and environment variables still override the base's file. [report] credit still comes from the working tree.
A gate change takes effect once it is merged. A pull request that loosens, tightens or adds a gate is compared under the base's gate, and the report adds a line below the verdict saying the change takes effect once it is merged. This is a change from tak 0.0.13 and earlier, which read [gate] from the working tree. To let a specific regression through on the pull request itself, use --accept BENCH from the workflow rather than editing the gate.
BASE's file is found by searching upward from the current directory through BASE's tree. With none there, every series is held to tak's defaults plus any flags and environment variables, never to the working tree's file, and the report says so. A BASE tak.toml that doesn't parse is an error: fix it on the base branch. It is one of the errors --no-gate does not waive: --no-gate never fails on the comparison, but errors still fail. The working tree's tak.toml has to parse as well. The CI guide covers the git objects this needs.
A series in the notes whose benchmark tak.toml no longer declares is held to [gate]. A series whose subject the benchmark no longer declares is held to the benchmark's gate.
tak detect reads gates from the working tree's tak.toml, which in a main-branch job is the merged commit, and applies each series' gate to its steps and its drift. A report-only benchmark never fails it. It reads allow_empty from there too, as does tak run --baseline --gate.
Environment filtering
tak removes known sources of non-determinism from the measured command's environment. The project setting replaces the default deny list, so repeat the defaults when adding project-specific credentials or configuration:
[env]
deny = ["GITHUB_TOKEN", "GH_TOKEN", "MYCLI_UPDATE_CHECK"]allow can opt a listed name back in without restating the deny list. It subtracts names from the deny list; it does not add variables to the environment. Passing credentials or network configuration through to a subject makes its measurements depend on state outside the repository.
Report credit
Generated comparison reports name tak in their footer by default. Disable that line when the surrounding report already provides the context:
[report]
credit = falseEvery setting follows the same precedence: command-line flag, environment variable, tak.toml, then the built-in default. tak settings --docs prints the resolved value, its source, every supported source, and the full setting documentation.