Skip to content

Enable PGO for macOS ARM64 ty releases - #4216

Merged
charliermarsh merged 4 commits into
charlie/ty-linux-pgo-prototypefrom
charlie/ty-pgo-macos-arm64
Aug 11, 2026
Merged

charliermarsh merged 4 commits into
charlie/ty-linux-pgo-prototypefrom
charlie/ty-pgo-macos-arm64

Conversation

@charliermarsh

@charliermarsh charliermarsh commented Aug 7, 2026 •

Copy link
Copy Markdown
Member

Summary

This PR enables PGO for ty's macOS ARM64 releases, following the approach outlined in #4213. The existing linker configuration and safe identical-code folding are preserved.

Across seven held-out macOS projects, ty check finishes 5.0% faster and uses 7.1% less CPU; incremental language-server edits are 7.3% faster. The production executable is 3.6% smaller than without PGO, although the wheel and release archive are each 2.8% larger.

The macOS ARM64 release job uses the same Namespace runner as Ruff, reducing its full-stack PGO build from 23m 47s to 6m 39s (72%). Three Namespace builds finished in 6m 38s to 7m 59s; the previous Depot runner took 11m 20s.

@charliermarsh
charliermarsh force-pushed the charlie/ty-pgo-macos-arm64 branch 3 times, most recently from 3752fd8 to 7f55272 Compare August 10, 2026 14:56
@charliermarsh charliermarsh added the release Related to the release process label Aug 10, 2026
@charliermarsh
charliermarsh force-pushed the charlie/ty-pgo-macos-arm64 branch 2 times, most recently from df5d59c to 6be935d Compare August 10, 2026 22:57
@charliermarsh
charliermarsh marked this pull request as ready for review August 10, 2026 23:16
@MichaReiser

Copy link
Copy Markdown
Member

Can you please add benchmark results that go beyond wheel size?

@charliermarsh

Copy link
Copy Markdown
Member Author

Added.

@charliermarsh
charliermarsh force-pushed the charlie/ty-pgo-macos-arm64 branch from 6be935d to 81e36e5 Compare August 11, 2026 13:41
@charliermarsh
charliermarsh force-pushed the charlie/ty-pgo-macos-arm64 branch from 8b344a9 to 6eacc9a Compare August 11, 2026 19:48
@charliermarsh
charliermarsh force-pushed the charlie/ty-pgo-macos-arm64 branch from 6eacc9a to b4ed7d7 Compare August 11, 2026 20:04
@charliermarsh
charliermarsh merged commit 8988165 into main Aug 11, 2026
34 checks passed
@charliermarsh
charliermarsh deleted the charlie/ty-pgo-macos-arm64 branch August 11, 2026 22:23
charliermarsh added a commit that referenced this pull request Aug 11, 2026
## Summary

This PR enables PGO for ty releases, starting with Linux x86-64. The
approach follows that outlined in
astral-sh/ruff#27570.

### Design

The release pipeline is modified as follows:

- We build an instrumented release binary using the Rust toolchain
pinned by the Ruff submodule.
- We run `ty check` on the same ten pinned ecosystem projects used by
Ruff. In total, the corpus contains 618 Python and stub files.
- We also exercise `ty server` against a temporary project, covering
diagnostics, hover, go-to-definition, completion, and incremental
cross-file edits.
- We merge the profiles via `llvm-profdata`.
- We derive LLVM's hot-code threshold from the profile's 95th
percentile, matching its size-optimization threshold and avoiding
unnecessary expansion of moderately hot functions. (Without this,
performance was marginally better, but wheel size _increased_ by 5-10%!)
- We feed the result back into the existing `maturin` build.

### Results

Here's the initial, untuned PGO performance on seven held-out projects
(and language-server performance on three held-out projects):

| Held-out project | `ty check` | Incremental edits |
| --- | ---: | ---: |
| Black | 7.1% faster | 17.1% faster |
| isort | 13.3% faster | 18.1% faster |
| Jinja | 5.5% faster | 13.7% faster |
| Django | 25.5% faster | — |
| pandas | 15.9% faster | — |
| scikit-learn | 19.6% faster | — |
| SymPy | 21.5% faster | — |
| Geometric mean | **15.8% faster** | **16.3% faster** |

As an aside, profiling the language server explicitly (i.e., including
the language server commands in the PGO training) improved
incremental-edit latency by another 7.6% over CLI-only PGO; without that
training, incremental edit performance was _still_ up (but not quite as
much).

Beyond raw performance, with the tuned PGO settings, we see the
following results:

- **Binary size decreased by 7.98%** (28.10 MB to 25.86 MB compared with
non-PGO).
- **Wheel size decreased by 1.32%** (12.71 MB to 12.54 MB compared with
non-PGO), and by 14.11% compared with untuned PGO.
- **Release archive size decreased by 1.53%** (12.35 MB to 12.16 MB
compared with non-PGO), and by 14.25% compared with untuned PGO.
- **Incremental language-server edits are 3.30% faster than untuned
PGO**.
- **Linux x86-64 release build finishes in 11m28s on eight-core Depot**,
compared with 12m59s on four-core Depot and 15m45s on the previous
GitHub runner. The full-stack bottleneck shifts to Windows at 12m24s;
other Linux targets retain their four-core runners.

### Stack

Additional platforms are covered in subsequent PRs:

- macOS ARM64: #4216
- Windows x86-64: #4217
- Linux ARM64: #4218

Ruff and uv follow the same approach; see
astral-sh/ruff#27570 and
astral-sh/uv#21001.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

release Related to the release process

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants