Industry Partner

Harvey's Legal Agent Benchmark

Updated 10/7/2026

Tests an agent's ability to complete legal work using documents, spreadsheets, presentations, and file-system tools.

As of October 7, 2026, Muse Spark 1.2 ranks first on Harvey's Legal Agent Benchmark with 25.42%, followed by Muse Spark 1.3 Max (23.75%) and Muse Spark 1.3 (22.92%).

Harvey's Legal Agent BenchmarkAgentic legal work with files
ACCURACY

Harvey's Legal Agent Benchmark leaderboard

Rank Model Task Pass Rate Cost / Test Input / Output Cost Duration
1 Muse Spark 1.2 25.42% $2.09 $1.25 / $4.25 24m16s
2 Muse Spark 1.3 Max 23.75% $2.26 $1.25 / $4.25 14m43s
3 Muse Spark 1.3 22.92% $2.76 $1.25 / $4.25 17m04s
4 Muse Spark 1.1 20.00% $0.80 $1.25 / $4.25 12m21s
5 Gemini 4 Argon 19.58% $10.55 $4 / $20 43m55s
6 Mistral Large 4 15.83% $4.18 $1.36 / $4.18 39m25s
7 Grok 4.6 15.83% $4.01 $2 / $6 45m05s
8 Grok 4.5 12.92% $2.02 $2 / $6 10m15s
9 Kimi K3 12.92% $3.83 $3 / $15 13m58s
10 Grok 4.7 12.50% $11.13 $2 / $6 42m05s
11 MiMo V2.6 Flash 11.25% $0.09 $0.14 / $0.28 17m12s
12 Qwen 3.8 27B 11.25% $4.03 $0.5 / $3 18m39s
13 Claude Fable 5 11.25% $19.23 $10 / $50 26m53s
14 MiMo V2.6 Pro 10.83% $0.22 $0.435 / $0.87 25m53s
15 Qwen 3.8 Max 10.42% $2.37 $2 / $6 49m12s
16 Gemini 3.8 Flash 10.00% $3.66 $1.5 / $7.5 29m21s
17 Claude Opus 4.8 9.58% $10.22 $5 / $25 24m08s
18 Step 5 Preview 9.17% $1.00 $1 / $2.7 29m02s
19 Hy4 Preview 9.17% $0.98 $0.834 / $2.501 29m25s
20 Ember-1 9.17% $3.33 $3 / $15 44m26s
21 Gemini 3.7 Flash 8.75% $2.56 $1.5 / $7.5 7m32s
22 DeepSeek V4 Flash 0731 8.33% $0.06 $0.44 / $1.32 16m54s
23 GLM 5.3 8.33% $4.30 $1.4 / $4.4 42m38s
24 DeepSeek V4 Pro 0813 7.50% $0.17 $1.32 / $3.96 21m53s
25 GLM 5.2 7.08% $2.06 $1.4 / $4.4 20m51s
26 DeepSeek V4.1 Flash 6.67% $0.17 $0.3 / $1.2 9m39s
27 Claude Opus 4.7 6.67% $10.81 $5 / $25 24m44s
28 GLM 5.3 Flash 6.67% $0.57 $0.075 / $0.25 52m17s
29 Claude Opus 5 6.67% $23.67 $5 / $25 55m38s
30 Claude Fable 5.1 6.67% $46.21 $10 / $50 1h52m
31 GPT-6 Astra 5.42% $26.16 $10 / $50 24m49s
32 GPT-6.1 Sol 5.42% $3.93 $2 / $10 43m14s
33 Claude Sonnet 4.6 5.00% $3.04 $3 / $15 18m43s
34 Claude Sonnet 5 5.00% $8.95 $2 / $10 38m52s
35 MiniMax-M3 4.17% $1.46 $0.6 / $2.4 22m31s
36 DeepSeek V4 3.75% $0.74 $1.32 / $3.96 6m42s
37 GPT 5.5 3.75% $4.60 $5 / $30 12m14s
38 Claude Opus 5.5 3.75% $21.38 $4 / $20 52m55s
39 Gemini 3.6 Flash 3.33% $1.90 $1.5 / $7.5 16m24s
40 GPT-6 Luna 2.92% $0.30 $0.1 / $0.5 17m32s
41 Claude Sonnet 5.5 2.92% $16.61 $2 / $10 57m44s
42 Gemini 3.5 Flash 2.50% $2.12 $1.5 / $9 7m58s
43 GPT-5.6 Sol 2.50% $10.37 $4 / $20 22m50s
44 MiMo V2.5 Pro 2.08% $0.11 $0.435 / $0.87 9m04s
45 Inkling 2.08% $1.25 $1 / $4.05 18m24s
46 MiMo V2.5 1.67% $0.04 $0.14 / $0.28 6m59s
47 GPT-6 Sol 1.67% $3.36 $2 / $10 13m40s
48 Kimi K2.6 1.67% $0.82 $0.95 / $4 23m25s
49 Inkling Small 1.67% $0.17 $0.3 / $1.2 31m11s
50 Qwen 3.7 Max 1.67% $1.58 $2.5 / $7.5 32m03s
51 Ling 3.0 Flash 1.25% $0.05 $0.075 / $0.22 2m41s
52 GPT-5.6 Luna 1.25% $0.38 $0.2 / $1.2 10m56s
53 Qwen 3.6 Plus 1.25% $3.94 $0.5 / $3 14m48s
54 Claude Haiku 5.5 1.25% $3.39 $0.1 / $0.5 35m14s
55 Claude Haiku 4.5 (Thinking) 0.83% $0.42 $1 / $5 6m01s
56 GPT-5.6 Terra 0.83% $3.20 $2 / $12 17m05s
57 Grok 4.3 0.42% $0.39 $1.25 / $2.5 2m49s
58 Nemotron 3 Ultra 0.42% N/A N/A 34m09s
59 Mistral Medium 3.5 0.42% $6.07 $1.5 / $7.5 1h02m
60 Gemini 3.1 Flash Lite Preview 0.00% $0.08 $0.25 / $1.5 60.55s
61 Mercury 2.5 0.00% $0.19 $0.2 / $0.75 2m12s
62 Gemini 3 Flash (12/25) 0.00% $0.26 $0.5 / $3 3m23s
63 Gemini 3.5 Flash Lite 0.00% $0.33 $0.3 / $2.5 3m53s
64 Ling 3.0 Flash Fin 0.00% $0.05 $0.06 / $0.18 5m58s
65 GPT 5.4 Nano 0.00% $0.18 $0.2 / $1.25 6m03s
66 MiniMax-M2.7 0.00% $0.49 $0.3 / $1.2 6m14s
67 Gemini 3.1 Pro Preview (02/26) 0.00% $1.15 $2 / $12 7m17s
68 Laguna M.1 0.00% N/A N/A 7m45s
69 GLM 5.1 0.00% $0.56 $1 / $3.2 8m18s
70 Grok 4.20 (Reasoning) 0.00% $0.99 $2 / $6 8m23s
71 Qwen 3.7 Plus 0.00% $0.23 $0.4 / $1.6 8m49s
72 GPT 5.4 Mini 0.00% $1.00 $0.75 / $4.5 11m21s
73 GPT 5.4 (xhigh) 0.00% $3.60 $2.5 / $15 14m36s
74 Laguna XS.2 0.00% N/A N/A 19m54s
75 Kimi K2.5 0.00% $0.35 $0.6 / $3 25m08s
76 Nemotron 3.5 Lightning 0.00% $0.03 $0.05 / $0.2 28m56s

Partners in Evaluation


Key Takeaways


Benchmark

The Legal Agent Benchmark is a benchmark recently released by Harvey to test the ability of models to support legal work in an agentic setting. There are two datasets as of today: the public set and the held-out test set. The results here are from the held-out set, and initial results have already been released by Harvey.

Each task asks an agent to produce legal work against a set of task-specific criteria. The agent is provided with six tools: Read File, Edit File, Write File, Glob, Bash, and Grep. It also has three skills: docx, pptx, and xlsx.

Reported results use the same methodology as Harvey’s initial leaderboard. Criteria pass rate is included to show how often models satisfy individual requirements, even when they do not fully resolve the task.


Results

Among models with a nonzero final score, MiMo V2.5, DeepSeek V4 Flash 0731, Muse Spark 1.1, and Muse Spark 1.2 form the cost/performance frontier.

Muse Spark 1.2 leads on the overall Harvey final score at 25.42%, with Muse Spark 1.3 Max second at 23.75%, Muse Spark 1.3 third at 22.92%, Muse Spark 1.1 fourth at 20.00%, and Grok 4.6 fifth at 15.83%. Claude Fable 5, tied for seventh at 11.25%, fell back to Claude Opus 4.8 on 4 tasks; counting those as failures gives a no-fallback score of 10.42%. The criteria pass rates are much higher: Muse Spark 1.3 Max reaches 94.74%, Muse Spark 1.2 94.52%, Muse Spark 1.1 92.86%, and Grok 4.6 92.52%.

Criteria Pass Rate by Task TypePercent of criteria passed per task type
Task type
Muse Spark 1.3 Max94.74% avg
Muse Spark 1.294.52% avg
Gemini 4 Argon94.04% avg
Muse Spark 1.192.86% avg
Grok 4.692.52% avg
Intellectual Property95.1%95.8%96.5%95.4%94.7%
Corporate M&A96.2%95.8%96.2%96.2%95.2%
Data Privacy/Cybersecurity99.3%95.8%97.5%93.7%97.9%
Banking Finance97.4%97.8%95.9%94.8%94.8%
Capital Markets96.1%96.9%95.7%92.5%94.1%
Corporate Governance97.3%96.2%97.3%96.9%93.2%
Trusts & Estates/Private Client93.9%93.5%93.5%91.6%92.7%
International Trade Sanctions94.1%93.1%93.3%90.0%93.3%
Real Estate96.5%96.5%95.9%95.6%95.0%
Healthcare/Life Sciences94.1%95.4%92.7%92.7%93.1%
0%99% criteria passed

The leaderboard can be filtered by task type and switched between task pass rate and criteria pass rate. Models perform best on task resolution in Energy/Natural Resources and Healthcare/Life Sciences, while criteria pass rates averaged across all models are highest in Intellectual Property, Corporate M&A, and Data Privacy/Cybersecurity.

Criteria Pass Rate vs. Task Resolution
CRITERIA PASS RATETASK RESOLUTION
Muse Spark 1.3 Max
94.7%/23.8%
Muse Spark 1.2
94.5%/25.4%
Gemini 4 Argon
94.0%/19.6%
Muse Spark 1.1
92.9%/20.0%
Grok 4.6
92.5%/15.8%
Muse Spark 1.3
92.2%/22.9%
Kimi K3
91.1%/12.9%
MiMo V2.6 Pro
90.7%/10.8%

Harvey grades a task as resolved only if every criterion passes. A model can satisfy most individual criteria and still miss task resolution credit.

There is a clear trend: strong models and agents satisfy most criteria, around 90% for top models. The remaining gaps are large enough that task resolution stays low even when criterion-level performance looks strong.

Tool Calling Statistics5/76 models

Average tool count per item, across the six tools in the benchmark harness.

Models heavily prefer Bash and Read File. Write File appears regularly for some models, while Edit File, Glob, and Grep are lower-volume.

The available skills rely on shell commands in their instructions and to run their scripts. The tools also often encourage models to read files through the harness.

The skills do not directly emphasize Edit File, Write File, Glob, or Grep. Those tools still appear in traces, but less consistently than Bash and Read File. Grep is not well-suited to the binary format of .docx, .pptx, and .xlsx files. Likewise, Edit is useful for text-based files such as .md files, not for these filetypes.

Skill Invocation Statistics5/76 models

Average skill invocations per item, across the three skills available to the agent.

Skill usage is dominated by docx, followed by xlsx. pptx is used less often, but it is not absent.

Methodology

We use Harvey’s generation and grading protocol in the same environment, with internet access disabled.

Harvey grades each submission with two LLM judges. Each judge computes a task pass rate. A task passes only if 100% of its criteria pass, and Harvey’s final score is the average of the two judge task pass rates.

The two judges were GPT 5.5 and Claude Sonnet 4.6. GPT 5.5 used medium reasoning and Claude Sonnet 4.6 was not modified.

While running the benchmark, we found that redline criteria needed DOCX tracked changes preserved when reading submitted files so judges could see inserted and deleted text. We fixed that bug and merged it upstream in harveyai/harvey-labs#76. The scoring rubric is unchanged.

The benchmark was modified to use our model library, an abstraction over various LLM provider APIs, and to run on Valkyrie, our framework for running agentic benchmarks. These are infrastructure changes and do not impact model performance.

To improve judge performance and reduce cost, we split the instruction prompt provided to each judge so common elements could be cached. This did not modify prompt content outside of caching.