Prove the lift · AI model performance evaluation

llama-4-maverick vs claude-sonnet-5

-1-point lift for llama-4-maverick on code-review-bench

Test summary

This benchmark evaluates llama-4-maverick against claude-sonnet-5 on the code-review-bench v2 dataset (7 rows), scored by an LLM-as-judge (llama-3.1-8b). llama-4-maverick reached a 60% pass rate versus 61% for claude-sonnet-5 — a -1-point lift.

Base model
claude-sonnet-5
via anthropic · 61% pass rate
Refined model
llama-4-maverick
via meta · 60% pass rate · 1.5s p95
Base pass rate
61%
Refined pass rate
60%
Refined p95
1.5s
Est. tokens saved
1.2K
0pts
no lift — the base model held its ground
0%
Base model
0%
Invoked-refined
2 improved2 regressed3 unchanged

Per-row results

1Spot a race in a read-modify-write on shared state100%0%-100
2Flag an off-by-one in a pagination loop100%100%0
3Find a resource leak on an early-return path100%100%0
4Detect a missing await on an async DB write100%100%0
5Catch an unsanitized input reaching a SQL string100%0%-100
6Identify an N+1 query in an ORM call site0%100%+100
7Note a broken null-check after a refactor0%100%+100
Prove the lift on your own data.

Run any two models over your tasks, judged automatically, and see the improvement — in seconds.

Try Invoked
Generated with Invoked · Browse all benchmarks