Prove the lift · AI model performance evaluation

claude-sonnet-4-6 vs claude-haiku-4-5

+33-point lift for claude-sonnet-4-6 on Insurance Claims Reasoning

Test summary

This benchmark evaluates claude-sonnet-4-6 against claude-haiku-4-5 on the Insurance Claims Reasoning v1 dataset (6 rows), scored by an LLM-as-judge (claude-haiku-4-5). claude-sonnet-4-6 reached a 83% pass rate versus 50% for claude-haiku-4-5 — a +33-point lift.

Base model
claude-haiku-4-5
via anthropic · 50% pass rate
Refined model
claude-sonnet-4-6
via anthropic · 83% pass rate · 5.8s p95
Base pass rate
50%
Refined pass rate
83%
Refined p95
5.8s
Est. tokens saved
214
+0pts
refined beats base on 6 rows
0%
Base model
0%
Invoked-refined
2 improved0 regressed4 unchanged

Per-row results

1Extract the policy number from: 'Policy No. HMX-4820193, effective 2026.' Reply with only the policy number.100%100%0
2What peril is described? 'A kitchen fire caused smoke damage throughout the first floor.' Reply with only the single word.100%100%0
3A commercial property policy has a $10,000 per-occurrence deductible AND a $25,000 annual aggregate deductible. Three separate losses occur in the year: $8,000,…0%0%0
4A policy has an 80% coinsurance clause. The building is worth $500,000 but insured for only $300,000. A $100,000 loss occurs with a $2,500 deductible. Using the…0%100%+100
5A claim occurred 03/15/2026. The policy has a 2-year suit limitation and a 60-day proof-of-loss requirement. The insured filed proof of loss on 06/01/2026 and w…0%100%+100
6Two policies cover the same $60,000 loss: Policy A has a $100,000 limit, Policy B has a $200,000 limit, both with 'other insurance' pro-rata clauses. Under pro-…100%100%0
Prove the lift on your own data.

Run any two models over your tasks, judged automatically, and see the improvement — in seconds.

Try Invoked
Generated with Invoked · Browse all benchmarks