Prove the lift · AI model performance evaluation
gemini-2.5-pro vs kimi-k2-thinking
-2-point lift for gemini-2.5-pro on insurance-extraction-eval
Test summary
This benchmark evaluates gemini-2.5-pro against kimi-k2-thinking on the insurance-extraction-eval v2 dataset (6 rows), scored by an LLM-as-judge (llama-3.1-8b). gemini-2.5-pro reached a 54% pass rate versus 56% for kimi-k2-thinking — a -2-point lift.
- Base model
- kimi-k2-thinking
- via moonshot · 56% pass rate
- Refined model
- gemini-2.5-pro
- via google · 54% pass rate · 0.8s p95
Base pass rate
56%
Refined pass rate
54%
Refined p95
0.8s
Est. tokens saved
1.4K
−0pts
no lift — the base model held its ground
0%
Base model0%
Invoked-refined▲2 improved▼1 regressed●3 unchanged
Per-row results
1Business-interruption period-of-restoration calculation0% → 100%+100
2Named-peril coverage determination under a basic HO form0% → 100%+100
3Aggregate deductible across a loss sequence under a per-occurrence + annual cap100% → 100%0
4Pro-rata distribution across layered carriers on a covered loss100% → 0%-100
5Coinsurance 80% clause recovery on an under-insured building0% → 0%0
6Replacement cost vs ACV on a depreciated roof100% → 100%0
Prove the lift on your own data.
Run any two models over your tasks, judged automatically, and see the improvement — in seconds.
Try InvokedGenerated with Invoked · Browse all benchmarks