Prove the lift · AI model performance evaluation

mixtral-8x7b vs gpt-4o-mini

+16-point lift for mixtral-8x7b on insurance-extraction-eval

Test summary

This benchmark evaluates mixtral-8x7b against gpt-4o-mini on the insurance-extraction-eval v1 dataset (6 rows), scored by an LLM-as-judge (llama-3.1-8b). mixtral-8x7b reached a 67% pass rate versus 51% for gpt-4o-mini — a +16-point lift.

Base model
gpt-4o-mini
via openai · 51% pass rate
Refined model
mixtral-8x7b
via mistral · 67% pass rate · 1.3s p95
Base pass rate
51%
Refined pass rate
67%
Refined p95
1.3s
Est. tokens saved
1.3K
+0pts
refined beats base on 6 rows
0%
Base model
0%
Invoked-refined
2 improved2 regressed2 unchanged

Per-row results

1Coinsurance 80% clause recovery on an under-insured building0%100%+100
2Business-interruption period-of-restoration calculation0%100%+100
3Suit-limitation clause: is a late-filed claim contractually barred?100%100%0
4Named-peril coverage determination under a basic HO form100%0%-100
5Replacement cost vs ACV on a depreciated roof100%0%-100
6Subrogation waiver interaction with additional-insured status100%100%0
Prove the lift on your own data.

Run any two models over your tasks, judged automatically, and see the improvement — in seconds.

Try Invoked
Generated with Invoked · Browse all benchmarks