v0.4.0
LatestNew
- Datasets and Prove the Lift turn a set of real tasks into a scored benchmark, then run two models head to head and see the measured improvement on your own data. An LLM judge grades every answer with a reason, and the result is an honest number: a regression shows as a negative lift, never a faked win
- The judge defaults to the cheapest model you have already connected, with a local Ollama judge as a first-class, fully private option. Invoked never hosts a judge; grading runs on your keys or your machine
- Publish any lift result as a public benchmark page a shareable link anyone can open, with the same animated reveal, the per-row breakdown, and the headline lift. Publishing is opt-in and discloses exactly what becomes public
- A public benchmark directory at invoked.ai/lift browse and filter shared results by model. Listing your own result is your choice, offered on publish and never required