Ori Eval: # Find the best model for what you're building
01──Find the best model, with proof
Just ask. Your agent will hand the request to Ori Eval, which explores your codebase, runs evaluations on models, and comes back with an answer.
02──New to evals? We figure out what to test with you
Never written an eval? Ori Eval scans your codebase for every place a model runs — which surface, where the code is, what model it uses today — and checks which one you want evals on. It guides you like an experienced engineering friend.
03──Consistent results you can trust, run after run
Keeping the test bench identical, run after run, makes evaluations repeatable and consistent. That's why Ori Eval is an agent: unlike a skill, it's able to pin model and effort. It comes pre-tuned - we've already chosen the harness and model that works best, so you don't have to.
04──Catch a bug before you fix it
Every bug becomes a test: it fails now, passes once you fix the agent, and keeps the bug from coming back.
05──Stop regressions before they ship
In CI, a bad change fails the build, so bugs never reach your users.
06──Re-run as new models drop
Your evals are just code. Manually re-run them when a new model drops, or schedule them monthly. You can sleep well knowing you always have the best model for what you're building.
07──Know it did the right thing
An eval checks three things: the tools the agent called, the tools it avoided, and the quality of the answer — so you catch the wrong move before your users do.