How-to seriesSeptember 28, 2026
How to choose a model for a specific task
Compare models on your actual inputs with a fixed rubric, bounded trial budget and a record of output quality, elapsed time and observed charges.
Choose a model by giving a few candidates the work you need done. A polished answer to an easy prompt can hide the failure that matters in production: a missing field, invented evidence or a refusal to leave an unknown value blank.
With Vaaya, your agent can discover model routes, run a bounded comparison and keep the call records together. Start with a small test you can review by hand. This tutorial supplies a test design, not a claim that one model won a benchmark.
1. Define an answer you would accept
Connect the agent through Vaaya's installation guide. Write down the input, required output and unacceptable errors before looking at model names.
For example, suppose the task is to turn a support ticket into a short summary and a category. Choose five fictional or appropriately redacted tickets: one ordinary request, one ambiguous request, one long thread, one with missing details and one outside the supported categories. Write the expected category and required summary facts yourself.
Treat these as separate checks: valid structure, correct category, preserved facts and no invented resolution. If a model returns attractive prose but fabricates a refund, it fails the factual requirement. Keep that requirement fixed across candidates.
2. Shortlist routes that fit the task
Ask consult which current models can handle the input and output you need, then inspect the catalog. Vaaya's openrouter/chat route requires a model and messages; other routes can have different fields or output limits. Resolve exact model identifiers before making calls.
For classification-only work, Jev is another possible route. Vaaya exposes openrouter/decisions for typed choices, probabilities and scores. Jev does not write a support reply. If the task needs both a routing label and prose, evaluate those jobs separately before deciding whether to combine them.
Use public benchmarks to narrow the list when their tasks resemble yours. Record the benchmark source and date. Do not import a score from a different task as your acceptance result.
3. Fix the inputs and bound the trial
Paste this prompt and fill in the missing details:
Use Vaaya to compare [two or three candidate model IDs] on
[this task]. Use these identical test inputs: [attach cases].
Acceptance rubric: [structure, facts and task-specific criteria].
My total trial budget is [amount], including retries.
Check each route's current schema and pricing. Show a call plan
before running it, set per-call ceilings, and track remaining spend.
Save raw outputs and grade every case against the same rubric.
Record actual charges and observed elapsed time. Mark anything
not measured as unknown. Recommend a model only for these tested cases.
Hold the prompt, evidence and requested output constant. Record settings that affect the comparison, including output limits and any supported sampling controls. If one model needs a different prompt to work, save that as a separate version so you can explain what changed.
4. Run, preserve and grade the results
Submit the planned calls within the budget. Keep raw responses alongside the parsed output and call record. If a response fails parsing, preserve it as a failed case before attempting a repair. Otherwise you may accidentally count an expensive retry as a clean first-pass success.
Grade without looking at the model name where practical. A human review is enough for a small set. An automated grader can help later, but check that grader on examples you already understand. A model-generated score alone is not proof of correctness.
The following is an illustrative worksheet, not measured performance:
| Model | Case | Valid shape | Required facts | Invented claim | Charge | Elapsed time |
|---|---|---|---|---|---|---|
| Candidate A | Ambiguous ticket | To review | To review | To review | Not run | Not run |
| Candidate B | Ambiguous ticket | To review | To review | To review | Not run | Not run |
5. Choose for the workload and record the exceptions
Compare accepted outputs, failure types and total observed spend. Cost per accepted result includes failed attempts and repairs. Keep an unmeasured value blank rather than treating it as zero.
Select the candidate that meets the required checks for the cases it will handle. If neither passes, revise the task boundaries, prompt or shortlist. A small test supports a limited deployment decision; expand the cases when new input patterns appear.
Save the test set, rubric, prompt versions and date with the chosen model. Re-run that same set before changing the model or routing rules. The tool reference is the place to check current call fields when you repeat the comparison.
Questions
Should I choose the model with the highest leaderboard rank?
Use a relevant, dated benchmark to shortlist candidates, then test your own task. A ranking on another dataset does not establish how well a model handles your inputs or output requirements.
Can Jev write the answers in this comparison?
Jev is a classification model for typed decisions such as choices, probabilities and scores. Use a text-generation model when the task requires prose, and validate any Jev-based grading against human-reviewed examples.
How should I compare model costs?
Record actual charges for the same test inputs and settings, along with failures and retries. Keep estimates labeled separately and compare cost per accepted result, not only the quoted price per token.