

Long description
Four gates ask whether the system can perform the task, whether facts are supported, whether actions are permitted, and whether the complete result is useful. A critical failure pauses expansion. Retest after a material change to the model, tools, sources, or scope.
Choose a model by checking how it performs on the work you plan to give it. A public ranking can suggest candidates, but it cannot decide how much review your team needs, whether a tool is available in your account, or whether the output meets your business rules.
This guide helps nontechnical owners organize a small, fair comparison. You will define requirements, test a shortlist, and make a decision that can be revisited. Official documentation was checked on October 10, 2026. No model is named as a universal winner.
Set requirements before reviewing candidates
Describe one task, the material it receives, and the result a person must approve. List the capabilities that are genuinely required: reading a certain input type, producing a consistent structure, using an approved tool, or working within a particular service environment.
Add the operational requirements. The workflow may need to finish while a reviewer is present, keep data within an approved arrangement, and stop when information is missing. Decide which conditions are mandatory and which are preferences. A model that fails a mandatory privacy or permission requirement should not win because its prose is appealing.
Verify features in the product and model version you will actually use. OpenAI's model-selection guidance notes that availability, tools, reasoning settings, and limits differ by product and version. Model selection.
Build a focused shortlist
Select a small set of plausible candidates rather than testing every model you can access. Include a candidate suitable for the task's complexity and another that offers a different cost or response-time tradeoff. Do not assume that a more expensive option is necessary before you have task evidence.
Record the exact model identifier, provider, product interface, relevant settings, and date. If an interface automatically chooses a model or changes its behavior, note that limitation. You are comparing the available workflow configurations, not conducting a controlled experiment on a hidden model.
Check service and data terms before sending material to each candidate. Use fictional or already approved examples first. Approval to use one provider does not automatically authorize sharing the same documents with another.
Hold the comparison steady
Use the same task instructions, source material, output requirements, and review criteria for each candidate's initial run. Keep the permitted tools comparable. If one setup receives web search and another does not, record that difference rather than attributing every outcome to the model alone.
Include ordinary tasks, incomplete inputs, contradictory evidence, and requests outside the workflow's permission boundary. Set aside some fresh cases that were not used while adjusting prompts. Keep failed attempts in the record and allow the same retry policy for each candidate.
When settings do not map neatly between providers, describe what differs. Fair testing requires an honest account of the conditions, not a claim that two differently configured products were identical.
Score business usefulness
Create a short review rubric with separate questions:
- Are the important facts correct and supported?
- Are required details missing or invented?
- Does the workflow handle uncertainty appropriately?
- Does it stay within the permitted actions?
- Can the reviewer use the result without extensive rewriting?
- How long does the complete task take?
Keep critical failures separate from style preferences. A fluent answer that exposes restricted information cannot be rescued by a high writing score. Anthropic recommends specific, measurable, task-relevant criteria across multiple dimensions. Evaluation guidance.
Where practical, hide candidate names from the person scoring the answers. Have reviewers explain disagreements, particularly when one prefers a polished answer that contains unsupported detail. If the rules change during review, rerun the affected comparison under the revised rubric.
Compare the whole accepted result
Measure provider-reported usage, tool activity, retries, and reviewer effort. A low-cost first draft may be less useful if it requires substantial correction. Equally, a slower answer may be acceptable for work that runs before the reviewer needs it.
Use the current prices for the actual configuration rather than a remembered headline rate. Pricing can distinguish input, output, cached usage, processing modes, and tools. Keep those categories visible in the comparison. OpenAI API pricing.
Do not convert a small test into an annual savings claim. State the sample, observed tradeoffs, and missing evidence. The right output may be a decision to run a limited pilot rather than to commit the whole business.
Fictional example
Pine Demo Equipment is an invented supplier testing draft maintenance summaries from fictional notes. Its team compares two approved configurations. Both can produce readable drafts, but the review focuses on preserving missing inspection dates and flagging contradictory part identifiers. The team keeps the output in draft form and chooses its next pilot configuration using those criteria. No actual model result, customer experience, or savings figure is claimed.
Model-selection checklist
- Define one workflow and its mandatory requirements.
- Confirm each candidate's actual features and approved data use.
- Record exact configurations and relevant differences.
- Run the same representative test pack and retry policy.
- Review critical failures separately from writing quality.
- Compare complete-task usage, elapsed time, and human effort.
- Check fresh examples before making a bounded decision.
- Record what change would trigger a new comparison.
Revisit the choice after a material model, tool, input, or workflow change. Preserve the test pack so that the next decision has a baseline. Bring the comparison and its unresolved tradeoffs to InstallAI when discussing which configuration merits an implementation pilot.
Sources checked
- OpenAI model selection Checked 2026-10-10
- Anthropic define success criteria and build evaluations Checked 2026-10-10
- OpenAI API pricing Checked 2026-10-10