Claude model selection
Anthropic’s model overview lists model capabilities and limits, and documents a Models API for querying capabilities programmatically.
The documentation lists model inputs and outputs, token limits, pricing, and capabilities. It explains using the Models API to query supported features.
Which model is worth testing on this workload?
Listed capabilities and comparative latency are not independent proof of reliability, cost per successful task, or performance on your workload.
Versions, effort settings, prompt length, retries, provider availability, and pricing changes affect a deployment decision.
Workload-specific model evaluation
Not run. This is a protocol, not a result.
- Define a narrow workload, a held-out sample, and an acceptance rubric before choosing candidates.
- Record model IDs, provider, effort, prompts, limits, and the pricing observed at test time.
- Run the same cases across candidates. Record failures, retries, output quality, latency, and total usage.
- Compare cost per accepted result and inspect failure modes. Do not generalize outside the measured workload.
Record these measurements
- Acceptance rate on held-out tasks
- Latency distribution
- Cost per accepted result
- Retry and refusal behavior
- Unsupported or incorrect claims
Keep all attempts, including failures. A ten-task trial is exploratory; report its limits rather than generalizing to other workloads.
One source. Clear boundaries.
Models overview
Model overview reviewed. No model API calls, benchmark runs, or model ranking performed.
Inspected 11 Oct 2026. Documentation can change after review.Have a useful independent source or a correction? Send it here.