01 / What’s documented

Claude model selection

Anthropic’s model overview lists model capabilities and limits, and documents a Models API for querying capabilities programmatically.

The documentation lists model inputs and outputs, token limits, pricing, and capabilities. It explains using the Models API to query supported features.

02 / What’s still open

Which model is worth testing on this workload?

Listed capabilities and comparative latency are not independent proof of reliability, cost per successful task, or performance on your workload.

Versions, effort settings, prompt length, retries, provider availability, and pricing changes affect a deployment decision.

03 / Proposed test

Workload-specific model evaluation

Not run. This is a protocol, not a result.

  1. Define a narrow workload, a held-out sample, and an acceptance rubric before choosing candidates.
  2. Record model IDs, provider, effort, prompts, limits, and the pricing observed at test time.
  3. Run the same cases across candidates. Record failures, retries, output quality, latency, and total usage.
  4. Compare cost per accepted result and inspect failure modes. Do not generalize outside the measured workload.

Record these measurements

  • Acceptance rate on held-out tasks
  • Latency distribution
  • Cost per accepted result
  • Retry and refusal behavior
  • Unsupported or incorrect claims

Keep all attempts, including failures. A ten-task trial is exploratory; report its limits rather than generalizing to other workloads.

04 / Source record

One source. Clear boundaries.

Original / Anthropic

Models overview

Model overview reviewed. No model API calls, benchmark runs, or model ranking performed.

Inspected 11 Oct 2026. Documentation can change after review.

Have a useful independent source or a correction? Send it here.

Next in the collectionClaude Code →