Skip to content
Luminesca.
Analysis · A closer look

How to Compare AI Models: A Repeatable Task and Cost Worksheet

Compare model candidates using fixed tasks, explicit acceptance criteria, failure records and cost per accepted result. Includes an illustrative calculation and a reusable worksheet.

Editorial correction — September 22, 2026. The former August model ranking mixed launch, price and benchmark claims without linking each claim to its supporting source. We have removed those rankings, prices and FAQ answers. This replacement is Luminesca's proposed evaluation method, not a report of a benchmark we ran or a list of current vendor offers.

Start with one decision and one task

Write the decision before choosing candidates: for example, “Which candidate can convert our approved test records into valid JSON with the least correction work?” Avoid combining writing style, coding, translation and research into a single score unless those are genuinely parts of the same job.

Use representative, non-sensitive examples. Include ordinary inputs and known failure cases. For structured extraction, that might mean an empty field, a quoted delimiter, an embedded newline, a duplicate key and a value that must remain a string. These cases have clear checks; an attractive-looking answer alone does not pass.

Define success before reading the answer

For each case, save the required output or an acceptance checklist. A JSON task can require valid syntax, a fixed set of keys, preserved values and no additional prose. A source-based answer can require a supporting passage for each factual assertion and an explicit “unknown” where the supplied material is insufficient.

Separate a failed format from a wrong fact. A parser can detect invalid JSON, but valid JSON can still contain invented values. For this site's own converter, the CSV repair case study documents concrete data-preservation cases. Those checks illustrate useful test design; they are not comparative model results.

Keep the comparison reproducible

  • Record the exact model identifier, provider, date and accessible interface.
  • Save the same prompt, attachments and tool permissions for each candidate.
  • Record settings, timeout, retry limit and any manual edits.
  • Separate a first attempt from a retry; retain failures as well as successes.
  • Use multiple runs when variation matters and report the number attempted.

If one system can browse and another cannot, you are comparing two configured workflows. That can be a useful product decision, but label it accordingly. A vendor's maximum context size also does not prove that your task will retain every important detail near that limit.

Compare cost per accepted result

Illustrative calculation, not a measured benchmark: candidate A costs $2.40 across 20 attempts, of which 18 pass your predefined checks. Its request cost per accepted result is $2.40 ÷ 18, approximately $0.133. Candidate B costs $1.80 for 20 attempts but has 10 accepted results, giving $0.180 per accepted result. The smaller total bill does not automatically produce the lower cost per useful result.

Include failed requests and retries in the numerator. Track reviewer time separately, and state how it is valued if you include it in the total. When no attempt passes, report zero accepted results and leave cost per accepted result undefined; do not divide by zero or display a misleading zero cost.

Before using a provider's pricing, check its current price page for the exact model, billing unit, input/output split, cached tokens and tool charges. Chat subscriptions and API usage are different purchasing contexts; this page supplies no current price quotations.

Copy this comparison worksheet

RecordWhat to enter
CandidateProvider, exact model identifier, settings, test date
CaseInput identifier, expected result, pass/fail criteria
OutcomePass, failure category, corrections and retry count
TimeTotal elapsed time and reviewer time, separately
CostActual request charges including failed attempts; pricing source and date
DecisionAccepted results / attempts, cost per accepted result, unresolved failures

What a benchmark can and cannot establish

A published benchmark is useful background, but its task and date matter. Stanford's 2025 AI Index describes a more than 280-fold reduction in the inference price of reaching GPT-3.5-level performance on MMLU between November 2022 and October 2024. This is a particular historical comparison, not a universal discount on every model or a measurement of your present workload. Source: Stanford AI Index 2025, research and development.

Keep a final set of cases out of prompt tuning, then check the chosen workflow on them. Small samples leave uncertainty; state the sample size and conditions instead of claiming an overall winner. Repeat the relevant checks when the model, prompt or provider behavior changes.

Separate investment, adoption and electricity statistics · A concrete browser reliability case: cancellable regular expressions