Benchmark investigation · July 2026

A five-model benchmark in quantitative finance

Two independently constructed task sets, five anonymous answers per task, and one model-blinded review protocol. This report presents the observed measurements without designating an overall winner.

01 / Study design

What was evaluated

The two task sets were constructed separately and use different taxonomies. Their results are presented side by side rather than treated as interchangeable samples.

Task set A

Original quantitative-finance benchmark

Broad coverage across quantitative-finance domains, response formats, difficulty tiers, and narrowly defined subdomains.

96tasks
480judgments
16domains
8formats
Task set B

Code and practical-methodology supplement

Two hundred single-turn tasks focused on code construction, diagnosis, and combined correction with practical methodology.

200tasks
1,000judgments
3categories
5answers each
Evaluation protocol. Each batch was reviewed in one Codex judging pass with five anonymous answers per task. The judge assigned a binary verdict and seven-level quality label to each answer, followed by a strict rank from 1 to 5. Model names were revealed only after each review was complete.
02 / Overall results

Observed performance by task set

The two panels use the same model order and metric definitions. Acceptance and rank-1 cells include their underlying counts; lower mean rank is better.

Quality-label distributions

Each bar contains every answer in the task set, ordered from fully correct to non-answer.

Exact counts and full distributions Includes the task-count-weighted combined view
Overall model statistics
ModelAcceptRejectMean qualityQuality-label distributionMean rankRank distributionRank 1
03 / Pairwise ranks

How often each model ranked above another

Each cell is the row model’s share of forced-rank wins against the column model. The two task sets share one color scale and model order.

Column model leadsRow model leads Near-neutral cells are approximately even; stronger color means a larger margin.

Original benchmark

96 paired decisions per model pair

Code and methodology supplement

200 paired decisions per model pair

Inspect one model pair in detail Acceptance, quality, mean rank, rank 1, and exact p
04 / Task variation

Observed results across task types

These figures show all five models at once across the benchmark’s principal task cuts. Color is normalized within each row, so it compares models for the same task type; exact values remain comparable across rows. Small samples remain visible but are not color-emphasized.

Better observed valueWorse observed valueLower mean rank is better.

Original benchmark · domains

Sixteen quantitative-finance domains; n is shown for every row.

Original benchmark · response formats

Eight answer formats.

Original benchmark · difficulty tiers

Four construction tiers.

Supplement · categories

Code construction, diagnosis, and combined correction/methodology.

Open the exact segment and task-evidence browser All domains, formats, tiers, subdomains, prompts, and judge rationales
Segment statistics for the selected batch, task cut, and model pair
Rank-1 Δ Evidence
05 / Provenance

Response collection metadata

Every campaign retains provider-reported token use, completion status, latency, and cost where the endpoint supplied it.

View the ten collection campaigns Run identifiers, tokens, latency, reported cost, and provider distribution
Cost coverage differs by endpoint. OpenRouter campaigns reported billed cost. The direct QuantBrain endpoints reported token use but no monetary cost, so those cells are marked “not reported” rather than estimated.
Canonical collection campaigns used by this evaluation
BatchModelRunCompleteErrorsPrompt tokensCompletion tokensReasoning tokensTotal tokensp50 / p95 latencyReported costProvider distribution
06 / Methods

Methods, definitions, and limitations

These notes are part of the report and should accompany any externally presented excerpt.

Binary verdict and quality

Accept/reject and the seven-level quality label were recorded as separate fields by the same judge. The report maps the ordered quality labels to equally spaced scores 7–1 for descriptive means; the full distributions are retained alongside them.

ScoreQuality labelProtocol meaning

Forced ranks and pairwise counts

Each task received a permutation of ranks 1–5. A model records a pairwise rank win when its rank is lower than the other model’s rank. Aggregate pairwise totals can tie; individual tasks cannot.

Strict ranks are relative and force differentiation even when two answers are substantively similar. Ranking reasons in the task audit preserve the judge’s stated basis.

Exact p-values

The two-sided exact binomial calculation compares paired rank wins with a 50% reference. Values are exploratory, are not adjusted for multiple comparisons, and do not account for task construction or reviewer uncertainty.

No p-value is presented as a standalone decision rule or as confirmatory evidence.

Limitations

    Subdomain rows are especially small: 82 contain one task and seven contain two. Treat those rows as task-level observations.

    Downloadable data

    All generated files come from deterministic builders. The analysis JSON and CSV files contain statistics, prompts, judgments, and rationales. The two ZIP files add the complete task inputs and canonical API response records, arranged task by task.

    Task evidence