Original quantitative-finance benchmark
Broad coverage across quantitative-finance domains, response formats, difficulty tiers, and narrowly defined subdomains.
Two independently constructed task sets, five anonymous answers per task, and one model-blinded review protocol. This report presents the observed measurements without designating an overall winner.
The two task sets were constructed separately and use different taxonomies. Their results are presented side by side rather than treated as interchangeable samples.
Broad coverage across quantitative-finance domains, response formats, difficulty tiers, and narrowly defined subdomains.
Two hundred single-turn tasks focused on code construction, diagnosis, and combined correction with practical methodology.
The two panels use the same model order and metric definitions. Acceptance and rank-1 cells include their underlying counts; lower mean rank is better.
Each bar contains every answer in the task set, ordered from fully correct to non-answer.
| Model | Accept | Reject | Mean quality | Quality-label distribution | Mean rank | Rank distribution | Rank 1 |
|---|
Each cell is the row model’s share of forced-rank wins against the column model. The two task sets share one color scale and model order.
96 paired decisions per model pair
200 paired decisions per model pair
These figures show all five models at once across the benchmark’s principal task cuts. Color is normalized within each row, so it compares models for the same task type; exact values remain comparable across rows. Small samples remain visible but are not color-emphasized.
Sixteen quantitative-finance domains; n is shown for every row.
Eight answer formats.
Four construction tiers.
Code construction, diagnosis, and combined correction/methodology.
| Rank-1 Δ | Evidence |
|---|
| Model | Accept | Reject | Mean quality | Quality distribution | Mean rank | Rank distribution | Rank 1 |
|---|
Every verdict, quality label, rank, prompt, and judge rationale is available below.
Every campaign retains provider-reported token use, completion status, latency, and cost where the endpoint supplied it.
| Batch | Model | Run | Complete | Errors | Prompt tokens | Completion tokens | Reasoning tokens | Total tokens | p50 / p95 latency | Reported cost | Provider distribution |
|---|
These notes are part of the report and should accompany any externally presented excerpt.
Accept/reject and the seven-level quality label were recorded as separate fields by the same judge. The report maps the ordered quality labels to equally spaced scores 7–1 for descriptive means; the full distributions are retained alongside them.
| Score | Quality label | Protocol meaning |
|---|
Each task received a permutation of ranks 1–5. A model records a pairwise rank win when its rank is lower than the other model’s rank. Aggregate pairwise totals can tie; individual tasks cannot.
Strict ranks are relative and force differentiation even when two answers are substantively similar. Ranking reasons in the task audit preserve the judge’s stated basis.
The two-sided exact binomial calculation compares paired rank wins with a 50% reference. Values are exploratory, are not adjusted for multiple comparisons, and do not account for task construction or reviewer uncertainty.
No p-value is presented as a standalone decision rule or as confirmatory evidence.
Subdomain rows are especially small: 82 contain one task and seven contain two. Treat those rows as task-level observations.
All generated files come from deterministic builders. The analysis JSON and CSV files contain statistics, prompts, judgments, and rationales. The two ZIP files add the complete task inputs and canonical API response records, arranged task by task.