One standard contour confirms
Mode B on 144 questions records FUS +7.915 with CI [+6.010, +9.714] and clears the configured decision rule.
Standalone benchmark investigation
A source-pinned review of three benchmark notebooks covering the QuantBrain 35B advisory endpoint and QuantBrain-code-397B-A17B. Results are reported by protocol and are not collapsed into one model score.
The advisory experiments show a positive product-relative result against GLM in one report-defined contour, but no demonstrated positive uplift against the Qwen3.5-35B-A3B comparator. The code experiments show opposing nominal, unadjusted signals—a debugging advantage and a closed-loop deficit—and no aggregate advantage.
Mode B on 144 questions records FUS +7.915 with CI [+6.010, +9.714] and clears the configured decision rule.
Four confidence intervals cross zero. The fifth is negative by interval and borderline under the paired sign-flip test.
Across 150 executable tasks, delta is +0.0100 with CI [−0.0566, +0.0787].
The judged tracks contain 6/150 critical errors for QuantBrain versus 1/150 for Qwen3.5-397B.
The three notebooks cover two distinct model sizes and evaluation systems. Their scales and outcomes answer different questions.
| Stratum | Candidate | Comparators | Task design | Score | Interpretation boundary |
|---|---|---|---|---|---|
| Financial advisory | QuantBrain 35B | GLM-4.7-Flash; Qwen3.5-35B-A3B | Paired Russian financial questions; 144- and 300-question views | Paired difference in 0–100 scores | Product-relative and probable-base comparisons; lineage is not checkpoint-attested |
| Code and applied methodology | QuantBrain-code-397B-A17B | Qwen3.5-397B-A17B | 300 tasks across six types, five tiers, and ten domains | Executable 0–1; judged 0–100 | The two evaluator families cannot support a single unified score |
FUS is candidate minus comparator. Positive values favor QuantBrain under the specified protocol.
GLM is a different model family, so that comparison does not isolate the effect of financial fine-tuning.
Qwen3.5-35B-A3B is the more relevant comparator, but an immutable pre-fine-tuning checkpoint identity was not supplied.
The Qwen retest combines candidate judgments from 14 July with comparator judgments from 16 July rather than rerunning both sides in one randomized wave. One advisory judge lane was not pinned across the early waves.
Mode-B OpenRouter answers were truncated with an estimated 3.2 characters per token, while QuantBrain used token-exact vLLM truncation.
Only one standard contour satisfies every configured guardrail
| Contour | n | FUS | 95% confidence interval | Critical rate · QB / GLM | Recorded decision |
|---|---|---|---|---|---|
| Mode A · 144 | 144 | +5.102 | [+2.991, +7.137] | 86.11% / 88.89% | Not confirmed |
| Mode B · 144 | 144 | +7.915 | [+6.010, +9.714] | 83.33% / 88.89% | Confirmed uplift |
| Mode A · 300 | 300 | +3.536 | [+1.879, +5.211] | 79.67% / 80.00% | Not confirmed |
| Mode B · 300 | 300 | +3.975 | [+2.206, +5.751] | 76.67% / 80.00% | Not confirmed |
| 2,048-token diagnostic · 144 | 144 | +17.118 | [+13.983, +20.257] | 41.67% / 73.61% | Diagnostic; not confirmed |
Saved decisions use product-bank thresholds of +6 for 144 questions and +5 for 300, plus interval, win-rate, critical-rate, and block guardrails.
No contour confirms positive uplift
| Contour | n | FUS | 95% confidence interval | Paired interpretation |
|---|---|---|---|---|
| Mode A · 144 | 144 | −0.259 | [−1.633, +1.104] | No detectable difference |
| Mode B · 144 | 144 | −1.139 | [−2.824, +0.550] | No detectable difference |
| Mode A · 300 | 300 | +0.830 | [−0.319, +2.003] | No detectable difference |
| Mode B · 300 | 300 | −1.902 | [−3.530, −0.247] | Negative CI; sign-flip p=0.0504 |
| 2,048-token diagnostic · 144 | 144 | +0.216 | [−2.229, +2.652] | No detectable difference |
The mode-B/300 result is fragile negative evidence: its bootstrap interval excludes zero, but the paired sign-flip p-value is 0.050395 and the recorded decision remains not_confirmed.
Why these views are not independent confirmations
126 of the smaller bank’s 144 IDs recur in the 300 bank. There are 319 unique prompts overall.
The candidate’s raw length-stop rate falls from 100% to 20.1% when the 144-question mode-A budget increases to 2,048 tokens.
Two advisory bank errors were corrected after older comparator judging without a complete rerun of that comparator.
The runs do not establish a practical advantage over Qwen3.5-35B-A3B; they do not prove that fine-tuning has no effect.
not_confirmed. The mode-B/300 view contains
only 234 pairs, with 66 missing. Claude Fable 5 and the GPT-5.6 Pro
reference are useful frontier anchors, not peer baselines; the GPT
answer was also visible to the judge.
The 300-task bank contains six types, five tiers, and ten domains. Objective execution and LLM-judged methodology tracks remain separate.
Generation and debugging use public fixtures and numeric checks in an eight-second subprocess. Closed-loop tasks use public and hidden fixtures over two turns.
The intended pinned Docker/no-network sandbox was unavailable. The lite evaluator can reject forbidden tokens and trim trailing code until the AST parses.
Analysis/recommendations, design specifications, and production operations are scored blindly on a 0–100 rubric.
The run uses one judge per answer. The release design calls for two judges plus adjudication, so this is a pilot rather than a release-grade verdict.
Score range 0–1
| Track | n | QuantBrain 397B | Qwen3.5-397B | Delta | 95% confidence interval | Inference |
|---|---|---|---|---|---|---|
| Generation + debug | 100 | 0.6008 | 0.5312 | +0.0697 | [−0.0170, +0.1580] | Not significant |
| Code generation | 50 | 0.6617 | 0.7043 | −0.0427 | [−0.1440, +0.0577] | Not significant |
| Code debugging | 50 | 0.5400 | 0.3580 | +0.1820 | [+0.0517, +0.3183] | Nominal candidate advantage |
| Closed loop | 50 | 0.6651 | 0.7743 | −0.1092 | [−0.2085, −0.0127] | Nominal comparator advantage |
| Combined objective | 150 | — | — | +0.0100 | [−0.0566, +0.0787] | Not significant |
The debugging advantage and closed-loop deficit are nominal findings without a preregistered multiplicity correction. The corresponding Wilcoxon/permutation p-values are 0.0177/0.0099 and 0.0229/0.0373.
Score range 0–100; FUS is QuantBrain minus Qwen
| Type | n | QuantBrain 397B | Qwen3.5-397B | FUS | 95% confidence interval |
|---|---|---|---|---|---|
| Analyze / recommend | 50 | 87.70 | 90.20 | −2.500 | [−5.250, −0.025] |
| Design specification | 50 | 81.40 | 82.58 | −1.175 | [−4.825, +2.450] |
| Production operations | 50 | 81.53 | 81.58 | −0.050 | [−3.625, +3.251] |
| Overall | 150 | 83.54 | 84.78 | −1.242 | [−3.133, +0.575] |
What the score does and does not establish
A solution specialized to the public case can receive full credit without being generally correct.
Grounding and JSON shape are checked more directly than the correctness of diagnosis values and conclusions.
Answers, extracted code, server identity, finish reasons, retries, and complete execution logs were not persisted.
The committed closed-loop script writes fewer fields than later notebook cells expect, so a clean rerun can break downstream analysis.
The retained outputs support a careful protocol-specific reading, but not a universal model ranking.
Only evidence belonging to this benchmark investigation is listed here.