Standalone benchmark investigation

Crucible test evidence audit

A source-pinned review of three benchmark notebooks covering the QuantBrain 35B advisory endpoint and QuantBrain-code-397B-A17B. Results are reported by protocol and are not collapsed into one model score.

Executive conclusion

The evidence is useful, protocol-specific, and mixed.

The advisory experiments show a positive product-relative result against GLM in one report-defined contour, but no demonstrated positive uplift against the Qwen3.5-35B-A3B comparator. The code experiments show opposing nominal, unadjusted signals—a debugging advantage and a closed-loop deficit—and no aggregate advantage.

01
35B · GLM comparison

One standard contour confirms

Mode B on 144 questions records FUS +7.915 with CI [+6.010, +9.714] and clears the configured decision rule.

35B · Qwen comparator

No positive uplift confirms

Four confidence intervals cross zero. The fifth is negative by interval and borderline under the paired sign-flip test.

397B · objective aggregate

Combined result is near zero

Across 150 executable tasks, delta is +0.0100 with CI [−0.0566, +0.0787].

397B · safety

Critical-error guardrail fails

The judged tracks contain 6/150 critical errors for QuantBrain versus 1/150 for Qwen3.5-397B.

Study boundary

Two benchmark strata, retained separately.

The three notebooks cover two distinct model sizes and evaluation systems. Their scales and outcomes answer different questions.

02
StratumCandidateComparatorsTask designScoreInterpretation boundary
Financial advisory QuantBrain 35B GLM-4.7-Flash; Qwen3.5-35B-A3B Paired Russian financial questions; 144- and 300-question views Paired difference in 0–100 scores Product-relative and probable-base comparisons; lineage is not checkpoint-attested
Code and applied methodology QuantBrain-code-397B-A17B Qwen3.5-397B-A17B 300 tasks across six types, five tiers, and ten domains Executable 0–1; judged 0–100 The two evaluator families cannot support a single unified score
The 35B and 397B results are not a common leaderboard. They involve different endpoints, comparators, question banks, budgets, scoring systems, and judges.
Advisory benchmark

QuantBrain 35B

FUS is candidate minus comparator. Positive values favor QuantBrain under the specified protocol.

03

Protocol

  • Mode A: thinking off, temperature 0, answer budget capped at 800 tokens.
  • Mode B: thinking on, with the final response targeted to the same task budget through different truncation paths.
  • Diagnostic: thinking off with a common 2,048-token answer budget.
  • Scores are paired, balanced across block × type cells, and summarized with stratified bootstrap intervals.

Attribution limits

GLM is a different model family, so that comparison does not isolate the effect of financial fine-tuning.

Qwen3.5-35B-A3B is the more relevant comparator, but an immutable pre-fine-tuning checkpoint identity was not supplied.

The Qwen retest combines candidate judgments from 14 July with comparator judgments from 16 July rather than rerunning both sides in one randomized wave. One advisory judge lane was not pinned across the early waves.

Mode-B OpenRouter answers were truncated with an estimated 3.2 characters per token, while QuantBrain used token-exact vLLM truncation.

1

QuantBrain 35B versus GLM-4.7-Flash

Only one standard contour satisfies every configured guardrail

ContournFUS95% confidence intervalCritical rate · QB / GLMRecorded decision
Mode A · 144144+5.102[+2.991, +7.137]86.11% / 88.89%Not confirmed
Mode B · 144144+7.915[+6.010, +9.714]83.33% / 88.89%Confirmed uplift
Mode A · 300300+3.536[+1.879, +5.211]79.67% / 80.00%Not confirmed
Mode B · 300300+3.975[+2.206, +5.751]76.67% / 80.00%Not confirmed
2,048-token diagnostic · 144144+17.118[+13.983, +20.257]41.67% / 73.61%Diagnostic; not confirmed

Saved decisions use product-bank thresholds of +6 for 144 questions and +5 for 300, plus interval, win-rate, critical-rate, and block guardrails.

2

QuantBrain 35B versus Qwen3.5-35B-A3B

No contour confirms positive uplift

ContournFUS95% confidence intervalPaired interpretation
Mode A · 144144−0.259[−1.633, +1.104]No detectable difference
Mode B · 144144−1.139[−2.824, +0.550]No detectable difference
Mode A · 300300+0.830[−0.319, +2.003]No detectable difference
Mode B · 300300−1.902[−3.530, −0.247]Negative CI; sign-flip p=0.0504
2,048-token diagnostic · 144144+0.216[−2.229, +2.652]No detectable difference

The mode-B/300 result is fragile negative evidence: its bootstrap interval excludes zero, but the paired sign-flip p-value is 0.050395 and the recorded decision remains not_confirmed.

3

Controls affecting interpretation

Why these views are not independent confirmations

Bank overlap

125 prompts repeat byte-for-byte

126 of the smaller bank’s 144 IDs recur in the 300 bank. There are 319 unique prompts overall.

Response budget

Truncation dominates

The candidate’s raw length-stop rate falls from 100% to 20.1% when the 144-question mode-A budget increases to 2,048 tokens.

Rubric stability

Two rows changed between waves

Two advisory bank errors were corrected after older comparator judging without a complete rerun of that comparator.

Conclusion boundary

No demonstrated positive uplift

The runs do not establish a practical advantage over Qwen3.5-35B-A3B; they do not prove that fine-tuning has no effect.

Earlier Qwen3-32B comparison and frontier anchors
The initial approximate comparator produced FUS −0.445, +2.897, −0.463, +0.904, and +4.768 across its five contours; every saved decision is not_confirmed. The mode-B/300 view contains only 234 pairs, with 66 missing. Claude Fable 5 and the GPT-5.6 Pro reference are useful frontier anchors, not peer baselines; the GPT answer was also visible to the judge.
Code benchmark

QuantBrain-code-397B-A17B

The 300-task bank contains six types, five tiers, and ten domains. Objective execution and LLM-judged methodology tracks remain separate.

04

Objective tracks · 150 tasks

Generation and debugging use public fixtures and numeric checks in an eight-second subprocess. Closed-loop tasks use public and hidden fixtures over two turns.

The intended pinned Docker/no-network sandbox was unavailable. The lite evaluator can reject forbidden tokens and trim trailing code until the AST parses.

LLM-judged tracks · 150 tasks

Analysis/recommendations, design specifications, and production operations are scored blindly on a 0–100 rubric.

The run uses one judge per answer. The release design calls for two judges plus adjudication, so this is a pilot rather than a release-grade verdict.

1

Executable results

Score range 0–1

TracknQuantBrain 397BQwen3.5-397BDelta95% confidence intervalInference
Generation + debug1000.60080.5312+0.0697[−0.0170, +0.1580]Not significant
Code generation500.66170.7043−0.0427[−0.1440, +0.0577]Not significant
Code debugging500.54000.3580+0.1820[+0.0517, +0.3183]Nominal candidate advantage
Closed loop500.66510.7743−0.1092[−0.2085, −0.0127]Nominal comparator advantage
Combined objective150+0.0100[−0.0566, +0.0787]Not significant

The debugging advantage and closed-loop deficit are nominal findings without a preregistered multiplicity correction. The corresponding Wilcoxon/permutation p-values are 0.0177/0.0099 and 0.0229/0.0373.

2

Applied-methodology results

Score range 0–100; FUS is QuantBrain minus Qwen

TypenQuantBrain 397BQwen3.5-397BFUS95% confidence interval
Analyze / recommend5087.7090.20−2.500[−5.250, −0.025]
Design specification5081.4082.58−1.175[−4.825, +2.450]
Production operations5081.5381.58−0.050[−3.625, +3.251]
Overall15083.5484.78−1.242[−3.133, +0.575]
Safety guardrail: QuantBrain records six critical errors (4.00%) versus one for Qwen (0.67%), exceeding the permitted one-percentage-point deterioration. The paired exact discordance test is p=0.125, so this is a failed operational guardrail and descriptive count, not evidence for a stable sixfold risk ratio.
3

Evaluator limitations

What the score does and does not establish

Public fixtures

Generation/debug has no hidden behavior set

A solution specialized to the public case can receive full credit without being generally correct.

Closed-loop grading

Turn two is weakly verified

Grounding and JSON shape are checked more directly than the correctness of diagnosis values and conclusions.

Auditability

Raw model outputs are absent

Answers, extracted code, server identity, finish reasons, retries, and complete execution logs were not persisted.

Rerun integrity

Saved summary schema is inconsistent

The committed closed-loop script writes fewer fields than later notebook cells expect, so a clean rerun can break downstream analysis.

Evidence quality

Strong paired structure, incomplete release controls.

The retained outputs support a careful protocol-specific reading, but not a universal model ranking.

05

What is comparatively strong

  • Common question IDs and paired differences.
  • Confidence intervals, guardrails, type-level tables, and retained negative findings.
  • Separate reporting of advisory and executable evaluator families.
  • Pinned notebook blobs and a machine-readable evidence record.

What constrains interpretation

  • Overlapping advisory banks and different judging dates.
  • Strong response-budget sensitivity and widespread truncation.
  • Two corrected rubric rows without a complete comparator rerun.
  • Approximate code sandbox, one judged lane, and missing raw outputs.
  • Several stale notebook narrative cells that conflict with saved tables.
Neutral reading: the runs provide no demonstrated positive advantage over the Qwen3.5-35B comparator, opposing nominal and unadjusted 397B debugging and closed-loop signals, and no aggregate 397B uplift. Those conclusions apply to these exact protocols and retained artifacts.
Evidence index

Crucible source artifacts

Only evidence belonging to this benchmark investigation is listed here.

06
Technical audit Full protocol reconstruction, exact tables, statistical caveats, and neutral interpretation. Open audit →
Machine-readable evidence Notebook identities, paired metrics, confidence intervals, p-values, safety rates, and audit findings. Open JSON →
Pinned benchmark source tree 16d155c5097c62b159109df1e8ac6c2b93cbd816 Open source →
35B Qwen-comparator notebook results_rebase_qwen35.ipynb · blob 85c60fed1277… Open notebook →
397B code benchmark notebook results_code_uplift.ipynb · blob 4ce9c6e28db8… Open notebook →