Aggregate result
Pareto frontier for financial judgement
Exact results for all 14 models
| Model | All-label accuracy | Recorded cost | Estimated context FP | Status |
|---|---|---|---|---|
| GPT-5.5 (medium) | 43/82 · 52.4% | $76 | 37/574 · 6.4% | 0 invalid · 0 failed |
| GPT-5.6 Sol (medium) | 42/82 · 51.2% | $84 | 6/574 · 1.0% | 0 invalid · 0 failed |
| GPT-5.6 Terra (medium) | 38/82 · 46.3% | $26 | 29/574 · 5.1% | 0 invalid · 0 failed |
| GPT-5.6 Luna (medium) | 35/82 · 42.7% | $12 | 84/574 · 14.6% | 0 invalid · 0 failed |
| Claude Opus 4.8 (medium) | 32/82 · 39.0% | $206 | 143/574 · 24.9% | 0 invalid · 0 failed |
| Claude Sonnet 4.6 (medium) | 30/82 · 36.6% | $111 | 185/574 · 32.2% | 0 invalid · 0 failed |
| GPT-5.4 (medium) | 33/82 · 40.2% | $42 | 33/574 · 5.7% | 0 invalid · 0 failed |
| GPT-5.4 mini (medium) | 25/82 · 30.5% | $15 | 55/567 · 9.7% | 0 invalid · 1 failed |
| DeepSeek V4 Flash | 28/82 · 34.1% | $0.97 | 108/567 · 19.0% | 0 invalid · 1 failed |
| DeepSeek V4 Pro (high) | 33/82 · 40.2% | $3.15 | 99/574 · 17.2% | 0 invalid · 0 failed |
| MiMo-V2.5 (medium-equiv.) | 27/82 · 32.9% | $1.54 | 117/567 · 20.6% | 1 invalid · 0 failed |
| Qwen3.7 Plus (medium-equiv.) | 17/82 · 20.7% | $6.52 | 134/539 · 24.9% | 1 invalid · 4 failed |
| Qwen3.6 35B-A3B | 18/82 · 22.0% | $3.70 | 84/546 · 15.4% | 3 invalid · 1 failed |
| Nemotron 3 Super | 7/82 · 8.5% | $8.54 | 31/196 · 15.8% | 31 invalid · 23 failed |
- Highest all-label accuracy
- 52.4%
- GPT-5.5 (medium)
- Highest atomic accuracy
- 71.1%
- GPT-5.5 (medium)
- Estimated context FP range
- 1.0%–32.2%
- Across released configurations
- Avg web searches per case
- 10.6
- Per model-case evaluation
Judgement, not sentiment
Financial news is easy to summarise and hard to judge. A headline can be positive while the financially relevant comparison is negative. A familiar number can hide a small change in wording that turns an approximate target into a ceiling. A genuine new fact can still be too small to matter.
Frontier Financial Judgement tests this harder decision. For each item, an agent has to decide whether the information is new, how important it is likely to be for valuation, and whether the expected effect is positive, negative, neutral, or unclear. These labels work together. Finding a new fact without understanding its significance is not enough, and sentiment is not a substitute for financial reasoning.
The benchmark was developed with professional equity analysts. It focuses on the part of news monitoring that depends on the prior public record, the right financial comparison, and an informed view of what could change expectations.
A signal hidden in noise
Each of the 82 cases contains eight supplied items. One is a controlled synthetic target built from an expert-authored event. Five are recently collected Live Articles and two are Historical Docs. Those seven real items are curated distractors: they are not materials used to author the target, and they do not establish the current answer.
The agent is not told which item is the target. It must assess all eight. This recreates an important part of real news flow, where the consequential disclosure sits beside repeated announcements, derivative coverage, older company documents, and stories with limited read-through.
Agents can use web search in the harness to investigate company background and the prior public record at the case cutoff. Synthetic targets make that research possible without letting the answer itself appear online. The result is a controlled test inside a realistic information environment.
Benchmark setup
How the realistic benchmark works
Scoring
Scoring keeps the joint decision visible. Atomic accuracy is the share of the three target labels answered correctly. All-label accuracy requires all three to match the expert assessment on a case. Invalid outputs and execution failures stay in the denominator and score as incorrect rather than quietly disappearing.
We also report an estimated context false-positive rate. It counts parseable distractor decisions that an agent marks as both new and more important than none. This is a useful measure of restraint, but it is not a conventional false-positive rate: the distractors are curated noise, not an independently expert-labelled negative dataset.