Sonto Research / Evals

Release
Version 1
Frozen

Frontier Financial Judgement: Can agents tell what might move a stock?

A benchmark of whether research agents can distinguish genuinely new, valuation-relevant financial information from realistic noise.

82

Cases

14

Models

656

Supplied items

246

Expert-assigned labels

Aggregate result

Pareto frontier for financial judgement

Pareto frontier for financial judgementFourteen models plotted by total recorded cost on a logarithmic horizontal axis and the proportion of 82 cases with all three target labels correct on the vertical axis. A line connects the observed non-dominated configurations on the cost and accuracy Pareto frontier.LOWER COST · HIGHER ACCURACYHIGHER COST · HIGHER ACCURACYLOWER COST · LOWER ACCURACYHIGHER COST · LOWER ACCURACY0%10%20%30%40%50%60%$1$3$10$30$100$300Total recorded cost (USD, log scale)Strict all-label accuracyGPT-5.5GPT-5.6 SolGPT-5.6 TerraGPT-5.6 LunaClaude Opus 4.8Claude Sonnet 4.6GPT-5.4GPT-5.4 miniDeepSeek V4 FlashDeepSeek V4 ProMiMo-V2.5Qwen3.7 PlusQwen3.6 35B-A3BNemotron 3 Super
Exact results for all 14 models
ModelAll-label accuracyRecorded costEstimated context FPStatus
GPT-5.5 (medium)43/82 · 52.4%$7637/574 · 6.4%0 invalid · 0 failed
GPT-5.6 Sol (medium)42/82 · 51.2%$846/574 · 1.0%0 invalid · 0 failed
GPT-5.6 Terra (medium)38/82 · 46.3%$2629/574 · 5.1%0 invalid · 0 failed
GPT-5.6 Luna (medium)35/82 · 42.7%$1284/574 · 14.6%0 invalid · 0 failed
Claude Opus 4.8 (medium)32/82 · 39.0%$206143/574 · 24.9%0 invalid · 0 failed
Claude Sonnet 4.6 (medium)30/82 · 36.6%$111185/574 · 32.2%0 invalid · 0 failed
GPT-5.4 (medium)33/82 · 40.2%$4233/574 · 5.7%0 invalid · 0 failed
GPT-5.4 mini (medium)25/82 · 30.5%$1555/567 · 9.7%0 invalid · 1 failed
DeepSeek V4 Flash28/82 · 34.1%$0.97108/567 · 19.0%0 invalid · 1 failed
DeepSeek V4 Pro (high)33/82 · 40.2%$3.1599/574 · 17.2%0 invalid · 0 failed
MiMo-V2.5 (medium-equiv.)27/82 · 32.9%$1.54117/567 · 20.6%1 invalid · 0 failed
Qwen3.7 Plus (medium-equiv.)17/82 · 20.7%$6.52134/539 · 24.9%1 invalid · 4 failed
Qwen3.6 35B-A3B18/82 · 22.0%$3.7084/546 · 15.4%3 invalid · 1 failed
Nemotron 3 Super7/82 · 8.5%$8.5431/196 · 15.8%31 invalid · 23 failed
Highest all-label accuracy
52.4%
GPT-5.5 (medium)
Highest atomic accuracy
71.1%
GPT-5.5 (medium)
Estimated context FP range
1.0%–32.2%
Across released configurations
Avg web searches per case
10.6
Per model-case evaluation

Judgement, not sentiment

Financial news is easy to summarise and hard to judge. A headline can be positive while the financially relevant comparison is negative. A familiar number can hide a small change in wording that turns an approximate target into a ceiling. A genuine new fact can still be too small to matter.

Frontier Financial Judgement tests this harder decision. For each item, an agent has to decide whether the information is new, how important it is likely to be for valuation, and whether the expected effect is positive, negative, neutral, or unclear. These labels work together. Finding a new fact without understanding its significance is not enough, and sentiment is not a substitute for financial reasoning.

The benchmark was developed with professional equity analysts. It focuses on the part of news monitoring that depends on the prior public record, the right financial comparison, and an informed view of what could change expectations.

A signal hidden in noise

Each of the 82 cases contains eight supplied items. One is a controlled synthetic target built from an expert-authored event. Five are recently collected Live Articles and two are Historical Docs. Those seven real items are curated distractors: they are not materials used to author the target, and they do not establish the current answer.

The agent is not told which item is the target. It must assess all eight. This recreates an important part of real news flow, where the consequential disclosure sits beside repeated announcements, derivative coverage, older company documents, and stories with limited read-through.

Agents can use web search in the harness to investigate company background and the prior public record at the case cutoff. Synthetic targets make that research possible without letting the answer itself appear online. The result is a controlled test inside a realistic information environment.

Benchmark setup

How the realistic benchmark works

How the realistic financial-judgement benchmark worksA human expert combines a synthetic financial event with generic primer paragraphs and an LLM article generator renders a realistic target article. Five related news articles and two historical documents are curated at the case cutoff. The target and seven context items form a locked bundle supplied unchanged to an agent harness with web search, command-line tools, a file system, and an LLM core. The agent assesses newness, importance, and direction.1Controlled target constructionHuman expertExpert-authoredsynthetic eventGeneric primerparagraphsLLM article generatorpublisher.example / company-newsSynthetic news articlerealistic presentationAbout  Media  Contact  Privacy2Real-world contextReal-world eventsSampled relatednews articles5 PER CASEHistoricaldocuments2 PER CASEcurated at the case cutoff3Frozen realistic caseLocked file bundlesupplied unchanged to every agent1 target5 articles2 documents4Agent environmentWEBSEARCHCLITOOLSFILESYSTEMAgent harnessLLM CORE5Agent assessmentNEWNESSIMPORTANCEDIRECTION

Scoring

Scoring keeps the joint decision visible. Atomic accuracy is the share of the three target labels answered correctly. All-label accuracy requires all three to match the expert assessment on a case. Invalid outputs and execution failures stay in the denominator and score as incorrect rather than quietly disappearing.

We also report an estimated context false-positive rate. It counts parseable distractor decisions that an agent marks as both new and more important than none. This is a useful measure of restraint, but it is not a conventional false-positive rate: the distractors are curated noise, not an independently expert-labelled negative dataset.