TASK 3 / PARTICIPANT HUB
Financial Audit Verification
Read an SEC XBRL filing and report two numbers: the value the filing states for a target concept, and the value its own calculation relationships imply. This is targeted numeric-fact verification, not a full financial-statement audit.
- Practice
- Open now
- Competition closes
- 15 Oct 2026 · 23:59 AoE
THE THREE PHASES
What each phase gives you back
All three phases accept submissions. They differ in what comes back: practice scores you against answers you already have, development scores and ranks you against answers you do not, and test takes your work and tells you only that it arrived.
| Phase | Status | Answers | What you get back | Ranking |
|---|---|---|---|---|
| Practice | Live | Public — the 332 FinMR cases, answers included | Scored the moment you upload | Never ranked |
| Development | Live | 680 held-out cases, answers withheld | Score, rank, and a public leaderboard | Ranked |
| Test | Live | 680 held-out cases, answers withheld | An acceptance receipt only | Decides the final result |
Practice is never ranked, and the reason is worth stating plainly: its answers are public, so any team could score 100% by reading them. It exists so you can rehearse the submission format against the real endpoint before it counts.
DATA
Where the data comes from
This site hosts no Task 3 files. The 332 public practice cases are the Financial Mathematical Reasoning task of the FinAuditing benchmark, released as TheFinAI/FinMR. The Task 3 starter kit downloads them for you, and ships the validator, the scorer, and one baseline per way of running a model.
The development and test questions are published separately, without their answers, as YanAdjeNole/FinReason-Task3— 680 cases each. They carry no rule label: for one of the three rules the label alone gives away the answer’s shape.
Every one of the 332 practice cases is built around one flagged Data Quality Committee rule — 110 for DQC_US_0015, 120 for DQC_US_0117, 102 for DQC_US_0126 — and in all of them the reported and calculated values disagree. A system that always predicted agreement would score zero.
NEXT
Run the practice set
The submission guide walks through producing predictions.jsonl and uploading it. The leaderboard page explains how the four rates are computed and shows the organizer baseline.