Official leaderboard · Multilingual document parsing
How well do models parse
documents in the real world?
A unified evaluation across 17 languages, native digital pages, and photographed documents. Scores are shown as percentages; higher is better.
3,400document images
17languages
3parsing tasks
2evaluation splits
Verified results
Leaderboard
How to read this board
One number, with the context still visible.
The public score is calculated over released samples. Private scores use the held-out split and are the basis for verified comparison. Digital and Photo expose the real-world robustness gap; Latin and Non-Latin avoid hiding language imbalance behind an average.
To add a model, run the released evaluation pipeline, then provide inference code, weights or API details, and predictions for official private-set evaluation.