MDPBench

Official leaderboard · Multilingual document parsing

How well do models parse
documents in the real world?

A unified evaluation across 17 languages, native digital pages, and photographed documents. Scores are shown as percentages; higher is better.

3,400document images
17languages
3parsing tasks
2evaluation splits

Verified results

Leaderboard

How to read this board

One number, with the context still visible.

The public score is calculated over released samples. Private scores use the held-out split and are the basis for verified comparison. Digital and Photo expose the real-world robustness gap; Latin and Non-Latin avoid hiding language imbalance behind an average.

To add a model, run the released evaluation pipeline, then provide inference code, weights or API details, and predictions for official private-set evaluation.