The runs of this benchmark
One row per CLI run: the audit writes one second statement per question per run, so a prediction file is a run. Every benchmark is indexed a level up.
| run | date | engine | question set | question set digest |
|---|---|---|---|---|
| hf | 2026-09-14T12:51:05.775721+00:00 | postgresql | https://huggingface.co/datasets/birdsql/bird_mini_dev/resolve/f65faf4ae3b638c1fa6df1d3370c8d92c8366301/data/mini_dev_pg-00000-of-00001.json (commit f65faf4a, downloaded 2026-09-14) | sha256:7fa740ef9225389cff6c34432120e8325d0ca3008d73db1ae38731234bc10da7 |
| zip | 2026-09-14T12:51:05.775721+00:00 | postgresql | https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json | sha256:d2731292f20b8d8569cd956dd747ffe1df13cd625076263e38ae9ebcef50b1ab |
Data and credit
The audited question sets are BIRD Mini-Dev and BIRD dev, both licensed under CC BY-SA 4.0. This site shows their question text, evidence text and gold SQL as they ship, beside results this project measured. The rest of the site is this project's own work under Apache-2.0, and NOTICE is the precise list of which files carry which.