The runs of this benchmark
One row per CLI run: the audit writes one second statement per question per run, so a prediction file is a run. Every benchmark is indexed a level up.
| run | date | engine | question set | question set digest |
|---|---|---|---|---|
| demo | 2026-09-14T13:13:54.071205+00:00 | sqlite | demo/questions.json | sha256:7ad58addc070b425cabfc9d67bbd9c95840253160cf3b73d672e9aeb4897f5ff |
Data and credit
The audited question sets are BIRD Mini-Dev and BIRD dev, both licensed under CC BY-SA 4.0. This site shows their question text, evidence text and gold SQL as they ship, beside results this project measured. The rest of the site is this project's own work under Apache-2.0, and NOTICE is the precise list of which files carry which.