minidev-sqlite

9 groups 99 runs

The groups of this benchmark

One row per prediction file, and one run per database it names, because a connection here is one file. The numbers are the sums of that group's runs, which they can be because no question is in two of them. Every benchmark is indexed a level up.

group runs audited compared EQUAL NOT_EQUAL ERROR probes fired credited by BIRD and NOT_EQUAL
gpt-35-turbo 11 498 409 164 245 89 53 27
gpt-35-turbo-instruct 11 498 372 144 228 126 49 26
gpt-4 11 498 472 209 263 26 57 32
gpt-4-32k 11 498 459 205 254 39 55 32
gpt-4-turbo 11 498 422 190 232 53 29 76
meta-llama-3-70b-instruct-2 11 498 440 177 263 58 49 29
meta-llama-3-8b-instruct-2 11 498 311 187 207 104 38 21
mistralai-mixtral-8x7b-instru-4 11 498 248 250 157 91 31 16
phi-3-medium-128k-instruct-1 11 498 366 235 131 132 45 25

Data and credit

The audited question sets are BIRD Mini-Dev and BIRD dev, both licensed under CC BY-SA 4.0. This site shows their question text, evidence text and gold SQL as they ship, beside results this project measured. The rest of the site is this project's own work under Apache-2.0, and NOTICE is the precise list of which files carry which.