gold
SELECT GSoffered FROM schools ORDER BY ABS(longitude) DESC NULLS LAST LIMIT 1
- abs(longitude)
- descending, nulls last
sha256:d4d87fb991696c450e12529c6c599b392de69debe7b966d28ef8cccdb2210f15
R-ORD NOT_EQUAL arbitrary-cut not-a-function-of-the-data
california_schools · mini_dev_postgresql from https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json
What is the grade span offered in the school with the highest longitude?
the hint the set supplies: the highest longitude refers to the school with the maximum absolute longitude value.
NOT_EQUAL states that these two statements disagree on this data under this rule. It does not state which of them is wrong.
SELECT GSoffered FROM schools ORDER BY ABS(longitude) DESC NULLS LAST LIMIT 1
sha256:d4d87fb991696c450e12529c6c599b392de69debe7b966d28ef8cccdb2210f15
SELECT schools.gsoffered
FROM schools
ORDER BY schools.longitude DESC
LIMIT 1;
sha256:ad3cdf827326e0095f8ec41c8da708e1998813bd07b0af516c415d3f697cd2ce
The marked tokens are where the two texts differ. Two statements that differ everywhere can return the same rows, and two that differ in one token can return other rows; the verdict above is read off the results.
from counterexample.json, 1 row, up to 25 shown per side
| side | gsofferedtext |
|---|---|
| gold | K-8 |
from counterexample.json, 1 row, up to 25 shown per side
| side | gsofferedtext |
|---|---|
| second | NULL |
set(second_rows) == set(gold_rows), float4/float8 cells as Python float as psycopg2 returns them, numeric as Decimal
https://github.com/bird-bench/mini_dev/blob/main/evaluation/evaluation_ex.py
result_eq: equal row counts and equal column counts, each row unordered as a quick rejection, then the two equal as a list when the gold text holds ORDER BY and as a multiset otherwise, under some permutation of the columns; DISTINCT is not stripped and re-executed, and the cells are PostgreSQL's as psycopg2 returns them
ruiqi-zhong/test-suite-sql-eval, exec_eval.py, result_eq, at commit 48cb78ec: https://github.com/ruiqi-zhong/test-suite-sql-eval/blob/48cb78ecf7f610620206283846c76751b18a1326/exec_eval.py
The class states what makes these two results unequal under this rule, read off the two results and nothing else. It does not state which of the two statements is wrong.
from counterexample.json, 1 row
| gsofferedtext |
|---|
| K-8 |
from counterexample.json, 1 row
| gsofferedtext |
|---|
| NULL |
A smell is a mechanical reason to read this gold statement again. It is a heuristic: it does not state that the statement is wrong, and a maintainer decides.
this statement orders by a text column holding only numbers, and ordering it as a number gives a different answer, so the gold may be sorting 9.5 above 10
no ORDER BY key resolves to a text column
{
"heuristic": true,
"reason": "no ORDER BY key resolves to a text column",
"keys": [
{
"key": "abs(longitude)",
"not_applicable": "the key is an expression and not a column"
}
]
}
this statement cuts its result at a LIMIT that does not decide which rows come back, so a different but equally correct statement can return other rows and score zero
from smells.json, 2 rows
| NULL null | 124.28481 |
| K-8 str | 124.28481 |
{
"heuristic": true,
"cut": 1,
"offset": 0,
"distinct_kept": false,
"unbounded_sql": "SELECT gsoffered, abs(longitude) AS attestql_ordering_key_0 FROM schools ORDER BY abs(longitude) DESC NULLS LAST",
"unbounded_rows": 17686,
"projected_columns": [
"gsoffered"
],
"ordering_key_columns": [
"attestql_ordering_key_0"
],
"ordering_keys": [
{
"key": "abs(longitude)",
"direction": "desc",
"nulls": "last",
"nulls_first_in_effect": false,
"returned_rows_null_in_this_key": 0,
"fires": false
}
],
"tied_at_the_cut": {
"positions": [
0,
1
],
"tied_rows": 2,
"distinct_projected_answers": 2,
"rows": [
[
{
"type": "null",
"value": null
}
],
[
{
"type": "str",
"value": "K-8"
}
]
]
},
"case": "tie-at-the-cut"
}
rerun over the same rows in another physical order this statement gives another answer, so its result depends on how the rows are stored and not only on the data
from smells.json, 1 row
| NULL |
{
"heuristic": true,
"rule": "R-ORD",
"baseline_result_hash": "sha256:d4d87fb991696c450e12529c6c599b392de69debe7b966d28ef8cccdb2210f15",
"baseline_result": {
"columns": [
{
"name": "gsoffered",
"declared_type": "text"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "str",
"value": "K-8"
}
]
],
"result_hash": "sha256:d4d87fb991696c450e12529c6c599b392de69debe7b966d28ef8cccdb2210f15"
},
"planner_statistics": {
"schools": {
"last_analyze": "2026-09-14 12:51:05.006896+00",
"last_autoanalyze": null,
"n_mod_since_analyze": 0
}
},
"shuffle": {
"seed": "1",
"row_limit": 300000,
"tables": [
"schools"
],
"tables_not_shuffled": [],
"tables_skipped_for_size": {
"laptimes": 400524,
"legalities": 427907,
"posthistory": 303155,
"trans": 1056320,
"yearmonth": 383282
},
"tables_not_reached_by_a_copy": {}
},
"shuffled_copies": {
"run": true,
"verdict": "not_equal",
"differs": true,
"result_hash": "sha256:ad3cdf827326e0095f8ec41c8da708e1998813bd07b0af516c415d3f697cd2ce",
"result": {
"columns": [
{
"name": "gsoffered",
"declared_type": "text"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "null",
"value": null
}
]
],
"result_hash": "sha256:ad3cdf827326e0095f8ec41c8da708e1998813bd07b0af516c415d3f697cd2ce"
}
},
"plan_variant": {
"run": false,
"reason": "the plan variant was not asked for"
}
}
this statement returns the same whole row more than once and never says DISTINCT, so a statement answering the same question once per row disagrees on multiplicity alone
{
"heuristic": true,
"rows": 1,
"distinct_rows": 1,
"repeated_rows": 0,
"largest_repeat": 1,
"result_bounded": false,
"distinct_stated": false,
"set_operation": false,
"not_applicable": "a result of fewer than two rows has nothing to repeat"
}
SELECT GSoffered FROM schools ORDER BY ABS(longitude) DESC NULLS LAST LIMIT 1
result_hash sha256:d4d87fb991696c450e12529c6c599b392de69debe7b966d28ef8cccdb2210f15 recomputed from this JSON: match
record_hash sha256:8a9d2f84c21b4e9e9f2043c7fd27c9869c0ce6ef2f199ba9812a18317fe401a5 recomputed from this JSON: match
from evidence-gold.json, 1 row
| gsofferedtext |
|---|
| K-8 |
SELECT schools.gsoffered
FROM schools
ORDER BY schools.longitude DESC
LIMIT 1;
result_hash sha256:ad3cdf827326e0095f8ec41c8da708e1998813bd07b0af516c415d3f697cd2ce recomputed from this JSON: match
record_hash sha256:a15cb8412e4f71584ae1a05577b6b757cc79ef3131e4c19ea9b02fcdee45f696 recomputed from this JSON: match
from evidence-second.json, 1 row
| gsofferedtext |
|---|
| NULL |
re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-ORD
re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-ORD