q1035

R-SET NOT_EQUAL

european_football_2 · mini_dev_postgresql from https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json

The question

Give the team_fifa_api_id of teams with more than 50 but less than 60 build-up play speed.

the hint the set supplies: teams with more than 50 but less than 60 build-up play speed refers to buildUpPlaySpeed >50 AND buildUpPlaySpeed <60;

NOT_EQUAL states that these two statements disagree on this data under this rule. It does not state which of them is wrong.

read by hand, 2026-09-04: A a wrong answer the benchmark credited: another table or projection; a scalar repeated per row of an unrelated table; a list at least twice the answer's length where the question asks for the things, not the rows

gold DISTINCT over 161 ids; the prediction returns 356, so the count of teams is misstated

a maintainer's reading of this question, out of classification.json, copied from plans/reports/prediction-mode-260904-real-predictions/classification.json. It is not a verdict and nothing above it was computed from it.

The statements

gold

SELECT DISTINCT team_fifa_api_id FROM Team_Attributes WHERE buildUpPlaySpeed > 50 AND buildUpPlaySpeed < 60

this statement states no ordering of its own

sha256:b5f316f44b8857af79c8d48c8c9f373862cffa9d41f817ff1f82b2f92a8884ad

second

SELECT team_fifa_api_id
FROM Team_Attributes
WHERE buildUpPlaySpeed > 50 AND buildUpPlaySpeed < 60

this statement states no ordering of its own

sha256:5dd8628b6284b80554776fe0f869af2d72b5c6acda1d534f5b754c724ef23caf

The marked tokens are where the two texts differ. Two statements that differ everywhere can return the same rows, and two that differ in one token can return other rows; the verdict above is read off the results.

The rows they differ in

in gold, not in the prediction, 0 rows

from counterexample.json, 0 rows, up to 25 shown per side

side team_fifa_api_idint8
no rows

in the prediction, not in gold, 195 rows

from counterexample.json, 195 rows, up to 25 shown per side, 25 on this page

sidetimes team_fifa_api_idint8
second4 110569
second4 13
second4 15005
second4 1715
second4 242
second4 31
second4 4
second4 462
second4 900
second3 100879
second3 110329
second3 110724
second3 110832
second3 111239
second3 15
second3 165
second3 1832
second3 203
second3 229
second3 232
second3 237
second3 449
second3 450
second3 459
second3 468

What the benchmark would have said

BIRD's own check: 1

set(second_rows) == set(gold_rows), float4/float8 cells as Python float as psycopg2 returns them, numeric as Decimal

https://github.com/bird-bench/mini_dev/blob/main/evaluation/evaluation_ex.py

gold_rows
161
second_rows
356
gold_distinct_rows
161
second_distinct_rows
161

the test-suite check: 0

result_eq: equal row counts and equal column counts, each row unordered as a quick rejection, then the two equal as a list when the gold text holds ORDER BY and as a multiset otherwise, under some permutation of the columns; DISTINCT is not stripped and re-executed, and the cells are PostgreSQL's as psycopg2 returns them

ruiqi-zhong/test-suite-sql-eval, exec_eval.py, result_eq, at commit 48cb78ec: https://github.com/ruiqi-zhong/test-suite-sql-eval/blob/48cb78ecf7f610620206283846c76751b18a1326/exec_eval.py

order_matters
false
gold_rows
161
second_rows
356
gold_columns
1
second_columns
1

the class: multiplicity

The class states what makes these two results unequal under this rule, read off the two results and nothing else. It does not state which of the two statements is wrong.

how many rows each result holds, and how many rows both hold, from evidence-gold.json and evidence-second.json161 gold356 prediction7 rows both results hold
The gold returned 161 rows and the prediction 356. 7 rows occur in both results at least once.
gold_types
["int8"]
second_types
["int8"]
multiset_equal
false
set_equal
true
order_equal
false
shorter_result_is_a_prefix
false

The results

gold, 161 rows

from counterexample.json, 161 rows, 25 on this page

team_fifa_api_idint8
485
110744
477
71
1889
68
229
481
52
70
80
1971
112225
1861
378
874
450
10030
1799
900
110745
1750
462
237
69
second, 356 rows

from counterexample.json, 356 rows, 25 on this page

team_fifa_api_idint8
434
77
77
77
614
614
614
614
1901
1901
650
650
650
1861
229
229
229
229
111989
1
1
240
240
100409
100409

The probes

A smell is a mechanical reason to read this gold statement again. It is a heuristic: it does not state that the statement is wrong, and a maintainer decides.

ordering-over-numeric-text not applicable

this statement orders by a text column holding only numbers, and ordering it as a number gives a different answer, so the gold may be sorting 9.5 above 10

the statement states no top level ORDER BY

what it measured
{
  "heuristic": true,
  "reason": "the statement states no top level ORDER BY"
}

arbitrary-cut not applicable

this statement cuts its result at a LIMIT that does not decide which rows come back, so a different but equally correct statement can return other rows and score zero

the statement states no LIMIT

what it measured
{
  "heuristic": true,
  "reason": "the statement states no LIMIT"
}

not-a-function-of-the-data quiet

rerun over the same rows in another physical order this statement gives another answer, so its result depends on how the rows are stored and not only on the data

what it measured
{
  "heuristic": true,
  "rule": "R-SET",
  "baseline_result_hash": "sha256:b5f316f44b8857af79c8d48c8c9f373862cffa9d41f817ff1f82b2f92a8884ad",
  "baseline_result": {
    "columns": [
      {
        "name": "team_fifa_api_id",
        "declared_type": "int8"
      }
    ],
    "row_count": 161,
    "truncated": false,
    "rows_shown": 10,
    "rows": [
      [
        {
          "type": "int",
          "value": 485
        }
      ],
      [
        {
          "type": "int",
          "value": 110744
        }
      ],
      [
        {
          "type": "int",
          "value": 477
        }
      ],
      [
        {
          "type": "int",
          "value": 71
        }
      ],
      [
        {
          "type": "int",
          "value": 1889
        }
      ],
      [
        {
          "type": "int",
          "value": 68
        }
      ],
      [
        {
          "type": "int",
          "value": 229
        }
      ],
      [
        {
          "type": "int",
          "value": 481
        }
      ],
      [
        {
          "type": "int",
          "value": 52
        }
      ],
      [
        {
          "type": "int",
          "value": 70
        }
      ]
    ],
    "result_hash": "sha256:b5f316f44b8857af79c8d48c8c9f373862cffa9d41f817ff1f82b2f92a8884ad"
  },
  "planner_statistics": {
    "team_attributes": {
      "last_analyze": "2026-09-14 12:51:04.628002+00",
      "last_autoanalyze": null,
      "n_mod_since_analyze": 0
    }
  },
  "shuffle": {
    "seed": "1",
    "row_limit": 300000,
    "tables": [
      "team_attributes"
    ],
    "tables_not_shuffled": [],
    "tables_skipped_for_size": {
      "laptimes": 400524,
      "legalities": 427907,
      "posthistory": 303155,
      "trans": 1056320,
      "yearmonth": 383282
    },
    "tables_not_reached_by_a_copy": {}
  },
  "shuffled_copies": {
    "run": true,
    "verdict": "equal",
    "differs": false,
    "result_hash": "sha256:55853fbc6e35cd7b4c6d6133c647830ab33d268315e267e9cc73fba5d258a1a1",
    "result": {
      "columns": [
        {
          "name": "team_fifa_api_id",
          "declared_type": "int8"
        }
      ],
      "row_count": 161,
      "truncated": false,
      "rows_shown": 10,
      "rows": [
        [
          {
            "type": "int",
            "value": 1
          }
        ],
        [
          {
            "type": "int",
            "value": 3
          }
        ],
        [
          {
            "type": "int",
            "value": 4
          }
        ],
        [
          {
            "type": "int",
            "value": 7
          }
        ],
        [
          {
            "type": "int",
            "value": 10
          }
        ],
        [
          {
            "type": "int",
            "value": 13
          }
        ],
        [
          {
            "type": "int",
            "value": 15
          }
        ],
        [
          {
            "type": "int",
            "value": 17
          }
        ],
        [
          {
            "type": "int",
            "value": 19
          }
        ],
        [
          {
            "type": "int",
            "value": 21
          }
        ]
      ],
      "result_hash": "sha256:55853fbc6e35cd7b4c6d6133c647830ab33d268315e267e9cc73fba5d258a1a1"
    }
  },
  "plan_variant": {
    "run": false,
    "reason": "the plan variant was not asked for"
  }
}

duplicate-full-row not applicable

this statement returns the same whole row more than once and never says DISTINCT, so a statement answering the same question once per row disagrees on multiplicity alone

what it measured
{
  "heuristic": true,
  "rows": 161,
  "distinct_rows": 161,
  "repeated_rows": 0,
  "largest_repeat": 1,
  "result_bounded": false,
  "distinct_stated": true,
  "set_operation": false,
  "not_applicable": "the statement states DISTINCT, so it cannot repeat a row"
}

The evidence records

gold: evidence-gold.json

SELECT DISTINCT team_fifa_api_id FROM Team_Attributes WHERE buildUpPlaySpeed > 50 AND buildUpPlaySpeed < 60
statement read from
data/questions/mini_dev_postgresql.json
digest
sha256:d2731292f20b8d8569cd956dd747ffe1df13cd625076263e38ae9ebcef50b1ab
origin
https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json, 2024-06-19

result_hash sha256:b5f316f44b8857af79c8d48c8c9f373862cffa9d41f817ff1f82b2f92a8884ad recomputed from this JSON: match

record_hash sha256:cf8704f3636499a9a930823ec3744723cd226d747cd834544ac6f4736e8eb0e8 recomputed from this JSON: match

the result this record holds, 161 rows, 50 of them here

from evidence-gold.json, 161 rows, 50 on this page

team_fifa_api_idint8
485
110744
477
71
1889
68
229
481
52
70
80
1971
112225
1861
378
874
450
10030
1799
900
110745
1750
462
237
69
1824
1905
236
680
59
1848
1862
1844
100632
110374
665
452
200
682
674
44
88
100879
100087
82
111083
286
240
480
100741

The whole result is in evidence-gold.json, beside this page.

what ran, and where
run
audit-9b355c99-9acf-40fa-bbae-77763a9af221
executed at
2026-09-14T12:53:47.527120+00:00
data as of
2026-09-14T12:53:37.264547+00:00
backend at checkout
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
backend that answered
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
database role
auditor
replay rule
R-SET
question set version
sha256:d2731292f20b8d8569cd956dd747ffe1df13cd625076263e38ae9ebcef50b1ab
validator
audit:libpg_query-parse
checks run
parses_as_exactly_one_statement, the_one_statement_is_a_select, no_placeholder_without_a_bound_parameter
statement timeout
120000 ms
rows
161 rows
the session it ran under
engine
postgresql
time_zone
Etc/UTC
date_style
ISO, MDY
interval_style
postgres
extra_float_digits
1
database_collation
en_US.utf8
work_mem
4096
hash_mem_multiplier
2

recorded beside them

statement_timeout
0
server_version
16.15 (Debian 16.15-1.pgdg13+2)
server_version_num
160015
transaction_read_only
on
max_parallel_workers_per_gather
2
server_encoding
UTF8
search_path
public
datlocprovider
c
daticulocale
datcollversion
2.41
the rendering and the data
version
attestql/audit/4
numeric_scale
6
timestamp_format
%Y-%m-%dT%H:%M:%S.%fZ
timezone
UTC
null_rendering
NULL
encoding
utf-8
schema digest
sha256:98bc0cba85d167faa0d5bb8e4db8801339559c93568b09cfd7c9a911f2c12123
source file sha256
sha256:31b1da211849d24a57c9af7636da46a5b82fc8a3ca1542bb3ebd8775e9a31cec
rows in public.team_attributes
1458

second: evidence-second.json

SELECT team_fifa_api_id
FROM Team_Attributes
WHERE buildUpPlaySpeed > 50 AND buildUpPlaySpeed < 60
statement read from
data/preds-pg/predict_mini_dev_gpt-4-turbo_postgresql.json
digest
sha256:428fd115a67a76d6312ea205aab64764b5a09444587588eda442a94cfee65211
origin
https://raw.githubusercontent.com/bird-bench/mini_dev/b3d4bcbbae9a96934ad812551eb400c7a3b23c12/llm/exp_result/sql_output_kg/predict_mini_dev_gpt-4-turbo_postgresql.json, 2024-06-19

result_hash sha256:5dd8628b6284b80554776fe0f869af2d72b5c6acda1d534f5b754c724ef23caf recomputed from this JSON: match

record_hash sha256:8dba9eee83d446addd3a9fede72a2208f9c644ae488f00b00bd24f4e880d0cf0 recomputed from this JSON: match

the result this record holds, 356 rows, 50 of them here

from evidence-second.json, 356 rows, 50 on this page

team_fifa_api_idint8
434
77
77
77
614
614
614
614
1901
1901
650
650
650
1861
229
229
229
229
111989
1
1
240
240
100409
100409
1906
1906
1848
1848
1848
32
21
675
1889
1889
234
88
88
88
88
3
3
3
160
4
4
4
4
4
59

The whole result is in evidence-second.json, beside this page.

what ran, and where
run
audit-9b355c99-9acf-40fa-bbae-77763a9af221
executed at
2026-09-14T12:53:47.530997+00:00
data as of
2026-09-14T12:53:37.264547+00:00
backend at checkout
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
backend that answered
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
database role
auditor
replay rule
R-SET
question set version
sha256:d2731292f20b8d8569cd956dd747ffe1df13cd625076263e38ae9ebcef50b1ab
validator
audit:libpg_query-parse
checks run
parses_as_exactly_one_statement, the_one_statement_is_a_select, no_placeholder_without_a_bound_parameter
statement timeout
120000 ms
rows
356 rows
the session it ran under
engine
postgresql
time_zone
Etc/UTC
date_style
ISO, MDY
interval_style
postgres
extra_float_digits
1
database_collation
en_US.utf8
work_mem
4096
hash_mem_multiplier
2

recorded beside them

statement_timeout
0
server_version
16.15 (Debian 16.15-1.pgdg13+2)
server_version_num
160015
transaction_read_only
on
max_parallel_workers_per_gather
2
server_encoding
UTF8
search_path
public
datlocprovider
c
daticulocale
datcollversion
2.41
the rendering and the data
version
attestql/audit/4
numeric_scale
6
timestamp_format
%Y-%m-%dT%H:%M:%S.%fZ
timezone
UTC
null_rendering
NULL
encoding
utf-8
schema digest
sha256:98bc0cba85d167faa0d5bb8e4db8801339559c93568b09cfd7c9a911f2c12123
source file sha256
sha256:31b1da211849d24a57c9af7636da46a5b82fc8a3ca1542bb3ebd8775e9a31cec
rows in public.team_attributes
1458

Running these again

gold

re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET

second

re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET

This question's run

run
audit-9b355c99-9acf-40fa-bbae-77763a9af221
server
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
question set
mini_dev_postgresql from https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json
replay rule
R-SET

the run this question belongs to

The JSON this page was rendered from