gold
SELECT player_api_id FROM Player ORDER BY weight DESC LIMIT 10
- weight
- descending
sha256:75325a400764432a13984d559881f0ab687f9cdedff538cc82c4dea6dadbd073
R-ORD GOLD-ONLY arbitrary-cut not-a-function-of-the-data
european_football_2 · dev_20251106-00000-of-00001 from https://huggingface.co/datasets/birdsql/bird_sql_dev_20251106/resolve/3c11fb193e5439b338e23677fa0aae11e8b85db9/data/dev_20251106-00000-of-00001.json (commit 3c11fb19, downloaded 2026-09-07)
What are the player api id of 10 heaviest players?
the hint the set supplies: heaviest refers to MAX(weight)
This question was audited without a prediction beside it, so there is nothing to compare the gold with. The probes below read the gold alone.
read by hand, 2026-09-07: wrong the gold does not answer its question on this data
three players tie at positions 8 to 10 of the weight order, so two of the ten ids are arbitrary
a maintainer's reading of this question, out of classification.json, copied from plans/reports/bird-dev-sqlite-260907/classification.json. It is not a verdict and nothing above it was computed from it.
SELECT player_api_id FROM Player ORDER BY weight DESC LIMIT 10
sha256:75325a400764432a13984d559881f0ab687f9cdedff538cc82c4dea6dadbd073
from evidence-gold.json, 10 rows
| player_api_idINTEGER |
|---|
| 148325 |
| 27313 |
| 5044 |
| 27267 |
| 101584 |
| 19020 |
| 210822 |
| 30669 |
| 40005 |
| 33060 |
A smell is a mechanical reason to read this gold statement again. It is a heuristic: it does not state that the statement is wrong, and a maintainer decides.
this statement orders by a text column holding only numbers, and ordering it as a number gives a different answer, so the gold may be sorting 9.5 above 10
no ORDER BY key resolves to a text column
{
"heuristic": true,
"reason": "no ORDER BY key resolves to a text column",
"keys": [
{
"key": "weight",
"column": "Player.weight",
"declared_type": "INTEGER",
"not_applicable": "the column is not declared as text"
}
]
}
this statement cuts its result at a LIMIT that does not decide which rows come back, so a different but equally correct statement can return other rows and score zero
from smells.json, 3 rows
| 40005 | 218 |
| 33060 | 218 |
| 93450 | 218 |
{
"heuristic": true,
"cut": 10,
"offset": 0,
"distinct_kept": false,
"unbounded_sql": "SELECT player_api_id, weight AS attestql_ordering_key_0 FROM Player ORDER BY weight DESC",
"unbounded_rows": 11060,
"projected_columns": [
"player_api_id"
],
"ordering_key_columns": [
"attestql_ordering_key_0"
],
"ordering_keys": [
{
"key": "weight",
"direction": "desc",
"nulls": "last",
"nulls_first_in_effect": false,
"returned_rows_null_in_this_key": 0,
"fires": false
}
],
"tied_at_the_cut": {
"positions": [
8,
9,
10
],
"tied_rows": 3,
"distinct_projected_answers": 3,
"rows": [
[
{
"type": "int",
"value": 40005
}
],
[
{
"type": "int",
"value": 33060
}
],
[
{
"type": "int",
"value": 93450
}
]
]
},
"case": "tie-at-the-cut"
}
rerun over the same rows in another physical order this statement gives another answer, so its result depends on how the rows are stored and not only on the data
from smells.json, 10 rows
| 148325 |
| 27313 |
| 5044 |
| 27267 |
| 101584 |
| 19020 |
| 30669 |
| 210822 |
| 33060 |
| 93450 |
{
"heuristic": true,
"rule": "R-ORD",
"baseline_result_hash": "sha256:75325a400764432a13984d559881f0ab687f9cdedff538cc82c4dea6dadbd073",
"baseline_result": {
"columns": [
{
"name": "player_api_id",
"declared_type": "INTEGER"
}
],
"row_count": 10,
"truncated": false,
"rows_shown": 10,
"rows": [
[
{
"type": "int",
"value": 148325
}
],
[
{
"type": "int",
"value": 27313
}
],
[
{
"type": "int",
"value": 5044
}
],
[
{
"type": "int",
"value": 27267
}
],
[
{
"type": "int",
"value": 101584
}
],
[
{
"type": "int",
"value": 19020
}
],
[
{
"type": "int",
"value": 210822
}
],
[
{
"type": "int",
"value": 30669
}
],
[
{
"type": "int",
"value": 40005
}
],
[
{
"type": "int",
"value": 33060
}
]
],
"result_hash": "sha256:75325a400764432a13984d559881f0ab687f9cdedff538cc82c4dea6dadbd073"
},
"planner_statistics": {},
"shuffle": {
"seed": "1",
"row_limit": 300000,
"tables": [
"Player"
],
"tables_not_shuffled": [],
"tables_skipped_for_size": {},
"tables_not_reached_by_a_copy": {}
},
"shuffled_copies": {
"run": true,
"verdict": "not_equal",
"differs": true,
"result_hash": "sha256:58670f0affcf787ca0dac7e1633bba6e9a526597e74b4e1339862a9edcdae72d",
"result": {
"columns": [
{
"name": "player_api_id",
"declared_type": "INTEGER"
}
],
"row_count": 10,
"truncated": false,
"rows_shown": 10,
"rows": [
[
{
"type": "int",
"value": 148325
}
],
[
{
"type": "int",
"value": 27313
}
],
[
{
"type": "int",
"value": 5044
}
],
[
{
"type": "int",
"value": 27267
}
],
[
{
"type": "int",
"value": 101584
}
],
[
{
"type": "int",
"value": 19020
}
],
[
{
"type": "int",
"value": 30669
}
],
[
{
"type": "int",
"value": 210822
}
],
[
{
"type": "int",
"value": 33060
}
],
[
{
"type": "int",
"value": 93450
}
]
],
"result_hash": "sha256:58670f0affcf787ca0dac7e1633bba6e9a526597e74b4e1339862a9edcdae72d"
}
},
"plan_variant": {
"run": false,
"reason": "the plan variant was not asked for"
}
}
this statement returns the same whole row more than once and never says DISTINCT, so a statement answering the same question once per row disagrees on multiplicity alone
{
"heuristic": true,
"rows": 10,
"distinct_rows": 10,
"repeated_rows": 0,
"largest_repeat": 1,
"result_bounded": false,
"distinct_stated": false,
"set_operation": false
}
SELECT player_api_id FROM Player ORDER BY weight DESC LIMIT 10
result_hash sha256:75325a400764432a13984d559881f0ab687f9cdedff538cc82c4dea6dadbd073 recomputed from this JSON: match
record_hash sha256:ec5c04a1cc61bf12b75911d1a7d7b1e4c75020919112511238563e0443c301f0 recomputed from this JSON: match
from evidence-gold.json, 10 rows
| player_api_idINTEGER |
|---|
| 148325 |
| 27313 |
| 5044 |
| 27267 |
| 101584 |
| 19020 |
| 210822 |
| 30669 |
| 40005 |
| 33060 |
re-run this statement read-only against SQLite 3.53.4 | file=/private/tmp/attestql-runs/data/dev/dev_databases/european_football_2/european_football_2.sqlite | size=597754880 under the session settings and over the data this record's fixture digest names, and compare the two results under R-ORD