askdata is scored on a gold set of 60 questions (53 answerable, 7 that should be refused)
with hand-written, cross-checked gold SQL. A question counts as correct when the generated query's result matches the gold result, or when an unanswerable question is refused.
The live site runs Gemini 2.5 Flash (thinking off) at ablation d, falling back to Gemini 2.5 Flash-Lite at ablation e (73% overall, 95% CI 62-83%) when its free quota runs out.
78%
overall accuracy, live configuration 95% CI 67-88%
77%
execution accuracy on the 53 answerable questions
86%
of the 7 unanswerable questions correctly refused
0.9s
median model time per question (all attempts)
Accuracy by ablation
Each step adds one thing to the one before: a Zero-shot, full schema; b + retrieved column docs and sample values;
c + retrieved few-shot examples; d + self-correction (max 2 retries); e + answerability rule (tuned on a separate dev set). Bars are overall accuracy over all 60 questions; whiskers are 95% bootstrap intervals over questions.
e + answerability rule (tuned on a separate dev set)
73%
62-83%
75%
74%
57%
0
0.13
0.6s
2,906
Gemini 2.5 Flash (thinking off)
a Zero-shot, full schema
25%
15-37%
17%
13%
86%
0
0.00
0.9s
1,021
Gemini 2.5 Flash (thinking off)
b + retrieved column docs and sample values
67%
55-78%
64%
60%
86%
0
0.00
0.9s
2,201
Gemini 2.5 Flash (thinking off)
c + retrieved few-shot examples
78%
67-88%
77%
75%
86%
0
0.00
0.9s
2,381
Gemini 2.5 Flash (thinking off)
d + self-correction (max 2 retries)
78%
67-88%
77%
75%
86%
0
0.07
0.9s
2,547
Gemini 2.5 Flash (thinking off)
e + answerability rule (tuned on a separate dev set)
70%
58-82%
68%
62%
86%
0
0.08
0.9s
2,766
Gemma 4 31B (open weights)
a Zero-shot, full schema
28%
17-40%
21%
19%
86%
0
0.00
29.0s
1,022
Gemma 4 31B (open weights)
b + retrieved column docs and sample values
75%
63-85%
72%
70%
100%
0
0.00
22.7s
2,202
Gemma 4 31B (open weights)
c + retrieved few-shot examples
92%
83-98%
91%
89%
100%
0
0.00
14.1s
2,382
Gemma 4 31B (open weights)
d + self-correction (max 2 retries)
92%
83-98%
91%
89%
100%
0
0.03
14.1s
2,472
Gemma 4 31B (open weights)
e + answerability rule (tuned on a separate dev set)
87%
77-95%
85%
83%
100%
0
0.08
14.1s
2,775
Exec. acc. = share of the 53 answerable questions whose result matched. "strict" counts only the gold definition; the other column also accepts the second definition of a listing count that the warehouse itself carries (see Method). Latency is model time summed over attempts. Run status: Gemini 2.5 Flash-Lite: complete (60 questions); Gemini 2.5 Flash (thinking off): complete (60 questions); Gemma 4 31B (open weights): complete (60 questions).
Does each step help?
Model
Step
Change
95% CI (paired bootstrap)
Gemini 2.5 Flash-Lite
a → b
+21.7 pts *
+10.0 to +33.3
Gemini 2.5 Flash-Lite
b → c
+26.7 pts *
+13.3 to +40.0
Gemini 2.5 Flash-Lite
c → d
+0.0 pts
+0.0 to +0.0
Gemini 2.5 Flash-Lite
d → e
+5.0 pts
+0.0 to +11.7
Gemini 2.5 Flash (thinking off)
a → b
+41.7 pts *
+28.3 to +55.0
Gemini 2.5 Flash (thinking off)
b → c
+11.7 pts
+0.0 to +23.3
Gemini 2.5 Flash (thinking off)
c → d
+0.0 pts
+0.0 to +0.0
Gemini 2.5 Flash (thinking off)
d → e
-8.3 pts
-16.7 to +0.0
Gemma 4 31B (open weights)
a → b
+46.7 pts *
+33.3 to +60.0
Gemma 4 31B (open weights)
b → c
+16.7 pts *
+6.7 to +26.7
Gemma 4 31B (open weights)
c → d
+0.0 pts
+0.0 to +0.0
Gemma 4 31B (open weights)
d → e
-5.0 pts
-11.7 to +1.7
* the interval excludes zero. With 60 questions, a one-question change is 1.7 points, so many step-to-step differences are within noise.
By question type
Model, ablation
lookup
aggregation
join
window
date
ambiguous
unanswerable
Gemini 2.5 Flash-Lite, a
50% (5/10)
18% (2/11)
10% (1/10)
0% (0/7)
25% (2/8)
0% (0/7)
29% (2/7)
Gemini 2.5 Flash-Lite, d
100% (10/10)
100% (11/11)
80% (8/10)
43% (3/7)
62% (5/8)
29% (2/7)
29% (2/7)
Gemini 2.5 Flash (thinking off), a
40% (4/10)
27% (3/11)
10% (1/10)
0% (0/7)
12% (1/8)
0% (0/7)
86% (6/7)
Gemini 2.5 Flash (thinking off), d
100% (10/10)
82% (9/11)
70% (7/10)
57% (4/7)
88% (7/8)
57% (4/7)
86% (6/7)
Gemma 4 31B (open weights), a
60% (6/10)
18% (2/11)
10% (1/10)
0% (0/7)
25% (2/8)
0% (0/7)
86% (6/7)
Gemma 4 31B (open weights), d
100% (10/10)
100% (11/11)
90% (9/10)
71% (5/7)
88% (7/8)
86% (6/7)
100% (7/7)
What goes wrong
Failures at ablation d, classified by a rule-based tagger (first matching rule wins: refusal errors, execution errors, empty results, missing the documented default filter, wrong tables or join type, date logic, a gold filter literal missing, wrong row count, otherwise wrong calculation). The tagger is a heuristic; the examples below are the evidence.
Error class
Gemini 2.5 Flash-Lite
Gemini 2.5 Flash (thinking off)
Gemma 4 31B (open weights)
missing default market filter
4
4
0
wrong table or join
3
3
1
answered an unanswerable question
5
1
0
date logic
3
1
1
wrong grouping or result shape
1
1
1
wrong aggregation or calculation
1
1
1
empty result (wrong filter value)
1
2
0
other execution error
1
0
1
Examples
missing default market filter M01: How many AI jobs are there? (Gemini 2.5 Flash-Lite)
Model SQL:
SELECT count(posting_key) FROM fct_posting WHERE family = 'ai_engineer'
Gold SQL:
SELECT count(*) FROM fct_posting WHERE family = 'ai_engineer' AND is_open AND remote_us AND is_canonical_listing
Verdict: no predicted column matches gold column 'count_star()'
missing default market filter M05: How many entry-level jobs are open? (Gemini 2.5 Flash-Lite)
Model SQL:
SELECT count(posting_key) FROM fct_posting WHERE seniority = 'entry' AND is_open
Gold SQL:
SELECT count(*) FROM fct_posting WHERE seniority = 'entry' AND is_open AND remote_us AND is_canonical_listing
Verdict: no predicted column matches gold column 'count_star()'
wrong table or join J10: What share of open remote-US ML engineer listings mention at least one LLM / GenAI skill? (Gemini 2.5 Flash-Lite)
Model SQL:
SELECT SUM(CASE WHEN ps.cluster = 'llm_genai' THEN 1 ELSE 0 END) * 1.0 / COUNT(DISTINCT f.posting_key) FROM fct_posting f LEFT JOIN posting_skills ps ON f.posting_key = ps.posting_key WHERE f.family = 'ml_engineer' AND f.is_open AND f.remote_us AND f.is_canonical_listing
Gold SQL:
SELECT count(*) FILTER (WHERE posting_key IN (SELECT posting_key FROM posting_skills WHERE cluster = 'llm_genai')) * 1.0 / count(*) AS share
FROM fct_posting WHERE family = 'ml_engineer' AND is_open AND remote_us AND is_canonical_listing
Verdict: no predicted column matches gold column 'share'
wrong table or join W06: How many open remote-US listings pay more than 1.5 times the median pay of their own role family? (Gemini 2.5 Flash-Lite)
Model SQL:
SELECT count(*) FROM fct_posting f LEFT JOIN mart_comp_by_family m ON f.family = m.family WHERE f.is_open AND f.remote_us AND f.is_canonical_listing AND f.comp_mid_usd > m.p50 * 1.5
Gold SQL:
SELECT count(*) AS listings FROM (
SELECT comp_mid_usd, median(comp_mid_usd) OVER (PARTITION BY family) AS fam_median
FROM fct_posting WHERE is_open AND remote_us AND is_canonical_listing)
WHERE comp_mid_usd > 1.5 * fam_median
Verdict: no predicted column matches gold column 'listings'
answered an unanswerable question U01: How many people applied to each posting? (Gemini 2.5 Flash-Lite)
Model SQL:
SELECT count(*) AS postings FROM fct_posting WHERE lower(workplace) = 'hybrid'
Gold: refuse (the warehouse cannot answer this).
Verdict: answered an unanswerable question
answered an unanswerable question U03: Which companies offer equity or stock options with their data roles? (Gemini 2.5 Flash-Lite)
Model SQL (after 2 retries):
SELECT DISTINCT T1.company_name FROM dim_company AS T1 INNER JOIN fct_posting AS T2 ON T1.company_key = T2.company_key WHERE T2.equity_mentioned = TRUE
Error: Binder Error: Table "T2" does not have a column named "equity_mentioned"
Candidate bindings: : "phd_mentioned"
Gold: refuse (the warehouse cannot answer this).
Verdict: answered an unanswerable question
date logic D01: How many open remote-US listings were published in the 30 days up to and including the latest crawl? (Gemini 2.5 Flash-Lite)
Model SQL (after 1 retries):
SELECT count(*) AS listings FROM fct_posting WHERE is_open AND remote_us AND is_canonical_listing AND published_at BETWEEN DATE '2026-08-30' AND DATE '2026-09-29'
Gold SQL:
SELECT count(*) AS listings FROM fct_posting
WHERE is_open AND remote_us AND is_canonical_listing
AND published_at::DATE > (SELECT max(last_seen) FROM fct_posting) - INTERVAL 30 DAY
Verdict: no predicted column matches gold column 'listings'
date logic D04: What is the median age in days of open remote-US listings in each role family? (Gemini 2.5 Flash-Lite)
Model SQL (after 1 retries):
SELECT family, approx_quantile(age_days_at_last_seen, 0.5) AS median_age FROM fct_posting WHERE is_open AND remote_us AND is_canonical_listing GROUP BY family
Gold SQL:
SELECT family, median(age_days_at_last_seen) AS median_age_days FROM fct_posting
WHERE is_open AND remote_us AND is_canonical_listing GROUP BY family
Verdict: no predicted column matches gold column 'median_age_days'
wrong grouping or result shape W01: For each role family, what is the highest pay midpoint among open remote-US listings, and which company posted it? (Gemini 2.5 Flash-Lite)
Model SQL:
SELECT
fp.family,
MAX(fp.comp_mid_usd) AS highest_pay_midpoint,
dc.company_name
FROM fct_posting AS fp
JOIN dim_company AS dc
ON fp.company_key = dc.company_key
WHERE
fp.is_open AND fp.remote_us AND fp.is_canonical_listing AND fp.comp_mid_usd IS NOT NULL
GROUP BY
fp.family,
dc.company_name
ORDER BY
fp.family,
highest_pay_midpoint DESC
Gold SQL:
SELECT family, company_name, comp_mid_usd FROM (
SELECT f.family, d.company_name, f.comp_mid_usd,
rank() OVER (PARTITION BY f.family ORDER BY f.comp_mid_usd DESC) AS r
FROM fct_posting f JOIN dim_company d USING (company_key)
WHERE f.is_open AND f.remote_us AND f.is_canonical_listing AND f.comp_mid_usd IS NOT NULL)
WHERE r = 1
Verdict: row count 481 != gold 8
wrong grouping or result shape W01: For each role family, what is the highest pay midpoint among open remote-US listings, and which company posted it? (Gemini 2.5 Flash (thinking off))
Model SQL:
SELECT family, company_name, comp_mid_usd FROM fct_posting AS fp JOIN dim_company AS dc ON fp.company_key = dc.company_key WHERE is_open AND remote_us AND is_canonical_listing QUALIFY ROW_NUMBER() OVER (PARTITION BY family ORDER BY comp_mid_usd DESC) = 1
Gold SQL:
SELECT family, company_name, comp_mid_usd FROM (
SELECT f.family, d.company_name, f.comp_mid_usd,
rank() OVER (PARTITION BY f.family ORDER BY f.comp_mid_usd DESC) AS r
FROM fct_posting f JOIN dim_company d USING (company_key)
WHERE f.is_open AND f.remote_us AND f.is_canonical_listing AND f.comp_mid_usd IS NOT NULL)
WHERE r = 1
Verdict: row count 7 != gold 8
wrong aggregation or calculation W03: What are the top 3 skills by share of listings for each role family? (Gemini 2.5 Flash-Lite)
Model SQL:
WITH RankedSkills AS (SELECT family, cluster_label, share, ROW_NUMBER() OVER (PARTITION BY family ORDER BY share DESC) as rn FROM mart_skill_demand WHERE scope = 'all' AND family IS NOT NULL) SELECT family, cluster_label, share FROM RankedSkills WHERE rn <= 3 ORDER BY family, rn
Gold SQL:
SELECT family, skill, share FROM (
SELECT family, skill, share, row_number() OVER (PARTITION BY family ORDER BY share DESC, skill) AS r
FROM mart_skill_demand WHERE scope = 'all')
WHERE r <= 3
Verdict: no predicted column matches gold column 'skill'
wrong aggregation or calculation W06: How many open remote-US listings pay more than 1.5 times the median pay of their own role family? (Gemini 2.5 Flash (thinking off))
Model SQL:
SELECT count(f.posting_key) AS listings_above_1_5x_median FROM fct_posting AS f JOIN mart_comp_by_family AS m ON f.family = m.family AND f.seniority = m.seniority WHERE f.is_open AND f.remote_us AND f.is_canonical_listing AND f.comp_mid_usd > (1.5 * m.p50)
Gold SQL:
SELECT count(*) AS listings FROM (
SELECT comp_mid_usd, median(comp_mid_usd) OVER (PARTITION BY family) AS fam_median
FROM fct_posting WHERE is_open AND remote_us AND is_canonical_listing)
WHERE comp_mid_usd > 1.5 * fam_median
Verdict: no predicted column matches gold column 'listings'
empty result (wrong filter value) W05: For each seniority level, show the median pay midpoint of open remote-US listings and how far it is above or below the overall median. (Gemini 2.5 Flash-Lite)
Model SQL (after 2 retries):
WITH OverallMedian AS (SELECT p50 AS overall_median FROM mart_comp_by_family WHERE family IS NULL AND seniority IS NULL), SeniorityMedian AS (SELECT seniority, p50 AS seniority_median FROM mart_comp_by_family WHERE family IS NULL AND seniority IS NOT NULL AND p50 IS NOT NULL) SELECT sm.seniority, sm.seniority_median, sm.seniority_median - om.overall_median AS difference_from_overall FROM SeniorityMedian sm, OverallMedian om WHERE om.overall_median IS NOT NULL ORDER BY sm.seniority;
Gold SQL:
SELECT seniority, median(comp_mid_usd) AS median_pay,
median(comp_mid_usd) - (SELECT median(comp_mid_usd) FROM fct_posting WHERE is_open AND remote_us AND is_canonical_listing) AS diff_from_overall
FROM fct_posting WHERE is_open AND remote_us AND is_canonical_listing
GROUP BY seniority
Verdict: row count 0 != gold 7
empty result (wrong filter value) A01: How many open remote-US data listings are there right now? (Gemini 2.5 Flash (thinking off))
Model SQL (after 2 retries):
SELECT open_listings FROM mart_daily_market WHERE crawl_date = '2026-09-29' AND remote_us = true AND family IS NULL AND seniority IS NULL
Gold SQL:
SELECT count(*) AS listings FROM fct_posting WHERE is_open AND remote_us AND is_canonical_listing
Verdict: row count 0 != gold 1
other execution error J05: Compare the median pay midpoint of open remote-US listings that mention RAG with those that don't. (Gemini 2.5 Flash-Lite)
Model SQL (after 2 retries):
SELECT CASE WHEN SUM(CASE WHEN ps.skill = 'rag' THEN 1 ELSE 0 END) > 0 THEN 'Mentions RAG' ELSE 'Does not mention RAG' END AS rag_mention_status, AVG(f.comp_mid_usd) AS median_pay_midpoint FROM fct_posting AS f LEFT JOIN posting_skills AS ps ON f.posting_key = ps.posting_key WHERE f.is_open AND f.remote_us AND f.is_canonical_listing GROUP BY rag_mention_status
Error: Binder Error: GROUP BY clause cannot contain aggregates!
Gold SQL:
SELECT (posting_key IN (SELECT posting_key FROM posting_skills WHERE skill = 'rag')) AS mentions_rag,
median(comp_mid_usd) AS median_pay
FROM fct_posting WHERE is_open AND remote_us AND is_canonical_listing
GROUP BY mentions_rag
Verdict: execution error
other execution error M06: How many remote jobs are in the dataset? (Gemma 4 31B (open weights))
Model SQL (after 2 retries):
(no SQL: refused - )
Error: no JSON object in reply
Gold SQL:
SELECT count(*) FROM fct_posting WHERE is_open AND remote_us AND is_canonical_listing
Verdict: execution error
Method
Data. A frozen snapshot of the Data Jobs Observatory warehouse (crawl of 2026-09-29): 11 tables, the same Parquet files this site loads into your browser.
Gold set. 60 questions in seven types (lookup, aggregation, join, window, date, ambiguous phrasing, unanswerable). Every gold query is executed by make gold-check; 38 also carry an independently written second query (through a different table or construct) that must return the same result, and 29 pin a hand-verified value so a data refresh cannot silently move the target. Ties at ranking cut-offs were checked and avoided.
A finding from verifying the gold set. The warehouse carries two definitions of "open remote-US listings": the documented default filter (is_open AND remote_us AND is_canonical_listing, 969 listings at 441 companies) and the counters in mart_daily_market/dim_company (1,016 remote listing keys, 448 companies), because the canonical flag picks one posting per listing across all locations. Gold follows the documented filter; for five questions the mart definition is accepted as an alternative and reported separately ("strict" excludes it).
Scoring. Result sets are compared as multisets; column names and order are ignored and extra columns are allowed; row order is enforced only for ranking questions (and ties may come back in any order); numbers match within 0.1% or when the prediction is the gold value rounded to its own decimals; a fraction and the same value as a percent match. Unanswerable questions are correct only if the agent refuses. A write request is blocked by the read-only guard regardless, but only a refusal scores.
Ablations. Temperature 0 throughout. Ablation d uses the same prompt as c, so it reuses c's first reply and adds up to two retries when a query errors or returns no rows, feeding the error back. Ablation e adds an answerability rule to the system prompt. It was added after the first run showed refusal was the weakest skill, written from the schema, and checked on a separate 16-question dev set (13/16 correct refusal decisions without it, 15/16 with it for Flash-Lite), never on the gold questions.
Statistics. 95% intervals are percentile bootstraps over questions (2,000 resamples); step-to-step changes use a paired bootstrap on the same questions.
Cost. Free-tier Gemini API only, rate-limited and cached by prompt hash, so the whole eval re-scores offline from the committed replies with make eval-offline.
Parity. The browser app builds its prompts with a JavaScript port of the Python code the eval uses; a test asserts byte-identical prompts for all 60 questions and identical guard verdicts.
Limitations
60 questions is small: intervals are wide and a single question moves a score by 1.7 points.
The gold SQL, the data-dictionary notes, the retrieval synonyms, the few-shot pool and the answerability rule were all written by the same author. Few-shot examples never answer a gold question (tested), but several are template siblings of one (same shape, different entity), which flatters ablation c. The answerability rule was written after seeing run-1 failures, so treat e as optimistic.
The scorer was corrected after inspecting run 1 (an order-dependent tolerance bug, and a timestamp answering a which-date question); every model was re-scored from the reply cache and no reply changed.
One domain, one warehouse snapshot, English questions only. The "ambiguous" questions encode one reading (the documented default filter); a reasonable analyst could disagree on some.
Execution accuracy can credit a wrong query that happens to return the right numbers, and it does not grade the one-sentence answer the site writes on top of the result.
Every question
ID
Question
Type
Gemini 2.5 Flash-Lite a
Gemini 2.5 Flash-Lite b
Gemini 2.5 Flash-Lite c
Gemini 2.5 Flash-Lite d
Gemini 2.5 Flash-Lite e
Gemini 2.5 Flash a
Gemini 2.5 Flash b
Gemini 2.5 Flash c
Gemini 2.5 Flash d
Gemini 2.5 Flash e
Gemma 4 31B a
Gemma 4 31B b
Gemma 4 31B c
Gemma 4 31B d
Gemma 4 31B e
L01
How many company job boards does the crawler track?
lookup
no
yes
yes
yes
yes
no
yes
yes
yes
yes
yes
yes
yes
yes
yes
L02
Which skill cluster (display name) does dbt belong to?
lookup
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
L03
How many job boards failed to answer on the latest crawl?
lookup
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
L04
How many data-role postings have been collected from Lever boards?
lookup
no
yes
yes
yes
yes
no
yes
yes
yes
yes
no
no
yes
yes
yes
L05
What is the median posted pay midpoint for ML engineers?
lookup
no
no
yes
yes
yes
no
yes
yes
yes
yes
no
yes
yes
yes
yes
L06
What is the URL of Reddit's job board?
lookup
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
L07
What pattern is used to detect the 'rag' skill in job descriptions?
lookup
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
L08
How many postings of any role were scanned in the most recent crawl?
lookup
yes
yes
yes
yes
yes
no
yes
yes
yes
yes
yes
yes
yes
yes
yes
L09
What share of open remote-US data analyst listings mention SQL?
lookup
no
no
yes
yes
yes
no
yes
yes
yes
yes
no
no
yes
yes
yes
L10
What fraction of all open remote-US listings disclose a pay range?
lookup
no
no
yes
yes
yes
no
yes
yes
yes
yes
no
yes
yes
yes
yes
A01
How many open remote-US data listings are there right now?
aggregation
no
yes
yes
yes
yes
no
yes
no
no
yes
no
yes
yes
yes
yes
A02
Break down open remote-US listings by role family.
aggregation
no
yes
yes
yes
yes
yes
yes
yes
yes
yes
no
no
yes
yes
no
A03
How many distinct companies have at least one open remote-US data listing?
aggregation
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
A04
Among open remote-US AI engineer listings that state a years-of-experience requirement, what is the average number of years asked?
aggregation
no
no
yes
yes
yes
no
no
yes
yes
yes
no
yes
yes
yes
yes
A05
How many postings came from each ATS platform?
aggregation
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
A06
What percentage of open remote-US listings are doorway roles?
aggregation
no
yes
yes
yes
yes
no
no
no
no
no
no
yes
yes
yes
yes
A07
How many open remote-US listings originally posted their pay in a currency other than US dollars?
aggregation
no
no
yes
yes
yes
no
yes
yes
yes
yes
no
yes
yes
yes
yes
A08
Which role family has the highest median pay among open remote-US listings?
aggregation
no
yes
yes
yes
yes
no
no
yes
yes
no
no
no
yes
yes
yes
A09
How many open remote-US listings name a PhD as their lowest degree requirement?
aggregation
no
no
yes
yes
yes
no
no
yes
yes
yes
no
yes
yes
yes
yes
A10
What are the lowest and highest posted pay midpoints among open remote-US data analyst listings?
aggregation
no
yes
yes
yes
yes
no
yes
yes
yes
yes
no
yes
yes
yes
yes
A11
For data engineers, how many open remote-US listings are there at each seniority level?
aggregation
no
yes
yes
yes
yes
no
yes
yes
yes
yes
no
yes
yes
yes
yes
J01
Which three companies have the most open remote-US data listings? Give the company name and count.
join
no
no
yes
yes
yes
no
no
yes
yes
no
no
yes
yes
yes
yes
J02
What are the 10 most frequently mentioned skills across open remote-US listings, with the number of listings mentioning each?
join
no
yes
yes
yes
yes
no
yes
yes
yes
yes
no
yes
yes
yes
yes
J03
How many open remote-US listings mention both Python and SQL?
join
no
no
yes
yes
yes
no
yes
yes
yes
yes
no
no
yes
yes
yes
J04
Which companies hiring through Ashby have more than 5 open remote-US data listings?
join
no
no
yes
yes
yes
no
no
yes
yes
no
no
no
yes
yes
yes
J05
Compare the median pay midpoint of open remote-US listings that mention RAG with those that don't.
join
no
no
no
no
no
no
no
no
no
no
no
no
no
no
no
J06
For each skill cluster, how many open remote-US listings mention at least one skill in that cluster?
join
no
yes
yes
yes
yes
no
yes
yes
yes
yes
no
yes
yes
yes
yes
J07
List the titles and links of Reddit's open remote-US AI engineer listings.
join
no
no
yes
yes
yes
no
yes
yes
yes
yes
no
yes
yes
yes
yes
J08
Which company has the most open remote-US listings that mention evals?
join
yes
no
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
J09
Among companies whose job boards list more than 1,000 postings of any role, how many open remote-US data listings does each have?
join
no
yes
yes
yes
yes
no
no
no
no
no
no
yes
yes
yes
yes
J10
What share of open remote-US ML engineer listings mention at least one LLM / GenAI skill?
join
no
no
no
no
no
no
no
no
no
no
no
no
yes
yes
yes
W01
For each role family, what is the highest pay midpoint among open remote-US listings, and which company posted it?
window
no
no
no
no
no
no
no
no
no
no
no
no
no
no
no
W02
Rank role families by number of open remote-US listings and show each family's share of the total.
window
no
yes
yes
yes
yes
no
no
yes
yes
yes
no
no
yes
yes
no
W03
What are the top 3 skills by share of listings for each role family?
window
no
no
no
no
no
no
yes
yes
yes
no
no
no
yes
yes
yes
W04
Order role families from largest to smallest by open remote-US listings and show a running total.
window
no
no
yes
yes
yes
no
no
no
no
no
no
no
no
no
no
W05
For each seniority level, show the median pay midpoint of open remote-US listings and how far it is above or below the overall median.
window
no
no
no
no
no
no
no
yes
yes
no
no
yes
yes
yes
yes
W06
How many open remote-US listings pay more than 1.5 times the median pay of their own role family?
window
no
no
no
no
no
no
no
no
no
no
no
yes
yes
yes
yes
W07
Within each role family, what percentage of open remote-US listings with posted pay are at the senior level or above (senior, staff+, manager+)?
window
no
no
yes
yes
yes
no
yes
yes
yes
yes
no
no
yes
yes
yes
D01
How many open remote-US listings were published in the 30 days up to and including the latest crawl?
date
no
no
no
no
no
no
yes
yes
yes
yes
no
yes
yes
yes
yes
D02
How many open remote-US listings were published in each month of 2026?
date
no
no
yes
yes
yes
no
no
no
no
no
no
yes
yes
yes
yes
D03
What is the oldest open remote-US listing by publish date? Give its title, company and publish date.
date
yes
yes
yes
yes
yes
yes
no
yes
yes
yes
yes
yes
yes
yes
yes
D04
What is the median age in days of open remote-US listings in each role family?
date
no
yes
no
no
yes
no
yes
yes
yes
yes
no
no
yes
yes
no
D05
How many open remote-US listings have been up for more than a year?
date
no
no
yes
yes
yes
no
yes
yes
yes
yes
no
no
no
no
yes
D06
On which day of the week were the most open remote-US listings published?
date
yes
yes
no
no
no
no
no
yes
yes
yes
yes
yes
yes
yes
yes
D07
How many open remote-US listings were published in the third quarter of 2026?
date
no
no
yes
yes
yes
no
yes
yes
yes
yes
no
yes
yes
yes
yes
D08
How many open remote-US listings still up today were first published before 2026, by publication year?
date
no
yes
yes
yes
yes
no
yes
yes
yes
yes
no
yes
yes
yes
yes
M01
How many AI jobs are there?
ambiguous
no
no
no
no
no
no
yes
yes
yes
yes
no
yes
yes
yes
yes
M02
What do data scientists make?
ambiguous
no
no
yes
yes
yes
no
yes
yes
yes
yes
no
yes
yes
yes
yes
M03
Who is hiring the most data people remotely?
ambiguous
no
no
yes
yes
yes
no
no
yes
yes
yes
no
yes
yes
yes
yes
M04
Is Python or SQL more in demand?
ambiguous
no
no
no
no
no
no
yes
yes
yes
no
no
yes
yes
yes
yes
M05
How many entry-level jobs are open?
ambiguous
no
no
no
no
no
no
yes
no
no
no
no
yes
yes
yes
no
M06
How many remote jobs are in the dataset?
ambiguous
no
no
no
no
no
no
no
no
no
no
no
yes
no
no
no
M07
How many jobs mention dbt?
ambiguous
no
no
no
no
no
no
yes
no
no
no
no
yes
yes
yes
yes
U01
How many people applied to each posting?
unanswerable
yes
yes
no
no
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
U02
What is the gender breakdown of people hired into data roles?
unanswerable
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
U03
Which companies offer equity or stock options with their data roles?
unanswerable
no
no
no
no
no
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
U04
How did the number of open AI engineer listings change month over month since June 2026?
unanswerable
no
no
no
no
no
no
no
no
no
no
no
yes
yes
yes
yes
U05
Which of these listings offer visa sponsorship?
unanswerable
no
no
no
no
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
yes
U06
What is the average headcount of the companies that are hiring?