The problem
If you’ve used a SQL practice tool before, you’ve hit one of two walls. Either it grades your final query against an answer key, pass or fail, no in-between. Or, once you’re stuck, it just hands you the correct query.
Neither builds the skill that actually generalizes. The real gap isn’t syntax, it’s decomposition: breaking an ambiguous question into shape, joins, filters, and aggregation before you write a single line.

The approach
Picture a hint ladder: a fixed staircase of hints, from the smallest possible nudge at the bottom to the full answer at the top. You climb it one rung at a time. You never get boosted to the top early, no matter how stuck you are, because reaching for the top rung is the exact habit this skill is built to break.
A fixed reasoning frame runs before you write any SQL. The hint ladder is what carries you through once you get stuck.
How it works
- A four-step pre-flight runs first, before any SQL, like a pilot’s checklist: output shape and grain, source and joins, row filters, aggregation.
- Stuck? You climb the hint ladder, one rung at a time, least to most revealing. You write each clause yourself, never handed one.
- Once you have a working query, it’s explained back to you visually: table transformations, join diagrams, an annotated clause-by-clause breakdown.
- A pattern library and a visual toolkit sit in reserve, pulled in only when you actually need them.

Design decisions worth noting
- Even “just give me the answer” gets redirected. The skill treats that as a request to get unstuck, not a request to see the finished query. You get the next rung, not the top of the ladder.
- The reasoning frame is fixed, not adaptive, on purpose. Consistency is what turns it into a habit you can run on your own later, once the skill isn’t in the room.
How does SQL Instructor teach query decomposition instead of just giving the answer?
SQL Instructor runs a fixed four-step pre-flight (output shape, source and joins, filters, aggregation) before any SQL is written, then climbs a hint ladder one rung at a time when you’re stuck, never skipping to the full answer. Tested against a 143,613-row dataset, it scored 100% versus 50% without the skill, across three runs each.
The benchmark
Tested against a real exercise: for each hospital facility, find the top 3 APR-DRGs by total charges, against a 143,613-row real dataset used elsewhere in this portfolio. A genuine step up from single-table top-N, aggregating raw rows to facility-and-DRG totals first, then ranking within each facility with a window function. Both conditions ran 3 times, since a single pair can’t separate a real effect from run-to-run noise.
| Metric | With skill | Without skill |
|---|---|---|
| Pass rate | 100% | 50% |
| Time | 48.6s | 28.6s |
| Tokens | 48,466 | 27,937 |
Three checks landed exactly as predicted: the pre-flight question, the refusal to hand over the answer on first ask, and the visual walkthrough, all failing at baseline and passing with the skill every run. Each maps to a named, explicit behavior in the skill, not an incidental side effect.
Three checks didn’t differentiate, and that’s an honest finding, not a gap. A capable baseline model already writes a correct two-level query and explains window function mechanics reasonably well without any coaching. What the skill actually changes is process and pedagogy, not baseline correctness.
Grading is self-graded, unweighted checklist counts, so treat the time and token figures as directional, not a production cost estimate. Full run data in the case study.

What this demonstrates
This is designed instruction, not a chatbot with a personality bolted on: a named teaching philosophy, a fixed decision procedure, and an escalation ladder that holds even when you push back and ask for the answer directly.
Using this, you can build the same discipline into anything that has to teach instead of just do the work: a training tool, a support copilot, an onboarding flow, anywhere the job is getting someone to climb the ladder themselves.
The same refusal-by-design discipline shows up in Operator Brain Evaluation, which holds a hard scope boundary instead of a hint ladder.
Common questions
Will SQL Instructor just give you the answer if you ask directly?
No. Even “just give me the answer” gets redirected. The skill treats that as a request to get unstuck, not a request to see the finished query, and gives you the next rung on the hint ladder instead.
How was the benchmark structured?
Against a real exercise (top 3 APR-DRGs by total charges per hospital facility, on the same 143,613-row dataset used elsewhere in this portfolio), run 3 times per condition to separate a real effect from run-to-run noise.