Pre-push challenging: you cannot push what you cannot explain
An agent wrote the change. You read the diff, it looked right, you approved it. A week later someone asks what dropping that timeout from 30 seconds to 10 actually changed, and you genuinely do not know. You did not write it, and you never really understood it. You recognized it.
I built a small tool that closes that gap at the last cheap moment: a git pre-push hook that asks you questions about your own outgoing commits and refuses the push until you answer them yourself. Commits stay free. Publishing is what costs you an explanation.
Comprehension debt
Technical debt lives in the code. It is visible, shared, and anyone on the team can read it.
This is different. The code can be excellent and the debt still accumulates, because the missing thing is not in the repository. It is in you. Call it comprehension debt: the gap between what your name is on and what you could explain without opening the file.
Technical debt is public. Comprehension debt is private, which is what makes it dangerous. Nobody can see it, including you, until you need the knowledge and find it was never there.
Why reading the diff does not work
The obvious fix is to read the diff more carefully. That does not work, and the learning research is direct about why.
Re-reading is the most popular study technique there is, and one of the weakest. It produces fluency: the text gets easier each pass, and that ease feels like understanding. In one classroom study cited in Make It Stick, students who reviewed material three times averaged a C+. Students who took three quizzes on the same material averaged an A−. Same content, same exam.
The harder it is to recall something from memory, the more the act of recalling it strengthens your hold on it. Learning that feels easy fades fast.
AI-generated code is the purest fluency trap I have met. It is well named and consistently formatted. It reads better than most human code, so it triggers the feeling of understanding faster and with less justification.
Reading the diff
Recognition. The answer is on screen, so you never find out whether you could have produced it.
Answering a question
Retrieval. The answer is not on screen. You either have the model in your head, or you find out now that you do not.
The gate
Commits are thinking out loud, so gating them would make the tool something you disable within a day. Pushing is different: it is the moment the code becomes everyone else's problem.
The tempting shortcut is to quiz on the working-tree diff or on HEAD. Both are wrong: a clean tree can still hold ten unpushed commits. A pre-push hook receives the truth on stdin, one line per ref, so that is what gets used.
One bug I shipped and had to fix: my first version excluded commits already on any remote. Pushing existing commits to a different remote still publishes them there, so the exclusion has to be scoped to the destination. Otherwise you can walk the whole quiz around the gate by adding a second remote.
The questions
Questions are generated by whichever coding agent you already use, from the real diffs. Changes are grouped by risk, and the number of questions scales with complexity, capped at ten.
| Change | Churn | Critical topics | Questions |
|---|---|---|---|
| One-line doc fix | 1 | 0 | 1 |
| Medium feature | 22 | 2 | 3 |
| Large refactor | 259 | 4 | 8 |
| Small but risky | 5 | 4 | 4 |
That last row matters most: the budget can never fall below the number of critical topics, so a five-line change touching four of them still gets four questions. Size is a bad proxy for risk.
Questions may only ask what the code does, what breaks, and who is affected. Never why the author chose it, because intent is not in the diff. Answers are judged on meaning, not wording, and there is no averaging: every topic has to stand on its own. "I don't know" is a first-class answer that triggers a teaching loop rather than a failure.
See it run
One real push, from the branch that contained the tool itself.
client:- timeout_seconds: 30+ timeout_seconds: 10- retries: 0+ retries: 2
config/settings.yml changed timeout_seconds 30 → 10 and retries 0 → 2. What is the practical consequence today?
config/settings.yml changed timeout_seconds 30 → 10 and retries 0 → 2. What is the practical consequence today?
client.py defines REQUEST_TIMEOUT = 10 as a module constant and never loads settings.yml. The two values agree today by coincidence.Ops edits settings.yml to timeout_seconds: 60 and restarts the service. What changes for a slow request?
settings.yml. It hardcodes its own timeout and implements no retry, so tuning those values in production would do nothing at all.Run A is the one that actually happened to me. The diff was clean and did exactly what it said: it changed two values in a file nothing reads. The defect was in me: I had approved a change whose effect I had not worked out, and I would have pushed it with complete confidence. The gate does not decide whether an inert change ships. It makes sure I know it is inert when I decide.
That is why "I don't know" is a button and not a failure. Punishing it just teaches you to guess, and a plausible guess is exactly the thing that hides the gap instead of closing it.
Limits
It is bypassable. The hook lives on your machine. You can delete it or run git push --no-verify and be done in three seconds, and so can an agent that runs git push for you. That is fine, because the threat model is not an adversary, it is you on a Thursday moving fast. For real enforcement you want branch protection and CI.
It fires once. The same research says retrieval works better spaced out. This quizzes you once, at push time. That is a real limitation, and the first of the next steps below.
It costs time and needs your agent. Every push adds a few model calls plus your own typing, and the questions only exist if your coding agent is installed, signed in and reachable.
The scope is strict. The record is bound to commit IDs, so an amend or a rebase asks everything again, even code you already explained. Pushing to a fork or a fresh remote counts every commit that remote has never seen, including other people's.
The grader is a model. It can pass a fluent, vague answer and fail a terse, correct one. That is the same fluency trap, one level up.
Next steps
- Spaced follow-ups. Bring a closed topic back on a later push, when recall is hard enough to count.
- Records bound to patches, not commits. Match on
git patch-id, so a rebase only asks about what actually changed. - An eval for the grader. Correct, partial and fluent-but-wrong answers, each graded several times, with the false-pass rate published.
None of this is about distrusting the model. The code it writes is often better than mine, and that is exactly the problem: quality is what makes it easy to accept without understanding. The quiz is not quality control on the AI. It is quality control on me.