Pre-push challenging: you cannot push what you cannot explain

· 6 min read

Pre-push challenging: a config diff changing a timeout and retries, a 5 / 5 quiz score, and an answer box

An agent wrote the change. You read the diff, it looked right, you approved it. A week later someone asks what dropping that timeout from 30 seconds to 10 actually changed, and you genuinely do not know. You did not write it, and you never really understood it. You recognized it.

I built a small tool that closes that gap at the last cheap moment: a git pre-push hook that asks you questions about your own outgoing commits and refuses the push until you answer them yourself. Commits stay free. Publishing is what costs you an explanation.

Comprehension debt

Technical debt lives in the code. It is visible, shared, and anyone on the team can read it.

This is different. The code can be excellent and the debt still accumulates, because the missing thing is not in the repository. It is in you. Call it comprehension debt: the gap between what your name is on and what you could explain without opening the file.

Technical debt is public. Comprehension debt is private, which is what makes it dangerous. Nobody can see it, including you, until you need the knowledge and find it was never there.

Why reading the diff does not work

The obvious fix is to read the diff more carefully. That does not work, and the learning research is direct about why.

Re-reading is the most popular study technique there is, and one of the weakest. It produces fluency: the text gets easier each pass, and that ease feels like understanding. In one classroom study cited in Make It Stick, students who reviewed material three times averaged a C+. Students who took three quizzes on the same material averaged an A−. Same content, same exam.

The harder it is to recall something from memory, the more the act of recalling it strengthens your hold on it. Learning that feels easy fades fast.

AI-generated code is the purest fluency trap I have met. It is well named and consistently formatted. It reads better than most human code, so it triggers the feeling of understanding faster and with less justification.

Reading the diff

Recognition. The answer is on screen, so you never find out whether you could have produced it.

Answering a question

Retrieval. The answer is not on screen. You either have the model in your head, or you find out now that you do not.

Reading the diff is review. Answering a question about it is retrieval. That is the whole borrowed idea.

The gate

Commits are thinking out loud, so gating them would make the tool something you disable within a day. Pushing is different: it is the moment the code becomes everyone else's problem.

git push pre-push hook
Read stdin
The real ref updates git hands the hook
→
Compute scope
Commits new to this destination
→
Check record
Bound to destination + old/new OIDs
Valid → allow
You explained exactly these commits, to exactly this remote.
Absent or stale → block
Amended, rebased, new commit, different remote. The quiz runs again on the new scope.
The hook never guesses the scope. It recomputes the scope from the ref updates and checks the record against that.

The tempting shortcut is to quiz on the working-tree diff or on HEAD. Both are wrong: a clean tree can still hold ten unpushed commits. A pre-push hook receives the truth on stdin, one line per ref, so that is what gets used.

One bug I shipped and had to fix: my first version excluded commits already on any remote. Pushing existing commits to a different remote still publishes them there, so the exclusion has to be scoped to the destination. Otherwise you can walk the whole quiz around the gate by adding a second remote.

The questions

Questions are generated by whichever coding agent you already use, from the real diffs. Changes are grouped by risk, and the number of questions scales with complexity, capped at ten.

ChangeChurnCritical topicsQuestions
One-line doc fix101
Medium feature2223
Large refactor25948
Small but risky544

That last row matters most: the budget can never fall below the number of critical topics, so a five-line change touching four of them still gets four questions. Size is a bad proxy for risk.

Questions may only ask what the code does, what breaks, and who is affected. Never why the author chose it, because intent is not in the diff. Answers are judged on meaning, not wording, and there is no averaging: every topic has to stand on its own. "I don't know" is a first-class answer that triggers a teaching loop rather than a failure.

See it run

One real push, from the branch that contained the tool itself.

grill-before-push Run A · I don’t know
$
Scope · 3 commits to origin/feature/grill-before-push
f90c0f7Cap questions and scale to complexity
0ec70beAdd pre-push hook and session server
a31d4c2Tune client timeout and retries
behavior failure config tests
Evidence · config/settings.yml
@@ -1,4 +1,4 @@ client:-  timeout_seconds: 30+  timeout_seconds: 10-  retries: 0+  retries: 2
grounded in the real diff
Question 7 of 10 · configuration

config/settings.yml changed timeout_seconds 30 → 10 and retries 0 → 2. What is the practical consequence today?

Your explanation…
Review code ▾ I don't know
Answer · human, not agent

config/settings.yml changed timeout_seconds 30 → 10 and retries 0 → 2. What is the practical consequence today?

…
Teaching loop · topic not closed
Missing concept: the config is never read
client.py defines REQUEST_TIMEOUT = 10 as a module constant and never loads settings.yml. The two values agree today by coincidence.
explained, then re-asked fresh scenario, not a reworded repeat
Follow-up · same topic

Ops edits settings.yml to timeout_seconds: 60 and restarts the service. What changes for a slow request?

…
Result
Config is inert
The client never reads settings.yml. It hardcodes its own timeout and implements no retry, so tuning those values in production would do nothing at all.
all topics closed push allowed

Two runs of the same question, replayed. Run A is the honest one: "I don't know" opens the teaching loop, and the topic only closes on a fresh scenario. Run B is what it looks like when you already have the answer. The agent writes and judges; it never fills the form.

Run A is the one that actually happened to me. The diff was clean and did exactly what it said: it changed two values in a file nothing reads. The defect was in me: I had approved a change whose effect I had not worked out, and I would have pushed it with complete confidence. The gate does not decide whether an inert change ships. It makes sure I know it is inert when I decide.

That is why "I don't know" is a button and not a failure. Punishing it just teaches you to guess, and a plausible guess is exactly the thing that hides the gap instead of closing it.

Limits

It is bypassable. The hook lives on your machine. You can delete it or run git push --no-verify and be done in three seconds, and so can an agent that runs git push for you. That is fine, because the threat model is not an adversary, it is you on a Thursday moving fast. For real enforcement you want branch protection and CI.

It fires once. The same research says retrieval works better spaced out. This quizzes you once, at push time. That is a real limitation, and the first of the next steps below.

It costs time and needs your agent. Every push adds a few model calls plus your own typing, and the questions only exist if your coding agent is installed, signed in and reachable.

The scope is strict. The record is bound to commit IDs, so an amend or a rebase asks everything again, even code you already explained. Pushing to a fork or a fresh remote counts every commit that remote has never seen, including other people's.

The grader is a model. It can pass a fluent, vague answer and fail a terse, correct one. That is the same fluency trap, one level up.

Next steps

None of this is about distrusting the model. The code it writes is often better than mine, and that is exactly the problem: quality is what makes it easy to accept without understanding. The quiz is not quality control on the AI. It is quality control on me.