Zeeshan Khan

A Maker Cannot Grade Itself

Essay

~7 min read · 2026

The first essay was about earning the standing to judge — staying close enough to the work that your own instruments still read true. This one is about what you do with that standing, which is the actual job: deciding what, and whom, to vouch for.

There is one rule underneath all of it. A maker cannot grade its own work. Not because makers are dishonest, but because the grader and the made thing share the same blind spots — the same misread of the requirement, the same case that never occurred to anyone, the same confidence that the plausible is the correct. This has always been true of people, which is why we invented review. What's new is that our most productive maker now grades itself by default, reports that it passed, and is believed.

The bold engineer

The people who take to these tools fastest are not the cautious ones. They are bold — they treat the whole lifecycle as delegable, ship at a cadence that would have been reckless three years ago and often isn't now, and they are frequently right. This is an asset, not a problem, and the instinct to slow them to a safe crawl is how you turn your best people into your most frustrated ones.

But boldness has a blind spot, and it is not carelessness. It is that a fast, capable engineer early in their judgment cannot yet see the trust line — the point past which a change stops being reversible, where a shipped mistake is not a rollback but a customer who no longer believes you. Seeing that line isn't a matter of talent or speed. It comes from having been on the wrong side of it before, which the bold newcomer by definition has not.

The wrong fix is to make them cautious. Caution isn't judgment, and a team trained into it loses the velocity that made it valuable without gaining the discernment that would have made the velocity safe. The right move is quieter: keep enough proximity to price the risk they can't yet price, hold the irreversible calls yourself, and extend trust that is fast and specific rather than blanket. Not block everyone and not wave everyone through — those are the two moves of a leader who has stopped paying attention, and both are ruinous, one slowly and one at once. The third option, earned and selective, is available only to a leader close enough to tell the change that's about to be brilliant from the one that's about to be an incident.

The point is not to contain the boldness. It is to be the grader it lacks, so it doesn't have to lose its nerve to stay safe. Speed is theirs. The trust line is yours.

The machine that grades itself

The same rule, aimed at the tool, is where the fashionable model of AI work breaks.

That model runs on self-verification: the machine builds, checks its output against the criteria, and reports done. It's efficient, and for reversible low-stakes work it's fine. But a machine checking its own work has not produced trust. It has produced a claim. It is vouching for its output with its own word — the exact thing that needed an independent check — and a maker grading itself always passes itself.

Worse, self-certification fails silently. A wrong answer that flags itself wrong is a manageable problem; the alarm sounds, someone responds. The dangerous one reports itself correct. Nothing fails loudly. Every dashboard stays green while the work quietly rots, because the warning was written by the same process it was supposed to warn about — and that process has neither the incentive nor the capacity to indict itself. This is how a codebase fills with slop under attentive-looking oversight: not because anyone ignored a red flag, but because the flag never went up. The loop reports success right until the customer discovers otherwise, and by then the trust is spent.

So the leader's job is to place the independent check where it counts — not on everything, which is the bottleneck-with-your-name-on-it again, but on the irreversible and the customer-facing, where a silent failure is expensive and permanent. The rule for where to insist on a human grader is the same rule as before: wherever being wrong can't be undone.

The graders you haven't hired

The last place this bites is the strangest, because it's about people who don't exist yet.

The short-term math is close to unarguable. A capable engineer with good tools now does what took a small team, and the marginal junior — months from contributing, slower and rougher than the machine — looks like cost with a distant, uncertain payoff. Entry-level hiring has contracted hard, and the reasoning isn't stupid. A company measuring itself over four quarters is behaving rationally when it won't pay today for value that arrives in three years.

But the rationality is horizon-bound, and the industry has run this experiment before: freeze the bottom rung in a downturn and discover, a few years on, that there's no one with three-to-five years of experience, because the people who'd have had it were never hired into the role that produces it. It is worse now, because the tasks that were a junior's training ground are exactly the ones handed to the machine. The rung isn't frozen. It's being removed.

What the removal misses is what junior work was ever for. It was never the code. It was the person: the debugging under real uncertainty, the change that looked fine and took down production at 3am, the slow formation of the instinct that tells a senior engineer — without being able to fully say why — that a working solution is a time bomb. That instinct is the grader's eye, in its first formation. It's learned by exposure and supervised mistakes, and the machine cannot teach it, because it cannot let someone make a consequential error inside a boundary and then help them understand what it cost.

The self-interested version, for any leader who files the rest under sentiment: mentoring is how seniors stay sharp, because explaining a decision forces you to actually understand it, and it's how tacit knowledge gets out of one head and into the group before that head leaves. Kill the pipeline and your seniors don't just lose apprentices — they become irreplaceable in the worst way: overloaded, un-leavable, the sole hosts of knowledge that lives nowhere else. An org of only seniors isn't strong. It's fragile and hasn't been tested yet.

The practical version isn't complicated, though it costs more than the alternative. Measure a junior by how fast their judgment is growing, not what they ship. Keep deliberate spaces where they solve things without the machine, so the fundamentals form. And invert the old apprenticeship: let the machine write the first draft and put the junior on the harder, more instructive work of finding what's wrong with it. You are not hiring hands anymore. You are raising the next people who can grade what the machine makes — and you cannot buy that back later at any price.

What it comes down to

Cheap generation didn't make judgment less necessary. It made judgment the whole job, by removing everything that used to stand in for it. When the plausible answer is free, the scarce thing is the ability to tell whether it's true — and that never lived in a credential, an alarm, a dashboard, or a machine's confident report about itself. It lives in a person who stayed close enough to still tell.

Value migrates, as it always does when a capability gets cheap, to whoever can still do the thing that didn't. Here that thing is grading — vouching for work the work can't vouch for itself. A leader's task is to be that grader, to place that trust where it can't be faked, and to raise the people who will do it after them. The trust has to be placed somewhere, by someone who stands behind it. Being worthy of that is the part of the job that doesn't get cheaper.