Every rejection note has a house style. We’ve thought more about what this role needs, and it isn’t the right match. Sometimes that’s the honest summary of a real finding. More often it’s a verdict from a process that never collected the evidence for it.

Fit is the most load-bearing word in startup hiring and the least disciplined one. It decides more outcomes than the take-home and the panel debrief combined, and unlike those it goes unwritten. Nobody signs their name to it. It shows up at the end, in the summary email, after a loop that tested something else entirely.

Nobody is a fit in the abstract

The claim always has four parts: this person, in this seat, under these constraints, against this bar. Drop one and you’re describing a personality, which is neither testable nor any of your business.

Two versions of the same criterion. First: “works well with others.” Second: “will carry the pager alone through a weekend for a system they didn’t build, and will wake a colleague when the runbook is wrong rather than guess.”

The first is a vibe. The second can be wrong, which means it can be checked. Build a round that produces evidence against it and you’ll find out. There’s no round you can build for “works well with others,” which is why teams that lean on it end up grading candidates against whichever colleague they happen to resemble.

The rewrite costs an afternoon. Most startups skip it because the founders think they already agree on what the seat is. They don’t, and the loop is where that surfaces.

If no round can test it, it isn’t a reason

Every criterion in your fit definition gets a named round attached to it — the round that could produce evidence against it. If nothing on the schedule can do that, the criterion is an impression, and impressions don’t get to reject people.

Interviewers aren’t neutral instruments. Jason Dana, Robyn Dawes, and Nathanial Peterson ran the cleanest demonstration I know of. In the first of three studies, participants predicting other students’ semester GPAs formed impressions just as confidently when the interviewee was answering with a random response system as when the answers were real, and they predicted worse than if they’d skipped the interview and used prior GPA alone. Watching an interview didn’t hurt. Conducting one did. Their abstract ends on a single sentence: “Our simple recommendation for those making screening decisions is not to use them.”

That’s students predicting GPAs, not managers predicting job performance, and the gap matters. What transfers is the mechanism: a confident read built on noise, crowding out the information you already had.

Structure is the fix. Google reports that its structured loops predicted on-the-job performance across functions and levels, saved about 40 minutes per interview in interviewer time, and left rejected candidates 35% happier than those who went through an unstructured one. That’s a company grading its own homework, so take it as a ceiling. The independent version is stronger than Google’s marketing needs it to be. When Sackett and colleagues re-ran the personnel-selection meta-analyses and corrected for decades of overstated range-restriction adjustments, most selection methods lost validity. Structured interviews came out top-ranked anyway.

Sit with the rejected-candidates number for a second. People can accept losing to a bar. Losing to a mood is what they don’t forgive.

You didn’t fail by interviewing. You failed by grading a criterion nothing in the loop assessed. Reject someone on collaboration style after four rounds that only ever probed technical depth and you’ve written a conclusion with no data under it. “Fit” is the word that covers the gap.

The seat has to exist before anyone can miss it

The most common fit failure is a role-definition failure wearing a candidate’s face.

Watch a fast-moving company open a seat nobody has held yet. Call it the first data engineer. One interviewer describes someone who builds the warehouse and owns the models end to end. Another describes a partner to the analysts, mostly unblocking their queries. A third thinks the job is finally making the dashboards trustworthy. All three are senior. All three are sincere. All three are describing different jobs. The candidate answers each of them accurately and then gets measured against a composite that exists in nobody’s head.

This isn’t only a hiring problem. Gallup’s engagement indicator puts the share of U.S. employees who say they know what’s expected of them at work at 49% as of May 2026. Role clarity is scarce for people who already have the job. Expecting it to show up spontaneously for a job that doesn’t exist yet is optimistic.

Here’s a cheap diagnostic, and a slightly brutal one. Before the req goes out, have every interviewer write one paragraph, separately: what does this person do in week six, who do they answer to day to day, what do they own that nobody else touches. Read the paragraphs side by side. If they don’t converge, there’s no seat yet. There’s a wish, and nobody can fit a wish.

Your loop is the honest demo

Culture is behavior under constraint. A hiring loop is a constraint, so the loop is the most honest thing a candidate sees.

They read it as an operating model, because that’s what it is. Rounds that appear and disappear mid-process. Headcount numbers that don’t agree between two conversations a week apart. A decision-maker named early who never materializes. A timeline that quietly doubles. None of that reads as a scheduling problem — it reads as how you behave when you’re busy. The people who pick up on it fastest are the ones with other options, which is a filter you didn’t design and wouldn’t choose.

If your loop can’t hold its shape for six weeks under normal load, better to learn that before a real deadline tests the same thing.

Write the fit spec

The artifact is small: one table, agreed before the req goes live, a row per criterion. Here’s one filled in for that first data engineer. The rows change per role. The shape doesn’t.

Claim about the seatWhat would disprove itRound that tests itDisqualifying signal
Ships against a charter nobody has written yetNeeds the spec settled before startingScenario round where the requirements contradict each otherAsks for the org chart before the problem
Makes the analysts self-sufficient instead of becoming their queueOptimizes for doing the work themselvesLive teach-back on something they builtCan’t explain the work to a non-specialist
Tolerates a role whose shape changes every quarterWants a fixed remit before committingDirect question, asked the same way of every candidateTreats a scope change as a broken promise
Picks a schema under incomplete information and defends it laterWaits for more data before decidingDesign round with a deliberately thin briefWon’t commit to a call under uncertainty

Then one governing rule. Rejection reasons must come off the table. If someone in the debrief raises a reason that isn’t on it, either the spec was incomplete — add the row, re-run the round that tests it — or the reason wasn’t real, and it doesn’t get to decide.

I build this kind of thing for a living, though usually for machines. rubric-bench exists because an LLM scoring free-text answers is a production dependency: you pin a golden set of answers to known verdicts, re-score after every prompt edit or model bump, and fail the build when the verdicts move. Nobody would ship a grader that way and skip the tests. A hiring panel is the same object with worse instrumentation — a grader nobody regression-tests, running a rubric nobody wrote down.

The spec isn’t free. Four rows with a round attached to each adds a week to a loop that a twenty-person company may not have. Published criteria are rehearsable: say “we test for scope pushback” out loud and you’ll meet candidates who’ve practiced pushing back on scope. And a spec tight enough to be testable will screen out the atypical hire who would have redefined the seat. Pay it anyway. The alternative isn’t “no spec.” It’s a spec living in four people’s heads, applied at the end, where nobody has to defend it.

The counter-argument

Fit isn’t fake, and treating it as pure bias overcorrects.

Lauren Rivera’s study of elite professional service firms found that evaluators sought candidates culturally similar to themselves in leisure pursuits, experiences, and self-presentation, and that concerns about shared culture “often outweighed concerns about absolute productivity”. Read closely, that’s a finding about unstated fit rather than fit as such. Her setting was law, banking, and consulting, not early-stage startups, and it’s a case study built from 120 interviews rather than a controlled trial. The mechanism travels anyway. A criterion nobody wrote down collapses into similarity-to-me, because that’s the only yardstick left in the room.

Genuine style mismatches exist too. Someone who does their best work in a quiet queue will be miserable in a seat that’s five customer calls a day, and hiring them is unkind to both parties. So style can be a real reason. What it can’t be is unstated, because unstated fit is unfalsifiable, and unfalsifiable reasons attract every bias you have. Name the constraint in the job post and you’ve been fair. Discover it in round four and call it fit, and you haven’t.

Back to the rejection note

Write the spec before the req, attach a round to every row, and forbid off-spec reasons in the debrief. That part’s mechanical.

The sentence at the end is the harder part. Tell people what you actually decided against. “We wanted someone who’s built a warehouse from nothing before, and we couldn’t establish that from the loop” is a sentence a person can act on, and one you can only write if you did the work to know it. “Not the right match on either side” is what you send when you didn’t.