Every rejection note has a house style. We’ve thought more about what this role needs, and it isn’t the right match. Sometimes that’s the honest summary of a real finding. More often it’s a verdict from a process that never collected the evidence for it.
Fit is the most load-bearing word in startup hiring and the least disciplined one. It decides more outcomes than the take-home and the panel debrief combined, and unlike those it goes unwritten. Nobody signs their name to it. It shows up at the end, in the summary email, after a loop that tested something else entirely.
Nobody is a fit in the abstract
The claim always has four parts: this person, in this seat, under these constraints, against this bar. Drop one and you’re describing a personality, which is neither testable nor any of your business.
Two versions of the same criterion. First: “works well with others.” Second: “will carry the pager alone through a weekend for a system they didn’t build, and will wake a colleague when the runbook is wrong rather than guess.”
The first is a vibe. The second can be wrong, which means it can be checked. Build a round that produces evidence against it and you’ll find out. There’s no round you can build for “works well with others,” which is why teams that lean on it end up grading candidates against whichever colleague they happen to resemble.
The rewrite costs an afternoon. Most startups skip it because the founders think they already agree on what the seat is. They don’t, and the loop is where that surfaces.
If no round can test it, it isn’t a reason
Every criterion in your fit definition gets a named round attached to it — the round that could produce evidence against it. If nothing on the schedule can do that, the criterion is an impression, and impressions don’t get to reject people.
Interviewers aren’t neutral instruments. Jason Dana, Robyn Dawes, and Nathanial Peterson ran the cleanest demonstration I know of. In the first of three studies, participants predicting other students’ semester GPAs formed impressions just as confidently when the interviewee was answering with a random response system as when the answers were real, and they predicted worse than if they’d skipped the interview and used prior GPA alone. Watching an interview didn’t hurt. Conducting one did. Their abstract ends on a single sentence: “Our simple recommendation for those making screening decisions is not to use them.”
That’s students predicting GPAs, not managers predicting job performance, and the gap matters. What transfers is the mechanism: a confident read built on noise, crowding out the information you already had.
Structure is the fix. Google reports that its structured loops predicted on-the-job performance across functions and levels, saved about 40 minutes per interview in interviewer time, and left rejected candidates 35% happier than those who went through an unstructured one. That’s a company grading its own homework, so take it as a ceiling. The independent version is stronger than Google’s marketing needs it to be. When Sackett and colleagues re-ran the personnel-selection meta-analyses and corrected for decades of overstated range-restriction adjustments, most selection methods lost validity. Structured interviews came out top-ranked anyway.
Sit with the rejected-candidates number for a second. People can accept losing to a bar. Losing to a mood is what they don’t forgive.
You didn’t fail by interviewing. You failed by grading a criterion nothing in the loop assessed. Reject someone on collaboration style after four rounds that only ever probed technical depth and you’ve written a conclusion with no data under it. “Fit” is the word that covers the gap.
The seat has to exist before anyone can miss it
The most common fit failure is a role-definition failure wearing a candidate’s face.
Watch a fast-moving company open a seat nobody has held yet. Call it the first data engineer. One interviewer describes someone who builds the warehouse and owns the models end to end. Another describes a partner to the analysts, mostly unblocking their queries. A third thinks the job is finally making the dashboards trustworthy. All three are senior. All three are sincere. All three are describing different jobs. The candidate answers each of them accurately and then gets measured against a composite that exists in nobody’s head.
This isn’t only a hiring problem. Gallup’s engagement indicator puts the share of U.S. employees who say they know what’s expected of them at work at 49% as of May 2026. Role clarity is scarce for people who already have the job. Expecting it to show up spontaneously for a job that doesn’t exist yet is optimistic.
Here’s a cheap diagnostic, and a slightly brutal one. Before the req goes out, have every interviewer write one paragraph, separately: what does this person do in week six, who do they answer to day to day, what do they own that nobody else touches. Read the paragraphs side by side. If they don’t converge, there’s no seat yet. There’s a wish, and nobody can fit a wish.
Your loop is the honest demo
Culture is behavior under constraint. A hiring loop is a constraint, so the loop is the most honest thing a candidate sees.
They read it as an operating model, because that’s what it is. Rounds that appear and disappear mid-process. Headcount numbers that don’t agree between two conversations a week apart. A decision-maker named early who never materializes. A timeline that quietly doubles. None of that reads as a scheduling problem — it reads as how you behave when you’re busy. The people who pick up on it fastest are the ones with other options, which is a filter you didn’t design and wouldn’t choose.
If your loop can’t hold its shape for six weeks under normal load, better to learn that before a real deadline tests the same thing.
Write the fit spec
The artifact is small: one table, agreed before the req goes live, a row per criterion. Here’s one filled in for that first data engineer. The rows change per role. The shape doesn’t.
| Claim about the seat | What would disprove it | Round that tests it | Disqualifying signal |
|---|---|---|---|
| Ships against a charter nobody has written yet | Needs the spec settled before starting | Scenario round where the requirements contradict each other | Asks for the org chart before the problem |
| Makes the analysts self-sufficient instead of becoming their queue | Optimizes for doing the work themselves | Live teach-back on something they built | Can’t explain the work to a non-specialist |
| Tolerates a role whose shape changes every quarter | Wants a fixed remit before committing | Direct question, asked the same way of every candidate | Treats a scope change as a broken promise |
| Picks a schema under incomplete information and defends it later | Waits for more data before deciding | Design round with a deliberately thin brief | Won’t commit to a call under uncertainty |
Then one governing rule. Rejection reasons must come off the table. If someone in the debrief raises a reason that isn’t on it, either the spec was incomplete — add the row, re-run the round that tests it — or the reason wasn’t real, and it doesn’t get to decide.
I build this kind of thing for a living, though usually for machines. rubric-bench exists because an LLM scoring free-text answers is a production dependency: you pin a golden set of answers to known verdicts, re-score after every prompt edit or model bump, and fail the build when the verdicts move. Nobody would ship a grader that way and skip the tests. A hiring panel is the same object with worse instrumentation — a grader nobody regression-tests, running a rubric nobody wrote down.
The spec isn’t free. Four rows with a round attached to each adds a week to a loop that a twenty-person company may not have. Published criteria are rehearsable: say “we test for scope pushback” out loud and you’ll meet candidates who’ve practiced pushing back on scope. And a spec tight enough to be testable will screen out the atypical hire who would have redefined the seat. Pay it anyway. The alternative isn’t “no spec.” It’s a spec living in four people’s heads, applied at the end, where nobody has to defend it.
The counter-argument
Fit isn’t fake, and treating it as pure bias overcorrects.
Lauren Rivera’s study of elite professional service firms found that evaluators sought candidates culturally similar to themselves in leisure pursuits, experiences, and self-presentation, and that concerns about shared culture “often outweighed concerns about absolute productivity”. Read closely, that’s a finding about unstated fit rather than fit as such. Her setting was law, banking, and consulting, not early-stage startups, and it’s a case study built from 120 interviews rather than a controlled trial. The mechanism travels anyway. A criterion nobody wrote down collapses into similarity-to-me, because that’s the only yardstick left in the room.
Genuine style mismatches exist too. Someone who does their best work in a quiet queue will be miserable in a seat that’s five customer calls a day, and hiring them is unkind to both parties. So style can be a real reason. What it can’t be is unstated, because unstated fit is unfalsifiable, and unfalsifiable reasons attract every bias you have. Name the constraint in the job post and you’ve been fair. Discover it in round four and call it fit, and you haven’t.
Back to the rejection note
Write the spec before the req, attach a round to every row, and forbid off-spec reasons in the debrief. That part’s mechanical.
The sentence at the end is the harder part. Tell people what you actually decided against. “We wanted someone who’s built a warehouse from nothing before, and we couldn’t establish that from the loop” is a sentence a person can act on, and one you can only write if you did the work to know it. “Not the right match on either side” is what you send when you didn’t.