We ran a four-person panel for a product manager role last year. After the final round we sat down to debrief and forty minutes later had decided nothing. One interviewer was in love with the candidate's strategic thinking. Another was nervous about their communication. A third said the candidate seemed "off" but couldn't say why. The fourth hadn't taken notes and was reconstructing impressions from memory. We ended up deferring to the most senior person in the room, which is exactly the kind of decision none of us would defend in daylight, and which — if I'm honest — is how most hiring decisions at most companies actually get made.
We weren't comparing candidates — we were comparing ourselves
The root cause was structural, not personal. Four people walked into four different interviews, asked four different sets of questions, and applied four private, unspoken definitions of "good." We weren't evaluating the candidate against the role. We were evaluating our own interviewers' taste against each other, and then resolving the tie by rank. No rubric, no shared questions, no shared definition of the bar — there was no chance of a clean decision and we'd never set ourselves up for one.
It's worse than just inefficient. An unstructured debrief is where bias does its quietest work, because "culture fit" and "seemed off" are exactly the kind of unfalsifiable impressions that let us hire people who remind us of ourselves and pass on people who don't. Nobody in that room would have described themselves as biased. We were four well-meaning people who genuinely wanted to hire the best candidate, and we'd built a process that made it nearly impossible to tell who that was.
What the rubric fixed
For the next PM hire I ran the Structured Interview Questions & Scoring Rubric Generator prompt through Claude Sonnet 4.6. I gave it four competencies, a 45-minute slot, and the stage. It returned one primary STAR question per competency, a follow-up probe, explicit strong-and-weak-answer signals, and a clean four-level rubric, plus an independent scorecard each interviewer fills out before anyone talks. Sonnet is the right model here — it's fast, it holds the XML structure tightly across a long output, and it's careful about the legal guardrails, so I never have to scrub a sketchy question about someone's family situation or age before sending the guide to the panel. Every interviewer got the same questions and the same rubric, and we agreed in advance that scorecards were due before the debrief.
- Debrief time: 40 minutes unstructured down to 18 minutes structured
- Everyone filled an independent scorecard before talking, so nobody anchored on the loudest or most senior opinion
- Disagreements became 'I scored Communication a 2, you scored a 3, here's the quote' — actually resolvable in seconds
- Time-to-decision: same day, not the 'let's sleep on it' two-day drift that usually means we forgot the details
Naming the 'off' feeling
The interviewer who said the candidate seemed "off" turned out to mean a clean 2 on the Communication row, which the rubric described as "answers are clear but lose focus under follow-up." Once it had a name and a definition, it was debatable instead of mystical. We talked about it for ninety seconds, agreed it was real but not disqualifying for this stage, scored it honestly, and moved on. Compare that to the previous debrief, where the same kind of feeling triggered fifteen minutes of circular discussion because nobody could pin it to anything. The rubric doesn't make feelings go away — it gives them a place to live where they can be examined.
We hired a different candidate that round and she's been one of the strongest PMs we've ever had — and crucially, I can point to the scorecard that says exactly why we picked her over the charismatic runner-up. That paper trail has saved me twice since, once in a calibration discussion where someone questioned the decision months later, and once when a rejected candidate politely asked for feedback and I could give them something specific and fair instead of a hand-wavy 'it wasn't quite a fit.'
There's a second benefit I didn't anticipate: the rubric made our junior interviewers braver. Before, a more junior person on the panel would defer to the senior voice almost automatically, because they had no language to push back with. Now a first-year teammate can say 'I scored that a 2 and here's the quote,' and the seniority in the room genuinely doesn't matter — the evidence is the evidence. That's quietly changed who feels allowed to have an opinion in our hiring, which is exactly the kind of thing you can't fix with a memo.
The honest caveat
The rubric only works if people actually fill the scorecard before they talk. The first time, two interviewers "forgot" and we slid right back into vibes for ten minutes before someone called it. Now I send the scorecard as a meeting pre-read and keep the debrief doc locked until everyone's submitted — boring process hygiene that turns out to be the whole ballgame. Grab the Structured Interview Questions & Scoring Rubric Generator prompt on Prompt Dock and run it before your next loop. I generate the rubric the same week I write the posting with my Inclusive Job Description Writer from a One-Line Role Brief prompt, so the whole funnel speaks the same language from the first click to the final yes.