Product Engineer: Interview Scorecard
Score 1–5 per question. Rubric anchors: Product Engineer question bank →
Stage 1 — Specification and Mental Modelling
Live spec writing exercise. Requirement provided verbally and in writing. 5 min think time, then discussion.
| Question | Weight | Score | Notes |
|---|---|---|---|
| P1-Q1 — Clarifying questions: what must be known before specifying? | 1.5× | ||
| P1-Q2 — Write the specification: behavioural, testable, complete (highest weight) | 2× | ||
| P1-Q3 — Deliberate exclusions: what was scoped out and why? | 1× | ||
| P1-Q4 — AI gap anticipation: what would AI exploit or miss in this spec? | 1.5× |
Stage 2 — Live AI-Assisted Build and Evaluation
Observed real-time work session. Candidate uses their own AI tool. 20 min build, then debrief. This is the highest-validity stage — observe carefully.
Observation checklist (fill during the live session)
| Observable behaviour | Observed | Notes |
|---|---|---|
| Provides context and specification in the prompt before generating | ||
| Reads output carefully before running it | ||
| Checks output against the specification, not just "looks right" | ||
| Iteration is diagnostic — changes are based on a hypothesis | ||
| Pauses to evaluate before generating again |
Planted issues
| Issue description | Category | Found? | Reasoning quality |
|---|---|---|---|
Debrief questions (post-session)
| Question | Weight | Score | Notes |
|---|---|---|---|
| P2-Q1 — Walk through function: including edge case behaviour? | 1.5× | ||
| P2-Q2 — Spec match: what's different between spec and output? (highest weight) | 2× | ||
| P2-Q3 — Iteration reasoning: how did you decide what to change? | 1.5× | ||
| P2-Q4 — Missing tests: specific cases the AI didn't write? | 1× |
Stage 3 — Output Review and Edge Case Hunting
50–80 lines of AI-generated code with planted issues. 10 min review, then 30 min discussion.
| Question | Weight | Score | Notes |
|---|---|---|---|
| P3-Q1 — Code review: find issues, explain failure behaviour (highest weight) | 2× | ||
| P3-Q2 — Silent failures: wrong but no error raised? | 1.5× | ||
| P3-Q3 — Adversarial: what would a malicious user do with this code? | 1× | ||
| P3-Q4 — Missing tests: specific cases by risk priority? | 1× |
Stage 4 — Structured Behavioural
Consistent questions, same order per candidate. Score on observed evidence only.
| Question | Weight | Score | Notes |
|---|---|---|---|
| P4-Q1 — Feature wrong in production: traced back to specification? | 1× | ||
| P4-Q2 — Pushed back on underspecified requirement: what and outcome? | 1× | ||
| P4-Q3 — AI output looked right but wasn't: how did you catch it? | 1.5× | ||
| P4-Q4 — AI use / no-use boundary: principled, not habitual? | 1× | ||
| P4-Q5 — Current experiments: specific, from real use, formed opinions (2×, disqualifier if 1) | 2× |
Compensation calibration — check if observed
Debrief notes
Complete before comparing with co-interviewer.
Most decisive positive signal Most decisive negative signal or absence If Hire with conditions — what specifically needs to develop One thing this candidate does that most candidates don'tFill in scores and notes above, then generate a structured prompt to paste into Claude, ChatGPT, or any LLM for a calibrated debrief summary.
✓ Copied to clipboard© Gabor Mayer. Licensed under Creative Commons Attribution 4.0 (CC BY 4.0). Free to share and adapt with attribution.