The AI-Era Engineering Playbook

Product Engineer: Interview Scorecard


Stage 1 — Specification and Mental Modelling

Live spec writing exercise. Requirement provided verbally and in writing. 5 min think time, then discussion.

QuestionWeightScoreNotes
P1-Q1 — Clarifying questions: what must be known before specifying?1.5×
P1-Q2 — Write the specification: behavioural, testable, complete (highest weight)
P1-Q3 — Deliberate exclusions: what was scoped out and why?
P1-Q4 — AI gap anticipation: what would AI exploit or miss in this spec?1.5×
Stage 1: / 30

Stage 2 — Live AI-Assisted Build and Evaluation

Observed real-time work session. Candidate uses their own AI tool. 20 min build, then debrief. This is the highest-validity stage — observe carefully.

Observation checklist (fill during the live session)

Observable behaviourObservedNotes
Provides context and specification in the prompt before generating
Reads output carefully before running it
Checks output against the specification, not just "looks right"
Iteration is diagnostic — changes are based on a hypothesis
Pauses to evaluate before generating again

Planted issues

Issue descriptionCategoryFound?Reasoning quality

Debrief questions (post-session)

QuestionWeightScoreNotes
P2-Q1 — Walk through function: including edge case behaviour?1.5×
P2-Q2 — Spec match: what's different between spec and output? (highest weight)
P2-Q3 — Iteration reasoning: how did you decide what to change?1.5×
P2-Q4 — Missing tests: specific cases the AI didn't write?
Stage 2 average: (auto-reject if below 2.0)
Stage 2: / 30

Stage 3 — Output Review and Edge Case Hunting

50–80 lines of AI-generated code with planted issues. 10 min review, then 30 min discussion.

QuestionWeightScoreNotes
P3-Q1 — Code review: find issues, explain failure behaviour (highest weight)
P3-Q2 — Silent failures: wrong but no error raised?1.5×
P3-Q3 — Adversarial: what would a malicious user do with this code?
P3-Q4 — Missing tests: specific cases by risk priority?
Stage 3: / 27.5

Stage 4 — Structured Behavioural

Consistent questions, same order per candidate. Score on observed evidence only.

QuestionWeightScoreNotes
P4-Q1 — Feature wrong in production: traced back to specification?
P4-Q2 — Pushed back on underspecified requirement: what and outcome?
P4-Q3 — AI output looked right but wasn't: how did you catch it?1.5×
P4-Q4 — AI use / no-use boundary: principled, not habitual?
P4-Q5 — Current experiments: specific, from real use, formed opinions (2×, disqualifier if 1)
Stage 4: / 32.5

/ 100
Score questions above to see recommendation

Compensation calibration — check if observed


Debrief notes

Complete before comparing with co-interviewer.

Most decisive positive signal Most decisive negative signal or absence If Hire with conditions — what specifically needs to develop One thing this candidate does that most candidates don't

Fill in scores and notes above, then generate a structured prompt to paste into Claude, ChatGPT, or any LLM for a calibrated debrief summary.

✓ Copied to clipboard

© Gabor Mayer. Licensed under Creative Commons Attribution 4.0 (CC BY 4.0). Free to share and adapt with attribution.