Product Methods Handbook

← All decks

Methods

21 methods

STRATEGY

1

Placing Bets

Treat each big choice as a bet: you can't be sure it'll work, so decide what you'll try, how much you'll spend finding out, and when you'll stop.

What it is

You almost never know in advance whether a big initiative will work. Placing bets just means being honest about that. Instead of dressing a plan up as a sure thing, you treat it as a bet — a sensible wager based on the best evidence you have right now. It's the same logic as an investor with a bit of money: you don't bet on a certainty (there isn't one), you make a smart wager, decide how much you're willing to risk, and know what would make you walk away.

Every bet has five plain parts: what you believe (your best read of the situation, from your research), what you'll try, what success looks like, the budget (how much time and people you'll spend before you judge it), and the stop line (the result that tells you to pull out). Two habits make betting work. You place several bets, not one — like spreading money across a few stocks, because some will lose. And you judge a bet by whether it was a smart call at the time, not only by whether it happened to win: a well-reasoned bet that didn't pay off was still a good bet, and you treat it as a lesson, not a failure. That's the difference between a team that keeps trying smart new things and one that only ever plays it safe.

When to use it

Strategy (Phase 1), when you turn your goals into the few things you'll actually back this year. And again at every review, to decide which bets to keep, change, or stop. Your bets become the themes on your Roadmap.

The four moves

Place — write each bet on one page.

  • Belief: what you think is true, and why (from your research).
  • Bet: what you'll do about it.
  • Success looks like: the measure that says it worked (linked to your North Star).
  • Budget: how much time and people you'll spend before you decide — a limit you set up front, not a guess at how long it'll take.
  • Stop line: the result that means you walk away (your kill criteria).

Prioritize — you can't back everything. Rank your bets by how much they'd move your North Star, and keep a healthy mix: some safe near-term wins, some bigger long shots. Only start as many as you can actually staff — a list of twenty "active" bets is a wish list, not a plan.

Track — keep them visible. Put your bets on a simple Now / Next / Later board and check them on a regular rhythm. As evidence comes in, update how confident you are. If a bet runs past its budget, it doesn't quietly carry on — it comes back to the table to be re-decided.

Validate — judge honestly. At each review, look at the evidence (is it real, recent, and enough?) and choose one of three things: keep going, adjust, or stop. Stopping a bet because it hit the stop line you set is a success — the system working as designed — not an embarrassment.

Diagram

BET CARD
Belief
new managers quit because year one is unsupported
Bet
a first-year manager support track
Success looks like
fewer regretted manager exits
Budget
1 quarter, 1 small team
Stop line
<30% finish, OR no change by week 8
NOW
Manager support
CONFIDENCE
NEXT
Skills map
CONFIDENCE
LATER
Internal mobility
CONFIDENCE

WORKED EXAMPLE (HR)

You notice new managers quit far more than average, and exit interviews say they felt "thrown in at the deep end." Your belief: the first year is unsupported. Your bet: a structured first-year support track. Success looks like: fewer regretted manager exits. Budget: one quarter, one small team. Stop line: fewer than 30% finish it, or the team's mood hasn't shifted by week eight. You run it, check the evidence at the quarter review, and either keep going, tweak it, or stop — without guilt either way. If it didn't work despite a sound decision, that's a lesson, not a black mark.

Pitfalls & how to run it well

Expecting every bet to win (then you only ever pick safe ones and never learn anything new); not setting a budget (the project drags on forever); not setting a stop line (nobody's willing to end it); and judging a bet only by whether it won rather than whether it was a smart call. Run it well by writing each bet on one page, setting the budget and stop line before you start, keeping the board visible, and treating a well-reasoned bet that lost as a lesson.

Connects to

Your beliefs come from Discovery and Assumption mapping; your placed bets become the themes on the Roadmap; you check them at the Exit gates using Evidence rating and Kill criteria.

2

Roadmapping

Translate strategy into a living, outcome-framed plan of what to learn and deliver next — never a fixed schedule of features.

What it is

A roadmap is the bridge between the Strategy hub and the actual work. It takes the year's few bets and the North Star and turns them into a rolling ~6-month plan expressed as outcomes and themes — problems to solve — rather than a locked list of features with delivery dates. That distinction is the whole point, and it's the place product and project cultures most visibly collide. A project plan commits to outputs ("ship A in Q1, B in Q2") and treats any deviation as failure. A product roadmap commits to outcomes ("cut time-to-productivity this half") and treats new evidence as a reason to re-aim — because the team doesn't yet know exactly which features will move the outcome. The roadmap is "living" because every exit gate (Stake) returns to Strategy and updates it; a roadmap that hasn't changed in six months is a sign the team stopped learning, not a sign of discipline.

When to use it

Strategy (Phase 1), after you've cascaded company strategy into your function's strategy and named the North Star and the year's bets. Re-touch it at every cycle as gates feed learnings back. Don't use a roadmap as a delivery contract or a Gantt chart of features — that's project thinking wearing a product label.

Process

  1. Anchor on the North Star and the year's bets — the roadmap exists to serve them, not to list everything anyone wants.
  2. Express every item as an outcome or a problem to solve, not a solution ("improve onboarding completion," not "build an onboarding portal").
  3. Group items into a few themes over a ~6-month horizon; sequence by priority and dependency, not by how confident you are of a delivery date.
  4. Attach a success measure to each theme (the outcome metric you'll judge it by), and ideally the kill criteria that would stop it.
  5. Make it visible to the whole team and key stakeholders — transparency is what turns a roadmap into shared commitment rather than a private plan.
  6. Stress-test it against viability — does leadership see the ROI?
  7. Revisit every cycle. As gates return learnings to Strategy, re-sequence, drop, or add. Expect it to change; that responsiveness is the feature.

Diagram

NORTH STAR · time-to-productivity ↓
MORE CERTAINLESS CERTAIN →
NOW
Onboarding
→ productive in 5 days
NEXT
Manager enablement
→ fewer escalations
LATER
Internal mobility
→ mobility rate ↑
✗ PROJECT PLAN
features + dates
✓ PRODUCT ROADMAP
outcomes + themes

WORKED EXAMPLE (HR)

North Star = internal mobility rate. Now: "Make open roles visible and findable" (outcome: % of employees who can name two current internal openings). Next: "De-risk the lateral move" (outcome: internal applications per role). Later: "Manager incentives for releasing talent." None of these is a feature; each carries the measure that tells you it worked.

Pitfalls & how to run it well

The dominant failure is listing projects-with-dates and calling it a roadmap. Others: over-specifying the Later column (it's the most uncertain — keep it loose); attaching no success measure to a theme; and never revisiting it. Run it well by keeping it visible, reviewing it at every Strategy return, and cutting ruthlessly to protect focus — if everything is on the roadmap, nothing is.

Connects to

Upstream: the Strategy cascade and North Star definition feed it; Prioritization decides what's Now versus Later. Downstream: each theme enters Discovery as a framed problem, and exit-gate learnings flow back to revise it.

DISCOVERY

3

Stakeholder Mapping (the circles / RACI)

Sort everyone connected to the work by their real relationship to the decision, so each person is involved correctly — no more, no less.

What it is

Stakeholder mapping places everyone touched by or touching an initiative into rings by their actual role in the decision. The circles model uses five: Core (the working team that decides and builds), Contributing (part-time SMEs, HRIS, IT), Consulted (e.g. managers, finance, legal — they give input, not control), Informed (kept in the loop), and Users (both B2C end users and B2B internal clients). It maps directly onto RACI — "consulted" and "informed" are literally RACI's C and I — and the underlying job is the same: separate who decides, who advises, and who is merely told. It's also where you name the product trio + sponsor explicitly inside the core: a problem owner (often the HRBP), a user advocate, a build/tech lead (People Ops / HRIS), and a sponsor for vision, resources, and air cover.

When to use it

At Discovery kickoff, before you plan research, and again whenever scope shifts. The cost of skipping it is role confusion that surfaces later as stalled decisions and blindsided stakeholders.

Process

  1. List everyone the initiative touches or who could touch it.
  2. Place each into a ring: core / contributing / consulted / informed / users.
  3. Make decision rights explicit — core decides, consulted advises, informed is told. Write it down so a "consulted" voice doesn't quietly behave like an approver.
  4. Name the trio + sponsor within the core.
  5. Make sure users are represented on both sides (B2C and B2B), not just executives.
  6. Agree how and how often each ring is engaged.

Diagram

Informed · told
Consulted · advises
Contributing
◎ CORE
trio + sponsor · decides
Users span all rings · B2C employees + B2B HRBPs / leaders

WORKED EXAMPLE (HR — BUDDY ONBOARDING)

Core: HRBP (problem owner), People Ops (build), HR/EX designer; sponsor = People director. Contributing: HRIS, IT. Consulted: legal (data), finance (budget), line managers. Informed: senior leadership. Users: new hires (B2C), managers and HRBPs (B2B).

Pitfalls & how to run it well

Putting everyone in "core" so every decision needs everyone (paralysis); forgetting users are stakeholders (you design for execs and lose adoption); letting a consulted stakeholder act as an approver. Run it well by keeping the center small, naming decision rights out loud, and re-mapping when scope changes.

Connects to

Feeds Dual discovery (it tells you who to research on each side) and clarifies the trio. It's the social map that the rest of Discovery runs on.

4

5W1H

A fast, structured interrogation that proves you actually understand a problem before you try to solve it.

What it is

5W1H forces you to answer, for any problem: Who has it (which segments, personas), Where / When does it occur (the moment of friction), Why does it happen (the root cause, not the symptom), and How will we know it's solved (the success metric, defined now). It's the lightweight front end to problem framing — the answers become the raw material for the problem statement. Two of the questions do most of the work. "Why," pushed several layers deep, guards against the most common product error: fixing a symptom while the cause persists. "How will we know it's solved" forces you to commit to an outcome measure before you fall in love with a solution, so you can later tell whether you actually changed anything.

When to use it

Early in Discovery, the moment a problem lands on the table — especially when it arrives pre-solutioned ("we need an X"). Pairs immediately with problem framing.

Process

  1. Who: name the specific segment or persona that feels it — not "employees" in general.
  2. Where / when: pin the moment of friction (which step, which context). Problems are rarely universal.
  3. Why: ask "why" repeatedly (the five-whys move) to get past the symptom to a root cause — which is itself a hypothesis to check, not a fact.
  4. How will we know: define the metric or behaviour that would tell you it's solved, in outcome terms ("productive in five days"), not completion terms ("form shipped").
  5. If you can't answer all of them, stop — you're not ready to design; you're ready to do more discovery.

Diagram

PROBLEM
Who?segments / personas
Where / When?moment of friction
Why?root cause (5 whys)
How will we know?outcome metric
★ carry the most weight

WORKED EXAMPLE (HR)

Problem as handed over: "managers don't give feedback." 5W1H: Who — people-managers of remote teams. Where/when — between formal review cycles, in distributed settings. Why (five whys) — no quick lightweight tool, plus fear of awkwardness, plus no established habit. How we'll know — share of 1:1s with documented feedback, and a pulse score on "I get useful feedback from my manager."

Pitfalls & how to run it well

Stopping at the first "why" (you fix a symptom); skipping "how will we know" (you ship with no way to judge success); treating it as bureaucracy rather than a ten-minute sharpening tool. Run it fast and out loud, and let the "why" get uncomfortable.

Connects to

Feeds Problem framing directly (its answers populate the problem statement) and Assumption mapping (the root-cause "why" is usually an assumption to test). Upstream of everything in Discovery.

5

Personas + JTBD

Turn raw research into a few distinct, research-grounded portraits of who you serve and the progress they're trying to make — in two layers, B2C and B2B.

What it is

A persona is a visualization of a segment — a composite, research-based portrait of one distinct type of user, built so the team designs for a specific person instead of a faceless "user." The operative word is distinct: a persona is only worth having if it's different enough to justify its existence. Personas come out of segmentation — dividing your users by a dimension that genuinely changes how you'd serve them (behaviour, needs, the job they're doing, their context) — not by demographics for their own sake. The test of a real persona is decision-relevance: would you make different product decisions for this persona than for that one? If two "personas" would lead to the same design, they're one segment wearing two faces — collapse them. A useful set is small (2–4) and sharply different.

Jobs-to-Be-Done (JTBD) captures the progress a persona is trying to make — the "job" they hire a solution to do — across three dimensions: functional (the practical task), emotional (how they want to feel), and social (how they want to be seen). JTBD is solution-agnostic and stable: features and demographics churn, but the job ("get up to speed in a new role") holds, which is what makes it a reliable target. It also has a direct line into delivery — a JTBD translates almost one-to-one into a user story, the format engineers use to describe backlog items, which is how a product-thinking HR team starts speaking the same language as its tech and product partners.

For internal functions you keep two layers (the B2C / B2B dual lens): end-user personas (employees, managers, new hires) and internal-stakeholder personas (HRBPs, leaders), each a distinct segment with its own jobs. Success means both — user adoption and operational alignment.

When to use it

Discovery, after the problem is framed and you have real research, to ground design in specific people. Always built from research (interviews, observation, data) — an invented persona just launders assumptions and is worse than none.

Process

  1. Segment first. Look across your research for the dimension that splits users into groups who need different things — behaviour, needs, jobs, context (demographics rarely do this alone).
  2. Test distinctness. For each candidate segment ask "would we design differently for this group?" If not, merge it. Keep the 2–4 that pass.
  3. Build the persona anatomy for each (below), grounded in what real people said and did — include a verbatim quote.
  4. Write the JTBD across functional, emotional, and social dimensions.
  5. Repeat for the stakeholder layer (B2B) — same anatomy, but the needs and jobs are governance and operational.
  6. Translate each JTBD into a user story so the work is ready to hand to delivery (see the bridge below).
  7. Name them, keep them few, and revisit as research deepens.

Persona anatomy

What a good persona contains:

  • Segment & distinctness — the dimension this persona represents, and why it's different enough to need its own card.
  • Demographics — role, tenure, team, location, and the few demographic facts that are actually relevant (not decoration).
  • Context / environment — where and how they work (remote, deskless, on the floor), which shapes everything else.
  • Needs — what they require to succeed.
  • Problems & frictions — what gets in their way today; the pain points (ties straight to the journey map).
  • Motivations — what drives them; why they'd care enough to change behaviour.
  • Habits & rituals — their routines and recurring behaviours; a solution has to fit these or it won't be adopted.
  • Current process / workarounds — how they solve this today, including the tools and hacks they already use. You're rarely competing with nothing — you're competing with the spreadsheet, the DM to a friend, the "I just wing it."
  • JTBD — functional / emotional / social (below).
  • Success metric — what "this worked" looks like for them.
  • Triggers — the event that kicks off the job (the "when" of the JTBD): first day, a reorg, a failed review.
  • Desired outcome — the gain they're after, in their words (the positive mirror of problems & frictions).
  • Influences & who they trust — whose opinion moves them; often decisive for B2B stakeholders.
  • Tech comfort — adoption-relevant: a low-comfort persona will quietly reject a clever-but-fiddly solution.
  • A verbatim quote — one real line from research; it anchors the persona in evidence and is the fastest tell of a grounded vs invented persona.
  • Frequency / volume — how often they hit this and how many of them exist (how much the segment matters).

JTBD — functional, emotional, social

  • Functional — the practical task: "get up to speed on my role and start contributing."
  • Emotional — how they want to feel: "feel confident and settled, not lost or anxious."
  • Social — how they want to be perceived: "be seen as competent and a good hire by my new team." A complete JTBD names all three, because solutions that nail the functional job but ignore the emotional and social ones get quietly rejected.

The bridge to delivery: a JTBD maps almost directly onto a user story — the one-line format engineers use to describe a backlog item.

  • JTBD: "When I start a new role, I want to know what's expected of me, so I can feel confident."
  • User story: "As a new hire, I want a clear week-one plan, so that I feel confident about what's expected." The "so that…" clause is the outcome, which keeps the backlog item tied to value rather than to a feature. This translation is how a product-thinking HR team hands work to product and engineering in the language they already use — instead of throwing a finished solution over the wall.

Diagram

All usersRemote graduate joinersExperienced lateralsHRBPs at scaleSenior leaders
N
B2C · New-hire Nina
Graduate analyst, week 1
segment: remote graduate joiners
Demographics: first corporate role, distributed team
Context: fully remote
Needs: clarity · someone to ask · early wins
Frictions: laptop late · “who do I ask?” · vague goals
Motivations: prove herself, belong
Habits: Slack first thing, to-do lists
Solves it today: DMs a friend, guesses, googles
I spent week one afraid to ask anything.
JTBD
Functional: get productive fast
Emotional: feel confident, not lost
Social: be seen as a strong hire
Success: knows who to ask + first task done by day 5
H
B2B · HRBP Henry
HR business partner
segment: HRBPs onboarding 20+/quarter
Demographics: HRBP, multiple teams
Context: many concurrent onboardings
Needs: consistency · fewer escalations
Frictions: uneven manager follow-through · compliance pressure
Motivations: a tight, trusted function
Habits: dashboards, checklists, weekly reviews
Solves it today: spreadsheets + chasing managers
I only find out it broke when a manager escalates.
JTBD
Functional: onboard consistently at scale
Emotional: trust nothing's slipping
Social: be seen as running a tight function
Success: zero week-one escalations across the cohort
JTBD → USER STORY
“When I start, I want to know what’s expected, so I can feel confident.”
“As a new hire, I want a clear week-one plan, so that I feel confident.”

WORKED EXAMPLE (HR)

New-hire Nina (B2C) — segment: remote graduate joiners, distinct from experienced lateral hires (who need a network, not orientation). Frictions: late laptop, no one to ask, unclear goals. Habits: Slack-first, list-maker. Solves it today by DMing a friend and guessing. JTBD — functional: get productive fast; emotional: feel confident, not lost; social: be seen as a strong hire. User story: "As a new hire, I want a clear week-one plan, so that I feel confident about what's expected." HRBP Henry (B2B) — segment: HRBPs onboarding 20+/quarter, distinct from senior leaders (who only want the dashboard). Needs: consistency, fewer escalations. JTBD — functional: onboard consistently at scale; emotional: trust nothing's slipping; social: be seen as running a tight function. The two jobs interact: Nina's confidence depends on Henry's consistency, so a solution has to serve both or it stalls.

Pitfalls & how to run it well

The big one: personas that aren't distinct — if two would drive the same decisions, you've got one segment and wasted effort, so collapse them. Others: demographics-only personas (age and location rarely change the design — behaviour and jobs do); inventing personas instead of grounding them in research; too many personas (4 is plenty); and writing JTBD as features ("wants an app") rather than progress ("wants to feel confident"). Run it well by segmenting on what actually changes the design, basing every line on real research, and naming all three JTBD dimensions.

Connects to

Built from Dual discovery research; feeds Journey mapping (whose journey?), Assumption mapping (each persona carries user assumptions to test), and the opportunities in the Opportunity-Solution Tree. Downstream, the JTBD → user story translation is the handoff into delivery backlogs — the shared language with product and engineering.

6

Journey Mapping

Lay out the user's end-to-end experience so you can see exactly where it breaks — functionally and emotionally.

What it is

A journey map visualises a user's experience of a process as a sequence of steps, and for each step captures four things: what the person does, the tools and touchpoints they use, the pain points they hit, and the emotion they feel — usually split across before / during / after. It turns a vague sense that "onboarding is rough" into a specific map of which moments fail and why. It builds directly on a persona (whose journey is this?) and is grounded in dual discovery research, not imagination. The layer teams skip is emotion, and it's the one that most often matters: a step that's functionally fine but emotionally awful — the anxious first day, the silent wait for an approval — is frequently where the real problem lives, and a purely functional map misses it. The output is a ranked set of moments of friction — the breakpoints worth solving.

When to use it

Discovery, after you have personas and some real research, to locate the highest-leverage moments to attack. Don't map an idealised "designed" journey; map the lived one.

Process

  1. Pick one persona and bound the scope — which journey, from where to where.
  2. Lay out the steps in sequence across before / during / after.
  3. For each step, capture actions, touchpoints/tools, pain points, and emotion.
  4. Draw the emotion as a line across the steps; the dips are your breakpoints.
  5. Ground every step in something a real person reported — not the official process diagram.
  6. Use the breakpoints to generate opportunities (this feeds ideation and the Opportunity-Solution Tree).

Diagram

BEFORE
DURING
AFTER
Step
offer
pre-board
day 1
week 1
month 1
Action
sign
setup
arrive
find footing
ramp up
Touchpoint
email
portal
laptop
Slack / mgr
reviews
Pain
forms
nothing
"who do I ask?"
unclear goals
Emotion
breakpoint
breakpoint

WORKED EXAMPLE (HR — ONBOARDING)

Steps: offer → pre-boarding → day 1 → week 1 → month 1. Emotion dips hard at day 1 (laptop not provisioned, no clear plan) and week 1 ("I don't know who to ask"). Those two dips are the problems worth solving — and notably, the day-1 dip is owned by IT, not HR, which the map makes impossible to ignore.

Pitfalls & how to run it well

Drawing the process as designed rather than as lived; dropping the emotion layer; mapping too broad a scope so nothing is specific. Run it well with one persona and a bounded scope, built from real stories, and pick one synthesis tool per project rather than over-documenting.

Connects to

Built on Personas + JTBD and Dual discovery. Its breakpoints feed the Opportunity-Solution Tree and surface user assumptions to map. Often the single most persuasive artifact for stakeholders, because the emotional dips are visceral.

7

Problem Framing

Define the problem precisely and testably before anyone proposes a solution — the highest-leverage, most-skipped step in product work.

What it is

Problem framing turns raw, messy signals — complaints, requests, dashboards, anecdotes, the answers from 5W1H — into a single clear, testable problem, separating symptom from root cause and bounding the scope. Its output is a problem statement, written to a template: "[User group] struggles with [problem] when [context] because [reason], which leads to [impact]." The core skill is resisting the pull toward solutions: requests almost always arrive pre-solutioned ("we need a new LMS"), which smuggles in an unexamined assumption about both the problem and its fix. A well-framed problem aligns the team on one target, makes discovery measurable, and keeps later ideation focused — which is why getting it right saves months that would otherwise be spent building the wrong thing beautifully.

When to use it

Discovery, right after 5W1H exploration. The resulting problem statement is exactly what the Discovery exit gate checks ("a framed problem the team agrees on, one sentence").

Process

  1. Gather the raw signals and the 5W1H answers.
  2. Strip the pre-loaded solution out of the request — turn "we need X" back into "what is X supposed to fix?"
  3. Separate symptom from root cause (the "why" from 5W1H).
  4. Bound the scope — a specific user group, context, and moment, not "everyone, everywhere, always."
  5. Write the statement using the template.
  6. Get the team to literally agree on the one sentence — agreement is part of the deliverable.

Diagram

raw request — “we need an LMS”
▼ strip solution, find root cause
framed problem
[user group]
[problem]
[context]
[reason]
[impact]

WORKED EXAMPLE (HR)

"Managers of remote teams struggle to give continuous feedback when they're between formal review cycles, because they lack a quick, lightweight tool and fear the conversation being awkward, which leads to missed development and unpleasant surprises at review time." Every slot is filled, it's bounded, and the "reason" is a hypothesis discovery can test.

Pitfalls & how to run it well

Framing the solution as the problem ("the problem is we don't have an LMS"); leaving it unbounded and universal; ending with three competing statements and no agreement. Run it well by forcing exactly one sentence, keeping the reason falsifiable, and re-framing without ego if discovery proves the first cut wrong.

Connects to

Built from 5W1H; the statement anchors Assumption mapping, the Opportunity-Solution Tree, and the entire solution search. It's the gate condition for leaving Discovery.

8

Assumption Mapping

Surface the beliefs your idea depends on, then decide what to test first.

What it is

Every plan rests on assumptions — beliefs treated as true without proof, "invisible decisions waiting to be tested." Assumption mapping pulls them into the open and plots each on two axes — importance (how badly the idea breaks if it's false) and certainty (how much evidence you already have) — so the team aims its scarce research at the beliefs that are both load-bearing and unproven. Assumptions come in four types worth forcing: user, problem, solution, and viability — and the riskiest are usually the viability ones ("leadership will fund this," "managers will allow the time"). The map is the bridge from believing to knowing: the dangerous assumptions become hypotheses, each handed to an experiment that designs the cheapest test that could disprove it. It also connects to the gate mechanics — a top-left assumption that fails becomes a kill criterion, and how confidently you can move it depends on evidence rating.

When to use it

Discovery, once the problem is framed and you have a candidate direction, before committing real research or build effort. Re-map as evidence arrives and new assumptions appear; it's a living artifact. Skip it only when the work is genuinely known and low-risk — then it's a project, not a discovery.

Process

  1. Brain-dump every assumption with the team, forced across the four types: user, problem, solution, viability. Phrase each as "We believe that…"
  2. Make each atomic and falsifiable — split compound beliefs, reword vague ones so they could be proven wrong.
  3. Rate each on importance and certainty.
  4. Plot them on the 2×2.
  5. Act by quadrant: high-importance + low-certainty → test first; high-importance + high-certainty → confirm cheaply with evidence you have; low-importance → park or ignore.
  6. Turn the top-left assumptions into hypotheses and hand them to Experiments.
  7. Re-map as you learn — tested assumptions move right, new ones appear.

Diagram

Low certainty
High certainty
High importance
Test first
high stakes, low evidence
Confirm cheaply
high stakes, have evidence
Low importance
Park
low stakes, low evidence
Ignore
low stakes, safe

WORKED EXAMPLE (HR — BUDDY ONBOARDING)

"New hires would actually engage an assigned buddy" → high importance, low certainty → test first. "Managers will free up buddies' time" → high importance, low certainty (viability) → test first. "New hires feel lost in week one" → high importance, medium certainty → confirm with existing survey/exit data. "We can match buddies sensibly" → medium importance → schedule after the leaps. "A Slack channel is the right format" → low importance (solution detail) → park. The two beliefs that could sink the idea are exactly the ones with the least evidence.

Pitfalls & how to run it well

Mapping only solution assumptions and skipping viability (usually the real killers); rating importance by how much you like the assumption rather than how much the idea depends on it; treating the map as a one-time artifact. Run it as a team (diverse views expose hidden assumptions), time-box 30–45 minutes, and force-rank if everything lands "high importance" — the point is to choose.

Connects to

Upstream: Problem framing, 5W1H, Personas/JTBD, and Journey mapping all surface assumptions. Downstream: each top-left assumption becomes a hypothesis for Experiments; the risk-ranked map is what the Discovery exit gate checks.

9

Dual Discovery

Research both sides of a two-sided product at once — B2C needs and B2B constraints — and reconcile them before designing.

What it is

Dual discovery is the operational execution of the B2C / B2B dual lens. You run end-user research (needs, friction, motivation) and stakeholder research (constraints, goals, success criteria) in parallel, then reconcile what the two sides need before committing to a design. It's a named discipline rather than just "do research" because teams reliably study only one side — usually the easier or louder one — and each omission has a predictable cost: skip the B2C side and you design something stakeholders approve but employees ignore; skip the B2B side and you design something employees love but leadership blocks or can't sustain. It draws on the full kit of research methods (in-depth interviews, observation, surveys, data analysis), and it's where the bias trap and the say-vs-do gap bite hardest — so you ask about real past behaviour, not hypotheticals, and recruit beyond friendly faces.

When to use it

Discovery, as the research engine, planned directly off the stakeholder map and aimed at the assumptions you most need to test. The Discovery exit gate explicitly checks that primary (not desk) research was done and that both sides were covered.

Process

  1. From the stakeholder map, identify who to talk to on each side (B2C users and B2B stakeholders).
  2. Match method to question: interviews for motivation and story, observation for actual behaviour, surveys for scale of a known issue, data for patterns.
  3. Run B2C research — employee/manager needs, pains, motivations — asking about real past behaviour and watching for bias.
  4. Run B2B research — stakeholder goals, constraints, success criteria, and what they'd need to approve and sustain the thing.
  5. Reconcile: where do the two sides' jobs align, and where do they conflict? Resolve the conflict before designing.
  6. Synthesise into personas, JTBD lines, the journey map, and the assumption map.

Diagram

B2C
interviews / observation
needs, friction
B2B
interviews / data
constraints, goals
Reconcile
shared insights → personas, journey, assumptions

WORKED EXAMPLE (HR)

B2C interviews with new hires surface "I didn't know who to ask in week one." B2B interviews with managers surface "I don't have time to onboard properly, and there's compliance pressure." Reconciled: a buddy system serves both Nina's need and Henry's consistency — but only if manager time is genuinely freed, which is the viability assumption to test before building.

Pitfalls & how to run it well

Researching only the easy or loud side; leading questions and friendly-faces selection bias; asking "would you use this?" instead of "tell me about the last time…"; gathering both sides but never reconciling them. Run it well by pairing every piece of user research with stakeholder research, splitting interviewer/note-taker roles, and recruiting deliberately diverse participants.

Connects to

Planned off Stakeholder mapping; feeds Personas + JTBD, Journey mapping, and Assumption mapping. It's the evidence engine that turns the Discovery assumptions from guesses into rankings.

SHAPE

10

Prioritization

Decide what to pursue first, because not every validated problem deserves resources now.

What it is

Prioritization is the discipline of choosing — making the trade-offs explicit and comparative rather than defaulting to whoever shouted loudest or whatever's newest. Its governing line is focus: "if everything is priority one, nothing is," and the hardest part of any strategy is what it refuses to do. Two lightweight tools do most of the work. The Impact/Effort matrix plots each candidate by value created against effort to deliver, instantly separating quick wins from long bets, fill-ins, and the money pit. ICE scoring rates each on Impact × Confidence × Ease, where the confidence term is the quiet genius — it docks ideas you believe in but have no evidence for, pushing you to either test them or drop them. Prioritization treats the work as a portfolio of bets, not a wish list.

When to use it

Shape, once Discovery has validated the problem and you have candidate opportunities or directions. It's how you pick the bet to de-risk and build. Use it again to sequence the roadmap (Now / Next / Later).

Process

  1. List the candidate opportunities or solutions (ideally from the Opportunity-Solution Tree, so you're comparing opportunities, not pet ideas).
  2. Score them — Impact/Effort for a quick read, ICE when confidence varies a lot.
  3. Be honest about the confidence term: an idea you love but can't evidence scores low until tested.
  4. Pick one defensible direction and be able to say why this, not that.
  5. Commit — one direction carried properly beats three half-bets.

Diagram

Low effort
High effort
High impact
Quick wins
Big bets
Low impact
Fill-ins
Money pit
ICE
I
C
E
Score
Buddy system
8
4
6
192
New LMS
6
3
2
36

WORKED EXAMPLE (HR)

Candidates to lift internal mobility: "internal job board" (high impact, low effort → quick win), "manager incentives to release talent" (high impact, high effort → big bet), "skills passport" (high impact, low confidence → test before committing). ICE pushes the skills passport down until a cheap test raises its confidence.

Pitfalls & how to run it well

Prioritising by loudness or recency; false precision in the scores (they're for sorting, not accounting); and producing a ranking but never actually cutting anything. Run it well by force-ranking, defending "this not that" out loud, and protecting focus.

Connects to

Takes candidates from the Opportunity-Solution Tree and ideation; the chosen bet then goes through Four-risk assessment and Experiments; the sequence feeds the Roadmap.

11

Opportunity-Solution Tree

Connect a desired outcome to the opportunities that could move it, the solutions for each, and the tests — so ideas stay tethered to the goal.

What it is

An opportunity-solution tree is a structure with four layers, top to bottom: a desired outcome at the root (laddered up to the North Star); the opportunities beneath it (the user needs, pains, and desires from Discovery that, if addressed, would move the outcome); the solutions under each opportunity; and the experiments under each solution. Read downward it says "to move this outcome, we could pursue these opportunities, for which we have these solution ideas, which we'll test like this." Its purpose is to prevent the most common failure in product work — jumping straight from a goal to a favourite solution, skipping the opportunity space entirely. By forcing every solution to hang off an explicit opportunity, it makes prioritization happen at the opportunity level (compare problems worth solving, not pet ideas), exposes tunnel vision, and makes the team's reasoning legible to stakeholders.

When to use it

Shape — it's the bridge from Discovery's insights to the solution search. Build it from journey breakpoints, JTBD, and the insights discovery produced.

Process

  1. Set the root: the outcome you're trying to move, tied to the North Star.
  2. Branch into opportunities — the distinct needs/pains from discovery that could move that outcome. Aim for a few real ones, not one.
  3. Under each opportunity, list candidate solutions (this is where ideation lives).
  4. Under each solution, name the experiment that would test it.
  5. Prioritise at the opportunity level first — which problem is most worth solving — then pick solutions within it.
  6. Prune relentlessly; a sprawling tree is a sign of avoiding the choice.

Diagram

Outcome: time-to-productivity ↓don’t know expectationsno one to askunclear goalsSol: buddySol: FAQ botExp: conciergeExp: pretotypeoutcome → opportunities → solutions → experiments

WORKED EXAMPLE (HR)

Outcome: new hires productive in five days. Opportunities (from the journey map's dips): "don't know what's expected," "no one to ask in week one," "unclear early goals." Under "no one to ask": solutions = buddy system, FAQ bot, office hours; experiments = a concierge buddy pilot, a fake-door for the bot. The tree makes clear you're choosing which problem first, not which toy.

Pitfalls & how to run it well

Starting from a beloved solution and reverse-engineering an opportunity to justify it; growing only one branch (tunnel vision); letting the tree sprawl into documentation. Run it well by prioritising opportunities before solutions and keeping it pruned to what you'll actually pursue.

Connects to

Built from Discovery insights (journey breakpoints, JTBD); feeds Prioritization (compare opportunities) and Experiments (the leaf nodes are tests). Sits at the seam between understanding the problem and choosing the solution.

12

6-3-5 Brainwriting

Six people silently write three ideas in five minutes, then pass the sheet — six rounds later you have up to 108 builds, and nobody got talked over.

What it is

6-3-5 brainwriting is structured, silent ideation: 6 participants each write 3 ideas in 5 minutes on a worksheet, then pass it to the neighbour, who reads what's there and adds three more — new ideas, variations, or builds on the rows above. After six rounds (~30 minutes) every sheet is full: up to 108 idea slots. Published by Bernd Rohrbach in 1968 (in the German magazine Absatzwirtschaft), it fixes the classic failures of shout-it-out brainstorming: dominant voices setting the agenda, anchoring on the first idea said aloud, and production blocking (only one person can talk at a time — but six can write at once). Writing is parallel, silent, and leaves a paper trail of who built on what. It's the best-known member of the brainwriting family, and the same write-first discipline behind note-and-vote facilitation.

When to use it

Shape — at the ideation seam, when the Opportunity-Solution Tree has named the opportunity and you need to fill its solutions layer with real options instead of the first pet idea. Reach for it any time the room needs volume and variety fast, or when the group has strong personalities and quiet experts. Remote works fine: a shared board where each "sheet" is a column and rotation is moving one column right.

Process

  1. Sharpen one prompt — a single "How might we…" tied to the chosen opportunity — and give everyone a worksheet of 3 columns × 6 rows.
  2. Round one: each person writes 3 ideas in their top row, silently, in 5 minutes.
  3. Pass the sheet one seat to the right. Read the rows above before writing — then add 3 more: new ideas, variations, or direct builds.
  4. Repeat until the sheets are full — 6 rounds, about 30 minutes, up to 108 slots.
  5. Cluster duplicates and near-duplicates, then converge: dot-vote or ICE-score the strongest candidates.
  6. Hang the survivors on the tree as solutions and send them into Prioritization.

Diagram

6-3-5 · SILENT BRAINWRITING
One worksheet — 3 ideas × 6 rounds
Idea 1
Idea 2
Idea 3
R1
buddy rota
FAQ bot
office hours
R2
↑ + build
new idea
↑ + build
R3
R6
Read the rows above, then add three — new, variation, or build.
Sheets rotate every 5 minutes
P1
P2
P3
P6
P5
P4
6 people × 3 ideas × 5 min6 rounds ≈ 30 min → 108 idea slots

WORKED EXAMPLE (HR)

Prompt: "How might we make sure a new hire has someone to ask in week one?" — the opportunity chosen on the Opportunity-Solution Tree. Six people, six silent rounds: buddy-of-the-day rota, FAQ bot trained on onboarding tickets, protected "office hours" block, manager check-in script, #ask-anything channel with an answer-SLA, first-week shadowing plan — plus builds, like pairing the rota with the office-hours block. The team clusters, dot-votes, and carries three candidates into ICE scoring.

Pitfalls & how to run it well

A vague prompt produces 108 vague slots — sharpen the "how might we" first. People stalling at one idea per round — push for quantity; the odd rows are where the builds come from. Skipping the read first step, which turns it into six parallel monologues and kills the building effect. And treating the number as the result — 108 slots are raw material for clustering, not 108 good ideas. Run it in true silence with an honest, visible 5-minute timer.

Connects to

Fills the solutions layer of the Opportunity-Solution Tree; its survivors flow into Prioritization (ICE) and on to Experiments. Kin to note-and-vote and silent-brainstorm facilitation — same write-first, speak-later muscle.

13

Four-Risk Assessment

Before building, check the chosen solution against the four risks — value, usability, feasibility, viability — because each threatens success differently and each needs a different test.

What it is

Every solution must survive four risks, assessed before the expensive build: Value (will people choose to use it?), Usability (can they figure out how to use it?), Feasibility (can we build it with what we have?), and Viability (does it work for the business — financially, legally, operationally, for the brand?). This is the same lens as desirability / feasibility / viability, broken one finer (desirability splits into value + usability). The framework's payoff is diagnostic: instead of vaguely "validating," you name which risk is largest and aim the right test at it. It maps cleanly onto the product trio — the PM carries value and viability, the designer carries usability and desirability, the tech lead carries feasibility — and onto the experiment types — value → pretotype, usability → prototype, feasibility → POC, viability → business/legal case. One thing teams routinely get wrong: legal belongs under viability, not feasibility (feasibility is "can we build it," viability is "are we allowed to and should we").

When to use it

Shape, on the chosen bet, before committing to build. It tells you what to test first.

Process

  1. State the solution clearly.
  2. Ask the four questions and rate where the risk is highest — honestly, not where it's easiest to test.
  3. Map each high risk to the cheapest test that attacks it (value → pretotype/concierge; usability → prototype; feasibility → POC; viability → business case / legal review).
  4. Test the biggest risk first.
  5. Assess it as the trio — each member surfaces the risk they own, together, not in sequence.

Diagram

VALUE
Will they choose it?
Test: pretotype · Owner: PM
USABILITY
Can they use it?
Test: prototype · Owner: designer
FEASIBILITY
Can we build it?
Test: POC · Owner: tech lead
VIABILITY
Works for the biz? (incl. legal)
Test: business case · Owner: PM

WORKED EXAMPLE (HR — BUDDY ONBOARDING)

Value (will new hires engage a buddy?) — uncertain, high. Usability (is the format easy?) — low risk. Feasibility (can we run matching?) — low. Viability (will managers give buddies the time; any policy issues?) — high. So the team tests value and viability first (a concierge buddy pilot + a manager-time agreement), and does not spend a sprint polishing the usability of something whose demand and viability are unproven.

Pitfalls & how to run it well

Polishing usability on something nobody wants (value unaddressed); proving feasibility of something the business can't sustain or isn't allowed to do (viability unaddressed); filing legal under feasibility. Run it well by naming the single biggest risk honestly and testing that, not the one that's most comfortable to test.

Connects to

Applied to the bet chosen in Prioritization / OST; routes each high risk to an Experiment type; mirrors the product trio. The Shape exit gate checks that all four were assessed.

14

Experiments

Run the smallest, cheapest test that could disprove your riskiest assumption — build to learn, not to launch.

What it is

"Experiments" is the family of build-to-learn tests, and the parent of the specific types that follow (pretotype, prototype, POC, MVP, A/B, concierge, Wizard-of-Oz). The discipline is constant regardless of type: take the assumption that's highest-importance and lowest-certainty, write it as a falsifiable hypothesis, pick the cheapest test that could disprove it, and define what counts as pass or fail before you run it. Two choices set the test: which of the four risks you're attacking, and how real the thing needs to be to learn (the build-to-learn ladder runs lightest to most real: pretotype → prototype → POC → MVP). A one-page Lean Experiment Canvas captures hypothesis, what you'll learn, the test, the success metric, and the next step if it works or doesn't — which is what stops "experiments" from being unfalsifiable demos. The mindset is "build to learn, not to launch": a test that disproves a hypothesis cheaply has succeeded.

When to use it

Shape (and on into Build), to move top-left assumptions from hope to evidence and to satisfy the gate that the riskiest assumption has been moved.

Process

  1. Take the riskiest assumption from the assumption map.
  2. Write it as a hypothesis: "We believe [change] will produce [measurable result] for [user]."
  3. Identify which risk it's really about (value / usability / feasibility / viability).
  4. Pick the cheapest test that fits that risk and the realness needed (see the ladder below).
  5. Fill the Lean Experiment Canvas — define the success threshold and the if-works / if-not next steps up front.
  6. Run it small, on real users where possible.
  7. Decide honestly: proceed, iterate, or kill against the pre-set threshold.

Diagram

most real / costly
MVPvalue in real use
POCfeasibility
Prototypeusability / desirability
Pretotypedemand — does anyone want it?
lightest / cheapest
Situational: A/B (which version) · Concierge (manual, open) · Wizard-of-Oz (manual, hidden)

WORKED EXAMPLE (HR)

Riskiest assumption: "new hires will engage a buddy." Hypothesis: "If we assign buddies to 10 new hires for two weeks, ≥70% will report knowing who to ask for help." Risk: value. Test: a concierge buddy pilot (manual matching). Canvas: success = ≥70% by end of week two; if met → scale, if not → redesign the format. Nothing is built until that pass/fail is known.

Pitfalls & how to run it well

Building to launch instead of to learn; running a "test" with no pre-defined success threshold (so any result can be spun); picking the test you find fun rather than the one that fits the risk; testing the easy risk rather than the biggest. Run it well with one canvas per experiment, the success metric and next steps defined before you start, and the smallest version that can still teach you something.

Connects to

Fed by Assumption mapping (the hypotheses) and Four-risk assessment (which risk to attack); each type is detailed below; results feed the Shape and Build exit gates and trigger kill criteria when a threshold is missed.

15

Pretotype (fake door)

The lightest test of all — measure demand before you build anything, even a prototype.

What it is

A pretotype tests one thing: does anyone actually want this?demand — at near-zero cost and before any real build. The classic form is the "fake door": a sign-up button, a landing page, or an announcement for a service that doesn't yet exist, where you measure how many people click or register. It sits at the bottom of the build-to-learn ladder because it challenges the assumption every other test takes for granted — that the thing is worth making at all. Where a prototype asks "can they use it?" and a POC asks "can we build it?", a pretotype asks the prior question: "should this exist?"

When to use it

Shape, against a value/demand assumption, when building even a prototype would be premature. It's the antidote to building something beautifully that nobody wanted.

Process

Create the lightest possible signal of the real thing (a button, a page, an email offer); expose it to real users; measure the action that indicates genuine interest (clicks, sign-ups, replies); compare against a pre-set threshold. Handle the ethics: tell sign-ups it's "coming soon," or route them to a manual fallback — never leave someone deceived or stranded.

Diagram

Try the new Buddy Match →
the door is fake
clicks: 47 / 200 sent
“Thanks — launching soon!” (ethical fallback)

WORKED EXAMPLE (HR)

Before building any buddy-matching system, email new hires a "Sign up for a buddy" link and count registrations. Low sign-up kills the idea in a day; high sign-up earns the next, more expensive test.

Pitfalls & how to run it well

Deceiving or stranding users (always provide a real fallback); reading vanity clicks as commitment (pick an action that costs the user something); pretotyping when you already know there's demand. Run it well with a clear pre-set threshold for "enough interest."

Connects to

A type of Experiment; attacks value risk; the lightest rung below Prototype, POC, and MVP.

16

Prototype

A representation built to test usability and desirability with users — without building the real thing.

What it is

A prototype is a mock-up, clickable screen, paper sketch, script, or staged walkthrough made to provoke real reactions and surface confusion early, when changing direction is cheap. It tests usability and desirability — "do they get it, and do they want it?" — and it is deliberately not functional underneath; the value is in what users do and say, not in the artifact. It comes in a fidelity range (paper → clickable → high-fidelity), and you use the lowest fidelity that answers the question. It sits above the pretotype on the ladder (it assumes demand and tests the experience) and below the POC and MVP.

When to use it

Shape, against a usability/desirability assumption, once demand is plausible. Use it to learn before a line of production code exists.

Process

Pick the fidelity that fits the question (don't polish a high-fi mock to test a flow); build the representation of just the slice you're testing; put it in front of real users; watch where they hesitate, misread, or reach for something that isn't there; capture the reactions, not applause.

Diagram

paper sketchclickable mockhigh-fi prototype
tests usability & desirability with users — not production, feasibility, or real value

WORKED EXAMPLE (HR)

A clickable mock-up of an onboarding dashboard shown to five new hires: you watch them hunt for "who's my buddy" and misread the week-one checklist — fixing the layout before anyone builds it.

Pitfalls & how to run it well

Over-investing in fidelity (you're testing reactions, not shipping); leading the user; mistaking polite praise for usability. Run it well by testing the lowest fidelity that answers the question and observing behaviour over opinion.

Connects to

A type of Experiment; attacks usability/desirability risk; distinct from POC (internal feasibility, no users) and MVP (real, live, delivers value).

17

POC (proof of concept)

A small, internal exercise to answer one question: can this even be built or made to work?

What it is

A proof of concept targets feasibility — the technical or practical "is this approach even possible with what we have?" It's allowed to be ugly, throwaway, and hidden from users; its only job is to retire a "we're not sure this can be done" doubt before the team bets real resources. The thing that sets it apart from every other test is the kind of uncertainty it resolves: not whether anyone wants it (value), not whether it works for the business (viability), only whether the approach is technically workable. That's why a POC is not an MVP — a POC proves can we, an MVP tests should we by putting real value in front of real users.

When to use it

Shape, against a feasibility assumption — when there's a genuine "can this be built / integrated / made to work at all?" doubt. Skip it when feasibility was never really in question.

Process

Isolate the single technical question; build the minimum wiring that answers it (a spike), with no interface and no users; confirm it works or doesn't; throw the code away. The output is a yes/no on feasibility, not a product.

Diagram

System A
just enough
System B
internal · no UI · throwaway — technically possible? yes / no
POC
can we build it? internal, throwaway
MVP
should we? real users, real value

WORKED EXAMPLE (HR)

Before promising a skills-matching feature, wire the HRIS to a matching script for a handful of records to confirm the data is clean and reachable enough to power it — no interface, thrown away after.

Pitfalls & how to run it well

Shipping a POC to users (it was never meant for them); over-building it into a half-product; running one when feasibility was never actually the risk. Run it well by keeping it isolated, internal, and disposable.

Connects to

A type of Experiment; attacks feasibility risk (the tech-lead's seat in the trio); distinct from Prototype (usability, with users) and MVP (value in real use).

18

MVP (minimum viable product)

The smallest real, live version that delivers genuine value and lets you learn from actual use.

What it is

The Minimum Viable Product is the smallest real version of a solution that delivers genuine value to real users and lets you learn from how they actually use it. Both words carry weight: minimum means stripped to the smallest thing that can stand on its own; viable means it must actually work and deliver value — a broken fragment is minimum but not viable. It's the top of the build-to-learn ladder and the most abused term in product, usually misused to mean "version one" or "the cut-down scope we could afford." It is none of those: not a prototype (not live, delivers no real value), not a POC (internal, proves only buildability), and not "everything minus polish." The discipline is the question: what is the smallest live thing that lets a real user get real value, so we learn whether this is worth scaling? The mindset is "build to learn, not to launch."

When to use it

The end of Shape into Build, against value and viability in real use, once cheaper tests have de-risked demand and usability. It's the first thing that's actually real.

Process

Define the single core value the user must get; strip everything not required to deliver that value once; build it for real but small; release it to a bounded group; measure the outcome (not just usage); learn and decide whether to scale.

Diagram

most real / live
v1 / fullthe complete product (not an MVP)
MVPREAL + LIVE + smallest slice that delivers value
POCworks internally, no users (feasibility)
Prototypelooks real, isn't (no value)
Pretotypedemand — does anyone want it?
lightest / cheapest

WORKED EXAMPLE (HR)

Not "the onboarding portal." The MVP is: assign buddies + a one-page week-one plan, to one incoming cohort, measured on day-7 confidence and time-to-first-task. Smallest live thing that delivers the value, instrumented to learn.

Pitfalls & how to run it well

Calling a feature-complete v1 an "MVP"; shipping something minimum but not viable (no real value); treating launch as the finish rather than the start of learning. Run it well by defining the one core value first and instrumenting the outcome from day one.

Connects to

A type of Experiment and the bridge into Build; attacks value/viability risk; the most-real rung above Pretotype, Prototype, and POC.

19

A/B Test

Compare two live versions with real users to settle "which version is better" with evidence instead of opinion.

What it is

An A/B test shows two (or more) versions of something to comparable groups of real users at the same time and measures which performs better on a pre-defined metric. Because the only meaningful difference between the groups is the change you're testing, a difference in results can be attributed to that change. It replaces taste and seniority with data on decisions that are genuinely uncertain and measurable — wording, layout, flow, timing. It is not a discovery tool: it optimises among options you already have, it won't tell you what to build.

When to use it

When you have two concrete variants of a live thing and a measurable outcome, and the decision between them is worth resolving with data rather than argument. Needs enough users for the difference to be real, not noise.

Process

Define the success metric before running; create the two variants differing in one thing; split comparable users between them; run until you have enough signal; pick the winner on the pre-set metric.

Diagram

users
Variant A
Variant B
measure
same metric
winner
metric defined first

WORKED EXAMPLE (HR)

Two versions of a benefits-enrolment reminder email sent to comparable employee groups; measure completion rate; roll out the winner. The metric (completion, not opens) is fixed before sending.

Pitfalls & how to run it well

Choosing the metric after seeing results (so you cherry-pick); too few users (noise read as signal); using it to "discover" rather than optimise. Run it well by fixing the metric up front and changing one variable at a time.

Connects to

A situational Experiment; mainly attacks value/usability at the margin; optimises among options that earlier methods (OST, prioritization) produced.

20

Concierge Test

Manually deliver the service to a few real users — openly — to learn what the experience needs before building anything.

What it is

In a concierge test the team does by hand, in the open, what a finished product would eventually automate. The user gets a real outcome; the team just provides it through human effort, and the user knows it's manual. The name comes from a hotel concierge personally handling each request. Its purpose is to learn the shape of the solution cheaply, on real cases — every step, exception, and judgement call the real service will have to handle — before committing to a build. It delivers genuine value (unlike a prototype) and builds nothing (unlike a pilot or MVP).

When to use it

Shape, against value and "what does this experience actually require" uncertainty, when you want to learn the service design before automating it.

Process

Pick a few real users; deliver the service to them by hand, openly; capture every step, decision, and edge case you hit; use that to design (or decide against) the eventual automated version.

Diagram

team member
hand-delivers
user 1
user 2
user 3
openly manual
logs every step, decision, edge case → informs the build

WORKED EXAMPLE (HR)

Personally match ten new hires to buddies yourself — choosing pairs, checking in, handling the awkward cases — before designing any matching system. You learn what makes a good match, which is exactly what an algorithm would need to encode.

Pitfalls & how to run it well

Automating before you understand the edge cases (the whole point is to find them by hand); doing it at a scale you can't sustain. Run it well with a deliberately tiny n and a disciplined log of what you learned.

Connects to

A type of Experiment; attacks value and surfaces feasibility/operational realities; distinct from Wizard-of-Oz (manual but hidden) and Prototype (delivers no real value).

21

Wizard-of-Oz

Users interact with what looks like a finished, automated product — but a human is doing the work behind the curtain.

What it is

In a Wizard-of-Oz test the user believes the system is real and automated, while a person quietly does the work behind the scenes — the wizard behind the curtain. That hidden-ness is the whole point and the difference from a concierge test: because the user thinks it's automated, you observe their genuine, unselfconscious behaviour with the "product," letting you test the real experience and demand for an automated solution without building the automation. It answers "if this worked automatically, would people use it and would it deliver value?" at a fraction of the cost.

When to use it

Shape, against value / the experience of an automated solution, when the cost or risk of building the automation is high and you want real behaviour, not stated intent.

Process

Build a convincing front (what the user sees as automated); have a human generate the responses/results live behind it; let real users interact believing it's real; measure whether they engage and benefit; only then decide whether the automation is worth building.

Diagram

FRONT STAGE · user sees
“automated” product
curtain
BACK STAGE · hidden
human generates results
the user believes it’s automated

WORKED EXAMPLE (HR)

A "smart" internal-mobility recommender whose suggestions are quietly hand-picked by a person, shown to employees as if generated by a system — to see whether people act on recommendations at all before anyone builds the matching engine.

Pitfalls & how to run it well

Letting the deception harm or mislead users in ways that matter (keep it honest and low-stakes — the test-ideas-responsibly principle applies); not being able to sustain the manual effort even briefly. Run it well at small scale, with a clear read on the behaviour you're testing.

Connects to

A type of Experiment; attacks value/experience risk for would-be automated solutions; distinct from Concierge (manual but open) and from POC (proves the automation is buildable, with no users).