What Is Qualitative Data?: Interviews, observations, and the evidence they provide in product decisions
A research readout can travel a long way from the material it rests on. Six interviews become “users hate setup,” which becomes a roadmap item, which becomes a sentence about the customer base. Each step can look reasonable on its own, and the finished claim can still assert far more than anyone observed. The discipline that prevents this is unglamorous: keep asking which record a claim is standing on.
“Said,” “did,” and “want” are three different claims
“Participants said” and “participants did” are different observations, and neither is the same as “customers want.” An interview supports a statement about what someone reported in that session. An observation supports a statement about what happened in the setting that was watched. Neither one produces a population rate, and neither, by itself, shows that a proposed change will move a metric.
Qualitative data can tell a product team that selected participants interpreted a label in a particular way, encountered a barrier during a task, or relied on a workaround in the contexts studied. It cannot, without an appropriate quantitative design, tell the team what percentage of the market shares that interpretation or how much a proposed change will improve activation.
What counts as qualitative data
In product work the material is concrete: interview transcripts, open-ended responses, observation notes, screenshots, photos, audio, or video. It is recorded evidence about qualities, meanings, experiences, and behavior. The term itself, though, carries two valid meanings that are easy to mix up:
| Context | What “qualitative data” usually means | Example | Appropriate analysis |
|---|---|---|---|
| Introductory statistics | Values of a categorical variable | Job role, plan type, or acquisition channel | Frequencies, proportions, mode, and category comparisons |
| Qualitative or product research | Naturalistic records interpreted for meaning and context | Interview transcript, field note, open-ended response, photo, audio, or video | Coding, case comparison, thematic or other method-appropriate interpretation |
The Australian Bureau of Statistics uses the first meaning: qualitative data describes categories, which may be names, symbols, or number codes. Those categories can be counted, although arithmetic such as a mean is generally not meaningful for an unordered category. Product researchers more often use the second meaning. The GOV.UK recording guidance lists notes, photos, audio, video, and copies of materials as records from a research session.
Qualitative data has no general-purpose formula. Coding a set of transcripts and counting each code can be useful, but the count is a derived summary of that particular material. It does not erase recruitment bias, make every coded instance equivalent, or convert a purposive interview sample into a market estimate. A sentiment score produced from text is likewise a new quantitative variable built from qualitative material; the score is not the original evidence.
Interviews and observations provide different evidence
An interview reaches what a participant can put into words: goals, explanations, expectations, vocabulary, remembered events, and how they frame the problem. The WHO describes the qualitative interview as a guided conversation for gathering in-depth information, and GOV.UK uses in-depth interviews to learn about users’ circumstances, service use, needs, and reported problems. The resulting record supports one kind of statement: what that participant reported in that session.
Contextual observation reaches what articulation misses. The GOV.UK Service Manual describes watching people perform an activity in their everyday environment, with their normal tools, documents, devices, and distractions. That record preserves order of actions, hesitation, handoffs, interruptions, barriers, and workarounds — the parts of a workflow people rarely reconstruct accurately from memory.
| Evidence source | It can support a finding about | It does not establish by itself |
|---|---|---|
| In-depth interview | Reported goals, beliefs, remembered experience, language, reasons, and perceived problems | Usual behavior, prevalence across customers, or the effect of a product change |
| Contextual observation | Actions, sequence, environment, tools, barriers, and workarounds in the observed setting | An unspoken motive, behavior in every other setting, or population frequency |
| Open-ended survey response or support message | The writer’s description, language, and self-selected concern | How common the concern is among non-respondents or what caused it |
| Analytics or a structured survey | Measured behavior or response distribution under the instrument and sample design | The meaning behind a pattern unless the design captures it |
| Randomized experiment | A treatment-effect estimate under a valid design and analysis | The full reason an effect occurred or how users understood the experience |
Silent observation preserves more of the activity’s natural flow but can leave the reason for an action unclear. Asking a question adds explanation while interrupting that flow. Interviews and observations therefore complement each other: the first can investigate meaning; the second can inspect conduct and context. Neither is automatically “truer.” Each is evidence for a different class of claim.
What qualitative data is good for in product work
Qualitative methods earn their place when the answer options are not yet known, when context matters, or when a team needs to understand why a process works differently from the diagram. The CDC’s field guide identifies exploratory questions, perceptions and subjective meaning, context, how-and-why questions, and how an intervention works in practice as appropriate uses. Those are public-health examples, but the decision logic transfers cleanly.
In product work, qualitative data is particularly useful for:
- discovering a problem the roadmap has not named yet;
- learning how users describe a task before the team fixes the taxonomy;
- mapping the real workflow around a product, including spreadsheets, approvals, handoffs, and offline steps;
- identifying why a metric pattern could have several plausible mechanisms;
- seeing where instructions, labels, or system feedback create the wrong mental model;
- generating hypotheses and candidate designs; and
- finding edge conditions that an average conceals.
The method should follow the decision. If the team needs to understand why selected administrators delay inviting colleagues, interviews may expose approval rules or perceived risks. Observation may reveal where they leave the workflow, which information they gather, and which tool they use next. Event data can then test how widespread that path is. A prototype study can test whether a redesigned flow is comprehensible, and a valid experiment can estimate whether it changes the target outcome.
This is a sequence of complementary questions, not a hierarchy in which numbers outrank words. The market research methods guide covers how to choose among desk research, interviews, surveys, tests, and pilots at different decision stages.
Turn raw material into a traceable finding
An interview transcript is data, not yet an insight. An observation note is data, not yet a product requirement. The move from record to decision contains judgment, and that judgment needs a visible trail.
A practical chain is:
record → observation → pattern → bounded finding → decision implication → check
Preserve the record before interpreting it
Record what was seen or heard with enough context to find it again. Label the session and participant consistently, retain the task or question that produced the material, and separate a verbatim statement or visible action from the note-taker’s conclusion. GOV.UK’s note-taking guidance recommends single-point notes tied to a participant and session, with observations kept distinct from interpretations.
“Participant returned to the account page twice before continuing” is an observation. “Navigation is confusing” is an interpretation. The interpretation may eventually be supported, but preserving the first statement lets another analyst inspect the reasoning and consider alternatives.
Code for the question, then compare cases
Coding labels a relevant segment so related material can be retrieved and compared. A code might mark an approval dependency, an unknown term, a manual workaround, or a trust concern. Codes can begin from the research question, emerge during analysis, or combine both approaches; the analysis method should be chosen deliberately rather than improvised after the strongest quote appears.
The CDC guidance emphasizes close reading, codebook development, coding, discussion, revision, and recoding until patterns become clear. It also makes an important limitation explicit: software can retrieve and organize material, but analysts still discern the pattern and reach the conclusion.
Compare within and across participants. Look for repetition, variation, sequence, conditions, and cases that contradict the emerging explanation. A theme is not merely a topic that appeared several times. It is a defensible pattern that helps answer the research question and whose boundaries can be stated.
Write the finding narrowly enough to be true
A decision-ready finding names the group, context, evidence, interpretation, and uncertainty. The following are illustrative formulations for an unnamed scenario, not company findings:
| Formulation | Problem |
|---|---|
| Users hate setup. | Users is unbounded, hate is an interpretation, and no task or evidence is named. |
| Setup was confusing in several sessions. | Repetition is implied, but the selected sample, source of confusion, and decision relevance remain unclear. |
| In this round, participants who needed an internal approval paused at the connection step and left the flow to obtain information not shown in the interface. | The observed group, condition, behavior, and barrier are visible; the finding remains bounded to the round. |
The last statement still does not prove that every customer has the problem or that one design will solve it. It is strong enough to justify a next action: inspect how often the condition occurs, prototype a deferral or explanation, recruit relevant participants, and test the revised flow.
Its purpose is to show the evidence chain. The raw observations support a bounded problem finding; analytics could test reach; a prototype session could test comprehension; and an experiment or before-and-after operational measure could address impact, depending on the decision.
Connect the finding to an action and a falsification check
The GOV.UK analysis method moves from discrete observations to themes, findings, and possible actions such as design changes, new research questions, prototype revisions, user stories, or roadmap inputs. Keep those levels separate in the research record:
- Evidence: what was captured and where.
- Finding: the bounded interpretation supported by that evidence.
- Implication: why the finding matters to the pending decision.
- Action: what the team will change, measure, or investigate.
- Check: what future observation would weaken or overturn the interpretation.
This prevents a common failure: presenting a proposed feature as though participants had requested and validated it. Participants provide evidence about their experience. The team owns the product response and the risk that its interpretation is wrong.
How many interviews are enough?
There is no universal participant count and no qualitative equivalent of a generic statistical-power formula. The required evidence depends on the decision, the diversity relevant to it, the method, the depth of each session, and the precision of the claim.
GOV.UK’s planning guide says teams will usually need four to eight participants per round for methods including contextual research, in-depth interviews, and usability testing, then recommends additional rounds rather than one large round when findings are unclear. That is a scoped practice guideline for iterative service research. It is not a promise of saturation or market representation.
Malterud, Siersma, and Guassora propose information power as a better way to reason about interview sample adequacy. A sample can require fewer participants when its information is highly relevant to a narrow aim, the participants are specific to that aim, dialogue quality is strong, established theory helps focus inquiry, and the analysis is appropriately bounded. A broad aim, varied population, thin conversations, or subgroup claims demand more evidence.
For product teams, stop asking only “Have we interviewed enough people?” Ask:
- Have we included the user types and operating conditions that could change this decision?
- Are new sessions still changing the explanation, not merely adding another example?
- Have we explored contradictions and plausible alternative causes?
- Is the claim no broader than the sample and context support?
- Would another method answer the unresolved part better?
A small round can be sufficient to expose a severe usability barrier in the tested flow. It is not sufficient to estimate what share of the customer base encounters it. Those are different decisions with different evidence requirements.
Make qualitative evidence credible enough to act on
Qualitative rigor is not a contest for the most transcripts or the largest wall of sticky notes. It comes from making the chain of evidence inspectable and limiting the conclusion to what the method can support.
Use these checks before a finding enters a roadmap or prioritization meeting:
- Decision fit: State the decision and research question before selecting the method. Do not run interviews when the unresolved question is a population rate.
- Relevant recruitment: Select participants because they can inform the question, and include meaningful differences in role, access needs, experience, account context, or workflow when those differences could change the answer.
- Neutral collection: Use open, non-leading prompts. In interviews, ask for recent concrete experience and probe the sequence; in observation, record what occurred before explaining it.
- Traceability: Keep session-level references from each finding back to notes, recordings, or artifacts, within the participant’s consent and the organization’s access rules.
- Comparative analysis: Examine all relevant cases, not only memorable quotes. Preserve contradictions and distinguish common patterns from isolated but consequential cases.
- More than one perspective: Involve observers or another analyst where the stakes justify it. The GOV.UK analysis guidance notes that group analysis can reduce the influence of one researcher or stakeholder.
- Bounded language: Write “participants in these sessions” when that is the evidence. Do not silently expand the subject to “customers” or “the market.”
- Triangulation by question: Use analytics, surveys, support data, experiments, or another qualitative method when the decision also needs reach, trend, causal effect, or a competing explanation tested.
- Counterevidence: Record what would falsify the finding and look for it. A persuasive quote is not permission to stop comparing cases.
Research material can contain identifiable behavior, recordings, account context, and sensitive disclosures. The GOV.UK consent guidance says participants should understand the purpose, data collected, use, sharing, recording, retention, and voluntary nature of participation. Apply the privacy, security, research-ethics, and legal requirements relevant to your organization and jurisdiction; a useful insight does not expand the consent under which its evidence was collected.
Pair the method with the product decision
The simplest way to avoid overclaiming is to name the missing verb in the decision:
| The team needs to… | Start with | Add when needed |
|---|---|---|
| Discover unknown problems or vocabulary | Interviews, contextual observation, open-ended feedback | Support logs and search data to locate adjacent evidence |
| Understand a workflow or workaround | Contextual observation plus probing | Event instrumentation or process data to test reach |
| Learn how people interpret a concept | In-depth interviews or concept evaluation | A structured survey if prevalence across a defined population matters |
| Find where a tested flow breaks | Moderated usability observation | Funnel data to measure scale; repeat studies across relevant user groups |
| Estimate how common a need is | A well-designed survey or behavioral measure | Interviews to explain response meaning and mechanisms |
| Determine whether a change caused an outcome | A valid experiment or causal design | Qualitative follow-up to understand why the effect did or did not occur |
Qualitative and quantitative evidence are not substitutes at opposite ends of one quality scale. They are instruments for different questions. The ABS notes that categorical and numeric data can describe different aspects of the same unit; product teams can extend that principle by pairing measured patterns with researched meaning. A drop-off rate says where a measurable event stopped. An interview or observation may reveal the competing reasons. A follow-up measurement tests whether the chosen response changed the outcome.
If you collect open-ended customer feedback, the customer feedback survey guide explains how to connect questions to decisions without turning every response into a roadmap item.
Sources
- Centers for Disease Control and Prevention, “Collecting and Analyzing Qualitative Data”
- Australian Bureau of Statistics, “Quantitative and qualitative data”
- World Health Organization Regional Office for the Western Pacific, “Applying communication for health: research tools: qualitative interviews”
- GOV.UK Service Manual, “Using in-depth interviews”
- GOV.UK Service Manual, “Contextual research and observation”
- GOV.UK Service Manual, “Taking notes and recording user research sessions”
- GOV.UK Service Manual, “Analyse a research session”
- GOV.UK Service Manual, “Plan user research for your service”
- Qualitative Health Research via PubMed, “Sample Size in Qualitative Interview Studies: Guided by Information Power”
- GOV.UK Service Manual, “Getting informed consent for user research”
Continue the evidence path
Related reading
Related
Qualitative vs. Quantitative Data: Key differences in evidence, question types, and limitations
Connect What Is Qualitative Data?: Interviews, observations, and the evidence they provide in product decisions with Qualitative vs. Quantitative Data: Key differences in evidence, question types, and limitations so a reader can see both halves of the comparison before choosing a method.
Next step
Market Research Methods for Lean B2B Teams: What to Use at Each Decision Stage
Connect What Is Qualitative Data?: Interviews, observations, and the evidence they provide in product decisions with Market Research Methods for Lean B2B Teams: What to Use at Each Decision Stage so a reader moves from the evidence type to selecting a design for the decision at hand.
Related
Customer Feedback Explained: Sources, Signal Quality, and Product-Decision Limits
Connect What Is Qualitative Data?: Interviews, observations, and the evidence they provide in product decisions with Customer Feedback Explained: Sources, Signal Quality, and Product-Decision Limits to separate research-session evidence from continuously collected feedback.