Foundations
What qualitative coding actually is, why it is interpretive work, and the small set of distinctions you have to get right before any of the later chapters help.
What coding is (and isn’t)
In qualitative research, coding is the act of attaching short labels to extracts of data — words, phrases, paragraphs, sometimes whole exchanges — that index something the analyst wants to track. The labels are the codes; the extracts are what they apply to. Coding is the way an analyst makes a large unstructured corpus interrogable.
It is worth being clear about what coding is not. It is not the same as the “coding” survey researchers do when they translate open-ended responses into a fixed taxonomy of categories. It is not a way of counting (although codes can be counted). It is not a substitute for analysis — it is the scaffolding on which analysis happens.
Saldaña (2021) puts the distinction sharply: a code is “a researcher-generated construct that symbolizes and thus attributes interpreted meaning to each individual datum for later purposes of pattern detection, categorization, theory-building, and other analytic processes.” The phrase that matters there is interpreted meaning. Coding is not data entry. It is interpretation, performed early and recorded so the interpretation can be argued with later.
Inductive, deductive, abductive
Three logics of inference, each producing a recognisably different relationship between codes and data.
Inductive coding
Codes are generated from the data. The analyst reads, notices, labels, and builds a codebook from the ground up. Inductive coding is the default for exploratory studies and for traditions (reflexive thematic analysis, grounded theory, IPA) that treat the participants’ framing as the analytical starting point. The risk is that “purely inductive” coding is a methodological fiction — every analyst arrives with priors, and the right move is to declare them rather than pretend they aren’t there.
Deductive coding
Codes come from theory, prior literature, or a pre-existing framework, and the analyst applies them to new data. Deductive coding is the default for framework analysis (Ritchie & Spencer, 1994), for studies replicating or extending prior work, and for survey-derived qualitative data. The risk is confirmation: a strong prior framework will see itself in the data even where it doesn’t fit.
Abductive coding
Codes alternate between inductive generation and deductive testing, organised around anomalies — data that doesn’t fit the working theory and forces the theory to revise. Timmermans and Tavory (2012) give the most-cited contemporary statement. In practice, most coding labelled “inductive” in methods sections is in fact abductive, and saying so is more accurate.
Codes, categories, themes
These three terms get used interchangeably in undergraduate methods textbooks, which is unfortunate because they are not interchangeable. A clean working set of definitions:
- A code is a label applied to a data extract. It is descriptive (“mother as primary carer”), in vivo (a phrase taken directly from a participant: “I had to be someone else”), or analytical (“the interrupted self”). A study typically generates dozens to low hundreds of codes.
- A category is a cluster of related codes that share a topic or descriptive domain. “Family relationships” is a category. Categories are organisational, not yet interpretive.
- A theme is an analytic claim about a pattern that runs through codes (and sometimes categories). A theme is an argument the data is making. “Family relationships are reframed after a caregiving role ends” is a theme.
A study can have 80 codes, organised into 8 categories, that support 3 themes, and that is a normal, well-shaped analysis. A study with 80 codes and 80 themes has skipped a step. A study with 3 codes and 12 themes is doing something else entirely. Chapter 2 goes into the codes-versus-themes distinction in depth for thematic analysis specifically; the same intuition transfers to grounded theory and IPA.
Saturation, honestly
Theoretical saturation — the point at which new data stops producing new codes — is the conventional stopping rule in grounded theory and has migrated into thematic analysis methods sections. The problem is that it is, as written, unfalsifiable: the analyst declares saturation reached, and the reader has no way to check.
Two pieces of recent work are worth knowing. Bowen (2008) walked through the operational ambiguity in the term and argued for a transparent reporting standard — how many interviews, what specifically stopped generating new codes, what was the threshold. Hennink, Kaiser and Marconi (2017) showed empirically that for in-depth interview studies, code saturation typically occurred by interview 9, butmeaning saturation (further development of existing codes) required 16–24 interviews.
The defensible practice in 2026: state your stopping rule in advance, report what you observed, and stop using “saturation” as if it were a single threshold. Code saturation, meaning saturation, and theoretical saturation are different things; the methods section should say which one you reached.
Before you start coding
Three small commitments you should make before opening the first transcript:
1. Decide which methodology you are doing — reflexive TA, codebook TA, grounded theory, IPA, framework analysis, narrative analysis — before any coding. The choice shapes the coding logic (inductive / deductive / abductive), whether intercoder agreement is required, how the codebook is structured, and what the final report looks like. Picking after the fact is the most common methodological mistake in qualitative theses.
2. Write a one-paragraph statement of positionality. Who is doing the coding, what relationship they have to the participants, what they expect to find, and what theoretical priors are in play. The point is not confessional; the point is that those priors will shape the codes and naming them early disciplines the analyst and helps the reader.
3. Pick a stopping rule — in advance. “We will code until two consecutive interviews produce no new codes against the working codebook” is a defensible stopping rule. “Until saturation” without further detail isn’t.
The five remaining chapters of this guide work through specific methodologies (thematic analysis, grounded theory, IPA), one common analytical question (intercoder reliability), and the ethics layer that runs through all of them. Start with the chapter that matches the methodology you’ve picked, not the one at the top.
Further reading
- Saldaña, J. (2021). The Coding Manual for Qualitative Researchers (4th ed.). SAGE.
- Boyatzis, R. E. (1998). Transforming Qualitative Information: Thematic Analysis and Code Development. SAGE.
- Timmermans, S., & Tavory, I. (2012). Theory construction in qualitative research: From grounded theory to abductive analysis. Sociological Theory, 30(3), 167–186.
- Bowen, G. A. (2008). Naturalistic inquiry and the saturation concept: a research note. Qualitative Research, 8(1), 137–152.
- Hennink, M. M., Kaiser, B. N., & Marconi, V. C. (2017). Code saturation versus meaning saturation: How many interviews are enough? Qualitative Health Research, 27(4), 591–608.
- Ritchie, J., & Spencer, L. (1994). Qualitative data analysis for applied policy research. In Analyzing Qualitative Data (pp. 173–194). Routledge.
This chapter is in draft. It has not yet been peer-reviewed by an external methodologist. Reviewer contact: [email protected].