Intercoder reliability
Cohen's κ, Krippendorff's α, and when each fits — without conceding more than the method actually requires to a positivist framing.
What you’re measuring (and not)
Intercoder reliability (sometimes intercoder agreement, or inter-rater reliability when the coding is numeric) is a statistic that summarises how often two or more coders apply the same code to the same extract. It is not a measure of whether the coding is correct. It is not a measure of whether the codes are good codes. It is a measure of consistency.
The conflation of consistency with correctness is the source of most of the trouble in this area. A codebook can produce κ = .85 across two coders and still be a bad codebook, if the categories are theoretically thin or the inclusion criteria are written so broadly that almost anything fits. The statistic guards against one specific failure mode — idiosyncratic, drifting, or non-replicable coding — and is silent on every other.
With that framing in place, the two statistics worth knowing are Cohen’s κ (kappa) and Krippendorff’s α (alpha).
Cohen’s κ
Cohen (1960) introduced κ as a chance-corrected agreement statistic for two raters and categorical data. The intuition is straightforward: if two raters agree on 80% of cases, but 60% agreement would happen by chance given the marginal frequencies, then the genuine agreement above chance is what κ captures.
The conventional thresholds (Landis & Koch, 1977) are widely cited and widely overstated:
- κ < .20 — slight agreement
- .21–.40 — fair
- .41–.60 — moderate
- .61–.80 — substantial
- .81–1.00 — almost perfect
In practice, .70 is the conventional minimum for publishable qualitative coding, and .80 is the bar in health-services research. The thresholds are heuristic; Landis and Koch invented them without empirical grounding and have been backed away from since (McHugh, 2012).
Cohen’s κ has two well-known failure modes. The first is the prevalence problem: with a heavily skewed code (say, 90% of extracts get the code, 10% don’t), even high raw agreement can produce a low κ because the chance-correction term is dominated by the marginals. The second is the bias problem: if the two raters apply the code at different overall rates (one rater 30% of extracts, the other 60%), κ penalises that bias even when their agreement on individual decisions is high. Both are well-documented (Feinstein & Cicchetti, 1990).
Krippendorff’s α
Krippendorff’s α (Krippendorff, 2004; Hayes & Krippendorff, 2007) generalises agreement statistics in three useful ways: any number of coders, any level of measurement (nominal, ordinal, interval), and tolerant of missing data. For nominal categorical coding with two coders, α and κ agree to three decimal places in most cases; the choice between them rarely changes the conclusion.
Where α earns its keep:
- Three or more coders. Cohen’s κ is defined for two raters; the multi-rater extensions (Fleiss’s κ, Light’s κ) have known instabilities. α handles arbitrary coder counts in one statistic.
- Partial coverage. When coders are randomly assigned to subsets of the dataset (a common design when coding is expensive), α handles the unbalanced design correctly; κ does not.
- Ordinal codes. If your code is severity-graded (low / moderate / high), α-ordinal credits near-misses appropriately; nominal κ treats low-vs-high as just as wrong as low-vs-moderate.
Conventional α thresholds are similar to κ: above .80 for tentative conclusions, above .67 for cautious conclusions about agreement, below .67 considered unacceptable for substantive claims (Krippendorff, 2004, p. 241).
The recurring debate
Whether qualitative researchers should report agreement statistics at all is a methodological argument that goes back to the 1980s and resurfaces every few years. The positivist position: any analytical claim should be replicable by another competent analyst; agreement statistics are how you demonstrate that. The interpretivist position: qualitative analysis is, by design, the situated interpretation of a knowledgeable analyst; demanding replicability concedes the wrong epistemology.
The most readable contemporary statement of the interpretivist position is Braun and Clarke (2019), which treats κ as conceptually incoherent for reflexive thematic analysis. The most readable consequence-oriented critique is McDonald, Schoenebeck and Forte (2019) on CSCW/HCI practice: they show that defaulting to κ across all qualitative work produces a fake-rigour aesthetic in journals that reviewers then enforce against work where the statistic doesn’t belong.
“In the absence of meaningful guidelines for when IRR is appropriate, the field has converged on a default that treats κ as a universal marker of qualitative rigour, regardless of whether the underlying methodology actually generates the kind of claim that IRR can support.”
— McDonald, Schoenebeck & Forte, 2019
A pragmatic position
A defensible practice, written so you can cite it in a methods section:
Report intercoder agreement when the codebook is intended to be applied beyond a single analyst — when other researchers will code further data using the same codebook, when coding is part of a structured content analysis with discrete categorical outputs, or when the analytic claim is about prevalence (“X% of participants reported Y”) rather than meaning (“participants narrated Y as a turning point”).
Do not report intercoder agreement when the analytic instrument is the analyst’s interpretive engagement. Reflexive thematic analysis, narrative analysis, IPA, and most phenomenological approaches fall here. The methods section should state explicitly that intercoder reliability is not the appropriate quality check for the approach being used, and what is — typically reflexive memos, an audit trail, member checking where appropriate, transparency about the analyst’s positionality.
When you do report κ or α, report it honestly. Include the per-code breakdown, not just the overall statistic. The overall number can hide a category where κ is .40 and the prose is (intentionally or not) directing the reader to assume the .85 average applies everywhere. The per-category breakdown is also where the actual codebook work shows up.
In QualCanvas
The intercoder reliability panel (Pro/Team) computes Cohen’s κ for any two researchers coding the same canvas and Krippendorff’s α when three or more researchers are present. The export gives you the overall statistic, the per-code breakdown, the confusion matrix between any two coders, and a disagreement queue you can step through to refine the codebook.
The disagreement queue is the useful part. The statistic tells you whether you have a problem; the queue tells you what to do about it. Most disagreements cluster on three or four boundary cases per code; an afternoon of joint review and a small codebook clarification usually moves κ from .65 to .80.
Further reading
- Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46.
- Krippendorff, K. (2004). Content Analysis: An Introduction to Its Methodology (2nd ed.). SAGE.
- Hayes, A. F., & Krippendorff, K. (2007). Answering the call for a standard reliability measure for coding data. Communication Methods and Measures, 1(1), 77–89.
- McHugh, M. L. (2012). Interrater reliability: the kappa statistic. Biochemia Medica, 22(3), 276–282.
- Feinstein, A. R., & Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543–549.
- McDonald, N., Schoenebeck, S., & Forte, A. (2019). Reliability and inter-rater reliability in qualitative research: Norms and guidelines for CSCW and HCI practice. Proceedings of the ACM on Human-Computer Interaction, 3(CSCW), 1–23.
- Braun, V., & Clarke, V. (2019). Reflecting on reflexive thematic analysis. Qualitative Research in Sport, Exercise and Health, 11(4), 589–597.
This chapter is in draft. It has not yet been peer-reviewed by an external methodologist. Reviewer contact: [email protected].