How to Calculate Intercoder Reliability: A Simple Guide (2026)

Intercoder reliability is calculated by double-coding a sample of your data, building a matrix that pairs each coder’s decision against the others, then comparing the observed agreement with the agreement you’d expect by chance alone. With two coders and categorical labels that means Cohen’s kappa. With three or more coders it usually means Krippendorff’s alpha. With numeric scores it means an intraclass correlation.

This guide walks through each of those calculations with the arithmetic shown, so you can reproduce every number by hand or check what your software produced.

Table of Contents
  1. 1What You Need
  2. 2Step-by-Step
  3. 3Step 1: Define the unit and write the category rules
  4. 4Step 2: Build the coding table
  5. 5Step 3: Compute observed agreement first
  6. 6Step 4: Calculate Cohen’s kappa for two coders
  7. 7Step 5: Read the same result off a two-category table
  8. 8Step 6: Calculate Krippendorff’s alpha for three or more coders
  9. 9Step 7: Use an intraclass correlation for numeric ratings
  10. 10Step 8: Calculate it in Excel, SPSS or R
  11. 11Step 9: Document and report the procedure
  12. 12What Counts as a Good Intercoder Reliability Score?
  13. 13Common Mistakes
  14. 14Why kappa sometimes misleads
  15. 15Frequently Asked Questions
  16. 16What is the difference between reliability and validity?
  17. 17When should I use Cohen’s kappa instead of Krippendorff’s alpha?
  18. 18Is percentage agreement enough on its own?
  19. 19What is a good intercoder reliability score?
  20. 20Should I calculate a separate kappa for each survey question?
  21. 21How much of my data should I double-code?
  22. 22Conclusion

What You Need

What You Need

Before you touch a formula, you need four things in place. If any one of them is missing, the coefficient you end up with will not mean much.

  • Two or more coders. One coder coding twice is a test-retest measure, not intercoder reliability. Same person, different days, does not count.
  • A shared set of documents or units. Exactly the same items must be coded by everyone, in the same form. If coder B reads a shortened version, you are measuring two different instruments.
  • A defined codebook. Mutually exclusive categories with written rules for edge cases. Every category needs to be reachable, including the awkward ones.
  • A coding table. One row per unit, one column per coder, one cell per decision. This table is the raw data for every calculation that follows.

Choose your unit of analysis before anything else. A paragraph, a whole interview, a tweet, a 10-second clip of video, a screening decision for a systematic review, a Likert score on a survey item. Each of those is a valid unit, but they give very different reliability numbers, so state yours explicitly in your methods section.

Once the table exists, the statistic choice collapses to two questions: how many coders, and what kind of data did they produce.

StatisticCodersData typeCorrects for chanceHandles missing dataRange
Percent agreement2 or moreNominal, ordinal, numericNoYes0 to 100%
Cohen’s kappaExactly 2Nominal (weighted version for ordinal)YesNo-1 to 1
Fleiss’ kappa3 or moreNominalYesNo-1 to 1
Krippendorff’s alpha2 or moreNominal, ordinal, intervalYesYes0 to 1
Scott’s pi2 or moreNominalYesNo-1 to 1
Intraclass correlation2 or moreNumeric ratingsYesYes-1 to 1

The rule of thumb most researchers settle on: two coders and nominal categories gives Cohen’s kappa, three or more coders gives Krippendorff’s alpha, and any numeric score from a rubric or rating scale gives an intraclass correlation coefficient. Fleiss’ kappa is a reasonable fallback for three or more nominal coders when your team is more comfortable with R’s irr package than with alpha, and Scott’s pi is worth knowing about because it is more forgiving than kappa when a category is very rare.

One caveat on that table. Kappa and Fleiss’ kappa assume every unit received the same number of ratings. Alpha assumes nothing about your design, which is why it is the safer default when coders skip items or a unit pulls out of the study.

Step-by-Step

Step-by-Step

Step 1: Define the unit and write the category rules

Write the codebook before you code anything, not after. For each category give a one-sentence definition, one clear example, one borderline case, and one thing that is explicitly out of scope. The borderline case line is the one that does the real work later, because that is where most disagreements come from.

If your coders are team members who have already talked about the topic, they are not independent in the statistical sense. That does not invalidate the study, but it does mean a high coefficient reflects shared training as much as a clear codebook. Say so in the methods section.

Step 2: Build the coding table

Double-code a sample rather than the whole corpus. A common target is 20% to 25% of units, with a floor of around 50 units for a stable estimate and a ceiling determined by how long discussion of disagreements will take. Below 30 or so units the confidence interval on kappa gets so wide that the number stops being informative, whatever it says.

Randomly select the subset. A convenience sample of the most obvious cases will produce a flattering coefficient that tells you nothing about the hard cases in the rest of your corpus.

Handle multi-coded units explicitly. If a coder may apply several codes to one paragraph, decide in advance whether each code counts as its own binary decision or whether you take the primary code only. Alpha handles the first design naturally; kappa does not, and most tools will silently collapse the multi-code into one cell unless you tell them not to.

Step 3: Compute observed agreement first

Start with percent agreement. It takes seconds, it is the baseline every other coefficient is compared against, and it is the number reviewers check by hand.

Percent agreement = (matching decisions ÷ total decisions) × 100

Using the coding table below, with 40 units coded by two people into three categories:

Coder A Coder BRelevantPartly relevantNot relevantRow total
Relevant173020
Partly relevant212317
Not relevant0213
Column total1917440

The diagonal cells are the agreements: 17 + 12 + 1 = 30 matching decisions out of 40 units.

Percent agreement = (30 ÷ 40) × 100 = 75%

Hold on to that 75%. You will need it in the next step, and it is also the figure that makes the limitations section later make sense.

Step 4: Calculate Cohen’s kappa for two coders

Cohen’s kappa adjusts observed agreement for the agreement two coders would reach by guessing, given how common each category is. The formula is:

kappa = (observed agreement − expected agreement) ÷ (1 − expected agreement)

Expected agreement is computed from the marginal totals only, not from the diagonal. Multiply each row total by its matching column total, add them up, and divide by the square of the total number of units.

Expected agreement = (20×19 + 17×17 + 3×4) ÷ 40²

Expected agreement = (380 + 289 + 12) ÷ 1600 = 681 ÷ 1600 = 0.4256

Observed agreement is the diagonal as a proportion: 30 ÷ 40 = 0.75.

kappa = (0.75 − 0.4256) ÷ (1 − 0.4256)

kappa = 0.3244 ÷ 0.5744 = 0.56

A kappa of 0.56 sits in the moderate band under the conventional thresholds discussed below. Your two coders agree well beyond chance, and they still disagree on 10 of 40 units, which is the number that matters for improving the codebook.

Step 5: Read the same result off a two-category table

When your scheme has only two categories, the table shrinks and the arithmetic gets short enough to do on the back of an envelope. Here are 100 units coded Yes or No.

Coder A Coder BYesNoRow total
Yes62567
No33033
Column total6535100

Observed agreement = (62 + 30) ÷ 100 = 0.92

Expected agreement = (67×65 + 33×35) ÷ 100² = (4355 + 1155) ÷ 10000 = 5510 ÷ 10000 = 0.551

kappa = (0.92 − 0.551) ÷ (1 − 0.551) = 0.369 ÷ 0.449 = 0.82

Notice how high the raw agreement is here, 92%, and how far kappa still sits below 1. That gap is not a flaw in the calculation. It is the whole point of the correction.

Step 6: Calculate Krippendorff’s alpha for three or more coders

Cohen’s kappa is defined only for two coders. With a team of three or more you have two sensible routes: alpha, which is the flexible one, and Fleiss’ kappa, which is the traditional one.

Alpha works by pooling all the pairable decisions any coder made on each unit into a coincidence matrix, then applying the same kind of disagreement-versus-chance logic to the whole thing. Every unit can be coded by a different number of coders, and units with no overlap are simply skipped, which is what makes alpha the right pick when someone skipped two paragraphs.

You do need to tell the software how far apart your categories are. For nominal categories, where Relevant and Not relevant are simply different with no ordering, use the nominal metric. For ordinal categories, where Partly relevant sits between Relevant and Not relevant, use the ordinal metric so that confusing Relevant with Not relevant costs more than confusing Relevant with Partly relevant. For interval scores, use the interval metric. Getting this wrong inflates or deflates alpha without any warning.

The same distance logic applies to Cohen’s kappa. Weighted kappa with linear or quadratic weights gives partial credit for near-misses instead of scoring them as flat disagreements. If your categories are ordered, unweighted kappa is throwing away information you already have.

Step 7: Use an intraclass correlation for numeric ratings

When coders assign a score rather than a category, such as a 1-to-5 rubric score or a 0-to-100 sentiment score, agreement is a correlation problem, not a matching problem. Percent agreement is the wrong tool because two scores of 3 and 4 are close but not identical, and unweighted kappa treats that pair as a flat failure.

An intraclass correlation coefficient (ICC) compares the between-unit variance to the total variance. If units differ a lot and coders score each unit almost identically, ICC approaches 1. The notation carries two decisions: the first number is whether coders are random samples from a wider population of possible coders, the second is whether you’re reporting a single coder’s reliability or the reliability of the average of all coders.

In practice, most teams report ICC(2,k): coders treated as a random sample, and the mean of k raters as the unit of measurement. That answers the question readers usually care about, which is whether the averaged score is stable. If you want the reliability of any single coder instead, use ICC(2,1). Report the model and the type explicitly, because the same data produces meaningfully different numbers under each choice.

Step 8: Calculate it in Excel, SPSS or R

No software is required for a two-coder nominal case; the arithmetic above is short enough to do by hand and auditable that way. Once you have more coders or ordinal data, use something reliable.

Excel. Build the agreement matrix in a block with the row labels down the left and the column labels across the top, then add row totals in the column to the right and column totals in the row underneath. Let N be the grand total in the corner cell.

Observed agreement: =SUMPRODUCT((matrix=TRANSPOSE(matrix))*matrix)/N^2, or more simply =SUM(diagonal_cells)/N

Expected agreement: =SUMPRODUCT(row_totals,column_totals)/N^2

Kappa: =(Po-Pe)/(1-Pe)

If you would rather not build formulas, pivot the coding table into a count matrix with Data > PivotTable, then reference those cells from the three formulas above. Excel has no native kappa function, which is why this route is a two-minute formula rather than a single click.

SPSS. Use Analyze > Descriptive Statistics > Crosstabs, put both coder columns into Cells, tick Row and Column under Statistics, and select Kappa in the Statistics box. For a linear or quadratic weighted kappa, use Data > Weight Cases first with a frequency variable, then request Linear and Quadratic in the same dialog. The syntax version:

CROSSTABS
  /VARIABLES=Coder1 Coder2
  /CELLS=COUNT ROW COLUMN
  /STATISTICS=KAPPA.

SPSS does not ship alpha or ICC for this purpose. For those, use R or one of the dedicated packages.

R. The irr package covers kappa, alpha and ICC in one place.

library(irr)

# Cohen's kappa, two coders, categorical data
kappa2(df$coder1, df$coder2)

# Cohen's kappa, two coders, ordinal data with linear weights
kappa2(df$coder1, df$coder2, weight = "linear")

# Krippendorff's alpha, matrix with units in rows and coders in columns
krippendorffsalpha(as.matrix(rating_matrix),
                   level_of_measurement = "nominal")

# Intraclass correlation for numeric ratings, two-way random, averaged raters
icc(rating_matrix, model = "2", type = "agreement", unit = "average")

Always report the confidence interval alongside the point estimate. A kappa of 0.56 on 40 units carries a much wider interval than the same kappa on 400 units, and a bare number hides that.

Step 9: Document and report the procedure

Reliability is a process, not a checkbox. Write down who coded, how they were trained, what proportion was double-coded, how disagreements were resolved, and whether the codebook changed after the reliability check. If you revised the codebook after seeing disagreements, say so and report the reliability figure from the final round, not the first attempt. That is honest, and it is normal practice.

What Counts as a Good Intercoder Reliability Score?

There is no official cut-off. The thresholds below are conventions published by named authors, and they disagree with each other at the margins. Use them as a shared language with your supervisor, not as a law.

Coefficient valueLandis & Koch (1977)Krippendorff (2004)McHugh (2012)
Below 0.00PoorUnreliableUnreliable
0.00 to 0.20SlightVery low, drop the measureWeak
0.21 to 0.40FairLowFair
0.41 to 0.60ModerateTentative onlyModerate
0.61 to 0.80SubstantialReliable for tentative conclusions at 0.667Strong
0.81 to 1.00Almost perfectReliable for confirmatory conclusions at 0.800Very strong

Two further cautions belong with that table. Hallgren (2012) notes that Landis and Koch’s bands were proposed without empirical support, so the labels are a convenience, not a validated scale. And the thresholds describe what a coefficient means, not whether your study is good enough. Journals in some fields will accept 0.67 for a pilot and reject 0.85 for a confirmatory analysis, because the same number carries different weight depending on the design.

Where thresholds matter most is exploratory work, where the codebook is still moving. Krippendorff’s own advice is to treat alpha below 0.667 as grounds for revising the measure rather than reporting the result, and to hold 0.667 to 0.800 back for tentative findings only.

Common Mistakes

Averaging kappa across items. This is the most common error in the field and it is wrong. If you have six survey items each rated by two people, run kappa once on the pooled coding table, not once per item and then averaged. Averaging kappas treats correlated decisions as independent and can move the reported value in either direction, usually in whichever direction flatters your data. On the r/statistics forum this question comes up constantly, and the consistent answer is: one coefficient, computed over the full cross-tabulation.

Running Cohen’s kappa with three coders. Some tools let you force it. The output is not a valid Cohen’s kappa. Use Krippendorff’s alpha or Fleiss’ kappa.

Using unweighted kappa on ordered categories. If your codebook has a natural sequence such as Strong, Mixed, Weak, the disagreement between Strong and Weak is not the same size as the disagreement between Strong and Mixed. Weighted kappa captures that; unweighted kappa treats both as complete failures and understates reliability.

Never defining the unit of analysis. Intercoder reliability without a stated unit is not comparable to any other study’s figure. Say “paragraph-level coding of 40 interview transcripts” rather than “reliability was assessed”.

Treating a small sample as definitive. Forty units is workable for a pilot. It is thin for a confirmatory claim. Report the confidence interval, and if it spans from moderate to excellent, describe it that way rather than rounding to a label.

Reporting kappa without the procedure behind it. A bare coefficient invites the reviewer question you wanted to avoid. Give the sample size, the unit, the category count, how disagreements were resolved, and the confidence interval. That is five sentences and it closes the whole conversation.

Letting software decide the measurement level. Most tools default to nominal. If your categories are ordinal and you leave the default, you get a number that is too low and no warning.

One more that is easy to miss: agreeing on the categories is not the same as applying them consistently. Some teams report near-perfect agreement on a five-category scheme where one category absorbs most of the decisions. Always report the marginal totals alongside the coefficient so the distribution is visible.

Why kappa sometimes misleads

The number can be low even when your coders look excellent, and understanding why saves a lot of needless worry. It comes down to what chance agreement looks like given how common your categories are.

Look at the 92% agreement example above. Expected agreement was 0.551, which is high, because with only two categories and 65 Yes decisions out of 100, two people guessing independently would agree about half the time no matter what. Kappa divides by what is left after that chance agreement is removed, so a high baseline pushes the coefficient down. This is the prevalence paradox: the rarer a category is, the more impressive raw agreement looks and the lower kappa reports.

The reverse case bites too. With eight balanced categories, random guessing produces almost no agreement, so expected agreement is near zero and kappa rises toward the raw agreement figure even if your coders are getting most units wrong.

This is why reporting percent agreement alongside kappa is not redundancy. The two numbers tell you different things: raw agreement describes your coders, kappa describes your coders relative to the difficulty of the scheme. Where they diverge sharply, the prevalence and bias indices for that table explain the gap in one line, and Artstein and Poesio (2008) is the standard reference for that discussion.

Qualitative researchers take a further step and question whether a single number can represent coding fidelity at all. Haslerig and Grummert’s recent work argues that chasing a coefficient can crowd out the conversation that actually improves a codebook. The practical compromise most teams land on is to do both: calculate the statistic because reviewers and grant reviewers expect one, and treat the disagreement discussion as the real method.

Which also answers the question that comes up more than any other on this topic: how do you handle disagreements? You do not average them away and you do not silently drop the unit. You discuss them, record what the disagreement was about, add or sharpen the rule that would have prevented it, and if the two coders still cannot agree on a unit, bring in a third person to adjudicate and record that decision. That record is more useful to a reader than the coefficient.

Frequently Asked Questions

What is the difference between reliability and validity?

Reliability asks whether two coders apply your categories the same way, which is agreement. Validity asks whether those categories capture what you claim they capture, which agreement cannot establish at any level. You can have two reliable coders agreeing consistently on a scheme that misses the construct entirely. Report reliability as evidence that your procedure is reproducible, and make your case for validity separately through the codebook design and the literature behind it.

When should I use Cohen’s kappa instead of Krippendorff’s alpha?

Use Cohen’s kappa when you have exactly two coders coding into nominal categories and every unit was coded by both. Switch to alpha as soon as a third coder joins, when units have uneven numbers of ratings, when some units were skipped, or when your categories are ordinal or interval and you want alpha to use the distance between them. Alpha handles all of those cases, so it is the safer default once your design stops being the simplest possible one.

Is percentage agreement enough on its own?

It is enough when the categories are few and balanced and you report it alongside the raw disagreement counts. It is not enough when one category dominates, because high raw agreement then tells you very little. In that situation a chance-corrected coefficient such as kappa or alpha separates genuine agreement from what the base rates produce on their own. Reporting percent agreement together with kappa costs you one extra line and pre-empts the most common reviewer objection.

What is a good intercoder reliability score?

Most fields treat 0.61 to 0.80 as substantial or strong and anything above 0.80 as excellent, while Krippendorff recommends 0.667 as the floor for reporting a measure at all and 0.800 before making confirmatory claims. Treat these as conventions rather than rules. A coefficient close to 1 on a two-category yes or no scheme means far less than the same value on an eight-category scheme, so always report the category count and the marginal totals next to the number.

Should I calculate a separate kappa for each survey question?

No. Build one coding table across all the items and compute a single coefficient from it. Computing kappa per item and then averaging those values is statistically wrong, because the per-item coefficients are not independent and averaging them either inflates or obscures the result depending on how the items differ. If you genuinely need item-level detail, report item-level coefficients side by side with the overall value and never summarise them with a single average.

How much of my data should I double-code?

Most teams double-code 20% to 25% of their units, choosing the sample at random rather than picking the easy cases. Aim for at least 50 units, and more if your category structure is complicated, because small samples produce confidence intervals too wide to interpret. Weigh that against the time cost of discussing every disagreement, which is usually the real constraint. A random 20% with a thorough disagreement discussion is more useful than a full second coding pass that nobody had time to think about.

Conclusion

Start by opening a spreadsheet and building the coding table: one row per unit, one column per coder, one cell per decision. Once it is in front of you, the rest is mechanical. Count the coders and name the measurement level, then pick the coefficient that matches: Cohen’s kappa for two coders and nominal categories, Krippendorff’s alpha for three or more or for missing data, an intraclass correlation for numeric ratings. Calculate percent agreement first so the chance-corrected number has something to sit against, report both with the sample size and the confidence interval, and use the disagreements to fix the codebook rather than to argue about the number.

Leave a Comment

Practical guides to statistics, surveys and research data

Read the latest guides