Free Tools › Cohen's Kappa Calculator
Measure agreement between two raters or screeners beyond chance. Enter an agreement table or paste two columns of ratings to get kappa with its 95% confidence interval, or weighted kappa for ordered categories.
FreeNo account neededYour data stays in your browser
| Rater 1 ↓ / Rater 2 → | Cat. 1 | Cat. 2 | Cat. 3 |
|---|---|---|---|
| Cat. 1 | |||
| Cat. 2 | |||
| Cat. 3 |
| Cohen's kappa | 0.606 (95% CI 0.462 to 0.751)substantial agreement (Landis & Koch benchmarks) |
|---|---|
| Standard error | 0.0736 |
| Test of κ = 0 | z = 7.69, p < 0.001 |
| Observed agreement | 73.8%59 of 80 items |
| Agreement expected by chance | 33.3% |
| Maximum possible kappa | 0.925Given the raters' marginal totals (unweighted) |
Report the observed percent agreement alongside kappa, and state which weights you used for ordered categories.
An agreement table, or two columns of ratings (e.g. include/exclude).
Unweighted for categories, linear or quadratic for ordered scales.
With its confidence interval and percent agreement.
Cohen's kappa compares the observed agreement between two raters with the agreement expected by chance. A kappa of 1 is perfect agreement and 0 is agreement no better than chance.
Weighted kappa gives partial credit for near-misses on ordered scales: linear weights penalise disagreements in proportion to their distance, and quadratic weights penalise large disagreements more.
Landis and Koch's benchmarks (slight, fair, moderate, substantial, almost perfect) are widely used, but some authors consider them lenient for health research, so report percent agreement too.
Paste the two screeners' include/exclude decisions as two columns, or enter the 2×2 table of agreements and disagreements.
Cohen's kappa is for two raters. For more raters, Fleiss' kappa or the ICC (for continuous ratings) are used.