Run this study#

What you are checking: whether two methods give the same category on the same specimens, and where they disagree.

Suggested starting plan: plan comparator-positive and comparator-negative specimens separately, after choosing the agreement you expect, the confidence level and the lowest acceptable lower confidence bound. For 95% expected agreement and a 90% lower-bound goal, test 200 positive and 200 negative specimens on both methods: 400 specimens, 800 results. With 190 of 200 agreeing in each group, the 95% Wilson lower bound is about 91.0%. Fewer agreements may miss the goal. Use the Wilson calculation to plan other goals. With three or more categories, include enough specimens in each category and set the limit on overall agreement.

Calculation minimum: one included pair gives a result. A limit is decided only when its group has at least 2 specimens: the reportable pairs for overall agreement, the comparator-positive specimens for positive agreement, the comparator-negative specimens for negative agreement. Otherwise it is Undecided.

  1. Decide first. Create the study. On Set up, declare the categories, the positive category, the nonreportable labels and whether the comparator is a true reference standard. Under Acceptance limits, choose Enter your own limit and set the smallest acceptable agreement or lower confidence bound. Choose specimens that represent your patients, including difficult ones near the cutoff.
  2. Test each specimen on both methods, following each method's instructions and stability limits. Keep the first result even if it leads to a repeat.
  3. Import the result pairs in the order you tested them, with any retest on a later row. On Data, choose Import a file and choose your file (start from Blank template (CSV) or qualitative_agreement.csv). Answer any question the dialog asks, then choose Import into this study. The first included row for each specimen is the one analyzed.
  4. Review both sides. Choose Calculate results. Check positive and negative agreement with their intervals, the denominators, the nonreportable counts and the disagreements. A nonreportable first result makes the affected limits Undecided. On Report, choose Download PDF.

Purpose#

Each specimen is tested by your test method and a comparative method. Each gives a category such as positive or negative, or a grade such as 1+.

What the percentages mean depends on the Comparative method type you choose:

  • Non-reference comparative method · percent agreement (default): positive percent agreement (PPA), negative percent agreement (NPA) and overall percent agreement. They describe agreement with that method only. See PPA and NPA.
  • Reference standard · sensitivity/specificity: for a true reference standard. The same formulas are then labelled Sensitivity and Specificity. See sensitivity and specificity and the FDA guidance.

When to use it#

Use it when each specimen has one category from each method. For repeated testing at known concentrations around a cutoff, use Near-cutoff precision. For numeric results, use Method comparison.

Study setup#

  1. In your project choose New study. Enter the Analyte (this study needs no unit), select Qualitative agreement under Qualitative tests, and choose Create study, or Create and import data if you already have results.
  2. On Set up, under Qualitative performance (Study mode is already Categorical agreement (test vs comparative)), fill in:
    • Result categories · declared order: default negative, positive. Two categories give a 2×2 table with PPA and NPA. Three or more give a larger table, listed lowest first when ordered.
    • Positive category: one of the declared categories, used for two-category tables.
    • Nonreportable labels: default equivocal, invalid, indeterminate. These results stay in the tables but leave the reportable block.
    • Comparative method type, and Category order: Unordered (nominal) or Ordered (semiquantitative). Ordered categories add adjacent agreement and Weighted kappa: None, Linear agreement weights or Quadratic agreement weights.
  3. Under Calculation settings, Confidence level (95 % by default, or 90 % or 99 %) sets the Wilson and kappa intervals.
  4. Under Acceptance limits, choose Enter your own limit (see Acceptance limits for this study).

Entering data#

Choose Import a file, or type or paste into the grid. Blank template (CSV) downloads the empty column template.

ColumnRequired?What to enter
Specimen IDYesCoded specimen ID, one per specimen
Comparative method resultYesA declared category or nonreportable label
Test method resultYesA declared category or nonreportable label
ReplicateNoNot used. The first included row for each specimen is analyzed, in entry order, and later rows are repeats.

Labels ignore case and surrounding spaces. Any other value, such as <0.5, appears on Set up under Labels found in the data. Choose Add to categories or Add to nonreportable there, or correct the row. To leave a row out, exclude it with a reason; see Corrections and exclusions.

Statistics#

  • Binary layout (rows test, columns comparative, FDA 2007 guidance Table 4): a = test+/comp+, b = test+/comp−, c = test−/comp+, d = test−/comp−. PPA (or sensitivity) = a/(a + c), NPA (or specificity) = d/(b + d), overall = (a + d)/n over reportable pairs.
  • Wilson score interval without continuity correction: centre (p + z²/2n)/(1 + z²/n), half-width z·√(p(1 − p)/n + z²/4n²)/(1 + z²/n), z = Φ⁻¹((1 + confidence)/2).
  • Cohen kappa κ = (pₒ − pₑ)/(1 − pₑ). Weighted kappa: linear weight 1 − |i − j|/(k − 1), quadratic 1 − (i − j)²/(k − 1)².
  • Kappa SE: Fleiss, Cohen & Everitt (1969) non-null large-sample variance (Eq. 8, or Eq. 13 unweighted). CI κ ± z·SE, without truncation. No SE with zero variance or fewer than 2 pairs.
  • McNemar exact p = min(1, 2·BinomCDF(min(b, c); b + c, ½)).
  • Paired difference 100·(b − c)/n with Newcombe (1998) method 10 interval.
  • Resolved results: each applied entry replaces that specimen's first pair in a copy of the data, which is analyzed the same way without acceptance limits.

See Methods and sources.

Acceptance limits for this study#

The dialog and its choices are in Acceptance limits. For this study:

  • Agreement percentages use reportable pairs. A nonreportable test result makes overall agreement and agreement in that specimen's comparator category Undecided. A nonreportable comparator result makes overall, positive and negative agreement Undecided. Limits on a lower confidence bound follow the same rules. The FDA guidance, section 6.2 explains why.
  • Positive and negative limits are Not applicable to tables with more than two categories.
  • Kappa, McNemar and the paired difference are for information.

Worked example#

Download qualitative_agreement.csv: 42 synthetic urine hCG rows for 40 specimens. Two specimens have a repeat row at the end. Three first results are nonreportable: Q015 (test equivocal), Q026 (comparative equivocal) and Q035 (test invalid).

Setup: default categories and labels, Positive category positive, Non-reference comparative method · percent agreement, 95 %. Acceptance limits: overall agreement at least 90 %, and positive agreement lower confidence bound at least 80 %.

Test ↓ / comparative →negativepositiveNonreportableTotal
negative17 (d)1 (c)119
positive2 (b)17 (a)019
Nonreportable1102
Total2019140
AgreementCount and percent95 % Wilson CI
Overall percent agreement34/37 = 91.9%78.7–97.2%
Positive percent agreement (PPA)17/18 = 94.4%74.2–99.0%
Negative percent agreement (NPA)17/19 = 89.5%68.6–97.1%
All-pairs agreement (nonreportable counted as non-agreement)34/40 = 85.0%70.9–92.9%
SupplementValue
Cohen kappa (unweighted)0.838, SE 0.090, 95 % normal CI 0.662 to 1.013
McNemar exact p (b = 2, c = 1)1.000
Positive-proportion difference (test − comparative)2.703 percentage points, 95 % paired CI −7.431 to 12.669, reportable pairs (n = 37)
Limit nameLimitObservedOutcome
Overall agreement≥ 90%91.9% among reportable pairsUndecided. Three specimens have a nonreportable first result.
Positive agreement / sensitivity, lower bound (%)≥ 80%74.2% among reportable pairsUndecided. Q015 (comparative positive, test equivocal) and Q026 (comparative equivocal) affect it.

The study is Undecided.

The discordant list shows Q011 (test negative / comparative positive), Q022 and Q030 (test positive / comparative negative) and the three nonreportable pairs. Q015's repeat reads positive, but the analysis uses its first result, equivocal.

With Reference standard · sensitivity/specificity, the numbers are the same and the rows read Sensitivity (17/18 = 94.4%) and Specificity (17/19 = 89.5%).

Ordered categories and weighted kappa#

Download qualitative_ordinal.csv: 30 synthetic urine protein grades negative, 1+, 2+, 3+. Setup: Ordered (semiquantitative), Linear agreement weights, acceptance limits overall agreement at least 80 % and positive agreement at least 90 %.

ResultSaved value
Exact (overall) agreement23/30 = 76.7% (95 % Wilson CI 59.1–88.2%)
Adjacent-category agreement (|i − j| ≤ 1)30/30 = 100.0% (88.6–100.0%)
Cohen kappa (unweighted)0.685, SE 0.104, 95 % normal CI 0.482 to 0.888
Linear weighted kappa0.805, SE 0.068, 95 % normal CI 0.671 to 0.939
Overall agreement (limit ≥ 80%)Not met (76.7%)
Positive agreement / sensitivity (%) (limit ≥ 90%)Not applicable. The table has 4 categories.

The study status is Criteria not met.

With Quadratic agreement weights, weighted kappa is 0.899 (95 % CI 0.823 to 0.975). Every disagreement here is one grade apart, so adjacent agreement is 100 %. The discordant list has a Distance column, the test grade minus the comparative grade.

Reading the results#

Contingency table. Counts each specimen by its first result. Percentages come from the reportable block. The Nonreportable row and column hold the rest.

Two overall figures. Overall percent agreement uses reportable pairs. All-pairs agreement counts nonreportable specimens as disagreement. A large gap means nonreportable results matter. Report both.

PPA and NPA (two categories only). Small denominators give wide intervals. Read the Wilson lower bound.

Kappa. Agreement beyond chance, on a scale from −1 to 1. Its large-sample CI can pass 1, as in 1.013 above, and is rough in small tables. With perfect agreement there is no SE or CI. Weighted kappa gives partial credit for near misses. Quadratic weights forgive one-grade differences more.

McNemar exact p (two categories). Tests whether the disagreements lean one way, b against c.

Positive-proportion difference (two categories). A positive value means the test calls more positives.

Repeat results lists later rows for the same specimen and whether they differ from the first. They are not counted. To use a retest, record it as a resolved result.

Look out for a low lower bound on PPA or NPA, many nonreportable results, disagreements all in one direction, or a "reference standard" that is not really one.

If the study fails, check for swapped columns, the wrong Positive category and undeclared labels. Then retest or adjudicate the discordant specimens. Against a comparative method that isn't a reference standard, a low PPA or NPA means the two methods disagree. Either one may be wrong.

Resolved results (retest or adjudication)#

Each specimen can have one resolved result, from a retest or from adjudication by a third method. It never changes the main table or the acceptance limits, which use first results.

  1. In the Discordant and nonreportable pairs list, choose Record resolved result on the specimen. Or choose Record resolved result in the Recorded resolved results panel and pick the specimen.
  2. Choose Resolved comparative result, Resolved test result and Method (Retest or Adjudication). Choose Save resolved result. A new entry replaces the earlier one for that specimen.
  3. Choose Recalculate results to apply them.

The Resolved results — secondary panel shows the tables recalculated with them and marks each entry as applied or not applied, with a reason.

Resolving only discordant or nonreportable specimens biases the secondary table, and a warning says so (discrepant resolution, FDA 2007 guidance). It clears once you resolve at least one concordant specimen.

Example: retest Q011 as positive on both methods. The main table stays a/b/c/d 17/2/1/17. The secondary is 18/2/0/17, with overall agreement 35/37 = 94.6% (95 % Wilson CI 82.3–98.5%), PPA 18/18 = 100.0% (82.4–100.0%), NPA unchanged at 17/19 = 89.5%, and all-pairs agreement 35/40 = 87.5% (73.9–94.5%). The discrepant-resolution warning appears.

The report shows them in section 1.2, Resolved results (secondary).

Common mistakes#

SymptomCorrection
Report says PPA/NPA but you meant sensitivity/specificityChoose Reference standard · sensitivity/specificity only for a true reference standard.
High agreement but a lower-bound limit is not metThe sample is too small. Test more specimens.