Run this study#

What you are checking: whether the new method agrees with your comparative method closely enough across the concentrations your laboratory reports.

Suggested starting plan: 40 patient specimens across low, middle and high results, including your medical decision levels, each measured once on both methods: 80 results in 40 rows. To work out the Deming λ from replicates, measure each specimen twice on each method: 160 results in 80 rows, still 40 specimens. Use more specimens for a tighter bias estimate or a wider range. A published CRP comparison used 40 specimens chosen across the range. The calculation minimum is 2 specimens (3 for slope and intercept intervals), and 5 to decide specimen agreement.

  1. Decide first. Choose the allowable difference for each specimen (in the study unit, as a percentage of the comparative result, or both; Use a CLIA / CAP limit also works if it fits your analyte, fluid, method and unit, see CLIA / CAP limits), your medical decision levels, the primary regression and, for Deming, how you will support λ. Create the study and enter these on Set up (Study setup).
  2. Prepare the specimens. Give each a coded ID. Check calibration and QC on both methods, measure each specimen on both within its stability time, and keep every replicate.
  3. Enter the results side by side, repeating the specimen ID on replicate rows. On Data, choose Import a file and choose your file (start from Blank template (CSV) or comparison.csv, replacing its demonstration data). Answer any question the dialog asks, then choose Import into this study.
  4. Calculate and review. Choose Calculate results. Results opens. Check the specimen count and range covered, open Specimen decisions under Results and acceptance to see which specimens are within their limits, then check decision-level bias and its interval. Check zero comparator results under combined limits. Check unexpected results against the original records; excluding a result asks for a reason. On Report, choose Download PDF.

The specimen limit is met when at least 95% of specimens, and at least 5, are within their allowable difference. Bias at a decision level outside your specimens' range stays Undecided.

Purpose#

A method comparison measures the same patient specimens on a new or changed method and on the method you already trust, to show how far apart the results are and whether the difference changes with concentration. Use it for a new analyzer, a move to another platform or checking a referral laboratory's method.

The test method is the y-axis and the comparative method the x-axis. Every difference is test minus comparative: a positive bias means the test method reads higher. The template and exported files call them candidate and comparator.

When to use it#

Use it for two methods, with both results for a specimen on one row. Use a different study for:

Study setup#

The experimental unit is the specimen (definition). Rows with the same specimen ID are averaged into one pair. Replicates can support the Deming ratio.

CountWhat happens
Fewer than 2 specimens with included pairsNo calculation.
Fewer than 3 specimensNo slope or intercept intervals.
Fewer than 5 specimensThe specimen limit is Undecided.
Fewer than 40 specimensA warning suggests at least 40 across the measuring interval.
  1. Choose New study. Enter Analyte and Unit, select Method comparison under Comparison and bias, and fill in Comparative method or instrument (x-axis) and Test method or instrument. Choose Create study, or Create and import data if you already have results.
  2. On Set up, enter up to 10 Medical decision levels separated by semicolons, for example 70; 126; 200. With a decimal comma, write 5,5; 12.
  3. Under Acceptance limits, choose Enter your own limit. Under What to limit, pick Specimen difference or Bias at a medical decision level (with its own Medical decision level), then enter the limit in the study unit, as a percentage, or both (both parts). More statistics adds slope, intercept, mean bias and mean relative bias. See Acceptance limits.
  4. In Method comparison, choose the Primary regression. A Deming fit reveals How λ is supported (below).

Entering data#

On Data, choose Import a file, or Add row and type or paste into the grid. Blank template (CSV) downloads the columns this study needs.

ColumnRequired?What to enter
Specimen IDYesCoded ID, repeated on every replicate row of that specimen
Comparative method (x)YesComparative result, in the study unit
Test method (y)YesTest result for the same specimen
Specimen group (level)NoA group label such as outpatient, the same on every replicate row. It appears in each plot point's label
ReplicateNoFor your records
Excluded, Exclusion reasonNoIn an imported file: yes/no, true/false or 1/0, with a reason

Import recognizes common headers such as comparator, reference or x, and candidate, test or y. See Importing a file. Correct results such as <5, or exclude them with a reason. See Missing and censored values.

These problems stop the calculation:

ProblemWhat to do
The primary regression can't run, for example weighted Deming with a zero value, or a Deming ratio that is not a positive numberChange the Primary regression, or correct how λ is supported in Method comparison on Set up.
A medical decision level is too far from your specimens for a bias intervalCorrect or remove that level.
A precision study linked for λ has changed since the last calculationSee Linked precision studies.

Statistics#

  • Specimen agreement: d = test − comparative for each specimen mean pair. The allowance at comparative result X is A, or P × X / 100, or with both limits the larger, max(A, P × X / 100) (Pass when within either (the larger allowance)) or the smaller, min(A, P × X / 100) (Pass only when within both), as chosen under With both limits. A specimen is within its limit when |d| ≤ allowance, inclusive, compared in exact decimal arithmetic on unrounded values.
  • Ordinary least squares of y on x, with Student t intervals on n − 2 df. Pearson r is descriptive.
  • Deming: slope = (Syy − λSxx + √((Syy − λSxx)² + 4λSxy²)) / (2Sxy), with λ as defined under How the Deming ratio is supported. Weighted Deming follows Linnet (1990) with constant-CV weights, as in the R package mcr. It needs positive values and stops after 30 iterations or at tolerance 1e-6. Both Deming fits use jackknife standard errors with t(n − 2) intervals.
  • Passing-Bablok (1983, 1984): shifted median of pairwise slopes, dropping slopes of exactly −1. Intercept = median of y − slope × x. The slope interval is rank-based. There is no intercept interval when comparative values span zero.
  • Replicate-derived λ = pooled within-specimen variance of y ÷ that of x, over specimens with at least 2 pairs. Linked λ = (test repeatability SD ÷ comparative repeatability SD)².
  • Differences: mean bias with a t CI on n − 1 df. Limits of agreement = mean ± 1.96 SD, each with a CI using SE = √(3s²/n) (Giavarina 2015).
  • Bias at decision level X = intercept + (slope − 1) × X, from the primary fit. Its interval is a Student t interval of the fitted response (OLS), a jackknife interval (Deming, weighted Deming) or the range over the corners of the slope and intercept intervals (Passing-Bablok). A level outside the range of comparative specimen means is extrapolated, and any limit at that level is Undecided.
  • Fitted-response band: OLS ŷ ± t × s × √(1/n + (x − x̄)²/Sxx). Deming fits use the jackknife. The band is pointwise.

Sources: Methods and sources.

The four regressions and what each needs#

All four fits use the specimen means. The limits, decision-level table, residual plot and band use the Primary regression. A fit that can't run shows why in the table.

Primary regressionUse it whenNeeds
Passing-Bablok · robust (default)Both methods have similar error and you want resistance to a few odd specimens.A positive, roughly linear relationship
Deming · constant SDBoth methods have roughly constant error SD.λ
Weighted Deming · constant CVError grows in proportion to concentration.λ, and all values above zero
Ordinary least squaresThe comparative method's error is negligible.Nothing extra

How the Deming ratio is supported#

For a Deming or weighted Deming primary fit, λ is the test method's error variance divided by the comparative method's. λ = 1 means both methods are equally imprecise.

How λ is supportedWhat happens
Computed from replicate pairs in this studyNeeds at least 2 specimens with at least 2 replicate pairs each, or neither Deming fit runs.
Taken from linked precision studiesUses multi-day precision or repeatability studies of the same analyte and unit in this project.
Stated assumptionUses the Assumed error-variance ratio λ you enter.
Not chosenUses Deming error-variance ratio λ, default 1.

With a Deming primary fit and Not chosen, or with a stale link, the slope, intercept and decision-level bias limits are Undecided and no band is drawn. Mean-bias and specimen limits are decided as usual.

Linked precision studies#

The Test method (y) and Comparative method (x) pickers list precision results for this analyte and unit in this project, one entry per level. Refresh the list of saved precision results reloads them. Link a different study to each method, at the level nearest your decision concentrations. The level needs an SD above 0.

A link goes stale when the linked study gets new data or a new calculation. The link keeps the result you chose and never moves to the newer one by itself. A new calculation still uses that earlier result: it warns that the linked study changed, the Deming limits above are Undecided and no band is drawn. A calculation with a stale link can't be signed off, and a calculation saved while the link was current can't be returned again once it goes stale. To decide the Deming limits, choose the current result in Method comparison on Set up, save and calculate again.

Acceptance limits for this study#

The specimen rule decides agreement: at least 95% of specimens, rounded up, must be within their limits. With 5–19 specimens all must be, with 20 specimens 19, with 25 specimens 24. A percentage limit needs a comparative result above zero on a ratio scale. A specimen that can't be evaluated leaves the limit Undecided unless enough specimens already fail: with 20 specimens, 1 failure and 1 not evaluated is Undecided, 2 failures and 1 not evaluated is Not met. See how outcomes are decided.

Without a specimen rule, a built-in Specimen agreement check stays Undecided and the study can't pass. The other limits check point estimates from the primary fit.

  • Slope limits are in ratio, Mean relative bias (%) limits in %, and every other limit in the study unit. A limit in another unit is Undecided. Specimen rules, including percentage rules, use the study unit.
  • A Mean relative bias (%) limit needs Show relative differences and allow relative-bias limits ticked under Calculation settings. Otherwise it is Undecided.
  • A decision-level limit sets its own Medical decision level, separate from the list on Set up.
  • A concentration-dependent specimen rule picks its segment by each specimen's comparative result, and decision-level bias by the decision level. A specimen or level in a gap is not evaluated. See Concentration-dependent limits.

Worked example#

Download comparison.csv: 20 synthetic glucose specimens, one pair each, comparative results 52.1 to 365.9 mg/dL.

SettingValue
Analyte, unitGlucose, mg/dL
Comparative method (x) / Test method (y)Laboratory analyzer B / Chemistry analyzer A
Confidence level95%
Primary regressionPassing-Bablok
How λ is supportedNot chosen (λ = 1)
Show relative differencesTicked
Medical decision levels70; 126; 200; 400

The specimen rule allows 6 mg/dL or 8% of the comparative result, whichever is greater. It becomes the segments 0,75,absolute,6 and 75,,relative,8. The 200 mg/dL limit uses the segments 0,100,absolute,4 and 100,,relative,4. All limits are inclusive.

Results as the screen shows them:

EstimateValue95% confidence interval
Mean bias, test minus comparative+3.80 mg/dL3.14–4.47
SD of differences1.415 mg/dL
95% limits of agreement1.03–6.58 mg/dLlower −0.12 to 2.18; upper 5.43–7.73
Mean relative bias (|comparative result| denominator)+2.4%2.0–2.7
RegressionSlope (interval)Intercept, mg/dL (interval)
Ordinary least squares (y on x)1.014 (1.012–1.016)+1.281 (0.887–1.676)
Deming1.014 (1.012–1.016)+1.280 (0.924–1.636)
Weighted Deming (constant CV)1.013 (1.011–1.014)+1.482 (1.259–1.704)
Passing-Bablok · primary1.014 (1.012–1.016)+1.390 (1.074–1.688)

Pearson r is > 0.9999.

Medical decision level (mg/dL)Estimated bias (mg/dL)IntervalInside studied range
70+2.351.88–2.77Yes
126+3.122.52–3.64Yes
200+4.143.37–4.79Yes
400+6.895.67–7.89No (extrapolated, can't pass)
Limit nameLimitObservedOutcome
Specimen differencesEach specimen within ±6 mg/dL or 8%, whichever is greater20/20 (100%)Met
Slope0.95–1.051.014Met
Intercept±5 mg/dL+1.39 mg/dLUndecided: 0 mg/dL is outside the studied range of 52.1 to 365.9 mg/dL
Bias at 126 mg/dL±6 at 126 mg/dL+3.12 mg/dLMet
Bias at 200 mg/dL±8.00 at 200 mg/dL (concentration-dependent: ±4 mg/dL below 100, ±4% from 100)+4.14 mg/dL · |Bias| 2.1%Met
Mean bias±5 mg/dL+3.80 mg/dLMet

The study is Undecided because the intercept limit is undecided. Warnings flag the 20-specimen count and the extrapolated 400 mg/dL level.

The test method reads about 3.8 mg/dL higher, and the bias grows with concentration.

Deming contrast. The same data with a Deming primary fit, decision level 126 mg/dL, the specimen rule above and three more limits: Slope 0.95 to 1.05, Bias at 126 mg/dL within ±6 mg/dL (estimate +3.04, interval 2.83–3.24) and Mean bias within ±5 mg/dL. With Not chosen (λ = 1), the slope and bias limits are Undecided and so is the study. With Stated assumption (λ = 1), every limit is Met and the study status is Criteria met.

Replicate-derived λ. Three specimens with three replicate pairs each:

specimen_id,comparator_result,candidate_result
A,10,10
A,11,12
A,12,14
B,19,21
B,21,23
B,20,25
C,30,33
C,32,34
C,34,35

Computed from replicate pairs in this study gives λ = 1.5 from 3 specimens, with 6 df per method. The within-specimen SDs are 1.7321 mg/dL for the test method and 1.4142 mg/dL for the comparative method. The 20-specimen file has no replicates, so this basis stops a Deming calculation on it.

Reading the results#

Regression plot. One point per specimen, mean test against mean comparative. The dashed line is the primary fit across the observed range. The dotted line is y = x, and points above it read higher on the test method. Select a point to jump to its row. The strip under it gives the specimen count and the range covered.

Slope and intercept by regression method lists all four fits, each with its interval under it (for example 95% CI 1.012–1.016), marks the primary one and gives Pearson r, with the λ used under it. Pearson r prints as > 0.9999 when it rounds to 1. The slope has no unit. Bias at medical decision levels gives each level's estimated bias, its interval and whether the level is inside the studied range.

Results and acceptance. Each limit shows its observed value and outcome, as in Why under the verdict. Specimen decisions below the table lists each specimen's results, difference, allowed difference, error index and outcome. The error index is the difference divided by its allowance, so −1 to +1 is within a symmetric limit.

Difference plot and summary. Test minus comparative for each specimen, against the comparative result, with the mean bias and 95% limits of agreement. The key gives their values in the plot's unit, for example Dashed line: mean bias +3.80 mg/dL and Dotted lines: 95% limits of agreement +1.03 to +6.58 mg/dL. The CI of the mean bias shows how well the average is known. The limits show where about 95% of individual differences fall, so a narrow CI with wide limits means a well-known average but large scatter. Tick Show relative differences and allow relative-bias limits under Calculation settings to plot 100 × (test − comparative) ÷ |comparative result| instead, skipping exact zeros. Save and recalculate to apply it.

Allowable-difference marks. Short marks at each specimen show mean-bias, mean-relative-bias and concentration-dependent decision-level limits. In the example they show the 200 mg/dL limit (±4 mg/dL below 100 mg/dL, ±4 % from 100 mg/dL up) and the ±5 mg/dL mean-bias limit. The specimen rule's limits are in Specimen decisions.

Residual diagnostic. Test result minus the primary fitted value, against the comparative result. Curvature or a widening spread means the fit or its error model is a poor match for the data. A pointwise band at your confidence level, shaded on the regression plot, shows how well the line is known. It appears for an OLS fit, or a Deming or weighted Deming fit with a current λ basis, with at least 3 specimens.

If the study fails or shows warning signs (bias near or beyond your limit, limits of agreement wide compared with clinical tolerance, a trend or fan shape in the difference or residual plot, slopes that differ between fits, isolated points), check for swapped columns, a wrong unit, a reused specimen ID or typing errors. If the data are right, the difference is real. Where the range is thin, add specimens and recalculate.

Common mistakes#

SymptomCorrection
Slope is the reciprocal of what you expected, or bias has the wrong signThe methods are swapped. Comparative is x, test is y.
Some specimens can't be evaluatedUse an absolute or greater-of limit, or extend the concentration segments to cover every specimen.
A copied limit shows Not applicableOpen it with Edit and choose Update limit.
Intercept limit always UndecidedThere are no specimens near zero. Check decision-level bias instead, or add low specimens.
Percentage statistics or weighted Deming can't be calculatedThey need a ratio scale with a meaningful zero (Measurement scale under Calculation settings). Use absolute differences on an interval scale.