Every estimate comes from a limited number of results, so it carries uncertainty.

Confidence intervals in plain terms#

A confidence interval (CI) shows how precisely an estimate is known. A 95% CI of 2.1–9.9 mg/dL for a mean of 6.0 mg/dL means the data fit a true mean anywhere in that range. More independent results usually narrow it.

A CI describes the estimate. In method comparison, the CI of the mean bias says how well the mean bias is known, and the limits of agreement say where 95% of individual differences fall.

Choosing the confidence level#

Set the confidence level on Set up, in Calculation settings: 90%, 95% or 99%. It applies to every interval in the study.

descriptive.csv holds five results (2, 4, 6, 8 and 10 mg/dL) from five specimens. As a descriptive summary:

Confidence levelMeanSample SDCI for the mean
90%6.0 mg/dL3.16 mg/dL3.0–9.0 mg/dL
95%6.0 mg/dL3.16 mg/dL2.1–9.9 mg/dL
99%6.0 mg/dL3.16 mg/dL−0.5 to 12.5 mg/dL

Only the width changes. The 99% interval goes below zero and is reported as calculated.

Choose the level before you calculate. Changing it later means recalculating.

When an interval can't be calculated#

The interval shows Not estimated, and the study notes and report give the reason. With the same five results:

SituationReason given
Only one resultOne result gives no SD, CV or mean CI
Several rows share a specimen ID (in a bias check: across days or runs, or without a recorded day and run)Results from one specimen are not independent
All results identicalThe SD is zero

Other study types have their own reasons. For example, at 95% confidence a limit of blank study needs at least 72 blank results for the upper confidence limit of the 95th percentile, and method comparison needs at least 2 specimens.

A criterion that depends on a missing interval is Undecided. See Criterion outcomes.

Intervals from counts or ranks#

Intervals from counts or ranks also use the chosen confidence level:

  • Wilson score intervals for agreement percentages, for example 34/37 = 91.9%, 95% CI 78.7–97.2%.
  • Exact rank-based intervals for nonparametric reference limits.

Confidence and population coverage#

Reference interval studies use two different percentages:

  • Target population coverage: the share of the healthy population the interval should contain. With 0.95 it runs from the 2.5th to the 97.5th percentile. See population coverage.
  • Confidence level: how precisely each limit is known, as a CI around each limit.

The example white blood cell study (female partition, 120 subjects), coverage 0.95, confidence 90%:

QuantityValue on screen
Lower reference limit4.21 10⁹/L
90% two-sided CI on the lower limit3.70–5.00 10⁹/L
Upper reference limit11.49 10⁹/L
90% two-sided CI on the upper limit10.30–14.00 10⁹/L

The interval 4.21 to 11.49 should hold 95% of healthy individuals. The CIs show how far each limit could move with another sample of 120.

How n is counted#

Observations are rows. They come from independent units: specimens, subjects, materials, pools, days or runs. Repeats add observations but no units.

Study typeCounted unit
Descriptive summaryEstimates use every included row. A criterion counts distinct specimen IDs.
RepeatabilityEach replicate of the one material at that level
Bias against an assigned valueEach replicate of the one material at that level when all share one recorded day and run; otherwise one material, so no mean or bias interval
Multi-day precisionReplicates within runs within days. The mean CI uses day means.
Method comparison, instrument comparisonSpecimens, with replicates averaged first. In instrument comparison, every specimen counts toward each analyzer's specimen rule, even one with a missing result.
Reagent lot comparisonPatient specimens, with replicates averaged per lot. QC material is excluded. A specimen measured on one lot only still counts toward the specimen rule, as not evaluated.
LinearityReplicates at each level. The fit uses level means.
Reference intervalSubjects. Repeat draws from one subject count once, as their mean.
Qualitative agreementSpecimens, one initial result each
Near-cutoff precisionReplicate results at each concentration
Limit of blankBlank results. Blank materials are counted separately.
StabilityMaterials (paired design) or specimens (independent design) at each time

The count strip and study notes show which count was used. For example, the example bias file has 10 results from 1 control material, all on day 1, run 1, so its mean CI is calculated for that run. The same 10 results spread over five days get no interval.

Why replicates are averaged#

In method, instrument and lot comparison, and for reference interval subjects, the specimen is the experimental unit. Replicates are averaged first, and each specimen mean is one point. Replicates share the specimen's matrix and interferences. Counting them separately would overstate precision and give extra weight to specimens with more replicates.

Which denominator a percentage uses#

PercentageDenominatorNot calculated when
CVThe mean: CV% = 100 × SD / meanInterval scale; fewer than 2 results; mean zero or negative; any negative result
Percent bias against an assigned valueThe assigned value: 100 × (mean − assigned) / assignedInterval scale; assigned value zero or negative
RecoveryThe assigned value: 100 × mean / assignedAs for percent bias
Relative difference, method comparison|comparative result|The comparative result is exactly zero
Mean relative bias criterion, method comparison|comparative result|The comparative result is exactly zero
Percentage specimen limit, method and instrument comparisonEach specimen's comparative (reference-analyzer) result: allowance = P × result / 100The result is zero or negative, or the scale is interval. That specimen is not evaluated.
Percentage limit, reagent lot comparisonEach specimen's previous-lot result: allowance = P × previous / 100As for the specimen limit
Descriptive percentage difference, reagent lot comparisonThe mean of the two lot results: 100 × (new − previous) / ((new + previous) / 2)Interval scale; a negative lot result; pair mean zero
Bias at a medical decision level, percentage limitThe decision levelInterval scale; decision level 0
Overall percent agreementReportable pairsAlways calculated. Nonreportable pairs are listed separately.
All-pairs agreementAll paired specimens, nonreportable counted as non-agreementAlways calculated

In the example qualitative agreement study, 40 specimens were paired and 37 were reportable on both methods. Overall agreement is 34/37 = 91.9% and all-pairs agreement 34/40 = 85.0%. Check the denominator before comparing studies.

See also: Study notes and unavailable results, Set up: study details and settings.