Every estimate comes from a limited number of results, so it carries uncertainty.
Confidence intervals in plain terms#
A confidence interval (CI) shows how precisely an estimate is known. A 95% CI of 2.1–9.9 mg/dL for a mean of 6.0 mg/dL means the data fit a true mean anywhere in that range. More independent results usually narrow it.
A CI describes the estimate. In method comparison, the CI of the mean bias says how well the mean bias is known, and the limits of agreement say where 95% of individual differences fall.
Choosing the confidence level#
Set the confidence level on Set up, in Calculation settings: 90%, 95% or 99%. It applies to every interval in the study.
descriptive.csv holds five results (2, 4, 6, 8 and 10 mg/dL) from five specimens. As a descriptive summary:
| Confidence level | Mean | Sample SD | CI for the mean |
|---|---|---|---|
| 90% | 6.0 mg/dL | 3.16 mg/dL | 3.0–9.0 mg/dL |
| 95% | 6.0 mg/dL | 3.16 mg/dL | 2.1–9.9 mg/dL |
| 99% | 6.0 mg/dL | 3.16 mg/dL | −0.5 to 12.5 mg/dL |
Only the width changes. The 99% interval goes below zero and is reported as calculated.
Choose the level before you calculate. Changing it later means recalculating.
When an interval can't be calculated#
The interval shows Not estimated, and the study notes and report give the reason. With the same five results:
| Situation | Reason given |
|---|---|
| Only one result | One result gives no SD, CV or mean CI |
| Several rows share a specimen ID (in a bias check: across days or runs, or without a recorded day and run) | Results from one specimen are not independent |
| All results identical | The SD is zero |
Other study types have their own reasons. For example, at 95% confidence a limit of blank study needs at least 72 blank results for the upper confidence limit of the 95th percentile, and method comparison needs at least 2 specimens.
A criterion that depends on a missing interval is Undecided. See Criterion outcomes.
Intervals from counts or ranks#
Intervals from counts or ranks also use the chosen confidence level:
- Wilson score intervals for agreement percentages, for example 34/37 = 91.9%, 95% CI 78.7–97.2%.
- Exact rank-based intervals for nonparametric reference limits.
Confidence and population coverage#
Reference interval studies use two different percentages:
- Target population coverage: the share of the healthy population the interval should contain. With 0.95 it runs from the 2.5th to the 97.5th percentile. See population coverage.
- Confidence level: how precisely each limit is known, as a CI around each limit.
The example white blood cell study (female partition, 120 subjects), coverage 0.95, confidence 90%:
| Quantity | Value on screen |
|---|---|
| Lower reference limit | 4.21 10⁹/L |
| 90% two-sided CI on the lower limit | 3.70–5.00 10⁹/L |
| Upper reference limit | 11.49 10⁹/L |
| 90% two-sided CI on the upper limit | 10.30–14.00 10⁹/L |
The interval 4.21 to 11.49 should hold 95% of healthy individuals. The CIs show how far each limit could move with another sample of 120.
How n is counted#
Observations are rows. They come from independent units: specimens, subjects, materials, pools, days or runs. Repeats add observations but no units.
| Study type | Counted unit |
|---|---|
| Descriptive summary | Estimates use every included row. A criterion counts distinct specimen IDs. |
| Repeatability | Each replicate of the one material at that level |
| Bias against an assigned value | Each replicate of the one material at that level when all share one recorded day and run; otherwise one material, so no mean or bias interval |
| Multi-day precision | Replicates within runs within days. The mean CI uses day means. |
| Method comparison, instrument comparison | Specimens, with replicates averaged first. In instrument comparison, every specimen counts toward each analyzer's specimen rule, even one with a missing result. |
| Reagent lot comparison | Patient specimens, with replicates averaged per lot. QC material is excluded. A specimen measured on one lot only still counts toward the specimen rule, as not evaluated. |
| Linearity | Replicates at each level. The fit uses level means. |
| Reference interval | Subjects. Repeat draws from one subject count once, as their mean. |
| Qualitative agreement | Specimens, one initial result each |
| Near-cutoff precision | Replicate results at each concentration |
| Limit of blank | Blank results. Blank materials are counted separately. |
| Stability | Materials (paired design) or specimens (independent design) at each time |
The count strip and study notes show which count was used. For example, the example bias file has 10 results from 1 control material, all on day 1, run 1, so its mean CI is calculated for that run. The same 10 results spread over five days get no interval.
Why replicates are averaged#
In method, instrument and lot comparison, and for reference interval subjects, the specimen is the experimental unit. Replicates are averaged first, and each specimen mean is one point. Replicates share the specimen's matrix and interferences. Counting them separately would overstate precision and give extra weight to specimens with more replicates.
Which denominator a percentage uses#
| Percentage | Denominator | Not calculated when |
|---|---|---|
| CV | The mean: CV% = 100 × SD / mean | Interval scale; fewer than 2 results; mean zero or negative; any negative result |
| Percent bias against an assigned value | The assigned value: 100 × (mean − assigned) / assigned | Interval scale; assigned value zero or negative |
| Recovery | The assigned value: 100 × mean / assigned | As for percent bias |
| Relative difference, method comparison | |comparative result| | The comparative result is exactly zero |
| Mean relative bias criterion, method comparison | |comparative result| | The comparative result is exactly zero |
| Percentage specimen limit, method and instrument comparison | Each specimen's comparative (reference-analyzer) result: allowance = P × result / 100 | The result is zero or negative, or the scale is interval. That specimen is not evaluated. |
| Percentage limit, reagent lot comparison | Each specimen's previous-lot result: allowance = P × previous / 100 | As for the specimen limit |
| Descriptive percentage difference, reagent lot comparison | The mean of the two lot results: 100 × (new − previous) / ((new + previous) / 2) | Interval scale; a negative lot result; pair mean zero |
| Bias at a medical decision level, percentage limit | The decision level | Interval scale; decision level 0 |
| Overall percent agreement | Reportable pairs | Always calculated. Nonreportable pairs are listed separately. |
| All-pairs agreement | All paired specimens, nonreportable counted as non-agreement | Always calculated |
In the example qualitative agreement study, 40 specimens were paired and 37 were reportable on both methods. Overall agreement is 34/37 = 91.9% and all-pairs agreement 34/40 = 85.0%. Check the denominator before comparing studies.
See also: Study notes and unavailable results, Set up: study details and settings.