Run this study#
What you are checking: the reference limits for a clearly defined healthy population, measured with your laboratory's method.
Suggested starting plan: for the default two-sided 95% nonparametric interval with 90% confidence on each limit, plan at least 120 eligible subjects per partition, one result each. Two partitions need 240 people and 240 results. This follows Horowitz's reference-interval guidance. With 95% confidence, plan at least 146 per partition.
Calculation minimums: nonparametric two-sided limits need 39 subjects per partition, and their confidence intervals need 119 at 90% confidence. The table under Study setup lists the rest, including 95% confidence.
- Decide first. Create the study. Define the population, eligibility rules, preparation and collection conditions, and any partitions. Enter the analyte and exact unit. On Set up, choose two-sided or one-sided limits, target population coverage, estimator and confidence. Add an acceptance limit only if your laboratory set one in advance.
- Recruit and measure. Collect eligible subjects with the same handling and measurement procedure. Give each person one coded ID.
- Import the cohort. Enter subject ID, partition and result. On Data, choose Import a file and choose your file (start from Blank template (CSV) or reference_establish_120.csv). Answer any question the dialog asks, then choose Import into this study. Check the source records behind flagged values, and exclude only for a documented reason.
- Review the limits. Choose Calculate results. For each partition, check the count, distribution, limits and confidence intervals. A wide confidence interval means the limit is poorly pinned down. On Report, choose Download PDF.
What you get: reference limits with confidence intervals, cohort details and a list of exclusions. Your laboratory decides whether to adopt the limits.
Purpose#
Calculate new reference limits from your own healthy subjects: a two-sided reference interval or a one-sided lower or upper limit. Each limit gets a confidence interval showing how precisely your subjects pin it down.
Two settings are easy to confuse:
- Target population coverage is the share of the healthy population inside the interval, for example 0.95. A two-sided 95 % interval leaves 2.5 % in each tail. A one-sided upper 95 % limit leaves 5 % above it. See population coverage.
- Confidence level, for example 90 %, applies to the interval around each limit. See confidence level.
A one-sided limit still gets a two-sided confidence interval.
When to use it#
Use it when you need new limits from healthy subjects. To check existing limits, use Reference interval verification.
Study setup#
- In your project choose New study. Enter Analyte and Unit, and select Establish a reference interval under Reference intervals.
- Choose Reference limits: two-sided, one-sided lower or one-sided upper. Then choose Create study, or Create and import data if you already have results.
- On Set up, set:
- Target population coverage (0–1), default
0.95. - Estimator: see Estimators and transformations.
- Target population coverage (0–1), default
- Check Confidence level in Calculation settings: 90 % (default), 95 % or 99 %.
- Most new limits need no acceptance limit. To add one, see Acceptance limits.
The design needs:
- Independent healthy subjects who meet your eligibility rules. Each subject counts once. Repeat draws are averaged and flagged.
- Optionally partitions, for example
femaleandmale, each with its own limits. - Enough subjects per partition:
| Setting | Subjects needed | With fewer |
|---|---|---|
| Nonparametric two-sided limits at 95% coverage | 39 | No limits |
| Confidence interval on a nonparametric two-sided 95% limit, 90% confidence | 119 | Limits, but no CI |
| Confidence interval on a nonparametric two-sided 95% limit, 95% confidence | 146 | Limits, but no CI |
| Nonparametric one-sided limit (lower or upper) at 95% coverage and its CI, 90% confidence | 59 | No limit and no CI |
| The same at 95% confidence | 72 | No limit and no CI |
| Nonparametric or parametric limits, guidance count | 120 | Limits, with a caution |
| Parametric or log-normal limits | 2 | No limits |
Other coverage and confidence settings need other counts. The result gives the number.
Entering data#
| Your data | Column in the study |
|---|---|
| Coded subject ID, one per person (no names) | Subject ID, required |
Reference group, for example female | Partition |
| Measured value | Result, required |
| Draw number, for your records | Replicate |
Without a partition column, all subjects are in partition 1. Choose Import a file; see Importing and mapping.
The analysis needs exact values. Correct a censored value such as <0.01, or exclude it with a reason; see Corrections and exclusions.
Fix these before you can calculate:
- a missing analyte or unit;
- a subject entered in two partitions;
- a cohort with no eligible subjects;
- rows marked with errors.
With too few subjects you still get a result, with the count and the reason the limits weren't calculated.
Estimators and transformations#
- Nonparametric percentiles (default): limits at rank p(n + 1) with exact confidence intervals.
- Parametric (normal theory): assumes a normal distribution.
- Log-normal (approximate CIs): assumes the log10 values are normal. It needs two-sided limits, coverage 0.95, confidence 90 %, 95 % or 99 %, and positive values. Limit CIs are approximate.
Data transformation is shown beside the estimator and set by it: None, or Base-10 logarithm for log-normal. Results are reported in the study unit.
Statistics#
- Subject value = mean of that subject's included draws.
- Nonparametric limits: order statistic at rank
p(n + 1)with linear interpolation (Hyndman–Fan type 6). Two-sided:p = (1 − coverage)/2and1 − (1 − coverage)/2. One-sided upper:p = coverage; lower:p = 1 − coverage. Two-sided limits needn ≥ 1/p − 1(39 at 95 %). - Limit CI: exact binomial order-statistic interval, ranks
randswith each tail at mostα/2, achieved coverageP(r ≤ B ≤ s − 1),B ~ Binomial(n, p). If a rank falls outside1…n, there is no CI. Horowitz 2008 gives ranks 1–7 and 114–120 at n = 120, 90 %. - Parametric limits:
mean ± z·SD,z = Φ⁻¹((1 + coverage)/2)two-sided orΦ⁻¹(coverage)one-sided. CI:limit ± t(n−1) × SE,SE = SD·√(1/n + z²/(2(n − 1)))(Bland). Not calculated when all values are identical. - Log-normal limits:
m ± z·son the log10 scale withz = Φ⁻¹(0.975), back-transformed as10^x. Each limit CI has log-scale half-widthΦ⁻¹((1 + confidence)/2) · s · √((1 + z²/2)/n). Achieved coverage is unknown. Values at or below zero, or identical log values, stop the analysis. - Review fences and diagnostics are as in Reference interval verification.
See Methods and sources.
Acceptance limits#
The dialog is in Acceptance limits. For this study:
- Enter your own limit opens the full editor. Under What to check, choose Lower reference limit or Upper reference limit, then enter a Lower limit, an Upper limit or both.
- The limit is compared with the reference limit itself, in the study unit. The CI width doesn't enter.
- A reference limit that wasn't calculated gives Undecided.
See Criterion outcomes.
Worked example#
reference_establish_120.csv: 240 synthetic white blood cell counts (10^9/L), 120 subjects each in female and male, one draw each.
Setup: two-sided, coverage 0.95, Nonparametric percentiles, 90 % confidence. One acceptance limit: Upper reference limit at most 12.5 10^9/L.
What the saved result shows#
| Partition | Subjects | Lower limit (90 % CI) | Upper limit (90 % CI) | Upper reference limit (≤ 12.5) |
|---|---|---|---|---|
| female | 120 | 4.21 (3.70–5.00) | 11.49 (10.30–14.00) | Met (11.49) |
| male | 120 | 4.60 (3.50–5.10) | 10.99 (10.00–11.50) | Met (10.99) |
The study status is Criteria met.
- CI on estimated limits: the lower limit's CI uses order statistics 1 and 7, the upper's 114 and 120. Because ranks are whole numbers, achieved coverage is 92 % against the nominal 90 %.
- Tail estimates: 6 of 120 female values fall outside the limits, against about 6 expected. Male: 5 of 120.
- The review list flagged 3 female values (
F08011.5,F00111.8,F11214; upper fence 11.25) and 1 male value (M03311.5; upper fence 11.2625). They stay in the result.
Excluding a flagged subject#
In a second study, row F112 (14.0) was excluded with a reason. Female result:
| Actual result (included only) | Including excluded (descriptive only) | |
|---|---|---|
| Independent subjects | 119 | 120 |
| Measurement rows | 119 | 120 |
| lower reference limit (10⁹/L) | 4.2 | 4.205 |
| upper reference limit (10⁹/L) | 11.2 | 11.492 |
The upper 90 % CI became 10 to 11.8, and 119 subjects is below the 120 recommended.
A cohort that is too small#
reference_establish.csv has 41 rows: 20 female subjects (F019 has two draws, averaged) and 20 male. With the nonparametric estimator, every limit and CI reads Not estimated, because two-sided limits need 39 subjects. Parametric (normal theory) gives, for example, female 3.23 to 10.76 10^9/L, with 90 % CIs of 1.94–4.52 on the lower limit and 9.47–12.05 on the upper.
On the 120-subject file, a one-sided upper limit at 90 % confidence gave female 10.40, CI 9.60–11.80, and male 10.48, CI 9.70–11.00. Each had 6 of 120 values above.
Reading the results#
Results show one partition at a time. With several, choose the Partition at the top.
Established limits shows Eligible independent subjects and each limit with its two-sided CI. The reason for any "Not estimated" is below the table.
Estimator details (expand it) holds the assumptions, sample-size and tail checks, exclusions and CI ranks.
The CI on a limit. A wide CI means your subjects don't pin the limit down. Nonparametric tails rest on a few subjects, so one extreme value can widen it a lot: the female upper CI, 10.3 to 14, runs to the largest value.
Parametric and log-normal limits assume a normal shape, of the values or of their log10, and this isn't tested. If the Q–Q plot curves or the box is lopsided, compare with the nonparametric result.
Subject results and established limits plots Raw points or a Histogram, with dotted lines at the limits.
Outside 1.5·IQR whiskers — review lists values beyond the Tukey fences, with View row links. Flagged values stay in the result. Check the record behind each flag. To remove a subject, exclude the row with a reason.
Including excluded rows (for information only) appears when a partition has excluded rows. It shows what the exclusions changed.
Distribution diagnostics · partition needs at least 3 subjects whose values are not all the same:
- Normal Q–Q: a straight line suggests a roughly normal shape, a curve suggests skew, and isolated end points are tail values.
- Box and raw-point strip: the box runs from Q1 to Q3 with the median, whiskers reach the most extreme values within 1.5 × IQR, and the strip shows every subject.
Watch for a limit resting on one or two extreme values, many flags on one side, or very different partition limits you intend to report as one.
What to do when the study fails#
- Every limit Not estimated: recruit more subjects in the thin partition. Use parametric limits only if the Q–Q plot and box show no skew.
- Needs ≥ N subjects where a CI should be: the exact CI needs N subjects in that partition. Add subjects.
- Criteria not met: recheck the unit and the acceptance limit.
- Undecided: usually the reference limit wasn't calculated. Add subjects.
- Flags on values that move the limits: check the source records. If a subject doesn't belong, exclude the row with a reason and recalculate.