Nested Case-Control Studies: Efficient Cohort Sampling for Observational Research

On this page
อ่านฉบับภาษาไทย (Thai version)
Abstract
A nested case-control study samples cases and controls from a defined cohort. Standard usage samples controls from each case's risk set. Cohort-based case-control sampling can also use exclusive or inclusive controls, but those schemes should be named explicitly. Controls provide the denominator. Exclusive controls are non-cases at follow-up's end; their exposure odds ratio estimates a disease odds ratio and approaches the risk ratio when disease is rare. Inclusive controls come from baseline and target the risk ratio. Concurrent controls come from each case's risk set. Matched risk-set analysis with conditional logistic regression targets the corresponding hazard ratio; under constant rates, this may be called an incidence-rate ratio. No rare-outcome approximation is required. A pooled unmatched odds ratio is not generally valid. In the example, the cohort risk ratio is 2.50 while illustrative exclusive, inclusive, and crude concurrent cross-products are 3.14, 2.50, and 2.77. Name the control scheme and analysis before interpretation.
Name the control-sampling scheme
Case-control sampling from a defined cohort can use exclusive, inclusive, or concurrent control sampling.
In standard epidemiologic usage, a nested case-control design selects controls from the risk set at each case time. Name the other schemes explicitly, and always report the actual sampling rule.
Controls stand for the denominator
A sampling scheme is the rule for selecting controls from the cohort. Its sampling fraction is the proportion selected from each eligible control group.
Exclusive controls represent survivors at the end. Inclusive controls represent people at baseline. Concurrent controls represent person-time at risk when each case occurs.
Three control-sampling schemes
| Scheme | Controls come from | Common literature name | Interpretation with design-aware analysis |
|---|---|---|---|
| Exclusive or cumulative | Non-cases at the end | Cumulative case-control | Odds ratio; approximates risk ratio when disease is rare |
| Inclusive or case-base | Everyone at baseline | Case-base; a standard case-cohort uses a baseline subcohort sampled without regard to future case status | Risk ratio |
| Concurrent or risk-set | People at risk at each case time | Standard nested case-control | Hazard ratio from matched risk-set analysis; incidence-rate ratio under constant rates |
One cohort, three sampled estimates
The cohort has 10,000 people. Among 2,000 exposed people, 600 become cases. Among 8,000 unexposed people, 960 become cases.
-
Step 1. Calculate the cohort risk ratio
\[ (600 / 2000) / (960 / 8000) = 0.30 / 0.12 = 2.50 \]
The exposed risk is 30%, and the unexposed risk is 12%.
-
Step 2. Sample exclusive controls
\[ (600 × 834) / (960 × 166) = 3.14 \]
A 1,000-person sample of end-of-follow-up non-cases has 166 exposed and 834 unexposed controls.
-
Step 3. Sample inclusive controls
\[ (600 × 800) / (960 × 200) = 2.50 \]
A baseline sample has 200 exposed and 800 unexposed controls.
-
Step 4. Sample concurrent controls
\[ (600 × 816) / (960 × 184) = 2.77 \]
The illustrative controls are constructed in proportion to 8,500 exposed and 37,600 unexposed person-years. Their cross-product ratio is 2.77 after rounding, close to the cohort rate ratio (600 / 8500) / (960 / 37600) = 2.76. In an actual nested case-control study, preserve the matched risk sets and use design-aware analysis. Do not treat a pooled 2 by 2 odds ratio as automatically valid.
Result: The example yields cross-products of 3.14, 2.50, and 2.77 because each control row represents a different denominator. The concurrent value is illustrative; matched risk-set analysis is required in a real nested case-control study.
The illustrative 2.77 crude ratio is close to 2.76 because the controls were constructed from the total person-time proportions.
When case-cohort is the right name
A standard case-cohort design includes a random subcohort selected from the baseline cohort without regard to future case status, plus the incident cases of interest. Its analysis must account for the subcohort sampling.
The same subcohort can be reused across outcomes, but reuse is an advantage rather than a defining requirement. Risk-set sampling instead rebuilds eligible controls at each case time.
Common nested-design errors
-
Using nested as a complete method description
The label alone does not identify the control-sampling scheme.
Fix: State when controls were sampled and from which eligible group.
-
Calling every sampled odds ratio a risk ratio
Exclusive, inclusive, and concurrent controls represent different denominators.
Fix: Map the control source to odds, risk, or rate before interpretation.
-
Calling inclusive sampling standard case-cohort
A standard case-cohort design uses a baseline subcohort sampled without regard to future case status and analysis that accounts for sampling.
Fix: Use case-base for the sampling scheme, and reserve case-cohort for the full design.
References
- Vandenbroucke JP, Pearce N. Case-control studies: basic concepts. International Journal of Epidemiology 2012;41:1480 to 1489. https://doi.org/10.1093/ije/dys147
- Wacholder S, McLaughlin JK, Silverman DT, Mandel JS. Selection of controls in case-control studies. I. Principles. American Journal of Epidemiology 1992;135:1019 to 1028. https://doi.org/10.1093/oxfordjournals.aje.a116396
- Knol MJ, Vandenbroucke JP, Scott P, Egger M. What do case-control studies estimate? Survey of methods and assumptions in published case-control research. American Journal of Epidemiology 2008;168:1073 to 1081. https://doi.org/10.1093/aje/kwn217
- Prentice RL. A case-cohort design for epidemiologic cohort studies and disease prevention trials. Biometrika 1986;73:1 to 11. https://doi.org/10.1093/biomet/73.1.1
- Rodrigues L, Kirkwood BR. Case-control designs in the study of common diseases: updates on the demise of the rare disease assumption and the choice of sampling scheme for controls. International Journal of Epidemiology 1990;19:205 to 213. https://doi.org/10.1093/ije/19.1.205
- Greenland S, Thomas DC. On the need for the rare disease assumption in case-control studies. American Journal of Epidemiology 1982;116:547 to 553. https://doi.org/10.1093/oxfordjournals.aje.a113439
Key takeaways
- Standard nested case-control design uses risk-set sampling; other cohort-based control schemes should be named explicitly.
- Controls represent a denominator, not merely people without disease.
- Exclusive sampling targets an odds ratio and needs rare disease to approximate a risk ratio.
- Inclusive sampling targets a risk ratio; matched concurrent risk-set analysis targets a hazard ratio without a rare-outcome approximation.
- A standard case-cohort uses a baseline subcohort sampled without regard to future case status and appropriate analysis.
Related in the wiki: [[cross-sectional-analogue-designs]] [[balanced-vs-imbalanced-diagnostic]]