Muzaffer BİLGİN, Ertuğrul ÇOLAK
Eskisehir Medical Journal - 2026;7(3):345-355
Introduction: In clinical research, sample size and power are usually calculated from closed -form analytic formulas based on a single assumed effect size. This study maps where analytic and Monte Carlo simulation power converge and diverge across clinical designs, and c ompares assurance (expected power) with classical nominal power. Materials and Methods: Six designs were examined: two -group mean difference, two proportions, one -way ANOVA, logistic regression, survival (log -rank/Cox), and diagnostic accuracy (ROC -AUC). For each, analytical power (pwr, power.prop.test , Hsieh, Schoenfeld, Hanley -McNeil) and simulation power (10,000 repetitions/cell) were compared under matched assumptions, together with assumption -violation scenarios for every design (unequal variance, skewed distribution, non -proportional hazards, non -binormal ROC, rare -event logistic regression, and heteroscedastic ANOVA). For each divergent scenario we additionally computed the correctly -specified analytical method (Welch; Lakatos -type average log -hazard -ratio; Welch ANOVA). The simulation engine's va lidity was verified via Type -I error calibration under H0, and the effect of optimistic effect -size specification on sample size was quantified. All powers are reported with Monte Carlo standard error (MCSE). Analyses used R 4.6.1 with per -cell determinist ic seeds and sessionInfo. Results: Type -I error matched the nominal level in all primary tests (0.046 -0.052; continuity -corrected proportion test conservative at 0.030), with two exceptions identified in this revision: heteroscedastic ANOVA was mildly liberal (0.068) and, under 2:1 allocati on with unequal variances, the pooled t -test was markedly liberal (0.127). Where assumptions held, analytical and simulation power closely overlapped (e.g., two -group d=0.50: 0.801 vs 0.806; logistic OR=1.50: 0.896 vs 0.888; survival HR=0.70: 0.734 vs 0.73 1). Under assumption violation, divergence was marked: with unequal variance, simulation power was 34 -64% below the naïve analytical estimate (d=0.50: 0.801 vs 0.315), and with a delayed treatment effect 79 -83% below the value assuming proportional hazards (HR=0.60: 0.959 vs 0.181). In both cases the correctly -specified method reproduced the simulated power to within Monte Carlo error (Welch 0.312; Lakatos -type 0.187). Under effect -size uncertainty, assurance remained below nominal power, and the gap widene d with uncertainty (prior SD=0.15: 0.801->0.740; SD=0.30: 0.801->0.674), a pattern reproduced in the survival and logistic designs. Overestimating the planning effect by 30% dropped a targeted 0.80 power to ~0.50. Conclusion: When assumptions are satisfied, analytical power calculation is sufficient and simulation -verifiable. When applied outside their stated assumptions, analytical formulas overestimate power in most of the violation scenarios examined - with the exception of the skewed -distribution case at small effect sizes, where the direction of the divergence reversed. The divergences reflect misapplication rather than a deficiency of analytical theory: in each case the correctly -specified formula recovered the simulated p ower. The study provides a reproducible R framework for practitioners.