Which Reporting Standard Applies to Your Study?
Four frameworks govern how medical statistics are written up, and they are not competitors — they stack. APA 7th edition governs the numeric notation itself (italics, decimal places, symbol order) and applies to virtually every manuscript regardless of design. On top of that, one design-specific checklist applies depending on what kind of study you ran.
| Framework | Governs | Applies To |
|---|---|---|
| APA 7th Edition | Statistical notation style: italics, decimals, symbol order, effect size and CI reporting | Every manuscript |
| CONSORT | Trial reporting: primary/secondary outcome effect sizes with CIs, per-group results, participant flow | Randomized controlled trials |
| STROBE | Observational reporting: unadjusted and adjusted estimates, confounder handling, missing data | Cohort, case-control, cross-sectional studies |
| PRISMA | Synthesis reporting: pooled effect sizes, heterogeneity statistics, study flow diagram | Systematic reviews and meta-analyses |
Every template in this guide follows APA numeric notation by default. Where CONSORT or STROBE adds a design-specific expectation for a particular test — reporting both the unadjusted and adjusted effect in a STROBE cohort study, for instance — it is called out in that test's explanation. PRISMA is included for completeness because many thesis committees and journals still ask which guideline you consulted even when your own study is a single trial or cohort rather than a synthesis; if any part of your work involves pooling results across studies (a meta-analysis chapter, for example), the same statistic-df-p-effect size-CI structure used throughout this guide still applies to each pooled estimate, on top of PRISMA's flow-diagram and heterogeneity-reporting requirements.
All four frameworks are maintained or indexed by the EQUATOR Network, the standard reference point medical journals point authors to when reporting requirements are not spelled out in their own author instructions. Many journals now require the relevant checklist (a completed CONSORT or STROBE checklist, for instance) to be uploaded as a supplementary file alongside the manuscript, so it is worth identifying which framework applies before you start drafting your Results section, not after.
Reporting Descriptive Statistics
Descriptive statistics open every Results section and describe your sample before any hypothesis test appears — this becomes your Table 1. Use normality testing to decide between the two template forms below.
Normally Distributed Continuous Variables
"The mean age of participants was 44.2 ± 13.6 years (range 19–74)."
"Age: 44.2 (13.6)" — reported with no label identifying which figure is the mean and which is the SD.
"Mean age was 44.2 ± 13.6 years." The ± symbol (or explicit "SD =") removes any ambiguity about what the second number represents.
Non-normal continuous variables should never be summarized with mean and SD alone — report median and interquartile range (IQR) instead: "Median length of stay was 4 days (IQR 3–7)." APA 7 permits either format but requires you to state explicitly which one you used and why (i.e., the result of your normality test). For a STROBE-compliant cohort or cross-sectional study, this Table 1 should also report the number and percentage of missing values for every variable, since undisclosed missingness is one of the most frequently cited STROBE deficiencies in peer review. Categorical baseline variables follow the same logic in reverse: report n (%) for every category, using Valid Percent once any missing data exists, exactly as covered in our SPSS output interpretation guide.
Comparing Two Groups
These four tests compare an outcome between two groups — the choice between them depends on whether groups are independent or paired, and whether the outcome is normally distributed.
Independent Samples T-Test
"Mean hemoglobin was significantly higher in males (13.6 ± 1.7 g/dL) than females (11.8 ± 1.6 g/dL); mean difference 1.8 g/dL, 95% CI [1.06, 2.54], t(83) = 4.87, p < 0.001, d = 1.07."
"There was a significant difference between males and females (p = 0.000)." No test statistic, no df, no effect size, and a p-value SPSS never actually produces.
Report the full statement above: t(83) = 4.87, p < 0.001, with the mean difference, its 95% CI, and Cohen's d, not a bare p-value.
The p-value alone cannot tell a reader whether a 1.8 g/dL difference in hemoglobin is clinically meaningful — the mean difference, its CI, and Cohen's d together do that job, which is why all three appear in the template above rather than the p-value in isolation. Use the independent t-test calculator to generate this output directly. For an RCT comparing two treatment arms, CONSORT explicitly requires the between-group difference and its CI to be the headline figure for both the primary and secondary outcomes — not the p-value alone, and not merely "significant" or "not significant" as a verdict.
Paired Samples T-Test
"Systolic blood pressure decreased significantly after the intervention, from 148.2 ± 14.1 mmHg to 135.8 ± 13.3 mmHg (mean decrease 12.4 mmHg, 95% CI [9.85, 14.95]), t(59) = 9.80, p < 0.001, d = 1.26."
"Blood pressure improved significantly after treatment (r = 0.71, p < 0.001)" — quoting the Paired Samples Correlations table instead of the actual Paired Samples Test table.
Report the mean difference and t-test from the Paired Samples Test table, as shown above — the correlation between pre and post scores is a diagnostic aside, not your result.
This is the single most common reporting error in before-after studies, and it matters because the two numbers tell entirely different stories: a correlation of 0.71 only says patients kept roughly the same relative ranking from pre- to post-treatment, while the paired t-test is what actually establishes whether the mean level changed. See our paired vs unpaired t-test guide for choosing the correct design, and the paired t-test calculator to compute this directly.
Mann-Whitney U Test
"Pain scores were significantly lower in the intervention group (median 3, IQR 2–4) than the control group (median 6, IQR 5–7), U = 412.5, z = -3.84, p < 0.001, r = 0.42."
"Pain was lower in the intervention group (p < 0.001)" with means reported instead of medians for a variable explicitly analyzed non-parametrically.
Report median and IQR (not mean and SD) alongside U, z, exact p, and an effect size r = Z/√N, matching the non-parametric test actually used.
Report medians consistently throughout — mixing means with a non-parametric test signals the analysis and the write-up do not match, which is exactly the kind of inconsistency a statistical reviewer is trained to catch on a first read. Because the Mann-Whitney U test does not directly estimate a mean difference, the effect size r (calculated as Z divided by the square root of N) is the standard way to convey magnitude, and should be interpreted using the same small/medium/large bands used for Pearson's r. Use the Mann-Whitney U calculator, and see Mann-Whitney vs t-test for when this test applies.
Wilcoxon Signed-Rank Test
"Anxiety scores decreased significantly from baseline (median 14, IQR 11–17) to post-intervention (median 9, IQR 7–12), Z = -4.21, p < 0.001, r = 0.55."
"Anxiety improved after the intervention (Z = -4.21)" — the Z statistic reported with no p-value, no effect size, and no medians to show the actual direction and magnitude of change.
Always pair Z with its exact p, an effect size r, and the pre/post medians — a test statistic alone tells a reader nothing interpretable.
The Wilcoxon signed-rank test is the non-parametric counterpart to the paired t-test, and the same principle applies: the p-value only tells you a change occurred, while the medians and effect size tell the reader how large that change was. Compute this directly with the Wilcoxon signed-rank calculator.
Comparing Three or More Groups
One-Way ANOVA
"Fasting glucose differed significantly across BMI categories, F(2, 82) = 18.73, p < 0.001, η² = 0.31. Tukey post-hoc comparisons showed the obese group had significantly higher glucose than both other groups (both p < 0.001), with no significant difference between normal-weight and overweight groups (p = 0.188)."
"ANOVA showed a significant difference between groups (p < 0.001)" with no statement of which specific groups differed.
Always follow a significant omnibus F with the post-hoc pairwise results, exactly as shown in the real example — a significant ANOVA alone only proves a difference exists somewhere.
Eta-squared (η²) is not produced automatically by most software and must be calculated as the between-groups sum of squares divided by the total sum of squares — worth the extra step, since a large F with a tiny η² in a very large sample tells a different clinical story than the same F with a large η². Use the one-way ANOVA calculator. For a three-arm CONSORT trial, report each pairwise comparison's own effect size and CI, not only the omnibus F.
Repeated Measures ANOVA
"Pain scores changed significantly across the three visits (Mauchly's test indicated sphericity was violated, χ²(2) = 6.87, p = 0.032; Greenhouse-Geisser corrected F(1.71, 100.9) = 27.4, p < 0.001, partial η² = 0.32). Pairwise comparisons showed significant reductions from baseline to week 4 and baseline to week 8 (both p < 0.001)."
"F(2, 118) = 27.4, p < 0.001" reported without ever mentioning Mauchly's test, when sphericity was in fact violated in this dataset.
State the Mauchly's test result first; if violated, report the Greenhouse-Geisser corrected F with its non-integer degrees of freedom, as shown above.
Reporting the sphericity check is not optional decoration — a violated-but-unreported sphericity assumption means the uncorrected F-test's p-value cannot be trusted, and an examiner who spots the omission will reasonably question every other result in the same analysis. Use the repeated measures ANOVA calculator for the full breakdown, including the automatically corrected degrees of freedom.
Kruskal-Wallis Test
"Symptom severity scores differed significantly across the three disease stages, H(2) = 15.62, p < 0.001. Post-hoc Dunn's tests with Bonferroni correction showed Stage III scores were significantly higher than Stage I (p < 0.001) and Stage II (p = 0.008), with no difference between Stage I and II (p = 0.412)."
"H = 15.62, significant" with no degrees of freedom, no exact p-value, and no post-hoc breakdown of which stages differed.
Report H with its df in parentheses, the exact p-value, and Bonferroni- or Dunn-corrected pairwise comparisons, mirroring how a significant ANOVA is followed up.
Epsilon-squared (ε²) is the recommended effect size for Kruskal-Wallis because, unlike eta-squared, it is derived from the rank-based H statistic itself rather than assuming normally distributed sums of squares. Run this with the Kruskal-Wallis calculator.
Friedman Test
"Quality-of-life scores differed significantly across the three follow-up time points, χ²(2) = 19.4, p < 0.001, Kendall's W = 0.27, indicating a small-to-moderate degree of concordance in ranking across time."
"Friedman's test was significant (p < 0.001)" with the chi-square statistic and Kendall's W both omitted entirely.
Report χ² with its df and Kendall's W as the effect size, exactly as in the real example, and follow with pairwise Wilcoxon comparisons if the omnibus result is significant.
Kendall's W ranges from 0 (no agreement in ranking across time points) to 1 (perfect agreement), and functions as the Friedman test's effect size in the same way η² does for ANOVA — report it whenever the omnibus result is significant. Use the Friedman test calculator for repeated ordinal or non-normal measurements across 3+ time points.
Categorical Association Tests
Chi-Square Test
"There was a statistically significant association between treatment group and clinical response, χ²(1, N = 100) = 7.84, p = 0.005, Cramér's V = 0.28."
"χ² = 7.84 (p < 0.05)" with no degrees of freedom, no sample size, and threshold notation instead of the exact p-value.
Report df and N inside the parentheses, the exact p, and Cramér's V — the full form shown in the real example above.
Cramér's V is preferred over the raw phi coefficient for any table larger than 2×2, and gives readers a bounded 0-to-1 sense of association strength that the chi-square statistic alone (which scales with sample size) cannot provide. Use the chi-square calculator. For a STROBE-compliant cohort study, also report the raw counts and row percentages in an accompanying table, not only the summary statistic in the text.
Fisher's Exact Test
"A rare adverse event occurred in 4 of 22 patients (18.2%) on the study drug versus 0 of 24 (0.0%) on placebo; this difference did not reach significance (Fisher's exact p = 0.081, OR could not be reliably estimated due to zero cell count)."
"χ² = 4.02, p = 0.045" reported for a 2×2 table where 50% of cells had an expected count below 5 — Pearson Chi-Square is invalid here.
Report Fisher's Exact p-value instead whenever any expected cell count is below 5, as shown in the real example — notice the conclusion changes from "significant" to "not significant."
Notice that the underlying counts and percentages are still reported in full even though the difference did not reach significance — a non-significant result is still a result, and CONSORT and STROBE both expect adverse-event and safety-outcome data to be reported completely regardless of statistical significance. Run this with the Fisher's Exact Test calculator, and see Chi-Square vs Fisher's Exact for the decision rule.
Correlation
Pearson Correlation
"Age was moderately, positively correlated with systolic blood pressure, r(108) = 0.41, 95% CI [0.24, 0.55], p < 0.001, r² = 0.17."
"Age and SBP were significantly correlated, so aging causes higher blood pressure" — treating a significant r as proof of causation, and omitting r² entirely.
Describe association language only ("was correlated with," not "caused"), and report r² alongside r to convey how much variance is actually shared.
A correlation coefficient can never establish direction of causation on its own, no matter how small the p-value — only a study design built to isolate cause (an RCT, or a longitudinal design with appropriate confounder adjustment) can support causal language, and journal reviewers routinely reject manuscripts that overstate a correlational finding this way. Use the correlation calculator to compute r with its 95% CI.
Spearman Correlation
"Pain score was moderately, negatively correlated with satisfaction score, ρ(94) = -0.49, 95% CI [-0.63, -0.32], p < 0.001."
"r = -0.49, p < 0.001" — using the Pearson symbol r for a coefficient that was actually calculated as Spearman's rho.
Use ρ (rho), not r, whenever the coefficient reported was Spearman's — the symbol itself tells a statistically literate reader which method was used.
Because Spearman's rho is calculated on ranks rather than raw values, it is also the appropriate choice whenever one or both variables are ordinal (Likert-type items, disease stage) rather than truly continuous, even if both happen to be normally distributed. See Pearson vs Spearman correlation for choosing between the two, computed via the same correlation calculator.
Regression Models
Linear Regression
"Age (B = 0.71, 95% CI [0.48, 0.94], β = 0.47, p < 0.001) and BMI (B = 1.12, 95% CI [0.31, 1.93], β = 0.24, p = 0.007) were independent significant predictors of systolic blood pressure. The model explained 32.2% of variance (adjusted R² = 0.322), F(2, 107) = 26.9, p < 0.001."
"Age was a stronger predictor than BMI because its B was larger" — comparing unstandardized B coefficients across predictors measured in different units.
Compare predictors using standardized β, not B, and always report the overall model fit (F, df, adjusted R²) alongside individual coefficients.
Adjusted R², not plain R², is the figure to headline once more than one predictor is in the model, because plain R² mechanically increases every time a predictor is added — even a clinically meaningless one — while the adjustment penalizes model complexity and gives a more honest sense of how much variance the model genuinely explains. Use the linear regression calculator, and see multivariate analysis for building multi-predictor models.
Logistic Regression
"Diabetes was associated with more than double the odds of 30-day readmission (aOR 2.64, 95% CI [1.36, 5.12], p = 0.004), adjusting for age (aOR 1.03 per year, 95% CI [1.00, 1.06], p = 0.045). The model showed good calibration (Hosmer-Lemeshow p = 0.635) and explained 24% of outcome variance (Nagelkerke R² = 0.24)."
"B = 0.97, p = 0.004" — reporting the raw log-odds coefficient instead of the exponentiated odds ratio a clinical reader can actually interpret.
Always report Exp(B), i.e., the odds ratio, with its 95% CI — never the raw B coefficient — plus model calibration and fit statistics.
The Hosmer-Lemeshow test follows the same reversed logic as a normality test: a non-significant result (p > 0.05) is the desired outcome, meaning the model's predicted probabilities are well calibrated against observed outcomes — reporting it without explaining this direction is a common source of reviewer confusion. Use the logistic regression calculator and see the full logistic regression guide.
Survival Analysis
Kaplan-Meier Survival Analysis
"Median progression-free survival was 18.4 months (95% CI [14.2, 22.6]) in the treatment group versus 11.1 months (95% CI [9.0, 13.2]) in the control group. The difference was statistically significant by log-rank test, χ²(1) = 12.8, p < 0.001."
"Survival was longer in the treatment group (p < 0.001)" with no median survival time, no CI, and no log-rank statistic reported.
Report each group's median survival with its 95% CI, plus the log-rank χ² and exact p, exactly as in the real example — and include the Kaplan-Meier curve with a number-at-risk table, as CONSORT's survival extension expects.
If the median survival time is not reached in one or both groups by the end of follow-up — common in trials with high cure rates or short follow-up — report the survival rate at a clinically meaningful fixed time point instead (e.g., "12-month survival was 78% vs 61%"), which CONSORT's survival-outcome extension treats as an acceptable and often more informative alternative. Use the Kaplan-Meier calculator and see the full Kaplan-Meier guide.
Cox Proportional Hazards Regression
"After adjusting for age and tumor stage, treatment with the study drug was associated with a significantly reduced hazard of disease progression (aHR 0.58, 95% CI [0.39, 0.85], p = 0.006). The proportional hazards assumption was confirmed via Schoenfeld residuals (p = 0.41)."
"HR = 0.58, significant" reported with no CI and no mention of whether the proportional hazards assumption was ever checked.
Report the adjusted HR with its 95% CI and exact p, and explicitly state that the proportionality assumption was tested (e.g., via Schoenfeld residuals), as required for a methodologically sound Cox model.
A hazard ratio describes the instantaneous risk of the event at any given moment, conditional on having survived to that moment — it is not the same quantity as a relative risk or an odds ratio, and using those terms interchangeably is a factual error a statistical reviewer will flag immediately. Use the Cox regression calculator to fit the model and check proportionality.
Diagnostic Accuracy and Agreement Statistics
ROC Curve Analysis
"Biomarker X showed good discrimination for disease presence, AUC = 0.81 (95% CI [0.73, 0.89], p < 0.001). At the optimal cutoff of ≥4.2 ng/mL (Youden index), sensitivity was 78% and specificity was 74%."
"AUC = 0.55, p = 0.03, confirming diagnostic usefulness" — treating a significant p-value (easy with a large sample) as proof of clinical usefulness despite a near-chance AUC.
Judge usefulness by AUC magnitude, not its p-value against 0.50, and always report the sensitivity/specificity pair at your chosen cutoff alongside the AUC.
The Youden index (sensitivity + specificity − 1) is the standard, defensible method for choosing a single cutoff to report, rather than eyeballing the curve or choosing whichever cutoff produces the most favorable-looking numbers after the fact. Use the ROC curve calculator and see Sensitivity, Specificity, PPV & NPV for the surrounding vocabulary.
Intraclass Correlation Coefficient (ICC)
"Inter-rater reliability for the pain scale was excellent, ICC = 0.91 (95% CI [0.86, 0.94]), based on a two-way random-effects model for absolute agreement between three raters."
"ICC = 0.91, p < 0.05" reported without stating which of the six ICC model variants (one-way/two-way, random/fixed, single/average measures) was used.
Always name the specific ICC model and its 95% CI — the p-value is rarely the meaningful figure here, since ICC model choice materially changes the number itself.
The same raw data can produce meaningfully different ICC values depending on whether a one-way or two-way model is chosen and whether consistency or absolute agreement is specified, so naming the exact model is not a formality — it is the only way a reader can judge whether your 0.91 is comparable to a 0.91 reported elsewhere. Use the ICC calculator and see the full ICC guide.
Bland-Altman Analysis
"Agreement between the new device and the reference standard was acceptable, with a mean bias of -0.3 mmHg and 95% limits of agreement of -8.2 to +7.6 mmHg."
"The two methods were significantly correlated (r = 0.94, p < 0.001), confirming agreement" — using a correlation coefficient to claim agreement between two measurement methods, a well-known methodological error.
Report mean bias and 95% limits of agreement from a Bland-Altman analysis — correlation measures association, not interchangeability, and is the wrong tool for a method-comparison question.
Two measurement methods can be very strongly correlated yet systematically biased relative to one another by a clinically important amount — correlation is blind to that kind of fixed offset, which is exactly the failure mode Bland-Altman analysis is designed to catch. Use the Bland-Altman calculator and see the full Bland-Altman guide.
Cohen's Kappa
"Inter-observer agreement for radiographic grading was substantial, κ = 0.72 (95% CI [0.58, 0.86]), p < 0.001."
"Agreement was 87% (p < 0.001)" — reporting raw percent agreement, which does not correct for the agreement expected by chance alone.
Report Cohen's kappa, not raw percent agreement, and interpret its value on a named agreement scale (e.g., Landis and Koch), as shown above.
Raw percent agreement systematically overstates true reliability whenever one category is much more common than another, because two raters can agree by chance alone a large fraction of the time on the dominant category — kappa corrects for exactly this, which is why journals in diagnostic and rater-reliability research require it over simple percent agreement. Use the Cohen's Kappa calculator and see the full Kappa guide.
Universal Statistical Reporting Checklist
Before submitting any thesis chapter or manuscript, run every result through this checklist regardless of which test produced it.
Name the test and the comparison
State which test was used and what was being compared, not just the outcome variable.
Report the test statistic with its df
t, F, χ², H, U, Z, or equivalent, in italics per APA style, with degrees of freedom in parentheses.
Report the exact p-value
Three decimal places, or "p < 0.001" for very small values — never "p = 0.000" or bare "p < 0.05."
Report an effect size
Cohen's d, η², r, Cramér's V, odds ratio, hazard ratio, or AUC — required by APA, CONSORT, and STROBE alike.
Report the 95% confidence interval
For the effect size or the mean/median difference, wherever your software provides one.
Match your summary statistic to the test
Mean ± SD with parametric tests, median (IQR) with non-parametric tests — never mix the two.
State the analyzed N
The sample size actually used in that specific test, which can differ from your total enrolled sample.
Follow a significant omnibus test with pairwise results
ANOVA, Kruskal-Wallis, Friedman, and Repeated Measures ANOVA all require post-hoc detail once significant.
Apply the correct design-specific checklist
CONSORT for an RCT, STROBE for an observational study, PRISMA for a systematic review or meta-analysis.
Distinguish statistical from clinical significance
State whether the magnitude of the effect, not just its p-value, would matter to a patient or clinician.
Cite your statistical software once in Methods
Name, version, and manufacturer — stated once, not repeated beside every individual result.
Cross-check tables, text, and figures
Every number in your Results narrative should match its corresponding table or figure exactly.
Downloadable Reporting Template
Every template shown in this guide is collected into a single fill-in-the-blank text file, organized by test, so you can paste your own numbers directly into your thesis or manuscript draft without retyping the structure.
Statistical Results Reporting Template
All 20 reporting templates in one plain-text file — descriptive statistics through Cohen's Kappa, ready to fill in with your own values.
⬇ Download the Template (.txt)Frequently Asked Questions
Let StatClinic Write Your Results Sentence For You
Run any of these 20 tests directly in StatClinic and get a journal-ready, APA-formatted results sentence generated automatically alongside your output. Free, no registration required.
Try StatClinic Free →