Launch StatClinic →
Research Writing

How to Write Statistical Results for Each Statistical Test: Complete Medical Research Guide

📖 32 min read 🗓 July 2026 ✓ Updated July 2026
S
StatClinic Editorial Team Statistical content for medical researchers and clinicians
Running the right statistical test is only half the job — a thesis committee or peer reviewer judges your work by how the Results section reads. This guide gives you a ready-to-use reporting template for 20 statistical tests used throughout medical research, each shown with a real worked example, a realistic incorrect version, its corrected fix, and a short explanation of why the correction matters under APA 7th edition notation and the CONSORT, STROBE, and PRISMA reporting guidelines. Keep this page open next to your manuscript and copy the template that matches your test.

Which Reporting Standard Applies to Your Study?

Four frameworks govern how medical statistics are written up, and they are not competitors — they stack. APA 7th edition governs the numeric notation itself (italics, decimal places, symbol order) and applies to virtually every manuscript regardless of design. On top of that, one design-specific checklist applies depending on what kind of study you ran.

FrameworkGovernsApplies To
APA 7th EditionStatistical notation style: italics, decimals, symbol order, effect size and CI reportingEvery manuscript
CONSORTTrial reporting: primary/secondary outcome effect sizes with CIs, per-group results, participant flowRandomized controlled trials
STROBEObservational reporting: unadjusted and adjusted estimates, confounder handling, missing dataCohort, case-control, cross-sectional studies
PRISMASynthesis reporting: pooled effect sizes, heterogeneity statistics, study flow diagramSystematic reviews and meta-analyses

Every template in this guide follows APA numeric notation by default. Where CONSORT or STROBE adds a design-specific expectation for a particular test — reporting both the unadjusted and adjusted effect in a STROBE cohort study, for instance — it is called out in that test's explanation. PRISMA is included for completeness because many thesis committees and journals still ask which guideline you consulted even when your own study is a single trial or cohort rather than a synthesis; if any part of your work involves pooling results across studies (a meta-analysis chapter, for example), the same statistic-df-p-effect size-CI structure used throughout this guide still applies to each pooled estimate, on top of PRISMA's flow-diagram and heterogeneity-reporting requirements.

All four frameworks are maintained or indexed by the EQUATOR Network, the standard reference point medical journals point authors to when reporting requirements are not spelled out in their own author instructions. Many journals now require the relevant checklist (a completed CONSORT or STROBE checklist, for instance) to be uploaded as a supplementary file alongside the manuscript, so it is worth identifying which framework applies before you start drafting your Results section, not after.

The Five Elements Every Result Needs Regardless of which test you ran, a complete, guideline-compliant result sentence contains: (1) the test statistic and its degrees of freedom, (2) the exact p-value, (3) an effect size or association measure, (4) the 95% confidence interval where available, and (5) the sample size actually analyzed. Missing any of these five is the single most common reason theses are sent back for revision.

Reporting Descriptive Statistics

Descriptive statistics open every Results section and describe your sample before any hypothesis test appears — this becomes your Table 1. Use normality testing to decide between the two template forms below.

Normally Distributed Continuous Variables

Reporting TemplateMean ± SD [or M = value, SD = value]; range = min–max
Real Example

"The mean age of participants was 44.2 ± 13.6 years (range 19–74)."

❌ Incorrect

"Age: 44.2 (13.6)" — reported with no label identifying which figure is the mean and which is the SD.

✅ Corrected

"Mean age was 44.2 ± 13.6 years." The ± symbol (or explicit "SD =") removes any ambiguity about what the second number represents.

Non-normal continuous variables should never be summarized with mean and SD alone — report median and interquartile range (IQR) instead: "Median length of stay was 4 days (IQR 3–7)." APA 7 permits either format but requires you to state explicitly which one you used and why (i.e., the result of your normality test). For a STROBE-compliant cohort or cross-sectional study, this Table 1 should also report the number and percentage of missing values for every variable, since undisclosed missingness is one of the most frequently cited STROBE deficiencies in peer review. Categorical baseline variables follow the same logic in reverse: report n (%) for every category, using Valid Percent once any missing data exists, exactly as covered in our SPSS output interpretation guide.

Take-Home Points Match your summary statistic to the distribution: mean ± SD for normal data, median (IQR) for skewed data. Always report the analyzed N, not just your total enrolled sample, since these can differ once missing data is excluded.

Comparing Two Groups

These four tests compare an outcome between two groups — the choice between them depends on whether groups are independent or paired, and whether the outcome is normally distributed.

Independent Samples T-Test

Reporting Templatet(df) = value, p = value, 95% CI [lower, upper], Cohen's d = value
Real Example

"Mean hemoglobin was significantly higher in males (13.6 ± 1.7 g/dL) than females (11.8 ± 1.6 g/dL); mean difference 1.8 g/dL, 95% CI [1.06, 2.54], t(83) = 4.87, p < 0.001, d = 1.07."

❌ Incorrect

"There was a significant difference between males and females (p = 0.000)." No test statistic, no df, no effect size, and a p-value SPSS never actually produces.

✅ Corrected

Report the full statement above: t(83) = 4.87, p < 0.001, with the mean difference, its 95% CI, and Cohen's d, not a bare p-value.

The p-value alone cannot tell a reader whether a 1.8 g/dL difference in hemoglobin is clinically meaningful — the mean difference, its CI, and Cohen's d together do that job, which is why all three appear in the template above rather than the p-value in isolation. Use the independent t-test calculator to generate this output directly. For an RCT comparing two treatment arms, CONSORT explicitly requires the between-group difference and its CI to be the headline figure for both the primary and secondary outcomes — not the p-value alone, and not merely "significant" or "not significant" as a verdict.

Paired Samples T-Test

Reporting Templatet(df) = value, p = value, 95% CI of the difference [lower, upper], Cohen's d = value
Real Example

"Systolic blood pressure decreased significantly after the intervention, from 148.2 ± 14.1 mmHg to 135.8 ± 13.3 mmHg (mean decrease 12.4 mmHg, 95% CI [9.85, 14.95]), t(59) = 9.80, p < 0.001, d = 1.26."

❌ Incorrect

"Blood pressure improved significantly after treatment (r = 0.71, p < 0.001)" — quoting the Paired Samples Correlations table instead of the actual Paired Samples Test table.

✅ Corrected

Report the mean difference and t-test from the Paired Samples Test table, as shown above — the correlation between pre and post scores is a diagnostic aside, not your result.

This is the single most common reporting error in before-after studies, and it matters because the two numbers tell entirely different stories: a correlation of 0.71 only says patients kept roughly the same relative ranking from pre- to post-treatment, while the paired t-test is what actually establishes whether the mean level changed. See our paired vs unpaired t-test guide for choosing the correct design, and the paired t-test calculator to compute this directly.

Mann-Whitney U Test

Reporting TemplateU = value, z = value, p = value, r = value (or Hodges-Lehmann median difference, 95% CI)
Real Example

"Pain scores were significantly lower in the intervention group (median 3, IQR 2–4) than the control group (median 6, IQR 5–7), U = 412.5, z = -3.84, p < 0.001, r = 0.42."

❌ Incorrect

"Pain was lower in the intervention group (p < 0.001)" with means reported instead of medians for a variable explicitly analyzed non-parametrically.

✅ Corrected

Report median and IQR (not mean and SD) alongside U, z, exact p, and an effect size r = Z/√N, matching the non-parametric test actually used.

Report medians consistently throughout — mixing means with a non-parametric test signals the analysis and the write-up do not match, which is exactly the kind of inconsistency a statistical reviewer is trained to catch on a first read. Because the Mann-Whitney U test does not directly estimate a mean difference, the effect size r (calculated as Z divided by the square root of N) is the standard way to convey magnitude, and should be interpreted using the same small/medium/large bands used for Pearson's r. Use the Mann-Whitney U calculator, and see Mann-Whitney vs t-test for when this test applies.

Wilcoxon Signed-Rank Test

Reporting TemplateZ = value, p = value, r = value
Real Example

"Anxiety scores decreased significantly from baseline (median 14, IQR 11–17) to post-intervention (median 9, IQR 7–12), Z = -4.21, p < 0.001, r = 0.55."

❌ Incorrect

"Anxiety improved after the intervention (Z = -4.21)" — the Z statistic reported with no p-value, no effect size, and no medians to show the actual direction and magnitude of change.

✅ Corrected

Always pair Z with its exact p, an effect size r, and the pre/post medians — a test statistic alone tells a reader nothing interpretable.

The Wilcoxon signed-rank test is the non-parametric counterpart to the paired t-test, and the same principle applies: the p-value only tells you a change occurred, while the medians and effect size tell the reader how large that change was. Compute this directly with the Wilcoxon signed-rank calculator.

Take-Home Points Parametric tests report mean ± SD with Cohen's d; non-parametric tests report median (IQR) with r. Never mix the two summary styles within the same result. Always include the CI or a directly comparable effect size, not the test statistic alone.

Comparing Three or More Groups

One-Way ANOVA

Reporting TemplateF(df_between, df_within) = value, p = value, η² = value; post-hoc: group A vs B, p = value
Real Example

"Fasting glucose differed significantly across BMI categories, F(2, 82) = 18.73, p < 0.001, η² = 0.31. Tukey post-hoc comparisons showed the obese group had significantly higher glucose than both other groups (both p < 0.001), with no significant difference between normal-weight and overweight groups (p = 0.188)."

❌ Incorrect

"ANOVA showed a significant difference between groups (p < 0.001)" with no statement of which specific groups differed.

✅ Corrected

Always follow a significant omnibus F with the post-hoc pairwise results, exactly as shown in the real example — a significant ANOVA alone only proves a difference exists somewhere.

Eta-squared (η²) is not produced automatically by most software and must be calculated as the between-groups sum of squares divided by the total sum of squares — worth the extra step, since a large F with a tiny η² in a very large sample tells a different clinical story than the same F with a large η². Use the one-way ANOVA calculator. For a three-arm CONSORT trial, report each pairwise comparison's own effect size and CI, not only the omnibus F.

Repeated Measures ANOVA

Reporting TemplateF(df, df) = value, p = value, partial η² = value [state Greenhouse-Geisser correction if sphericity was violated]
Real Example

"Pain scores changed significantly across the three visits (Mauchly's test indicated sphericity was violated, χ²(2) = 6.87, p = 0.032; Greenhouse-Geisser corrected F(1.71, 100.9) = 27.4, p < 0.001, partial η² = 0.32). Pairwise comparisons showed significant reductions from baseline to week 4 and baseline to week 8 (both p < 0.001)."

❌ Incorrect

"F(2, 118) = 27.4, p < 0.001" reported without ever mentioning Mauchly's test, when sphericity was in fact violated in this dataset.

✅ Corrected

State the Mauchly's test result first; if violated, report the Greenhouse-Geisser corrected F with its non-integer degrees of freedom, as shown above.

Reporting the sphericity check is not optional decoration — a violated-but-unreported sphericity assumption means the uncorrected F-test's p-value cannot be trusted, and an examiner who spots the omission will reasonably question every other result in the same analysis. Use the repeated measures ANOVA calculator for the full breakdown, including the automatically corrected degrees of freedom.

Kruskal-Wallis Test

Reporting TemplateH(df) = value, p = value, ε² = value; post-hoc pairwise, p = value (Bonferroni-adjusted)
Real Example

"Symptom severity scores differed significantly across the three disease stages, H(2) = 15.62, p < 0.001. Post-hoc Dunn's tests with Bonferroni correction showed Stage III scores were significantly higher than Stage I (p < 0.001) and Stage II (p = 0.008), with no difference between Stage I and II (p = 0.412)."

❌ Incorrect

"H = 15.62, significant" with no degrees of freedom, no exact p-value, and no post-hoc breakdown of which stages differed.

✅ Corrected

Report H with its df in parentheses, the exact p-value, and Bonferroni- or Dunn-corrected pairwise comparisons, mirroring how a significant ANOVA is followed up.

Epsilon-squared (ε²) is the recommended effect size for Kruskal-Wallis because, unlike eta-squared, it is derived from the rank-based H statistic itself rather than assuming normally distributed sums of squares. Run this with the Kruskal-Wallis calculator.

Friedman Test

Reporting Templateχ²(df) = value, p = value, Kendall's W = value
Real Example

"Quality-of-life scores differed significantly across the three follow-up time points, χ²(2) = 19.4, p < 0.001, Kendall's W = 0.27, indicating a small-to-moderate degree of concordance in ranking across time."

❌ Incorrect

"Friedman's test was significant (p < 0.001)" with the chi-square statistic and Kendall's W both omitted entirely.

✅ Corrected

Report χ² with its df and Kendall's W as the effect size, exactly as in the real example, and follow with pairwise Wilcoxon comparisons if the omnibus result is significant.

Kendall's W ranges from 0 (no agreement in ranking across time points) to 1 (perfect agreement), and functions as the Friedman test's effect size in the same way η² does for ANOVA — report it whenever the omnibus result is significant. Use the Friedman test calculator for repeated ordinal or non-normal measurements across 3+ time points.

Take-Home Points A significant omnibus test (ANOVA, Repeated Measures ANOVA, Kruskal-Wallis, or Friedman) only tells you a difference exists somewhere among groups — always report the follow-up pairwise comparisons by name. Check and report sphericity (Mauchly's test) whenever repeated measures are involved.

Categorical Association Tests

Chi-Square Test

Reporting Templateχ²(df, N = value) = value, p = value, Cramér's V = value
Real Example

"There was a statistically significant association between treatment group and clinical response, χ²(1, N = 100) = 7.84, p = 0.005, Cramér's V = 0.28."

❌ Incorrect

"χ² = 7.84 (p < 0.05)" with no degrees of freedom, no sample size, and threshold notation instead of the exact p-value.

✅ Corrected

Report df and N inside the parentheses, the exact p, and Cramér's V — the full form shown in the real example above.

Cramér's V is preferred over the raw phi coefficient for any table larger than 2×2, and gives readers a bounded 0-to-1 sense of association strength that the chi-square statistic alone (which scales with sample size) cannot provide. Use the chi-square calculator. For a STROBE-compliant cohort study, also report the raw counts and row percentages in an accompanying table, not only the summary statistic in the text.

Fisher's Exact Test

Reporting TemplateFisher's exact p = value, Odds Ratio = value, 95% CI [lower, upper]
Real Example

"A rare adverse event occurred in 4 of 22 patients (18.2%) on the study drug versus 0 of 24 (0.0%) on placebo; this difference did not reach significance (Fisher's exact p = 0.081, OR could not be reliably estimated due to zero cell count)."

❌ Incorrect

"χ² = 4.02, p = 0.045" reported for a 2×2 table where 50% of cells had an expected count below 5 — Pearson Chi-Square is invalid here.

✅ Corrected

Report Fisher's Exact p-value instead whenever any expected cell count is below 5, as shown in the real example — notice the conclusion changes from "significant" to "not significant."

Notice that the underlying counts and percentages are still reported in full even though the difference did not reach significance — a non-significant result is still a result, and CONSORT and STROBE both expect adverse-event and safety-outcome data to be reported completely regardless of statistical significance. Run this with the Fisher's Exact Test calculator, and see Chi-Square vs Fisher's Exact for the decision rule.

Take-Home Points Always include df and N inside the parentheses for chi-square. Switch to Fisher's Exact whenever expected cell counts are small — and say so explicitly in your Results, since it changes which p-value a reader should trust.

Correlation

Pearson Correlation

Reporting Templater(df) = value, p = value, 95% CI [lower, upper], r² = value
Real Example

"Age was moderately, positively correlated with systolic blood pressure, r(108) = 0.41, 95% CI [0.24, 0.55], p < 0.001, r² = 0.17."

❌ Incorrect

"Age and SBP were significantly correlated, so aging causes higher blood pressure" — treating a significant r as proof of causation, and omitting r² entirely.

✅ Corrected

Describe association language only ("was correlated with," not "caused"), and report r² alongside r to convey how much variance is actually shared.

A correlation coefficient can never establish direction of causation on its own, no matter how small the p-value — only a study design built to isolate cause (an RCT, or a longitudinal design with appropriate confounder adjustment) can support causal language, and journal reviewers routinely reject manuscripts that overstate a correlational finding this way. Use the correlation calculator to compute r with its 95% CI.

Spearman Correlation

Reporting Templateρ(df) = value [or ρ = value], p = value, 95% CI [lower, upper]
Real Example

"Pain score was moderately, negatively correlated with satisfaction score, ρ(94) = -0.49, 95% CI [-0.63, -0.32], p < 0.001."

❌ Incorrect

"r = -0.49, p < 0.001" — using the Pearson symbol r for a coefficient that was actually calculated as Spearman's rho.

✅ Corrected

Use ρ (rho), not r, whenever the coefficient reported was Spearman's — the symbol itself tells a statistically literate reader which method was used.

Because Spearman's rho is calculated on ranks rather than raw values, it is also the appropriate choice whenever one or both variables are ordinal (Likert-type items, disease stage) rather than truly continuous, even if both happen to be normally distributed. See Pearson vs Spearman correlation for choosing between the two, computed via the same correlation calculator.

Take-Home Points Use the correct Greek letter: r for Pearson, ρ (rho) for Spearman — they are not interchangeable notation. Report r² (or interpret the magnitude) alongside any correlation coefficient; a significant r says nothing about how strong the relationship actually is.

Regression Models

Linear Regression

Reporting TemplateB = value, 95% CI [lower, upper], β = value, p = value; model: F(df, df) = value, p = value, Adjusted R² = value
Real Example

"Age (B = 0.71, 95% CI [0.48, 0.94], β = 0.47, p < 0.001) and BMI (B = 1.12, 95% CI [0.31, 1.93], β = 0.24, p = 0.007) were independent significant predictors of systolic blood pressure. The model explained 32.2% of variance (adjusted R² = 0.322), F(2, 107) = 26.9, p < 0.001."

❌ Incorrect

"Age was a stronger predictor than BMI because its B was larger" — comparing unstandardized B coefficients across predictors measured in different units.

✅ Corrected

Compare predictors using standardized β, not B, and always report the overall model fit (F, df, adjusted R²) alongside individual coefficients.

Adjusted R², not plain R², is the figure to headline once more than one predictor is in the model, because plain R² mechanically increases every time a predictor is added — even a clinically meaningless one — while the adjustment penalizes model complexity and gives a more honest sense of how much variance the model genuinely explains. Use the linear regression calculator, and see multivariate analysis for building multi-predictor models.

Logistic Regression

Reporting TemplateAdjusted OR = value, 95% CI [lower, upper], p = value; model: χ²(df) = value, p = value, Nagelkerke R² = value
Real Example

"Diabetes was associated with more than double the odds of 30-day readmission (aOR 2.64, 95% CI [1.36, 5.12], p = 0.004), adjusting for age (aOR 1.03 per year, 95% CI [1.00, 1.06], p = 0.045). The model showed good calibration (Hosmer-Lemeshow p = 0.635) and explained 24% of outcome variance (Nagelkerke R² = 0.24)."

❌ Incorrect

"B = 0.97, p = 0.004" — reporting the raw log-odds coefficient instead of the exponentiated odds ratio a clinical reader can actually interpret.

✅ Corrected

Always report Exp(B), i.e., the odds ratio, with its 95% CI — never the raw B coefficient — plus model calibration and fit statistics.

The Hosmer-Lemeshow test follows the same reversed logic as a normality test: a non-significant result (p > 0.05) is the desired outcome, meaning the model's predicted probabilities are well calibrated against observed outcomes — reporting it without explaining this direction is a common source of reviewer confusion. Use the logistic regression calculator and see the full logistic regression guide.

Take-Home Points Linear regression: report unstandardized B with CI for magnitude, standardized β for comparing predictors, and adjusted R² for overall fit. Logistic regression: always report Exp(B) as an odds ratio with its CI, never the raw B. STROBE requires both the unadjusted and the fully adjusted estimate for your primary exposure in observational studies.

Survival Analysis

Kaplan-Meier Survival Analysis

Reporting TemplateMedian survival = value (95% CI [lower, upper]); log-rank χ²(df) = value, p = value
Real Example

"Median progression-free survival was 18.4 months (95% CI [14.2, 22.6]) in the treatment group versus 11.1 months (95% CI [9.0, 13.2]) in the control group. The difference was statistically significant by log-rank test, χ²(1) = 12.8, p < 0.001."

❌ Incorrect

"Survival was longer in the treatment group (p < 0.001)" with no median survival time, no CI, and no log-rank statistic reported.

✅ Corrected

Report each group's median survival with its 95% CI, plus the log-rank χ² and exact p, exactly as in the real example — and include the Kaplan-Meier curve with a number-at-risk table, as CONSORT's survival extension expects.

If the median survival time is not reached in one or both groups by the end of follow-up — common in trials with high cure rates or short follow-up — report the survival rate at a clinically meaningful fixed time point instead (e.g., "12-month survival was 78% vs 61%"), which CONSORT's survival-outcome extension treats as an acceptable and often more informative alternative. Use the Kaplan-Meier calculator and see the full Kaplan-Meier guide.

Cox Proportional Hazards Regression

Reporting TemplateAdjusted HR = value, 95% CI [lower, upper], p = value [state proportional hazards assumption check]
Real Example

"After adjusting for age and tumor stage, treatment with the study drug was associated with a significantly reduced hazard of disease progression (aHR 0.58, 95% CI [0.39, 0.85], p = 0.006). The proportional hazards assumption was confirmed via Schoenfeld residuals (p = 0.41)."

❌ Incorrect

"HR = 0.58, significant" reported with no CI and no mention of whether the proportional hazards assumption was ever checked.

✅ Corrected

Report the adjusted HR with its 95% CI and exact p, and explicitly state that the proportionality assumption was tested (e.g., via Schoenfeld residuals), as required for a methodologically sound Cox model.

A hazard ratio describes the instantaneous risk of the event at any given moment, conditional on having survived to that moment — it is not the same quantity as a relative risk or an odds ratio, and using those terms interchangeably is a factual error a statistical reviewer will flag immediately. Use the Cox regression calculator to fit the model and check proportionality.

Take-Home Points Always report median survival time with its CI per group, not just a p-value. Cox regression results are hazard ratios, not odds ratios or risk ratios — label them correctly, and confirm (and report) that the proportional hazards assumption held.

Diagnostic Accuracy and Agreement Statistics

ROC Curve Analysis

Reporting TemplateAUC = value, 95% CI [lower, upper], p = value; optimal cutoff, sensitivity = value%, specificity = value%
Real Example

"Biomarker X showed good discrimination for disease presence, AUC = 0.81 (95% CI [0.73, 0.89], p < 0.001). At the optimal cutoff of ≥4.2 ng/mL (Youden index), sensitivity was 78% and specificity was 74%."

❌ Incorrect

"AUC = 0.55, p = 0.03, confirming diagnostic usefulness" — treating a significant p-value (easy with a large sample) as proof of clinical usefulness despite a near-chance AUC.

✅ Corrected

Judge usefulness by AUC magnitude, not its p-value against 0.50, and always report the sensitivity/specificity pair at your chosen cutoff alongside the AUC.

The Youden index (sensitivity + specificity − 1) is the standard, defensible method for choosing a single cutoff to report, rather than eyeballing the curve or choosing whichever cutoff produces the most favorable-looking numbers after the fact. Use the ROC curve calculator and see Sensitivity, Specificity, PPV & NPV for the surrounding vocabulary.

Intraclass Correlation Coefficient (ICC)

Reporting TemplateICC = value, 95% CI [lower, upper], model type (e.g., two-way random, absolute agreement)
Real Example

"Inter-rater reliability for the pain scale was excellent, ICC = 0.91 (95% CI [0.86, 0.94]), based on a two-way random-effects model for absolute agreement between three raters."

❌ Incorrect

"ICC = 0.91, p < 0.05" reported without stating which of the six ICC model variants (one-way/two-way, random/fixed, single/average measures) was used.

✅ Corrected

Always name the specific ICC model and its 95% CI — the p-value is rarely the meaningful figure here, since ICC model choice materially changes the number itself.

The same raw data can produce meaningfully different ICC values depending on whether a one-way or two-way model is chosen and whether consistency or absolute agreement is specified, so naming the exact model is not a formality — it is the only way a reader can judge whether your 0.91 is comparable to a 0.91 reported elsewhere. Use the ICC calculator and see the full ICC guide.

Bland-Altman Analysis

Reporting TemplateMean bias = value, 95% limits of agreement [lower, upper]
Real Example

"Agreement between the new device and the reference standard was acceptable, with a mean bias of -0.3 mmHg and 95% limits of agreement of -8.2 to +7.6 mmHg."

❌ Incorrect

"The two methods were significantly correlated (r = 0.94, p < 0.001), confirming agreement" — using a correlation coefficient to claim agreement between two measurement methods, a well-known methodological error.

✅ Corrected

Report mean bias and 95% limits of agreement from a Bland-Altman analysis — correlation measures association, not interchangeability, and is the wrong tool for a method-comparison question.

Two measurement methods can be very strongly correlated yet systematically biased relative to one another by a clinically important amount — correlation is blind to that kind of fixed offset, which is exactly the failure mode Bland-Altman analysis is designed to catch. Use the Bland-Altman calculator and see the full Bland-Altman guide.

Cohen's Kappa

Reporting Templateκ = value, 95% CI [lower, upper], p = value (interpreted as [poor/fair/moderate/substantial/almost perfect] agreement)
Real Example

"Inter-observer agreement for radiographic grading was substantial, κ = 0.72 (95% CI [0.58, 0.86]), p < 0.001."

❌ Incorrect

"Agreement was 87% (p < 0.001)" — reporting raw percent agreement, which does not correct for the agreement expected by chance alone.

✅ Corrected

Report Cohen's kappa, not raw percent agreement, and interpret its value on a named agreement scale (e.g., Landis and Koch), as shown above.

Raw percent agreement systematically overstates true reliability whenever one category is much more common than another, because two raters can agree by chance alone a large fraction of the time on the dominant category — kappa corrects for exactly this, which is why journals in diagnostic and rater-reliability research require it over simple percent agreement. Use the Cohen's Kappa calculator and see the full Kappa guide.

Take-Home Points AUC magnitude, not its p-value, determines diagnostic usefulness. Never substitute a correlation coefficient for agreement — Bland-Altman (continuous measurements) and Kappa (categorical ratings) are the correct tools, and ICC always needs its specific model type named.

Universal Statistical Reporting Checklist

Before submitting any thesis chapter or manuscript, run every result through this checklist regardless of which test produced it.

1

Name the test and the comparison

State which test was used and what was being compared, not just the outcome variable.

2

Report the test statistic with its df

t, F, χ², H, U, Z, or equivalent, in italics per APA style, with degrees of freedom in parentheses.

3

Report the exact p-value

Three decimal places, or "p < 0.001" for very small values — never "p = 0.000" or bare "p < 0.05."

4

Report an effect size

Cohen's d, η², r, Cramér's V, odds ratio, hazard ratio, or AUC — required by APA, CONSORT, and STROBE alike.

5

Report the 95% confidence interval

For the effect size or the mean/median difference, wherever your software provides one.

6

Match your summary statistic to the test

Mean ± SD with parametric tests, median (IQR) with non-parametric tests — never mix the two.

7

State the analyzed N

The sample size actually used in that specific test, which can differ from your total enrolled sample.

8

Follow a significant omnibus test with pairwise results

ANOVA, Kruskal-Wallis, Friedman, and Repeated Measures ANOVA all require post-hoc detail once significant.

9

Apply the correct design-specific checklist

CONSORT for an RCT, STROBE for an observational study, PRISMA for a systematic review or meta-analysis.

10

Distinguish statistical from clinical significance

State whether the magnitude of the effect, not just its p-value, would matter to a patient or clinician.

11

Cite your statistical software once in Methods

Name, version, and manufacturer — stated once, not repeated beside every individual result.

12

Cross-check tables, text, and figures

Every number in your Results narrative should match its corresponding table or figure exactly.

Downloadable Reporting Template

Every template shown in this guide is collected into a single fill-in-the-blank text file, organized by test, so you can paste your own numbers directly into your thesis or manuscript draft without retyping the structure.

Statistical Results Reporting Template

All 20 reporting templates in one plain-text file — descriptive statistics through Cohen's Kappa, ready to fill in with your own values.

⬇ Download the Template (.txt)

Frequently Asked Questions

Do I need to report exact p-values or is p < 0.05 acceptable? +
Report exact p-values to three decimal places (e.g., p = 0.032) for every test. APA 7th edition, ICMJE, and virtually every major medical journal require exact values rather than threshold notation, because p = 0.049 and p = 0.001 both satisfy "p < 0.05" yet represent very different strength of evidence. The only accepted exception is very small p-values, reported as p < 0.001 rather than a string of zeros.
What is the difference between how CONSORT and STROBE want results reported? +
CONSORT applies to randomized controlled trials and requires the effect size for the primary outcome with its 95% CI, reported separately per group, plus a participant flow diagram. STROBE applies to cohort, case-control, and cross-sectional studies and additionally expects both unadjusted and adjusted effect estimates and a clear description of how confounders and missing data were handled. Both share the same core notation (statistic, df, exact p, effect size, CI) — they differ in study-design-specific elements around the numbers, not the numbers themselves.
Should effect sizes always be reported alongside a p-value? +
Yes. A p-value only indicates whether an effect is statistically distinguishable from zero; it says nothing about magnitude. APA 7, CONSORT, and STROBE all require an effect size (Cohen's d, η², odds ratio, hazard ratio, r, or equivalent) with its 95% CI alongside every primary result, because p-values in large samples can be significant for clinically trivial effects, and non-significant in small samples despite a potentially important effect.
How do I cite the statistical software used in my results section? +
State the software name, version, and manufacturer once, in the Statistical Analysis subsection of your Methods: for example, "Statistical analyses were performed using IBM SPSS Statistics, version 28.0 (IBM Corp., Armonk, NY)." This is required by APA 7th edition and expected by most medical journals. You do not need to repeat it beside every individual result in your Results section.
What if my target journal does not specify a reporting format? +
Default to APA 7th edition notation for the numeric formatting of every test, and layer on the design-specific checklist that matches your study: CONSORT for a randomized controlled trial, STROBE for a cohort, case-control, or cross-sectional study, or PRISMA for a systematic review or meta-analysis. Many journals implicitly expect this combination even when their author instructions do not name a guideline, and some explicitly require the relevant EQUATOR Network checklist submitted alongside the manuscript.
How do I report results for a statistical test not covered in this guide? +
Apply the same universal structure used throughout this guide regardless of the specific test: name the test and comparison, report the test statistic with its degrees of freedom, report the exact p-value, report an effect size or association measure, report the 95% CI where available, and state the sample size actually analyzed. This five-element structure generalizes to virtually any inferential test in the medical literature.
Do non-parametric tests need confidence intervals reported too? +
Yes, wherever they can be calculated. For Mann-Whitney U, report the Hodges-Lehmann estimate of the median difference with its 95% CI if your software provides it; for correlation coefficients (Spearman's rho), report the CI using Fisher's z-transformation. When a direct CI is not readily available, report the exact p-value, the test statistic, and a non-parametric effect size (r = Z/√N) so the magnitude of the result is still conveyed, not just its significance.

Let StatClinic Write Your Results Sentence For You

Run any of these 20 tests directly in StatClinic and get a journal-ready, APA-formatted results sentence generated automatically alongside your output. Free, no registration required.

Try StatClinic Free →