How to Test Data for Normality in Theses and Journal Articles
A definitive methodological decision framework for graduate theses and peer-reviewed journals: Shapiro-Wilk vs. Kolmogorov-Smirnov, skewness and kurtosis benchmarks, Q-Q plot diagnostics, 3-tier non-normality fallbacks, and APA 7 reporting.
Normality cannot be reduced to a single p-value. Follow this evidence-based hierarchy tailored to your sample size (N):
- Small and Moderate Samples (n < 50 and n ≤ 300): The Shapiro-Wilk test is the most powerful test available (Razali & Wah, 2011). Look for p > .05 and verify that data points in the Normal Q-Q plot adhere tightly to the 45-degree diagonal reference line.
- Skewness and Kurtosis Benchmarks: Across graduate theses and dissertations, the standard citation is George & Mallery (2010, 2020), where coefficients between [-2.0, +2.0] signify acceptable univariate normality. For high-impact Q1/Q2 journals, aim for Tabachnick & Fidell (2019) [-1.5, +1.5].
- Large Samples (N > 300): Shapiro-Wilk and Kolmogorov-Smirnov become hyper-sensitive, frequently yielding p < .05 on trivial imperfections. Under the Central Limit Theorem, prioritize skewness/kurtosis bounds and Q-Q plots to justify parametric testing.
Normality Decision Matrix & Methodological Standards
Comparative benchmark guidelines across prominent psychometric and statistical authorities.
| Standard / Tier | Acceptable Window | Primary Citation | Recommendation | Methodological Justification |
|---|---|---|---|---|
| Ideal Normal Symmetry | -1.00 ≤ Skewness & Kurtosis ≤ +1.00 | Hair et al. (2019) | Unconditional Parametric Acceptance | Data adhere closely to the theoretical normal distribution. Independent t-tests, One-Way ANOVA, Pearson correlation, and linear regression can proceed safely. |
| Conservative / High-Impact Standard | -1.50 ≤ Skewness & Kurtosis ≤ +1.50 | Tabachnick & Fidell (2019) | Recommended for Q1 / Q2 Journals | Strict univariate normality boundary widely accepted by leading medical, psychometric, and behavioral science editorial boards. |
| Graduate Dissertation Benchmark | -2.00 ≤ Skewness & Kurtosis ≤ +2.00 | George & Mallery (2010, 2020) | Standard Dissertation Threshold | The most universally cited rule of thumb across graduate thesis committees and institutional guidelines in social and behavioral sciences. |
| Z-Score Ratio Rule (Small Samples Only) | |Z| < 1.96 (α = .05) or |Z| < 2.58 (α = .01) | Field (2018); Tabachnick & Fidell (2019) | Valid Strictly for n < 50 | Calculated as Z = Statistic / SE. In moderate to large samples (N > 100), SE shrinks dramatically, causing harmless deviations to trigger p < .001. Avoid for N > 100. |
| Structural Equation Modeling (SEM) | Skewness < 3.00, Kurtosis < 10.00 | Kline (2016) | Multivariate Analysis Ceiling | Upper tolerance threshold for confirmatory factor analysis (CFA) and path models in AMOS, LISREL, or Mplus before requiring robust estimators (e.g., Satorra-Bentler / MLM). |
| Severe Non-Normality / Heavy Tails | |Skewness| > 2.00 or |Kurtosis| > 2.00 | Field (2018) | Reject Standard Parametric Test | Data exhibit marked asymmetry or extreme outliers. Transition to non-parametric equivalents (Mann-Whitney U, Kruskal-Wallis) or run 95% BCa Bootstrapping. |
6-Stage Normality Screening & Decision Protocol
From raw data screening to definitive dissertation and journal results reporting.
Research Design, Measurement Scale, and Group-Wise Testing
The single most prevalent conceptual error in empirical research is attempting to perform normality tests on individual ordinal Likert survey items (e.g., 1 = Strongly Disagree to 5 = Strongly Agree). Individual Likert responses are discrete, bounded integer categories that inherently cannot form a continuous Gaussian bell curve. Normality testing must be restricted to continuous variables measured on an interval or ratio scale, or to summated composite and mean scores derived from validated multi-item scales.
The second non-negotiable rule is the requirement for group-wise testing in comparative designs: When planning an independent samples t-test (e.g., Treatment vs. Control) or a One-Way ANOVA across multiple conditions, the distributional normality assumption applies to the dependent variable WITHIN EACH SUBGROUP SEPARATELY, not to the pooled aggregate sample. Pooling groups can artificially induce bimodal distributions or mask genuine within-group skewness, rendering inferential statistics invalid.
- Verified that the target variable is measured on a continuous (interval/ratio) scale or is a composite multi-item index
- Configured software (e.g., SPSS 'Explore > Factor List' or 'Split File') to test normality separately within each experimental or demographic subgroup
- Identified total sample size (N) and subgroup sizes (n) to select appropriate statistical tests and sensitivity criteria
Distributional Metrics: Skewness and Kurtosis Benchmarks
Two fundamental coefficients mathematically define the shape of your sample distribution: Skewness and Kurtosis. Skewness quantifies the asymmetry of the distribution relative to the mean. A coefficient of zero indicates perfect bilateral symmetry. Positive skewness reflects data concentrated on the lower end with a tail extending toward positive infinity; negative skewness reflects clustering on the high end with an elongated lower tail.
Kurtosis measures the tail weight and peakedness relative to a normal distribution. In modern statistical packages like SPSS and R, 'Excess Kurtosis' is computed where the normal curve equals zero (mesokurtic). Positive kurtosis indicates heavy tails and outlier proneness (leptokurtic), whereas negative kurtosis indicates light tails and platykurtic flatness. For graduate theses and dissertations, the standard citation is George & Mallery (2010, 2020), which deems skewness and kurtosis values between [-2.0, +2.0] as an acceptable approximation of univariate normality. For submissions to high-impact Q1/Q2 journals, Tabachnick & Fidell's (2019) stricter [-1.5, +1.5] benchmark is expected.
- Recorded both raw Skewness and Kurtosis point estimates alongside their respective Standard Errors (SE)
- For small samples (n < 50), calculated the Z-score ratio (Statistic / SE) to verify that |Z| < 1.96 at α = .05
- For large samples (N > 100), disregarded the hypersensitive Z-score ratio and evaluated absolute coefficients against George & Mallery's [-2.0, +2.0] window
Goodness-of-Fit Tests: Shapiro-Wilk (W) vs. Kolmogorov-Smirnov (D)
Statistical software generates two primary goodness-of-fit tests under normality screening: the Shapiro-Wilk test and the Kolmogorov-Smirnov test (with Lilliefors significance correction). In these procedures, the null hypothesis (H0) states that the empirical sample is drawn from a normally distributed population; the alternative hypothesis (H1) posits a significant departure from normality. Consequently, researchers seek a non-significant result (p > .05) to retain H0.
Rigorous Monte Carlo simulation studies by Razali & Wah (2011) demonstrate that the Shapiro-Wilk test possesses substantially greater statistical power than the Kolmogorov-Smirnov test across virtually all sample sizes and non-normal distribution geometries. For small to moderate samples (n < 50 and n ≤ 300), Shapiro-Wilk should always be preferred; Kolmogorov-Smirnov is prone to Type II errors, frequently failing to flag critical departures from normality.
- Evaluated whether the p-value (Sig.) exceeds the .05 significance threshold
- For small to medium samples (n ≤ 300), based the formal conclusion primarily on the Shapiro-Wilk (W) statistic
- For large samples (N > 300), accounted for test hypersensitivity by corroborating numerical p-values with Q-Q plots and distributional indices
Visual and Graphical Diagnostics: Q-Q Plots, Histograms, and Boxplots
Graphical diagnostics represent the most resilient safeguard against the dual risks of test underpowering in small samples and overpowering in large samples (Field, 2018). The Normal Q-Q (Quantile-Quantile) Plot plots observed sample quantiles against theoretical normal quantiles along a 45-degree diagonal reference line. Close adherence of points to this diagonal confirms univariate normality.
A bell-shaped curve overlay on the sample histogram verifies unimodality; distinct bimodal peaks signal that two unmodeled subpopulations (e.g., novices vs. experts) have been conflated. Furthermore, boxplots provide an objective screen for distribution symmetry and flag mild outliers (exceeding 1.5 × IQR, displayed as open circles) and extreme outliers (exceeding 3.0 × IQR, displayed as asterisks).
- Verified that data points in the Normal Q-Q plot trace the 45-degree diagonal reference line without systematic curvilinear trends
- Confirmed that the Detrended Normal Q-Q plot displays a random scatter of points centered around the horizontal zero line
- Inspected histograms to ensure absence of severe multimodality or severe ceiling/floor clustering
- Screened boxplots to identify and isolate data entry typographical errors
3-Tier Fallback Protocol When the Normality Assumption Fails
If numerical tests yield p < .05, skewness and kurtosis breach the [-2.0, +2.0] threshold, and Q-Q plots demonstrate pronounced departure from linearity, researchers must execute a structured fallback protocol. Three scientifically sound avenues exist: The first and most established path is transitioning to non-parametric equivalents: Mann-Whitney U for independent samples t-test, Wilcoxon Signed-Rank for paired samples, Kruskal-Wallis H for One-Way ANOVA, and Spearman's rho for Pearson correlation.
The second approach is mathematical data transformation (e.g., logarithmic log10, square root, or inverse transformation for positively skewed data; reflect-and-log for negative skewness). However, transformations carry a heavy interpretive cost: transformed variables lose their original physical units and become challenging to interpret substantively. The third and modern gold standard is Bootstrapping (1,000 to 5,000 resamples). By computing 95% Bias-Corrected and Accelerated (BCa) confidence intervals, researchers can retain their original parametric model without relying on the Gaussian distributional assumption.
- If transitioning to non-parametric tests, verified underlying distribution shape assumptions (e.g., shape similarity for Mann-Whitney U medians)
- If applying transformations, re-tested the transformed variables for normality and documented the mathematical formula transparently
- Considered 95% BCa Bootstrapping as a robust modern method to preserve original metric scales
- Maintained data integrity: never deleted genuine respondent observations to artificially force p > .05
Official APA 7th Edition Reporting Standards (JARS-Quant)
The Publication Manual of the American Psychological Association (7th ed., Section 6.44 and Chapter 12 JARS-Quant) specifies exact formatting requirements for reporting distributional assumptions in academic manuscripts. A critical stylistic rule is the 'Leading Zero Rule': statistical values that cannot exceed 1.00 by definition (such as p-values, Shapiro-Wilk W, and correlation coefficients) must be reported WITHOUT a leading zero before the decimal point (e.g., W = .98, p = .184, NOT W = 0.98, p = 0.184).
All statistical symbols (N, n, M, SD, W, D, p, Z) must be italicized in both in-text narrative and tabular headers. Your method section must explicitly justify the normality threshold employed, citing foundational literature (e.g., George & Mallery, 2010 or Tabachnick & Fidell, 2019) to ensure resilience during thesis defenses and peer review.
- Omitted leading zeros for all values bounded by 1.00 (e.g., W = .97, p = .214)
- Italicized all statistical symbols (N, n, M, SD, W, D, p, Z)
- Included formal citations for the specific normality cutoff criteria selected
- Clearly reported the final decision regarding parametric versus non-parametric test execution
Turnkey APA 7 Reporting Templates
Peer-reviewed templates for your thesis or manuscript methodology and results sections.
Frequently Asked Questions (FAQs)
Practical methodological solutions for challenging empirical scenarios and reviewer queries.
The Kolmogorov-Smirnov test is p < .05 while Shapiro-Wilk is p > .05. Which one should I report to my reviewers?↓
For sample sizes n < 50 and n ≤ 300, prioritize and report the Shapiro-Wilk test. Monte Carlo simulation evidence by Razali and Wah (2011) demonstrates that Shapiro-Wilk achieves substantially superior statistical power across diverse non-normal distributions, whereas Kolmogorov-Smirnov has lower power and is prone to Type II errors in small to moderate samples. In your methodology section, cite Razali and Wah (2011) to justify your choice.
My sample size is N = 450. The histogram looks like a bell curve, but Shapiro-Wilk yields p < .001. Can I still use parametric tests?↓
Yes. This is the well-known 'large sample paradox'. When sample sizes exceed 300, goodness-of-fit hypothesis tests become hyper-sensitive to trivial, inconsequential departures from symmetry. In accordance with Field (2018) and the Central Limit Theorem, evaluate skewness and kurtosis absolute values against George & Mallery's (2010) [-2.0, +2.0] window and inspect Normal Q-Q plots. If values remain within this threshold and points trace the diagonal line, parametric tests remain valid and robust.
My dissertation committee requested an authoritative citation for acceptable skewness and kurtosis values. Which source should I cite?↓
Cite George, D., & Mallery, P. (2010/2020). IBM SPSS Statistics Step by Step for the [-2.0, +2.0] threshold. If submitting to a high-impact (Q1/Q2) journal in psychology, medicine, or psychometrics, cite Tabachnick, B. G., & Fidell, L. S. (2019). Using Multivariate Statistics for the conservative [-1.5, +1.5] threshold.
Can I test normality on individual 5-point Likert items?↓
No. Individual Likert items are discrete ordinal categories (1, 2, 3, 4, 5) and cannot form a continuous bell-shaped distribution. Normality testing must be reserved for continuous variables (interval/ratio scale) or summated/mean composite scale scores.
When comparing two or more groups (e.g., Male vs. Female), should normality be tested across the whole sample or separately per group?↓
Normality must strictly be tested separately within each individual group. For between-subjects designs (e.g., independent samples t-test, ANOVA), the parametric assumption requires each population cell to be normally distributed. In SPSS, split your file (Data > Split File > Compare groups) or assign the grouping variable to the Factor List under Analyze > Descriptive Statistics > Explore.
Can I delete extreme outliers to achieve p > .05 on a Shapiro-Wilk test?↓
No. Deleting legitimate participant cases to manipulate p-values or artificially force normality constitutes scientific misconduct (cherry-picking and p-hacking). Outliers may only be removed or corrected if verified as data-entry errors (e.g., an age entered as 250). For genuine extreme scores, you must either retain them and report non-parametric tests (e.g., Mann-Whitney U), conduct 95% BCa bootstrapping, or report a transparent sensitivity analysis comparing results with and without the outlier.
Primary Methodological Literature & Authority Standards
- Shapiro, S. S., & Wilk, M. B. (1965): An analysis of variance test for normality (complete samples). Biometrika, 52(3/4), 591–611. https://doi.org/10.1093/biomet/52.3-4.591
- Razali, N. M., & Wah, Y. B. (2011): Power comparisons of Shapiro-Wilk, Kolmogorov-Smirnov, Lilliefors and Anderson-Darling tests. Journal of Statistical Modeling and Analytics, 2(1), 21–33.
- George, D., & Mallery, P. (2020): IBM SPSS Statistics 26 Step by Step: A Simple Guide and Reference (16th ed.). Routledge. https://doi.org/10.4324/9780429056765
- Tabachnick, B. G., & Fidell, L. S. (2019): Using Multivariate Statistics (7th ed.). Pearson.
- Field, A. (2018): Discovering Statistics Using IBM SPSS Statistics (5th ed.). SAGE Publications.
- Hair, J. F., Black, W. C., Babin, B. J., & Anderson, R. E. (2019): Multivariate Data Analysis (8th ed.). Cengage Learning.
- Kline, R. B. (2016): Principles and Practice of Structural Equation Modeling (4th ed.). Guilford Press.
- American Psychological Association (2020): Publication Manual of the APA (7th ed.). Section 6.44 (Leading Zero Rule) & Chapter 12 (JARS-Quant Standards).
Related Methodological & Statistical Guides
AcademicFix provides expert guidance on research design, parametric assumption validation, and transparent APA 7 reporting. We strictly refuse ghostwriting, conducting data analysis in place of the researcher, writing theses, fabricating data, or promising guaranteed committee/journal acceptance. To verify that your empirical findings and formatting conform to international publishing criteria prior to final submission, you may request an independent pre-submission technical review.