Methodology & Parametric Assumption Framework

How to Test Data for Normality in Theses and Journal Articles

A definitive methodological decision framework for graduate theses and peer-reviewed journals: Shapiro-Wilk vs. Kolmogorov-Smirnov, skewness and kurtosis benchmarks, Q-Q plot diagnostics, 3-tier non-normality fallbacks, and APA 7 reporting.

By: AcademicFix Editorial Team•Last Reviewed: October 2, 2026•Reading Time: 11 min read
⚡ Quick Decision Rule: Which Normality Metric Takes Precedence?

Normality cannot be reduced to a single p-value. Follow this evidence-based hierarchy tailored to your sample size (N):

  • Small and Moderate Samples (n < 50 and n ≤ 300): The Shapiro-Wilk test is the most powerful test available (Razali & Wah, 2011). Look for p > .05 and verify that data points in the Normal Q-Q plot adhere tightly to the 45-degree diagonal reference line.
  • Skewness and Kurtosis Benchmarks: Across graduate theses and dissertations, the standard citation is George & Mallery (2010, 2020), where coefficients between [-2.0, +2.0] signify acceptable univariate normality. For high-impact Q1/Q2 journals, aim for Tabachnick & Fidell (2019) [-1.5, +1.5].
  • Large Samples (N > 300): Shapiro-Wilk and Kolmogorov-Smirnov become hyper-sensitive, frequently yielding p < .05 on trivial imperfections. Under the Central Limit Theorem, prioritize skewness/kurtosis bounds and Q-Q plots to justify parametric testing.

Normality Decision Matrix & Methodological Standards

Comparative benchmark guidelines across prominent psychometric and statistical authorities.

Standard / TierAcceptable WindowPrimary CitationRecommendationMethodological Justification
Ideal Normal Symmetry-1.00 ≤ Skewness & Kurtosis ≤ +1.00Hair et al. (2019)Unconditional Parametric AcceptanceData adhere closely to the theoretical normal distribution. Independent t-tests, One-Way ANOVA, Pearson correlation, and linear regression can proceed safely.
Conservative / High-Impact Standard-1.50 ≤ Skewness & Kurtosis ≤ +1.50Tabachnick & Fidell (2019)Recommended for Q1 / Q2 JournalsStrict univariate normality boundary widely accepted by leading medical, psychometric, and behavioral science editorial boards.
Graduate Dissertation Benchmark-2.00 ≤ Skewness & Kurtosis ≤ +2.00George & Mallery (2010, 2020)Standard Dissertation ThresholdThe most universally cited rule of thumb across graduate thesis committees and institutional guidelines in social and behavioral sciences.
Z-Score Ratio Rule (Small Samples Only)|Z| < 1.96 (α = .05) or |Z| < 2.58 (α = .01)Field (2018); Tabachnick & Fidell (2019)Valid Strictly for n < 50Calculated as Z = Statistic / SE. In moderate to large samples (N > 100), SE shrinks dramatically, causing harmless deviations to trigger p < .001. Avoid for N > 100.
Structural Equation Modeling (SEM)Skewness < 3.00, Kurtosis < 10.00Kline (2016)Multivariate Analysis CeilingUpper tolerance threshold for confirmatory factor analysis (CFA) and path models in AMOS, LISREL, or Mplus before requiring robust estimators (e.g., Satorra-Bentler / MLM).
Severe Non-Normality / Heavy Tails|Skewness| > 2.00 or |Kurtosis| > 2.00Field (2018)Reject Standard Parametric TestData exhibit marked asymmetry or extreme outliers. Transition to non-parametric equivalents (Mann-Whitney U, Kruskal-Wallis) or run 95% BCa Bootstrapping.

6-Stage Normality Screening & Decision Protocol

From raw data screening to definitive dissertation and journal results reporting.

01

Research Design, Measurement Scale, and Group-Wise Testing

The single most prevalent conceptual error in empirical research is attempting to perform normality tests on individual ordinal Likert survey items (e.g., 1 = Strongly Disagree to 5 = Strongly Agree). Individual Likert responses are discrete, bounded integer categories that inherently cannot form a continuous Gaussian bell curve. Normality testing must be restricted to continuous variables measured on an interval or ratio scale, or to summated composite and mean scores derived from validated multi-item scales.

The second non-negotiable rule is the requirement for group-wise testing in comparative designs: When planning an independent samples t-test (e.g., Treatment vs. Control) or a One-Way ANOVA across multiple conditions, the distributional normality assumption applies to the dependent variable WITHIN EACH SUBGROUP SEPARATELY, not to the pooled aggregate sample. Pooling groups can artificially induce bimodal distributions or mask genuine within-group skewness, rendering inferential statistics invalid.

Never Pool Comparative Groups for Normality Testing: Testing an aggregate column across heterogeneous conditions violates the core assumption of between-subjects designs. If males are normally distributed but females exhibit severe skewness, the independent t-test assumption is breached.
Execution Checklist for This Step:
  • Verified that the target variable is measured on a continuous (interval/ratio) scale or is a composite multi-item index
  • Configured software (e.g., SPSS 'Explore > Factor List' or 'Split File') to test normality separately within each experimental or demographic subgroup
  • Identified total sample size (N) and subgroup sizes (n) to select appropriate statistical tests and sensitivity criteria
02

Distributional Metrics: Skewness and Kurtosis Benchmarks

Two fundamental coefficients mathematically define the shape of your sample distribution: Skewness and Kurtosis. Skewness quantifies the asymmetry of the distribution relative to the mean. A coefficient of zero indicates perfect bilateral symmetry. Positive skewness reflects data concentrated on the lower end with a tail extending toward positive infinity; negative skewness reflects clustering on the high end with an elongated lower tail.

Kurtosis measures the tail weight and peakedness relative to a normal distribution. In modern statistical packages like SPSS and R, 'Excess Kurtosis' is computed where the normal curve equals zero (mesokurtic). Positive kurtosis indicates heavy tails and outlier proneness (leptokurtic), whereas negative kurtosis indicates light tails and platykurtic flatness. For graduate theses and dissertations, the standard citation is George & Mallery (2010, 2020), which deems skewness and kurtosis values between [-2.0, +2.0] as an acceptable approximation of univariate normality. For submissions to high-impact Q1/Q2 journals, Tabachnick & Fidell's (2019) stricter [-1.5, +1.5] benchmark is expected.

Why the Z-Score Ratio Fails in Large Samples: Because Standard Error (SE) is inversely proportional to the square root of sample size (SE = sqrt(6/N)), as N grows, SE collapses toward zero. With N = 500, a negligible skewness of 0.20 yields Z = 3.65 (p < .001). Never rely on Z-score ratios when N > 100.
Execution Checklist for This Step:
  • Recorded both raw Skewness and Kurtosis point estimates alongside their respective Standard Errors (SE)
  • For small samples (n < 50), calculated the Z-score ratio (Statistic / SE) to verify that |Z| < 1.96 at α = .05
  • For large samples (N > 100), disregarded the hypersensitive Z-score ratio and evaluated absolute coefficients against George & Mallery's [-2.0, +2.0] window
03

Goodness-of-Fit Tests: Shapiro-Wilk (W) vs. Kolmogorov-Smirnov (D)

Statistical software generates two primary goodness-of-fit tests under normality screening: the Shapiro-Wilk test and the Kolmogorov-Smirnov test (with Lilliefors significance correction). In these procedures, the null hypothesis (H0) states that the empirical sample is drawn from a normally distributed population; the alternative hypothesis (H1) posits a significant departure from normality. Consequently, researchers seek a non-significant result (p > .05) to retain H0.

Rigorous Monte Carlo simulation studies by Razali & Wah (2011) demonstrate that the Shapiro-Wilk test possesses substantially greater statistical power than the Kolmogorov-Smirnov test across virtually all sample sizes and non-normal distribution geometries. For small to moderate samples (n < 50 and n ≤ 300), Shapiro-Wilk should always be preferred; Kolmogorov-Smirnov is prone to Type II errors, frequently failing to flag critical departures from normality.

The Large Sample Size Paradox (N > 300): With large sample sizes (N > 300 to 500), goodness-of-fit tests possess extreme power and will flag trivial, cosmetically imperfect deviations as p < .001. When N is large, rely on the Central Limit Theorem, Normal Q-Q plots, and absolute skewness/kurtosis bounds.
Execution Checklist for This Step:
  • Evaluated whether the p-value (Sig.) exceeds the .05 significance threshold
  • For small to medium samples (n ≤ 300), based the formal conclusion primarily on the Shapiro-Wilk (W) statistic
  • For large samples (N > 300), accounted for test hypersensitivity by corroborating numerical p-values with Q-Q plots and distributional indices
04

Visual and Graphical Diagnostics: Q-Q Plots, Histograms, and Boxplots

Graphical diagnostics represent the most resilient safeguard against the dual risks of test underpowering in small samples and overpowering in large samples (Field, 2018). The Normal Q-Q (Quantile-Quantile) Plot plots observed sample quantiles against theoretical normal quantiles along a 45-degree diagonal reference line. Close adherence of points to this diagonal confirms univariate normality.

A bell-shaped curve overlay on the sample histogram verifies unimodality; distinct bimodal peaks signal that two unmodeled subpopulations (e.g., novices vs. experts) have been conflated. Furthermore, boxplots provide an objective screen for distribution symmetry and flag mild outliers (exceeding 1.5 × IQR, displayed as open circles) and extreme outliers (exceeding 3.0 × IQR, displayed as asterisks).

Interpreting S-Shaped Curves in Q-Q Plots: If points in a Q-Q plot snake across the 45-degree diagonal in an S-shape, the distribution possesses heavy tails (leptokurtosis) and higher-than-expected outlier concentrations.
Execution Checklist for This Step:
  • Verified that data points in the Normal Q-Q plot trace the 45-degree diagonal reference line without systematic curvilinear trends
  • Confirmed that the Detrended Normal Q-Q plot displays a random scatter of points centered around the horizontal zero line
  • Inspected histograms to ensure absence of severe multimodality or severe ceiling/floor clustering
  • Screened boxplots to identify and isolate data entry typographical errors
05

3-Tier Fallback Protocol When the Normality Assumption Fails

If numerical tests yield p < .05, skewness and kurtosis breach the [-2.0, +2.0] threshold, and Q-Q plots demonstrate pronounced departure from linearity, researchers must execute a structured fallback protocol. Three scientifically sound avenues exist: The first and most established path is transitioning to non-parametric equivalents: Mann-Whitney U for independent samples t-test, Wilcoxon Signed-Rank for paired samples, Kruskal-Wallis H for One-Way ANOVA, and Spearman's rho for Pearson correlation.

The second approach is mathematical data transformation (e.g., logarithmic log10, square root, or inverse transformation for positively skewed data; reflect-and-log for negative skewness). However, transformations carry a heavy interpretive cost: transformed variables lose their original physical units and become challenging to interpret substantively. The third and modern gold standard is Bootstrapping (1,000 to 5,000 resamples). By computing 95% Bias-Corrected and Accelerated (BCa) confidence intervals, researchers can retain their original parametric model without relying on the Gaussian distributional assumption.

Ethical Safeguard: Prohibition of Arbitrary Outlier Pruning: Pruning valid participant cases solely to achieve p > .05 or force a normal distribution constitutes scientific misconduct (cherry-picking and p-hacking). Outliers may only be removed if proven to stem from verifiable data-entry or instrumentation errors.
Execution Checklist for This Step:
  • If transitioning to non-parametric tests, verified underlying distribution shape assumptions (e.g., shape similarity for Mann-Whitney U medians)
  • If applying transformations, re-tested the transformed variables for normality and documented the mathematical formula transparently
  • Considered 95% BCa Bootstrapping as a robust modern method to preserve original metric scales
  • Maintained data integrity: never deleted genuine respondent observations to artificially force p > .05
06

Official APA 7th Edition Reporting Standards (JARS-Quant)

The Publication Manual of the American Psychological Association (7th ed., Section 6.44 and Chapter 12 JARS-Quant) specifies exact formatting requirements for reporting distributional assumptions in academic manuscripts. A critical stylistic rule is the 'Leading Zero Rule': statistical values that cannot exceed 1.00 by definition (such as p-values, Shapiro-Wilk W, and correlation coefficients) must be reported WITHOUT a leading zero before the decimal point (e.g., W = .98, p = .184, NOT W = 0.98, p = 0.184).

All statistical symbols (N, n, M, SD, W, D, p, Z) must be italicized in both in-text narrative and tabular headers. Your method section must explicitly justify the normality threshold employed, citing foundational literature (e.g., George & Mallery, 2010 or Tabachnick & Fidell, 2019) to ensure resilience during thesis defenses and peer review.

Defense Strategy for Thesis Committees: If committee examiners ask why you chose Shapiro-Wilk over Kolmogorov-Smirnov, cite Razali & Wah (2011) for superior statistical power. If questioned on a [-2, +2] cutoff, cite George & Mallery (2010).
Execution Checklist for This Step:
  • Omitted leading zeros for all values bounded by 1.00 (e.g., W = .97, p = .214)
  • Italicized all statistical symbols (N, n, M, SD, W, D, p, Z)
  • Included formal citations for the specific normality cutoff criteria selected
  • Clearly reported the final decision regarding parametric versus non-parametric test execution

Turnkey APA 7 Reporting Templates

Peer-reviewed templates for your thesis or manuscript methodology and results sections.

Methodology Section Paragraph Template:
“Data distributions were formally screened for univariate normality using a combination of numerical tests, distributional metrics, and visual diagnostics prior to inferential testing. Continuous variable distributions were examined within each experimental condition. Given the sample size (N = 184), Kolmogorov-Smirnov (with Lilliefors significance correction) and Shapiro-Wilk tests were inspected alongside skewness and kurtosis indices. Skewness and kurtosis values fell within the recommended empirical thresholds of ±1.5 (Tabachnick & Fidell, 2019). Normal Q-Q plots were inspected to verify that observed data points adhered closely to the 45-degree diagonal reference line. Where severe non-normality was detected, 95% bias-corrected and accelerated (BCa) bootstrap intervals (1,000 resamples) or equivalent non-parametric tests were planned.”
Results Section Paragraph Template:
“Prior to conducting the independent samples t-test, the assumption of normality was evaluated for each group separately. For Group A (n = 92), scores exhibited acceptable symmetry (Skewness = .38, SE = .25; Kurtosis = -.28, SE = .50), and the Shapiro-Wilk test indicated no statistically significant departure from normality (W = .98, p = .184). Similarly, for Group B (n = 92), the distribution conformed to normality (Skewness = .44, SE = .25; Kurtosis = -.19, SE = .50; W = .97, p = .122). Inspection of Normal Q-Q plots confirmed that observed values closely traced the theoretical diagonal reference line. Because both groups satisfied the univariate normality assumption within the ±1.5 threshold (Tabachnick & Fidell, 2019), parametric inferential analyses were deemed robust.”

Frequently Asked Questions (FAQs)

Practical methodological solutions for challenging empirical scenarios and reviewer queries.

The Kolmogorov-Smirnov test is p < .05 while Shapiro-Wilk is p > .05. Which one should I report to my reviewers?↓

For sample sizes n < 50 and n ≤ 300, prioritize and report the Shapiro-Wilk test. Monte Carlo simulation evidence by Razali and Wah (2011) demonstrates that Shapiro-Wilk achieves substantially superior statistical power across diverse non-normal distributions, whereas Kolmogorov-Smirnov has lower power and is prone to Type II errors in small to moderate samples. In your methodology section, cite Razali and Wah (2011) to justify your choice.

My sample size is N = 450. The histogram looks like a bell curve, but Shapiro-Wilk yields p < .001. Can I still use parametric tests?↓

Yes. This is the well-known 'large sample paradox'. When sample sizes exceed 300, goodness-of-fit hypothesis tests become hyper-sensitive to trivial, inconsequential departures from symmetry. In accordance with Field (2018) and the Central Limit Theorem, evaluate skewness and kurtosis absolute values against George & Mallery's (2010) [-2.0, +2.0] window and inspect Normal Q-Q plots. If values remain within this threshold and points trace the diagonal line, parametric tests remain valid and robust.

My dissertation committee requested an authoritative citation for acceptable skewness and kurtosis values. Which source should I cite?↓

Cite George, D., & Mallery, P. (2010/2020). IBM SPSS Statistics Step by Step for the [-2.0, +2.0] threshold. If submitting to a high-impact (Q1/Q2) journal in psychology, medicine, or psychometrics, cite Tabachnick, B. G., & Fidell, L. S. (2019). Using Multivariate Statistics for the conservative [-1.5, +1.5] threshold.

Can I test normality on individual 5-point Likert items?↓

No. Individual Likert items are discrete ordinal categories (1, 2, 3, 4, 5) and cannot form a continuous bell-shaped distribution. Normality testing must be reserved for continuous variables (interval/ratio scale) or summated/mean composite scale scores.

When comparing two or more groups (e.g., Male vs. Female), should normality be tested across the whole sample or separately per group?↓

Normality must strictly be tested separately within each individual group. For between-subjects designs (e.g., independent samples t-test, ANOVA), the parametric assumption requires each population cell to be normally distributed. In SPSS, split your file (Data > Split File > Compare groups) or assign the grouping variable to the Factor List under Analyze > Descriptive Statistics > Explore.

Can I delete extreme outliers to achieve p > .05 on a Shapiro-Wilk test?↓

No. Deleting legitimate participant cases to manipulate p-values or artificially force normality constitutes scientific misconduct (cherry-picking and p-hacking). Outliers may only be removed or corrected if verified as data-entry errors (e.g., an age entered as 250). For genuine extreme scores, you must either retain them and report non-parametric tests (e.g., Mann-Whitney U), conduct 95% BCa bootstrapping, or report a transparent sensitivity analysis comparing results with and without the outlier.

Primary Methodological Literature & Authority Standards

  • Shapiro, S. S., & Wilk, M. B. (1965): An analysis of variance test for normality (complete samples). Biometrika, 52(3/4), 591–611. https://doi.org/10.1093/biomet/52.3-4.591
  • Razali, N. M., & Wah, Y. B. (2011): Power comparisons of Shapiro-Wilk, Kolmogorov-Smirnov, Lilliefors and Anderson-Darling tests. Journal of Statistical Modeling and Analytics, 2(1), 21–33.
  • George, D., & Mallery, P. (2020): IBM SPSS Statistics 26 Step by Step: A Simple Guide and Reference (16th ed.). Routledge. https://doi.org/10.4324/9780429056765
  • Tabachnick, B. G., & Fidell, L. S. (2019): Using Multivariate Statistics (7th ed.). Pearson.
  • Field, A. (2018): Discovering Statistics Using IBM SPSS Statistics (5th ed.). SAGE Publications.
  • Hair, J. F., Black, W. C., Babin, B. J., & Anderson, R. E. (2019): Multivariate Data Analysis (8th ed.). Cengage Learning.
  • Kline, R. B. (2016): Principles and Practice of Structural Equation Modeling (4th ed.). Guilford Press.
  • American Psychological Association (2020): Publication Manual of the APA (7th ed.). Section 6.44 (Leading Zero Rule) & Chapter 12 (JARS-Quant Standards).

Related Methodological & Statistical Guides

AcademicFix Editorial & Methodological Advisory Scope

AcademicFix provides expert guidance on research design, parametric assumption validation, and transparent APA 7 reporting. We strictly refuse ghostwriting, conducting data analysis in place of the researcher, writing theses, fabricating data, or promising guaranteed committee/journal acceptance. To verify that your empirical findings and formatting conform to international publishing criteria prior to final submission, you may request an independent pre-submission technical review.