The UNIVARIATE Procedure

Goodness-of-Fit Tests

When you specify the NORMAL option in the PROC UNIVARIATE statement or you request a fitted parametric distribution in the HISTOGRAM statement, the procedure computes goodness-of-fit tests for the null hypothesis that the values of the analysis variable are a random sample from the specified theoretical distribution. See Example 3.22.

When you specify the NORMAL option, these tests, which are summarized in the output table labeled "Tests for Normality," include the following:

  • Shapiro-Wilk test

  • Kolmogorov-Smirnov test

  • Anderson-Darling test

  • Cramér–von Mises test

The Kolmogorov-Smirnov D statistic, the Anderson-Darling statistic, and the Cramér–von Mises statistic are based on the empirical distribution function (EDF). However, some EDF tests are not supported when certain combinations of the parameters of a specified distribution are estimated. See Table 31 for a list of the EDF tests available. You determine whether to reject the null hypothesis by examining the p-value that is associated with a goodness-of-fit statistic. When the p-value is less than the predetermined critical value (alpha), you reject the null hypothesis and conclude that the data did not come from the specified distribution.

If you want to test the normality assumptions for analysis of variance methods, beware of using a statistical test for normality alone. A test’s ability to reject the null hypothesis (known as the power of the test) increases with the sample size. As the sample size becomes larger, increasingly smaller departures from normality can be detected. Because small deviations from normality do not severely affect the validity of analysis of variance tests, it is important to examine other statistics and plots to make a final assessment of normality. The skewness and kurtosis measures and the plots that are provided by the PLOTS option, the HISTOGRAM statement, the PROBPLOT statement, and the QQPLOT statement can be very helpful. For small sample sizes, power is low for detecting larger departures from normality that might be important. To increase the test’s ability to detect such deviations, you might want to declare significance at higher levels, such as 0.15 or 0.20, rather than the often-used 0.05 level. Again, consulting plots and additional statistics can help you assess the severity of the deviations from normality.

Shapiro-Wilk Statistic

If the sample size is less than or equal to 2000 and you specify the NORMAL option, PROC UNIVARIATE computes the Shapiro-Wilk statistic, W (also denoted as upper W Subscript n to emphasize its dependence on the sample size n). The W statistic is the ratio of the best estimator of the variance (based on the square of a linear combination of the order statistics) to the usual corrected sum of squares estimator of the variance (Shapiro and Wilk 1965). When n is greater than three, the coefficients to compute the linear combination of the order statistics are approximated by the method of Royston (1992). The statistic W is always greater than zero and less than or equal to one left-parenthesis 0 less-than upper W less-than-or-equal-to 1 right-parenthesis.

Small values of W lead to the rejection of the null hypothesis of normality. The distribution of W is highly skewed. Seemingly large values of W (such as 0.90) might be considered small and lead you to reject the null hypothesis. The method for computing the p-value (the probability of obtaining a W statistic less than or equal to the observed value) depends on n. For n = 3, the probability distribution of W is known and is used to determine the p-value. For n greater-than 4, a normalizing transformation is computed:

StartLayout 1st Row  upper Z Subscript n Baseline equals StartLayout Enlarged left-brace 1st Row 1st Column left-parenthesis minus log left-parenthesis gamma minus log left-parenthesis 1 minus upper W Subscript n Baseline right-parenthesis right-parenthesis minus mu right-parenthesis slash sigma 2nd Column if 4 less-than-or-equal-to n less-than-or-equal-to 11 2nd Row 1st Column left-parenthesis log left-parenthesis 1 minus upper W Subscript n Baseline right-parenthesis minus mu right-parenthesis slash sigma 2nd Column if 12 less-than-or-equal-to n less-than-or-equal-to 2000 EndLayout EndLayout

The values of sigma, gamma, and mu are functions of n obtained from simulation results. Large values of upper Z Subscript n indicate departure from normality, and because the statistic upper Z Subscript n has an approximately standard normal distribution, this distribution is used to determine the p-values for n greater-than 4.

EDF Goodness-of-Fit Tests

When you fit a parametric distribution, PROC UNIVARIATE provides a series of goodness-of-fit tests based on the empirical distribution function (EDF). The EDF tests offer advantages over traditional chi-square goodness-of-fit test, including improved power and invariance with respect to the histogram midpoints. For a thorough discussion, refer to D’Agostino and Stephens (1986).

The empirical distribution function is defined for a set of n independent observations upper X 1 comma ellipsis comma upper X Subscript n Baseline with a common distribution function upper F left-parenthesis x right-parenthesis. Denote the observations ordered from smallest to largest as upper X Subscript left-parenthesis 1 right-parenthesis Baseline comma ellipsis comma upper X Subscript left-parenthesis n right-parenthesis Baseline. The empirical distribution function, upper F Subscript n Baseline left-parenthesis x right-parenthesis, is defined as

StartLayout 1st Row 1st Column upper F Subscript n Baseline left-parenthesis x right-parenthesis equals 0 comma 2nd Column x less-than upper X Subscript left-parenthesis 1 right-parenthesis Baseline 2nd Row 1st Column upper F Subscript n Baseline left-parenthesis x right-parenthesis equals StartFraction i Over n EndFraction comma 2nd Column upper X Subscript left-parenthesis i right-parenthesis Baseline less-than-or-equal-to x less-than upper X Subscript left-parenthesis i plus 1 right-parenthesis Baseline 3rd Column i equals 1 comma ellipsis comma n minus 1 3rd Row 1st Column upper F Subscript n Baseline left-parenthesis x right-parenthesis equals 1 comma 2nd Column upper X Subscript left-parenthesis n right-parenthesis Baseline less-than-or-equal-to x EndLayout

Note that upper F Subscript n Baseline left-parenthesis x right-parenthesis is a step function that takes a step of height StartFraction 1 Over n EndFraction at each observation. This function estimates the distribution function upper F left-parenthesis x right-parenthesis. At any value x, upper F Subscript n Baseline left-parenthesis x right-parenthesis is the proportion of observations less than or equal to x, while upper F left-parenthesis x right-parenthesis is the probability of an observation less than or equal to x. EDF statistics measure the discrepancy between upper F Subscript n Baseline left-parenthesis x right-parenthesis and upper F left-parenthesis x right-parenthesis.

The computational formulas for the EDF statistics make use of the probability integral transformation upper U equals upper F left-parenthesis upper X right-parenthesis. If upper F left-parenthesis upper X right-parenthesis is the distribution function of X, the random variable U is uniformly distributed between 0 and 1.

Given n observations upper X Subscript left-parenthesis 1 right-parenthesis Baseline comma ellipsis comma upper X Subscript left-parenthesis n right-parenthesis Baseline, the values upper U Subscript left-parenthesis i right-parenthesis Baseline equals upper F left-parenthesis upper X Subscript left-parenthesis i right-parenthesis Baseline right-parenthesis are computed by applying the transformation, as discussed in the next three sections.

PROC UNIVARIATE provides three EDF tests:

  • Kolmogorov-Smirnov

  • Anderson-Darling

  • Cramér–von Mises

The following sections provide formal definitions of these EDF statistics.

Kolmogorov D Statistic

The Kolmogorov-Smirnov statistic (D) is defined as

upper D equals sup Subscript x Baseline StartAbsoluteValue upper F Subscript n Baseline left-parenthesis x right-parenthesis minus upper F left-parenthesis x right-parenthesis EndAbsoluteValue

The Kolmogorov-Smirnov statistic belongs to the supremum class of EDF statistics. This class of statistics is based on the largest vertical difference between upper F left-parenthesis x right-parenthesis and upper F Subscript n Baseline left-parenthesis x right-parenthesis.

The Kolmogorov-Smirnov statistic is computed as the maximum of upper D Superscript plus and upper D Superscript minus, where upper D Superscript plus is the largest vertical distance between the EDF and the distribution function when the EDF is greater than the distribution function, and upper D Superscript minus is the largest vertical distance when the EDF is less than the distribution function.

StartLayout 1st Row 1st Column upper D Superscript plus 2nd Column equals 3rd Column max Underscript i Endscripts left-parenthesis StartFraction i Over n EndFraction minus upper U Subscript left-parenthesis i right-parenthesis Baseline right-parenthesis 2nd Row 1st Column upper D Superscript minus 2nd Column equals 3rd Column max Underscript i Endscripts left-parenthesis upper U Subscript left-parenthesis i right-parenthesis Baseline minus StartFraction i minus 1 Over n EndFraction right-parenthesis 3rd Row 1st Column upper D 2nd Column equals 3rd Column max left-parenthesis upper D Superscript plus Baseline comma upper D Superscript minus Baseline right-parenthesis EndLayout

PROC UNIVARIATE uses a modified Kolmogorov D statistic to test the data against a normal distribution with mean and variance equal to the sample mean and variance.

Anderson-Darling Statistic

The Anderson-Darling statistic and the Cramér–von Mises statistic belong to the quadratic class of EDF statistics. This class of statistics is based on the squared difference left-parenthesis upper F Subscript n Baseline left-parenthesis x right-parenthesis minus upper F left-parenthesis x right-parenthesis right-parenthesis squared. Quadratic statistics have the following general form:

upper Q equals n integral Subscript negative normal infinity Superscript plus normal infinity Baseline left-parenthesis upper F Subscript n Baseline left-parenthesis x right-parenthesis minus upper F left-parenthesis x right-parenthesis right-parenthesis squared psi left-parenthesis x right-parenthesis d upper F left-parenthesis x right-parenthesis

The function psi left-parenthesis x right-parenthesis weights the squared difference left-parenthesis upper F Subscript n Baseline left-parenthesis x right-parenthesis minus upper F left-parenthesis x right-parenthesis right-parenthesis squared.

The Anderson-Darling statistic (upper A squared) is defined as

upper A squared equals n integral Subscript negative normal infinity Superscript plus normal infinity Baseline left-parenthesis upper F Subscript n Baseline left-parenthesis x right-parenthesis minus upper F left-parenthesis x right-parenthesis right-parenthesis squared left-bracket upper F left-parenthesis x right-parenthesis left-parenthesis 1 minus upper F left-parenthesis x right-parenthesis right-parenthesis right-bracket Superscript negative 1 Baseline d upper F left-parenthesis x right-parenthesis

Here the weight function is psi left-parenthesis x right-parenthesis equals left-bracket upper F left-parenthesis x right-parenthesis left-parenthesis 1 minus upper F left-parenthesis x right-parenthesis right-parenthesis right-bracket Superscript negative 1.

The Anderson-Darling statistic is computed as

upper A squared equals negative n minus StartFraction 1 Over n EndFraction sigma-summation Underscript i equals 1 Overscript n Endscripts left-bracket left-parenthesis 2 i minus 1 right-parenthesis log upper U Subscript left-parenthesis i right-parenthesis Baseline plus left-parenthesis 2 n plus 1 minus 2 i right-parenthesis log left-parenthesis 1 minus upper U Subscript left-parenthesis i right-parenthesis Baseline right-parenthesis right-bracket
Cramér–von Mises Statistic

The Cramér–von Mises statistic (upper W squared) is defined as

upper W squared equals n integral Subscript negative normal infinity Superscript plus normal infinity Baseline left-parenthesis upper F Subscript n Baseline left-parenthesis x right-parenthesis minus upper F left-parenthesis x right-parenthesis right-parenthesis squared d upper F left-parenthesis x right-parenthesis

Here the weight function is psi left-parenthesis x right-parenthesis equals 1.

The Cramér–von Mises statistic is computed as

upper W squared equals sigma-summation Underscript i equals 1 Overscript n Endscripts left-parenthesis upper U Subscript left-parenthesis i right-parenthesis Baseline minus StartFraction 2 i minus 1 Over 2 n EndFraction right-parenthesis squared plus StartFraction 1 Over 12 n EndFraction
Probability Values of EDF Tests

Once the EDF test statistics are computed, PROC UNIVARIATE computes the associated probability values (p-values).

For the Gumbel, inverse Gaussian, generalized Pareto, and Rayleigh distributions, PROC UNIVARIATE computes associated probability values (p-values) by resampling from the estimated distribution. By default 500 EDF test statistics are computed and then compared to the EDF test statistic for the specified (fitted) distribution. The number of samples can be controlled by setting EDFNSAMPLES=n. For example, to request Gumbel distribution Goodness-of-Fit test p-values based on 5000 simulations, use the following statement:

proc univariate data=test;
   histogram / gumbel(edfnsamples=5000);
run;

For the beta, exponential, gamma, lognormal, normal, power function, and Weibull distributions the UNIVARIATE procedure uses internal tables of probability levels similar to those given by D’Agostino and Stephens (1986). If the value is between two probability levels, then linear interpolation is used to estimate the probability value.

The probability value depends upon the parameters that are known and the parameters that are estimated for the distribution. Table 31 summarizes different combinations fitted for which EDF tests are available.

Table 31: Availability of EDF Tests

Distribution Parameters Tests Available
Threshold Scale Shape
Beta theta known sigma known alpha comma beta known All
theta known sigma known alpha comma beta less-than 5 unknown All
Exponential theta known, sigma known All
theta known sigma unknown All
theta unknown sigma known All
theta unknown sigma unknown All
Gamma theta known sigma known alpha known All
theta known sigma unknown alpha known All
theta known sigma known alpha unknown All
theta known sigma unknown alpha greater-than 1 unknown All
theta unknown sigma known alpha greater-than 1 known All
theta unknown sigma unknown alpha greater-than 1 known All
theta unknown sigma known alpha greater-than 1 unknown All
theta unknown sigma unknown alpha greater-than 1 unknown All
Lognormal theta known zeta known sigma known All
theta known zeta known sigma unknown upper A squared and upper W squared
theta known zeta unknown sigma known upper A squared and upper W squared
theta known zeta unknown sigma unknown All
theta unknown zeta known sigma less-than 3 known All
theta unknown zeta known sigma less-than 3 unknown All
theta unknown zeta unknown sigma less-than 3 known All
theta unknown zeta unknown sigma less-than 3 unknown All
Normal theta known sigma known All
theta known sigma unknown upper A squared and upper W squared
theta unknown sigma known upper A squared and upper W squared
theta unknown sigma unknown All
Power function theta known sigma known alpha known All
theta known sigma known alpha less-than 5 unknown All
Weibull theta known sigma known c known All
theta known sigma unknown c known upper A squared and upper W squared
theta known sigma known c unknown upper A squared and upper W squared
theta known sigma unknown c unknown upper A squared and upper W squared
theta unknown sigma known c greater-than 2 known All
theta unknown sigma unknown c greater-than 2 known All
theta unknown sigma known c greater-than 2 unknown All
theta unknown sigma unknown c greater-than 2 unknown All


Last updated: April 10, 2023