The UNIVARIATE Procedure

Descriptive Statistics

This section provides computational details for the descriptive statistics that are computed with the PROC UNIVARIATE statement. These statistics can also be saved in an OUT= data set by specifying keywords listed in Table 14 in the OUTPUT statement.

Standard algorithms (Fisher 1973) are used to compute the moment statistics. The computational methods used by the UNIVARIATE procedure are consistent with those used by other SAS procedures for calculating descriptive statistics.

The following sections give specific details on a number of statistics calculated by the UNIVARIATE procedure.

Mean

The sample mean is calculated as

x overbar Subscript w Baseline equals StartFraction sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i Baseline x Subscript i Baseline Over sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i Baseline EndFraction

where n is the number of nonmissing values for a variable, x Subscript i is the ith value of the variable, and w Subscript i is the weight associated with the ith value of the variable. If there is no WEIGHT variable, the formula reduces to

x overbar equals StartFraction 1 Over n EndFraction sigma-summation Underscript i equals 1 Overscript n Endscripts x Subscript i

Sum

The sum is calculated as sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i Baseline x Subscript i, where n is the number of nonmissing values for a variable, x Subscript i is the ith value of the variable, and w Subscript i is the weight associated with the ith value of the variable. If there is no WEIGHT variable, the formula reduces to sigma-summation Underscript i equals 1 Overscript n Endscripts x Subscript i.

Sum of the Weights

The sum of the weights is calculated as sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i Baseline, where n is the number of nonmissing values for a variable and w Subscript i is the weight associated with the ith value of the variable. If there is no WEIGHT variable, the sum of the weights is n.

Variance

The variance is calculated as

StartFraction 1 Over d EndFraction sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i Baseline left-parenthesis x Subscript i Baseline minus x overbar Subscript w Baseline right-parenthesis squared

where n is the number of nonmissing values for a variable, x Subscript i is the ith value of the variable, x overbar Subscript w is the weighted mean, w Subscript i is the weight associated with the ith value of the variable, and d is the divisor controlled by the VARDEF= option in the PROC UNIVARIATE statement:

d equals StartLayout Enlarged left-brace 1st Row 1st Column n minus 1 2nd Column if VARDEF equals DF left-parenthesis default right-parenthesis 2nd Row 1st Column n 2nd Column if VARDEF equals upper N 3rd Row 1st Column left-parenthesis sigma-summation Underscript i Endscripts w Subscript i Baseline right-parenthesis minus 1 2nd Column if VARDEF equals WDF 4th Row 1st Column sigma-summation Underscript i Endscripts w Subscript i Baseline 2nd Column if VARDEF equals WEIGHT vertical-bar WGT EndLayout

If there is no WEIGHT variable, the formula reduces to

StartFraction 1 Over d EndFraction sigma-summation Underscript i equals 1 Overscript n Endscripts left-parenthesis x Subscript i Baseline minus x overbar right-parenthesis squared

Standard Deviation

The standard deviation is calculated as

s Subscript w Baseline equals StartRoot StartFraction 1 Over d EndFraction sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i Baseline left-parenthesis x Subscript i Baseline minus x overbar Subscript w Baseline right-parenthesis squared EndRoot

where n is the number of nonmissing values for a variable, x Subscript i is the ith value of the variable, x overbar Subscript w is the weighted mean, w Subscript i is the weight associated with the ith value of the variable, and d is the divisor controlled by the VARDEF= option in the PROC UNIVARIATE statement. If there is no WEIGHT variable, the formula reduces to

s equals StartRoot StartFraction 1 Over d EndFraction sigma-summation Underscript i equals 1 Overscript n Endscripts left-parenthesis x Subscript i Baseline minus x overbar right-parenthesis squared EndRoot

Skewness

The sample skewness, which measures the tendency of the deviations to be larger in one direction than in the other, is calculated as in Table 28, depending on the VARDEF= option.

Table 28: Formulas for Skewness

VARDEF Formula
DF (default) StartFraction n Over left-parenthesis n minus 1 right-parenthesis left-parenthesis n minus 2 right-parenthesis EndFraction sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i Superscript 3 slash 2 Baseline left-parenthesis StartFraction x Subscript i Baseline minus x overbar Subscript w Baseline Over s Subscript w Baseline EndFraction right-parenthesis cubed
N StartFraction 1 Over n EndFraction sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i Superscript 3 slash 2 Baseline left-parenthesis StartFraction x Subscript i Baseline minus x overbar Subscript w Baseline Over s Subscript w Baseline EndFraction right-parenthesis cubed
WDF Missing
WEIGHT | WGT Missing


where n is the number of nonmissing values for a variable, x Subscript i is the ith value of the variable, x overbar Subscript w is the sample average, s is the sample standard deviation, and w Subscript i is the weight associated with the ith value of the variable. If VARDEF=DF, then n must be greater than 2. If there is no WEIGHT variable, then w Subscript i Baseline equals 1 for all i equals 1 comma ellipsis comma n.

The sample skewness can be positive or negative; it measures the asymmetry of the data distribution and estimates the theoretical skewness StartRoot beta 1 EndRoot equals mu 3 mu 2 Superscript negative three-halves, where mu 2 and mu 3 are the second and third central moments. Observations that are normally distributed should have a skewness near zero.

Kurtosis

The sample kurtosis, which measures the heaviness of tails, is calculated as in Table 29, depending on the VARDEF= option.

Table 29: Formulas for Kurtosis

VARDEF Formula
DF (default) StartFraction n left-parenthesis n plus 1 right-parenthesis Over left-parenthesis n minus 1 right-parenthesis left-parenthesis n minus 2 right-parenthesis left-parenthesis n minus 3 right-parenthesis EndFraction sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i Superscript 2 Baseline left-parenthesis StartFraction x Subscript i Baseline minus x overbar Subscript w Baseline Over s Subscript w Baseline EndFraction right-parenthesis Superscript 4 minus StartFraction 3 left-parenthesis n minus 1 right-parenthesis squared Over left-parenthesis n minus 2 right-parenthesis left-parenthesis n minus 3 right-parenthesis EndFraction
N StartFraction 1 Over n EndFraction sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i Superscript 2 Baseline left-parenthesis StartFraction x Subscript i Baseline minus x overbar Subscript w Baseline Over s Subscript w Baseline EndFraction right-parenthesis Superscript 4 minus 3
WDF Missing
WEIGHT | WGT Missing


where n is the number of nonmissing values for a variable, x Subscript i is the ith value of the variable, x overbar Subscript w is the sample average, s Subscript w is the sample standard deviation, and w Subscript i is the weight associated with the ith value of the variable. If VARDEF=DF, then n must be greater than 3. If there is no WEIGHT variable, then w Subscript i Baseline equals 1 for all i equals 1 comma ellipsis comma n.

The sample kurtosis measures the heaviness of the tails of the data distribution. It estimates the adjusted theoretical kurtosis denoted as beta 2 minus 3, where beta 2 equals StartFraction mu 4 Over mu 2 squared EndFraction, and mu 4 is the fourth central moment. Observations that are normally distributed should have a kurtosis near zero.

Coefficient of Variation (CV)

The coefficient of variation is calculated as

upper C upper V equals StartFraction 100 times s Subscript w Baseline Over x overbar Subscript w Baseline EndFraction

Geometric Mean

The geometric mean is calculated as

left-parenthesis product Underscript i equals 1 Overscript n Endscripts x Subscript i Superscript w Super Subscript i Superscript Baseline right-parenthesis Superscript 1 slash sigma-summation Underscript i equals 1 Overscript n Endscripts w Super Subscript i

where n is the number of nonmissing values for a variable, x Subscript i is the ith value of the variable, and w Subscript i is the weight associated with the ith value of the variable.

If there is no WEIGHT variable, the formula reduces to

left-parenthesis product Underscript i equals 1 Overscript n Endscripts x Subscript i Baseline right-parenthesis Superscript 1 slash n

If any x Subscript i is negative, the geometric mean is set to missing.

Last updated: April 10, 2023