The UNIVARIATE Procedure

Calculating Percentiles

The UNIVARIATE procedure automatically computes the 1st, 5th, 10th, 25th, 50th, 75th, 90th, 95th, and 99th percentiles (quantiles), as well as the minimum and maximum of each analysis variable. To compute percentiles other than these default percentiles, use the PCTLPTS= and PCTLPRE= options in the OUTPUT statement.

You can specify one of five definitions for computing the percentiles with the PCTLDEF= option. Let n be the number of nonmissing values for a variable, and let x 1 comma x 2 comma ellipsis comma x Subscript n Baseline represent the ordered values of the variable. Let the tth percentile be y, set p equals StartFraction t Over 100 EndFraction, and let

StartLayout 1st Row 1st Column n p 2nd Column equals 3rd Column j plus g 4th Column when PCTLDEF equals 1 comma 2 comma 3 comma or 5 2nd Row 1st Column left-parenthesis n plus 1 right-parenthesis p 2nd Column equals 3rd Column j plus g 4th Column when PCTLDEF equals 4 EndLayout

where j is the integer part of the quantity and g is the fractional part of the quantity. Then the PCTLDEF= option defines the tth percentile, y, as described in Table 30.

Table 30: Percentile Definitions

PCTLDEF Description Formula
1 Weighted average at x Subscript n p y equals left-parenthesis 1 minus g right-parenthesis x Subscript j Baseline plus g x Subscript j plus 1
where x 0 is taken to be x 1
2 Observation numbered closest to np StartLayout 1st Row 1st Column y equals x Subscript j Baseline 2nd Column if g less-than one-half 2nd Row 1st Column y equals x Subscript j Baseline 2nd Column if g equals one-half and j is even 3rd Row 1st Column y equals x Subscript j plus 1 Baseline 2nd Column if g equals one-half and j is odd 4th Row 1st Column y equals x Subscript j plus 1 Baseline 2nd Column if g greater-than one-half EndLayout
3 Empirical distribution function StartLayout 1st Row 1st Column y equals x Subscript j Baseline 2nd Column if g equals 0 2nd Row 1st Column y equals x Subscript j plus 1 Baseline 2nd Column if g greater-than 0 EndLayout
4 Weighted average aimed y equals left-parenthesis 1 minus g right-parenthesis x Subscript j Baseline plus g x Subscript j plus 1
at x Subscript left-parenthesis n plus 1 right-parenthesis p where x Subscript n plus 1 is taken to be x Subscript n
5 Empirical distribution function with averaging StartLayout 1st Row 1st Column y equals one-half left-parenthesis x Subscript j Baseline plus x Subscript j plus 1 Baseline right-parenthesis 2nd Column if g equals 0 2nd Row 1st Column y equals x Subscript j plus 1 Baseline 2nd Column if g greater-than 0 EndLayout


Weighted Percentiles

When you use a WEIGHT statement, the percentiles are computed differently. The 100pth weighted percentile y is computed from the empirical distribution function with averaging:

y equals StartLayout Enlarged left-brace 1st Row 1st Column x 1 2nd Column if w 1 greater-than p upper W 2nd Row 1st Column one-half left-parenthesis x Subscript i Baseline plus x Subscript i plus 1 Baseline right-parenthesis 2nd Column if sigma-summation Underscript j equals 1 Overscript i Endscripts w Subscript j Baseline equals p upper W 3rd Row 1st Column x Subscript i plus 1 Baseline 2nd Column if sigma-summation Underscript j equals 1 Overscript i Endscripts w Subscript j Baseline less-than p upper W less-than sigma-summation Underscript j equals 1 Overscript i plus 1 Endscripts w Subscript j Baseline EndLayout

where w Subscript i is the weight associated with x Subscript i and upper W equals sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i is the sum of the weights.

Note that the PCTLDEF= option is not applicable when a WEIGHT statement is used. However, in this case, if all the weights are identical, the weighted percentiles are the same as the percentiles that would be computed without a WEIGHT statement and with PCTLDEF=5.

Confidence Limits for Percentiles

You can use the CIPCTLNORMAL option to request confidence limits for percentiles, assuming the data are normally distributed. These limits are described in Section 4.4.1 of Hahn and Meeker (1991). When 0 less-than p less-than one-half, the two-sided 100 left-parenthesis 1 minus alpha right-parenthesis percent-sign confidence limits for the 100 pth percentile are

StartLayout 1st Row 1st Column lower limit 2nd Column equals 3rd Column upper X overbar minus g prime left-parenthesis StartFraction alpha Over 2 EndFraction semicolon 1 minus p comma n right-parenthesis s 2nd Row 1st Column upper limit 2nd Column equals 3rd Column upper X overbar minus g prime left-parenthesis 1 minus StartFraction alpha Over 2 EndFraction semicolon p comma n right-parenthesis s EndLayout

where n is the sample size. When one-half less-than-or-equal-to p less-than 1, the two-sided 100 left-parenthesis 1 minus alpha right-parenthesis percent-sign confidence limits for the 100 pth percentile are

StartLayout 1st Row 1st Column lower limit 2nd Column equals 3rd Column upper X overbar plus g prime left-parenthesis StartFraction alpha Over 2 EndFraction semicolon 1 minus p comma n right-parenthesis s 2nd Row 1st Column upper limit 2nd Column equals 3rd Column upper X overbar plus g prime left-parenthesis 1 minus StartFraction alpha Over 2 EndFraction semicolon p comma n right-parenthesis s EndLayout

One-sided 100 left-parenthesis 1 minus alpha right-parenthesis percent-sign confidence bounds are computed by replacing StartFraction alpha Over 2 EndFraction by alpha in the appropriate preceding equation. The factor g prime left-parenthesis gamma comma p comma n right-parenthesis is related to the noncentral t distribution and is described in Owen and Hua (1977) and Odeh and Owen (1980). See Example 3.10.

You can use the CIPCTLDF option to request distribution-free confidence limits for percentiles. In particular, it is not necessary to assume that the data are normally distributed. These limits are described in Section 5.2 of Hahn and Meeker (1991). The two-sided 100 left-parenthesis 1 minus alpha right-parenthesis percent-sign confidence limits for the 100 pth percentile are

StartLayout 1st Row 1st Column lower limit 2nd Column equals 3rd Column upper X Subscript left-parenthesis l right-parenthesis 2nd Row 1st Column upper limit 2nd Column equals 3rd Column upper X Subscript left-parenthesis u right-parenthesis EndLayout

where upper X Subscript left-parenthesis j right-parenthesis is the jth order statistic when the data values are arranged in increasing order:

upper X Subscript left-parenthesis 1 right-parenthesis Baseline less-than-or-equal-to upper X Subscript left-parenthesis 2 right-parenthesis Baseline less-than-or-equal-to midline-horizontal-ellipsis less-than-or-equal-to upper X Subscript left-parenthesis n right-parenthesis

The lower rank l and upper rank u are integers that are symmetric (or nearly symmetric) around left floor n p right floor plus 1, where left floor n p right floor is the integer part of n p and n is the sample size. Furthermore, l and u are chosen so that upper X Subscript left-parenthesis l right-parenthesis and upper X Subscript left-parenthesis u right-parenthesis are as close to upper X Subscript left floor n p right floor plus 1 as possible while satisfying the coverage probability requirement,

upper Q left-parenthesis u minus 1 semicolon n comma p right-parenthesis minus upper Q left-parenthesis l minus 1 semicolon n comma p right-parenthesis greater-than-or-equal-to 1 minus alpha

where upper Q left-parenthesis k semicolon n comma p right-parenthesis is the cumulative binomial probability,

upper Q left-parenthesis k semicolon n comma p right-parenthesis equals sigma-summation Underscript i equals 0 Overscript k Endscripts StartBinomialOrMatrix n Choose i EndBinomialOrMatrix p Superscript i Baseline left-parenthesis 1 minus p right-parenthesis Superscript n minus i

In some cases, the coverage requirement cannot be met, particularly when n is small and p is near 0 or 1. To relax the requirement of symmetry, you can specify CIPCTLDF(TYPE = ASYMMETRIC). This option requests symmetric limits when the coverage requirement can be met, and asymmetric limits otherwise.

If you specify CIPCTLDF(TYPE = LOWER), a one-sided 100 left-parenthesis 1 minus alpha right-parenthesis percent-sign lower confidence bound is computed as upper X Subscript left-parenthesis l right-parenthesis, where l is the largest integer that satisfies the inequality

1 minus upper Q left-parenthesis l minus 1 semicolon n comma p right-parenthesis greater-than-or-equal-to 1 minus alpha where 0 less-than u less-than-or-equal-to n

If you specify CIPCTLDF(TYPE = UPPER), a one-sided 100 left-parenthesis 1 minus alpha right-parenthesis percent-sign upper confidence bound is computed as upper X Subscript left-parenthesis u right-parenthesis, where u is the smallest integer that satisfies the inequality

upper Q left-parenthesis u minus 1 semicolon n comma p right-parenthesis greater-than-or-equal-to 1 minus alpha where 0 less-than u less-than-or-equal-to n

Note that confidence limits for percentiles are not computed when a WEIGHT statement is specified. See Example 3.10.

Last updated: April 10, 2023