The UNIVARIATE Procedure

Creating Summary Plots

When ODS Graphics is enabled, the PLOTS option in the PROC UNIVARIATE statement produces the following diagnostic plots that describe the data distribution:

  • horizontal histogram

  • box plot

  • normal probability plot

  • side-by-side box plots (when you specify a BY variable)

If you specify a WEIGHT statement, PROC UNIVARIATE provides a weighted histogram, a weighted box plot based on the weighted quantiles, and a weighted normal probability plot.

Horizontal Histogram

The vertical axis of the horizontal histogram defines intervals of data values, referred to as bins. Horizontal bars indicate the number of observations that lie within each bin. The number of bins is determined by using the method of Terrell and Scott (1985).

Box Plot

The box plot, also known as a schematic box plot, appears beside the horizontal bar chart and uses the same vertical scale. The box plot provides a visual summary of the data and identifies outliers. The bottom and top edges of the box correspond to the sample 25th (Q1) and 75th (Q3) percentiles. The box length is one interquartile range (Q3 – Q1). The center horizontal line corresponds to the sample median. The central marker corresponds to the sample mean. The vertical lines that project out from the box, called whiskers, extend as far as the data extend, up to and including a distance of 1.5 interquartile ranges. Values farther away are potential outliers, and are identified with markers.

Normal Probability Plot

The normal probability plot plots the empirical quantiles against the quantiles of a standard normal distribution. Markers that indicate the data values are overlaid with a straight reference line that is drawn by using the sample mean and standard deviation. If the data are from a normal distribution, the data values tend to fall along the reference line. The vertical coordinate is the data value, and the horizontal coordinate is normal upper Phi Superscript negative 1 Baseline left-parenthesis v Subscript i Baseline right-parenthesis where

StartLayout 1st Row 1st Column v Subscript i 2nd Column equals 3rd Column StartFraction r Subscript i Baseline minus three-eighths Over n plus one-fourth EndFraction 2nd Row 1st Column normal upper Phi Superscript negative 1 Baseline left-parenthesis dot right-parenthesis 2nd Column equals 3rd Column inverse of the standard normal distribution function 3rd Row 1st Column r Subscript i 2nd Column equals 3rd Column rank of the i th data value when ordered from smallest to largest 4th Row 1st Column n 2nd Column equals 3rd Column number of nonmissing observations EndLayout

For a weighted normal probability plot, the ith ordered observation is plotted against normal upper Phi Superscript negative 1 Baseline left-parenthesis v Subscript i Baseline right-parenthesis where

StartLayout 1st Row 1st Column v Subscript i 2nd Column equals 3rd Column StartStartFraction left-parenthesis 1 minus StartFraction 3 Over 8 i EndFraction right-parenthesis sigma-summation Underscript j equals 1 Overscript i Endscripts w Subscript left-parenthesis j right-parenthesis Baseline OverOver left-parenthesis 1 plus StartFraction 1 Over 4 n EndFraction right-parenthesis sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i Baseline EndEndFraction 2nd Row 1st Column w Subscript left-parenthesis j right-parenthesis 2nd Column equals 3rd Column weight associated with the j th ordered observation EndLayout

When each observation has an identical weight, w Subscript j Baseline equals w, the formula for v Subscript i reduces to the expression for v Subscript i in the unweighted normal probability plot:

v Subscript i Baseline equals StartFraction i minus three-eighths Over n plus one-fourth EndFraction

When the value of VARDEF= is WDF or WEIGHT, a reference line with intercept ModifyingAbove mu With caret and slope ModifyingAbove sigma With caret is added to the plot. When the value of VARDEF= is DF or N, the slope is StartFraction ModifyingAbove sigma With caret Over StartRoot w overbar EndRoot EndFraction where w overbar equals StartFraction sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i Baseline Over n EndFraction is the average weight.

When each observation has an identical weight and the value of VARDEF= is DF, N, or WEIGHT, the reference line reduces to the usual reference line with intercept ModifyingAbove mu With caret and slope ModifyingAbove sigma With caret in the unweighted normal probability plot.

If the data are normally distributed with mean mu and standard deviation sigma, and each observation has an identical weight w, then the points on the plot should lie approximately on a straight line. The intercept for this line is mu. The slope is sigma when VARDEF= is WDF or WEIGHT, and the slope is StartFraction sigma Over StartRoot w EndRoot EndFraction when VARDEF= is DF or N.

Note: You can also use the PROBPLOT statement to produce probability plots, see the section PROBPLOT Statement.

Side-by-Side Box Plots

When you use a BY statement with the PLOTS option, PROC UNIVARIATE produces side-by-side box plots, one for each BY group. The box plots (also known as schematic plots) use a common scale that enables you to compare the data distribution across BY groups. This plot appears after the univariate analyses of all BY groups. Use the NOBYPLOT option to suppress this plot.

Legacy Line Printer Plots

When ODS Graphics is disabled, the PLOTS option in the PROC UNIVARIATE statement produces diagnostic plots by using legacy line printer output.

In line printer output, the horizontal histogram is replaced by either a stem-and-leaf plot (Tukey 1977) or a horizontal bar chart. If any single interval contains more than 49 observations, a horizontal bar chart is produced; otherwise, a stem-and-leaf plot is produced. Both plots provide a method to visualize the overall distribution of the data. However, the stem-and-leaf plot provides more detail because each point in the plot represents an individual data value.

To change the number of stems that the plot displays, use the PLOTSIZE= option to increase or decrease the number of rows in the plot. Instructions that are displayed below the plot explain how to determine the values of the variable. If no instructions appear, you multiply Stem.Leaf by 1 to determine the values of the variable. For example, if the stem value is 10 and the leaf value is 1, then the variable value is approximately 10.1. For the stem-and-leaf plot, the procedure rounds a variable value to the nearest leaf. If the variable value is exactly halfway between two leaves, the value rounds to the nearest leaf with an even integer value. For example, a variable value of 3.15 has a stem value of 3 and a leaf value of 2.

In line printer box plots, extreme values between 1.5 and 3 interquartile ranges from the top or bottom edge of the box are plotted with a zero, and more extreme values are plotted with an asterisk (*).

In line printer probability plots, asterisks (*) indicate the data values and the reference line is drawn by using plus signs (+).

Last updated: April 10, 2023