The UNIVARIATE Procedure

PROC UNIVARIATE Statement

  • PROC UNIVARIATE <options>;

The PROC UNIVARIATE statement is required to invoke the UNIVARIATE procedure. If you do not specify any other statements, it produces a variety of statistics that summarize the data distribution of each analysis variable:

  • sample moments

  • basic measures of location and variability

  • confidence intervals for the mean, standard deviation, and variance

  • tests for location

  • tests for normality

  • trimmed and Winsorized means

  • robust estimates of scale

  • quantiles and related confidence intervals

  • extreme observations and extreme values

  • frequency counts for observations

  • missing values

In addition, you can use options in the PROC UNIVARIATE statement to do the following:

  • specify the input data set to be analyzed

  • specify a graphics catalog for saving traditional graphics output

  • specify rounding units for variable values

  • specify the definition to use to calculate percentiles

  • specify the divisor to use to calculate variances and standard deviations

  • suppress tables

  • save statistics in an output data set

You can specify the following options:

ALL

requests all statistics and tables that the FREQ, MODES, NEXTRVAL=5, PLOTS, and CIBASIC options generate. If the analysis variables are not weighted, this option also requests the statistics and tables that are generated by the CIPCTLDF, CIPCTLNORMAL, LOCCOUNT, NORMAL, ROBUSTSCALE, TRIMMED=0.25, and WINSORIZED=0.25 options. In producing the output, PROC UNIVARIATE also uses any values that you specify for the ALPHA=, MU0=, NEXTRVAL=, CIBASIC, CIPCTLDF, CIPCTLNORMAL, TRIMMED=, or WINSORIZED= options.

ALPHA=alpha

specifies the level of significance alpha for 100 left-parenthesis 1 minus alpha right-parenthesis percent-sign confidence intervals, where alpha must be between 0 and 1.

Specialized ALPHA= options are available for a number of confidence interval options. For example, you can specify CIBASIC(ALPHA=0.10) to request a table of basic confidence limits at the 90% level. The default value of these options is the value of this (ALPHA=) option. By default, ALPHA=0.05, which results in 95% confidence intervals.

ANNOTATE=SAS-data-set
ANNO=SAS-data-set

specifies an input data set that contains annotate variables as described in SAS/GRAPH: Reference. You can use this SAS-data-set to add features to your traditional graphics. PROC UNIVARIATE adds the features in this SAS-data-set to every graph that it produces. PROC UNIVARIATE does not use this SAS-data-set unless you create a traditional graph by using a plot statement. This option does not apply to ODS Graphics output. Use the ANNOTATE= option in the plot statement if you want to add a feature to a specific graph produced by that statement.

CIBASIC <(options )>

requests confidence limits for the mean, standard deviation, and variance based on the assumption that the data are normally distributed. If you specify this option, you must use the default value of the VARDEF= option, which is DF.

You can specify one or both of the following options within parentheses:

TYPE=LOWER |UPPER |TWOSIDED

specifies the type of confidence limit. By default, TYPE=TWOSIDED.

ALPHA=alpha

specifies the level of significance alpha for 100 left-parenthesis 1 minus alpha right-parenthesis percent-sign confidence intervals, where alpha must be between 0 and 1. The default value is the value of main ALPHA= option, which you can specify in the PROC statement.

CIPCTLDF <(options )>
CIQUANTDF <(options )>

requests confidence limits for quantiles based on a method that is distribution-free; that is, no specific parametric distribution, such as the normal distribution, is assumed for the data. PROC UNIVARIATE uses order statistics (ranks) to compute the confidence limits as described by Hahn and Meeker (1991). This option does not apply if you use a WEIGHT statement.

You can specify one or both of the following options within parentheses:

TYPE=LOWER |UPPER |SYMMETRIC

specifies the type of confidence limit. By default, TYPE=SYMMETRIC.

ALPHA=alpha

specifies the level of significance alpha for 100 left-parenthesis 1 minus alpha right-parenthesis percent-sign confidence intervals. The value alpha must be between 0 and 1. The default value is the value of main ALPHA= option, which you can specify in the PROC statement.

CIPCTLNORMAL <(options )>
CIQUANTNORMAL <(options )>

requests confidence limits for quantiles based on the assumption that the data are normally distributed. The computational method is described in Section 4.4.1 of Hahn and Meeker (1991) and uses the noncentral t distribution as given by Odeh and Owen (1980). This option does not apply if you use a WEIGHT statement

You can specify one or both of the following options within parentheses:

TYPE=LOWER |UPPER |TWOSIDED

specifies the type of confidence limit. By default, TYPE=TWOSIDED.

ALPHA=alpha

specifies the level of significance alpha for 100 left-parenthesis 1 minus alpha right-parenthesis percent-sign confidence intervals. The value alpha must be between 0 and 1. The default value is the value of main ALPHA= option, which you can specify in the PROC statement.

DATA=SAS-data-set

specifies the input SAS data set to be analyzed. If the DATA= option is omitted, the procedure uses the most recently created SAS data set.

EXCLNPWGT
EXCLNPWGTS

excludes observations that have nonpositive weight values (zero or negative) from the analysis. By default, PROC UNIVARIATE counts observations that have negative or zero weights in the total number of observations. This option applies only when you use a WEIGHT statement.

FORCEQN

forces calculation of the robust estimate of scale upper Q Subscript n. Because this calculation is very computationally intensive, upper Q Subscript n is not computed by default for a variable that has more than 65,526 nonmissing observations. On some hosts, upper Q Subscript n cannot be computed at all when there are more than 65,526 nonmissing observations.

FORCESN

forces calculation of the robust estimate of scale upper S Subscript n. Because this calculation is computationally intensive, upper S Subscript n is not computed by default for a variable that has more than 1 million nonmissing observations.

FREQ

requests a frequency table that consists of the variable values, frequencies, cell percentages, and cumulative percentages.

If you specify the WEIGHT statement, PROC UNIVARIATE includes the weighted count in the table and uses its value to compute the percentages.

GOUT=graphics-catalog

specifies the SAS catalog in which PROC UNIVARIATE saves its traditional graphics output. If you omit the libref in the name of the graphics-catalog, PROC UNIVARIATE looks for the catalog in the temporary library called WORK and creates the catalog if it does not exist. This option does not apply to ODS Graphics output.

IDOUT

includes ID variables in the output data that an OUTPUT statement creates. The value of an ID variable in the output data set is its first value from the input data set or BY group. By default, ID variables are not included in the output data sets that an OUTPUT statement creates.

LOCCOUNT

requests a table that shows the number of observations greater than, not equal to, and less than the value of MU0=. PROC UNIVARIATE uses these values to construct the sign test and the signed rank test. This option does not apply if you use a WEIGHT statement.

MODES
MODE

requests a table of all possible modes. By default, when the data contain multiple modes, PROC UNIVARIATE displays the lowest mode in the table of basic statistical measures. When all the values are unique, PROC UNIVARIATE does not produce a table of modes.

MU0=values
LOCATION=values

specifies the value of the mean or location parameter (mu 0) in the null hypothesis for tests of location, which are summarized in the table labeled "Tests for Location: Mu0=value." If you specify one value, PROC UNIVARIATE tests the same null hypothesis for all analysis variables. If you specify multiple values, a VAR statement is required, and PROC UNIVARIATE tests a different null hypothesis for each analysis variable, matching variables and location values by their order in the two lists. By default, MU0=0.

The following statement tests the hypothesis mu 0 equals 0 for the first variable and the hypothesis mu 0 equals 0.5 for the second variable.

proc univariate mu0=0 0.5;
NEXTROBS=n

specifies the number of extreme observations that PROC UNIVARIATE lists in the table of extreme observations. The table lists the n lowest observations and the n highest observations. You can specify NEXTROBS=0 to suppress the table of extreme observations. By default, NEXTROBS=5.

NEXTRVAL=n

specifies the number of extreme values that PROC UNIVARIATE lists in the table of extreme values. The table lists the n lowest unique values and the n highest unique values. By default, NEXTRVAL=0 and no table is displayed.

NOBYPLOT

suppresses side-by-side box plots that are created by default when you use the BY statement and either the ALL option or the PLOTS option in the PROC statement.

NOPRINT

suppresses all the tables of descriptive statistics that the PROC UNIVARIATE statement creates. NOPRINT does not suppress the tables that the HISTOGRAM statement creates. You can use the NOPRINT option in the HISTOGRAM statement to suppress the creation of its tables. Use NOPRINT when you want only to create an output data set that is produced by the OUT= or OUTTABLE= option.

NORMAL
NORMALTEST

requests tests for normality that include a series of goodness-of-fit tests based on the empirical distribution function. The table provides test statistics and p-values for the Shapiro-Wilk test (provided the sample size is less than or equal to 2,000), the Kolmogorov-Smirnov test, the Anderson-Darling test, and the Cramér–von Mises test. This option does not apply if you use a WEIGHT statement.

NOTABCONTENTS

suppresses the table of contents entries for tables of summary statistics that are produced by the PROC UNIVARIATE statement.

NOVARCONTENTS

suppresses grouping entries that are associated with analysis variables in the table of contents. By default, the table of contents lists results that are associated with an analysis variable in a group that has the variable name.

OUTTABLE=SAS-data-set

creates an output data set that contains univariate statistics arranged in tabular form, with one observation per analysis variable. For more information, see the section OUTTABLE= Output Data Set.

PCTLDEF=value
DEF=value

specifies the definition that PROC UNIVARIATE uses to calculate quantiles, where value can be 1, 2, 3, 4, or 5. You cannot use PCTLDEF= when you compute weighted quantiles. For more information, see the section Calculating Percentiles. By default, PCTLDEF=5.

PLOTS <(plot-options )>
PLOT <(plot-options )>

produces a panel of plots for each analysis variable. If ODS Graphics is enabled, the panel contains a horizontal histogram, a box plot, and a normal probability plot. Otherwise, the procedure produces a stem-and-leaf plot (or a horizontal bar chart), a box plot, and a normal probability plot by using legacy line printer output. If you specify a BY statement, side-by-side box plots of the data from the BY groups are displayed following the univariate output for the last BY group.

You can specify the following plot-options to produce titles and footnotes for the plots when ODS Graphics is enabled.

ODSFOOTNOTE=FOOTNOTE |FOOTNOTE1 |'string'

adds a footnote to ODS Graphics output. You can specify the following values:

FOOTNOTE

uses the value of the SAS FOOTNOTE statement as the graph footnote.

FOOTNOTE1

uses the value of the SAS FOOTNOTE statement as the graph footnote.

'string'

uses the specified string as the graph footnote. The string can contain either of the following escaped characters, which are replaced with the appropriate values from the analysis: set-minusn is replaced by the analysis variable name, or set-minusl is replaced by the analysis variable label (or name if the analysis variable has no label).

ODSFOOTNOTE2=FOOTNOTE2 |'string'

adds a secondary footnote to ODS Graphics output. You can specify the following values:

FOOTNOTE2

uses the value of the SAS FOOTNOTE2 statement as the secondary graph footnote.

'string'

uses the specified string as the secondary graph footnote. The string can contain either of the following escaped characters, which are replaced with the appropriate values from the analysis: set-minusn is replaced by the analysis variable name, or set-minusl is replaced by the analysis variable label (or name if the analysis variable has no label).

ODSTITLE=TITLE |TITLE1 |NONE |DEFAULT |LABELFMT |'string'

specifies a title for ODS Graphics output. You can specify the following values:

TITLE

uses the value of SAS TITLE statement as the graph title.

TITLE1

uses the value of SAS TITLE statement as the graph title.

NONE

suppresses all titles from the graph.

DEFAULT

uses the default ODS Graphics title (a descriptive title that consists of the plot type and the analysis variable name).

LABELFMT

uses the default ODS Graphics title with the variable label instead of the variable name.

'string'

uses the specified string as the graph title. The string can contain the following escaped characters, which are replaced with the appropriate values from the analysis: set-minusn is replaced by the analysis variable name, or set-minusl is replaced by the analysis variable label (or name if the analysis variable has no label).

ODSTITLE2=TITLE2 |'string'

specifies a secondary title for ODS Graphics output. You can specify the following values:

TITLE1

uses the value of SAS TITLE2 statement as the secondary graph title.

'string'

uses the specified string as the secondary graph title. The string can contain the following escaped characters, which are replaced with the appropriate values from the analysis: set-minusn is replaced by the analysis variable name, or set-minusl is replaced by the analysis variable label (or name if the analysis variable has no label).

SSPLOT (plot-options )

specifies plot-options that apply only to the side-by-side box plots of BY group data. You can specify any of the plot-options listed previously, with the exceptions of ODSTITLE=LABELFMT and the substitution of an analysis variable name or label in a quoted string.

PLOTSIZE=n

specifies the approximate number of rows to use in legacy line printer plots that are produced when ODS Graphics is disabled and you specify the ALL option or the PLOTS option in the PROC statement. If n is larger than the value of the SAS system option PAGESIZE=, PROC UNIVARIATE uses the value of PAGESIZE=. If n is less than eight, PROC UNIVARIATE uses eight rows to draw the plots.

ROBUSTSCALE

produces a table that contains robust estimates of scale. The statistics include the interquartile range, Gini’s mean difference, the median absolute deviation about the median (MAD), and two statistics proposed by Rousseeuw and Croux (1993): upper Q Subscript n, and upper S Subscript n. For more information, see the section Robust Estimates of Scale. This option does not apply if you use a WEIGHT statement.

ROUND=units

specifies the units to use to round the analysis variables prior to computing statistics. If you specify one unit, PROC UNIVARIATE uses this unit to round all analysis variables. If you specify multiple units, a VAR statement is required, and each unit rounds the values of the corresponding analysis variable. If ROUND=0, no rounding occurs. This option reduces the number of unique variable values, thereby reducing memory requirements for the procedure. For example, to make the rounding unit 1 for the first analysis variable and 0.5 for the second analysis variable, submit the following statements:

proc univariate round=1 0.5;
   var Yieldstrength tenstren;
run;

When a variable value is midway between the two nearest rounded points, the value is rounded to the nearest even multiple of the roundoff value. For example, with a roundoff value of 1, the variable values of –2.5, –2.2, and –1.5 are rounded to –2; the values of –0.5, 0.2, and 0.5 are rounded to 0; and the values of 0.6, 1.2, and 1.4 are rounded to 1.

SUMMARYCONTENTS='string'

specifies the table of contents entry to use for grouping the summary statistics. You can specify SUMMARYCONTENTS='' to suppress the grouping entry.

TRIMMED=values <(options )>
TRIM=values <(options )>

requests a table of trimmed means, where value specifies the number or the proportion of observations that PROC UNIVARIATE trims. If value is the number n of trimmed observations, n must be between 0 and half the number of nonmissing observations. If value is a proportion p between 0 and ½, the number of observations that PROC UNIVARIATE trims is the smallest integer that is greater than or equal to n p, where n is the number of observations. To include confidence limits for the mean and the Student’s t test in the table, you must use the default value of VARDEF=, which is DF. For more information about computing trimmed means, see the section Trimmed Means. This option does not apply if you use a WEIGHT statement.

You can specify one or both of the following options:

TYPE=LOWER |UPPER |TWOSIDED

specifies the type of confidence limit for the mean. By default, TYPE=TWOSIDED.

ALPHA=alpha

specifies the level of significance alpha for 100 left-parenthesis 1 minus alpha right-parenthesis percent-sign confidence intervals, where alpha must be between 0 and 1. By default, ALPHA=0.05, which results in 95% confidence intervals.

VARDEF=divisor

specifies the divisor to use in the calculation of variances and standard deviation. The following table shows the possible values for divisor and associated divisors, where n is the number of observations and w Subscript i is the weight for the ith observation.

Divisor Description Formula Notes
DF Degrees of freedom n minus 1 When you use the WEIGHT statement and VARDEF=DF, the variance is an estimate of sigma squared where the variance of the ith observation is normal v normal a normal r left-parenthesis x Subscript i Baseline right-parenthesis equals StartFraction sigma squared Over w Subscript i Baseline EndFraction. This yields an estimate of the variance of an observation with unit weight.
N Number of observations n
WDF Sum of weights minus one left-parenthesis normal upper Sigma Subscript i Baseline w Subscript i Baseline right-parenthesis minus 1
WEIGHT | WGT Sum of weights normal upper Sigma Subscript i Baseline w Subscript i When you use the WEIGHT statement and VARDEF=WGT, the computed variance is asymptotically (for large n) an estimate of StartFraction sigma squared Over w overbar EndFraction where w overbar is the average weight. This yields an asymptotic estimate of the variance of an observation with average weight.

The procedure computes the variance as StartFraction normal upper C normal upper S normal upper S Over normal d normal i normal v normal i normal s normal o normal r EndFraction where CSS is the corrected sums of squares and equals sigma-summation Underscript i equals 1 Overscript n Endscripts left-parenthesis x Subscript i Baseline minus x overbar right-parenthesis squared. When you weight the analysis variables, normal upper C normal upper S normal upper S equals sigma-summation Underscript i equals 1 Overscript n Endscripts w Subscript i Baseline left-parenthesis x Subscript i Baseline minus x overbar Subscript w Baseline right-parenthesis squared, where x overbar Subscript w is the weighted mean.

By default, VARDEF=DF, which computes the standard error of the mean, confidence limits, and Student’s t test.

WINSORIZED=values <(options )>
WINSOR=values <(options )>

requests of a table of Winsorized means, where value is the number or the proportion of observations that PROC UNIVARIATE uses to compute the Winsorized mean. If the value is the number n of Winsorized observations, n must be between 0 and half the number of nonmissing observations. If value is a proportion p between 0 and ½, the number of observations that PROC UNIVARIATE uses is equal to the smallest integer that is greater than or equal to n p, where n is the number of observations. To include confidence limits for the mean and the Student t test in the table, you must use the default value of the VARDEF= option, which is DF. For more information about computing Winsorized means, see the section Winsorized Means. This option does not apply if you use a WEIGHT statement.

You can specify one or both of the following options:

TYPE=LOWER |UPPER |TWOSIDED

specifies the type of confidence limit for the mean. By default, TYPE=TWOSIDED.

ALPHA=alpha

specifies the level of significance alpha for 100 left-parenthesis 1 minus alpha right-parenthesis percent-sign confidence intervals, alpha must be between 0 and 1. By default, ALPHA=0.05, which results in 95% confidence intervals.

Last updated: April 10, 2023