The UNIVARIATE Procedure
PROC UNIVARIATE Statement
PROC UNIVARIATE <options>;
The PROC UNIVARIATE statement is required to invoke the UNIVARIATE procedure. If you do not specify any other statements, it produces a variety of statistics that summarize the data distribution of each analysis variable:
sample moments
basic measures of location and variability
confidence intervals for the mean, standard deviation, and variance
tests for location
tests for normality
trimmed and Winsorized means
robust estimates of scale
quantiles and related confidence intervals
extreme observations and extreme values
frequency counts for observations
missing values
In addition, you can use options in the PROC UNIVARIATE statement to do the following:
specify the input data set to be analyzed
specify a graphics catalog for saving traditional graphics output
specify rounding units for variable values
specify the definition to use to calculate percentiles
specify the divisor to use to calculate variances and standard deviations
suppress tables
save statistics in an output data set
You can specify the following options:
- ALL
requests all statistics and tables that the FREQ, MODES, NEXTRVAL=5, PLOTS, and CIBASIC options generate. If the analysis variables are not weighted, this option also requests the statistics and tables that are generated by the CIPCTLDF, CIPCTLNORMAL, LOCCOUNT, NORMAL, ROBUSTSCALE, TRIMMED=0.25, and WINSORIZED=0.25 options. In producing the output, PROC UNIVARIATE also uses any values that you specify for the ALPHA=, MU0=, NEXTRVAL=, CIBASIC, CIPCTLDF, CIPCTLNORMAL, TRIMMED=, or WINSORIZED= options.
-
ALPHA=
-
specifies the level of significance
for
confidence intervals, where
must be between 0 and 1.
Specialized ALPHA= options are available for a number of confidence interval options. For example, you can specify CIBASIC(ALPHA=0.10) to request a table of basic confidence limits at the 90% level. The default value of these options is the value of this (ALPHA=) option. By default, ALPHA=0.05, which results in 95% confidence intervals.
-
ANNOTATE=SAS-data-set
ANNO=SAS-data-set specifies an input data set that contains annotate variables as described in SAS/GRAPH: Reference. You can use this SAS-data-set to add features to your traditional graphics. PROC UNIVARIATE adds the features in this SAS-data-set to every graph that it produces. PROC UNIVARIATE does not use this SAS-data-set unless you create a traditional graph by using a plot statement. This option does not apply to ODS Graphics output. Use the ANNOTATE= option in the plot statement if you want to add a feature to a specific graph produced by that statement.
- CIBASIC <(options )>
-
requests confidence limits for the mean, standard deviation, and variance based on the assumption that the data are normally distributed. If you specify this option, you must use the default value of the VARDEF= option, which is DF.
You can specify one or both of the following options within parentheses:
- TYPE=LOWER |UPPER |TWOSIDED
specifies the type of confidence limit. By default, TYPE=TWOSIDED.
-
ALPHA=
specifies the level of significance
for
confidence intervals, where
must be between 0 and 1. The default value is the value of main ALPHA= option, which you can specify in the PROC statement.
-
CIPCTLDF <(options )>
CIQUANTDF <(options )> -
requests confidence limits for quantiles based on a method that is distribution-free; that is, no specific parametric distribution, such as the normal distribution, is assumed for the data. PROC UNIVARIATE uses order statistics (ranks) to compute the confidence limits as described by Hahn and Meeker (1991). This option does not apply if you use a WEIGHT statement.
You can specify one or both of the following options within parentheses:
- TYPE=LOWER |UPPER |SYMMETRIC
specifies the type of confidence limit. By default, TYPE=SYMMETRIC.
-
ALPHA=
specifies the level of significance
for
confidence intervals. The value
must be between 0 and 1. The default value is the value of main ALPHA= option, which you can specify in the PROC statement.
-
CIPCTLNORMAL <(options )>
CIQUANTNORMAL <(options )> -
requests confidence limits for quantiles based on the assumption that the data are normally distributed. The computational method is described in Section 4.4.1 of Hahn and Meeker (1991) and uses the noncentral t distribution as given by Odeh and Owen (1980). This option does not apply if you use a WEIGHT statement
You can specify one or both of the following options within parentheses:
- TYPE=LOWER |UPPER |TWOSIDED
specifies the type of confidence limit. By default, TYPE=TWOSIDED.
-
ALPHA=
specifies the level of significance
for
confidence intervals. The value
must be between 0 and 1. The default value is the value of main ALPHA= option, which you can specify in the PROC statement.
- DATA=SAS-data-set
specifies the input SAS data set to be analyzed. If the DATA= option is omitted, the procedure uses the most recently created SAS data set.
-
EXCLNPWGT
EXCLNPWGTS excludes observations that have nonpositive weight values (zero or negative) from the analysis. By default, PROC UNIVARIATE counts observations that have negative or zero weights in the total number of observations. This option applies only when you use a WEIGHT statement.
- FORCEQN
forces calculation of the robust estimate of scale
. Because this calculation is very computationally intensive,
is not computed by default for a variable that has more than 65,526 nonmissing observations. On some hosts,
cannot be computed at all when there are more than 65,526 nonmissing observations.
- FORCESN
forces calculation of the robust estimate of scale
. Because this calculation is computationally intensive,
is not computed by default for a variable that has more than 1 million nonmissing observations.
- FREQ
-
requests a frequency table that consists of the variable values, frequencies, cell percentages, and cumulative percentages.
If you specify the WEIGHT statement, PROC UNIVARIATE includes the weighted count in the table and uses its value to compute the percentages.
- GOUT=graphics-catalog
-
specifies the SAS catalog in which PROC UNIVARIATE saves its traditional graphics output. If you omit the libref in the name of the graphics-catalog, PROC UNIVARIATE looks for the catalog in the temporary library called WORK and creates the catalog if it does not exist. This option does not apply to ODS Graphics output.
- IDOUT
-
includes ID variables in the output data that an OUTPUT statement creates. The value of an ID variable in the output data set is its first value from the input data set or BY group. By default, ID variables are not included in the output data sets that an OUTPUT statement creates.
- LOCCOUNT
-
requests a table that shows the number of observations greater than, not equal to, and less than the value of MU0=. PROC UNIVARIATE uses these values to construct the sign test and the signed rank test. This option does not apply if you use a WEIGHT statement.
-
MODES
MODE -
requests a table of all possible modes. By default, when the data contain multiple modes, PROC UNIVARIATE displays the lowest mode in the table of basic statistical measures. When all the values are unique, PROC UNIVARIATE does not produce a table of modes.
-
MU0=values
LOCATION=values -
specifies the value of the mean or location parameter (
) in the null hypothesis for tests of location, which are summarized in the table labeled "Tests for Location: Mu0=value." If you specify one value, PROC UNIVARIATE tests the same null hypothesis for all analysis variables. If you specify multiple values, a VAR statement is required, and PROC UNIVARIATE tests a different null hypothesis for each analysis variable, matching variables and location values by their order in the two lists. By default, MU0=0.
The following statement tests the hypothesis
for the first variable and the hypothesis
for the second variable.
proc univariate mu0=0 0.5; - NEXTROBS=n
-
specifies the number of extreme observations that PROC UNIVARIATE lists in the table of extreme observations. The table lists the n lowest observations and the n highest observations. You can specify NEXTROBS=0 to suppress the table of extreme observations. By default, NEXTROBS=5.
- NEXTRVAL=n
specifies the number of extreme values that PROC UNIVARIATE lists in the table of extreme values. The table lists the n lowest unique values and the n highest unique values. By default, NEXTRVAL=0 and no table is displayed.
- NOBYPLOT
-
suppresses side-by-side box plots that are created by default when you use the BY statement and either the ALL option or the PLOTS option in the PROC statement.
- NOPRINT
-
suppresses all the tables of descriptive statistics that the PROC UNIVARIATE statement creates. NOPRINT does not suppress the tables that the HISTOGRAM statement creates. You can use the NOPRINT option in the HISTOGRAM statement to suppress the creation of its tables. Use NOPRINT when you want only to create an output data set that is produced by the OUT= or OUTTABLE= option.
-
NORMAL
NORMALTEST -
requests tests for normality that include a series of goodness-of-fit tests based on the empirical distribution function. The table provides test statistics and p-values for the Shapiro-Wilk test (provided the sample size is less than or equal to 2,000), the Kolmogorov-Smirnov test, the Anderson-Darling test, and the Cramér–von Mises test. This option does not apply if you use a WEIGHT statement.
- NOTABCONTENTS
-
suppresses the table of contents entries for tables of summary statistics that are produced by the PROC UNIVARIATE statement.
- NOVARCONTENTS
-
suppresses grouping entries that are associated with analysis variables in the table of contents. By default, the table of contents lists results that are associated with an analysis variable in a group that has the variable name.
- OUTTABLE=SAS-data-set
-
creates an output data set that contains univariate statistics arranged in tabular form, with one observation per analysis variable. For more information, see the section OUTTABLE= Output Data Set.
-
PCTLDEF=value
DEF=value -
specifies the definition that PROC UNIVARIATE uses to calculate quantiles, where value can be 1, 2, 3, 4, or 5. You cannot use PCTLDEF= when you compute weighted quantiles. For more information, see the section Calculating Percentiles. By default, PCTLDEF=5.
-
PLOTS <(plot-options )>
PLOT <(plot-options )> -
produces a panel of plots for each analysis variable. If ODS Graphics is enabled, the panel contains a horizontal histogram, a box plot, and a normal probability plot. Otherwise, the procedure produces a stem-and-leaf plot (or a horizontal bar chart), a box plot, and a normal probability plot by using legacy line printer output. If you specify a BY statement, side-by-side box plots of the data from the BY groups are displayed following the univariate output for the last BY group.
You can specify the following plot-options to produce titles and footnotes for the plots when ODS Graphics is enabled.
- ODSFOOTNOTE=FOOTNOTE |FOOTNOTE1 |'string'
-
adds a footnote to ODS Graphics output. You can specify the following values:
- FOOTNOTE
uses the value of the SAS FOOTNOTE statement as the graph footnote.
- FOOTNOTE1
uses the value of the SAS FOOTNOTE statement as the graph footnote.
- 'string'
uses the specified string as the graph footnote. The string can contain either of the following escaped characters, which are replaced with the appropriate values from the analysis:
n is replaced by the analysis variable name, or
l is replaced by the analysis variable label (or name if the analysis variable has no label).
- ODSFOOTNOTE2=FOOTNOTE2 |'string'
-
adds a secondary footnote to ODS Graphics output. You can specify the following values:
- FOOTNOTE2
uses the value of the SAS FOOTNOTE2 statement as the secondary graph footnote.
- 'string'
uses the specified string as the secondary graph footnote. The string can contain either of the following escaped characters, which are replaced with the appropriate values from the analysis:
n is replaced by the analysis variable name, or
l is replaced by the analysis variable label (or name if the analysis variable has no label).
- ODSTITLE=TITLE |TITLE1 |NONE |DEFAULT |LABELFMT |'string'
-
specifies a title for ODS Graphics output. You can specify the following values:
- TITLE
uses the value of SAS TITLE statement as the graph title.
- TITLE1
uses the value of SAS TITLE statement as the graph title.
- NONE
suppresses all titles from the graph.
- DEFAULT
uses the default ODS Graphics title (a descriptive title that consists of the plot type and the analysis variable name).
- LABELFMT
uses the default ODS Graphics title with the variable label instead of the variable name.
- 'string'
uses the specified string as the graph title. The string can contain the following escaped characters, which are replaced with the appropriate values from the analysis:
n is replaced by the analysis variable name, or
l is replaced by the analysis variable label (or name if the analysis variable has no label).
- ODSTITLE2=TITLE2 |'string'
-
specifies a secondary title for ODS Graphics output. You can specify the following values:
- TITLE1
uses the value of SAS TITLE2 statement as the secondary graph title.
- 'string'
uses the specified string as the secondary graph title. The string can contain the following escaped characters, which are replaced with the appropriate values from the analysis:
n is replaced by the analysis variable name, or
l is replaced by the analysis variable label (or name if the analysis variable has no label).
- SSPLOT (plot-options )
specifies plot-options that apply only to the side-by-side box plots of BY group data. You can specify any of the plot-options listed previously, with the exceptions of ODSTITLE=LABELFMT and the substitution of an analysis variable name or label in a quoted string.
- PLOTSIZE=n
-
specifies the approximate number of rows to use in legacy line printer plots that are produced when ODS Graphics is disabled and you specify the ALL option or the PLOTS option in the PROC statement. If n is larger than the value of the SAS system option PAGESIZE=, PROC UNIVARIATE uses the value of PAGESIZE=. If n is less than eight, PROC UNIVARIATE uses eight rows to draw the plots.
- ROBUSTSCALE
-
produces a table that contains robust estimates of scale. The statistics include the interquartile range, Gini’s mean difference, the median absolute deviation about the median (MAD), and two statistics proposed by Rousseeuw and Croux (1993):
, and
. For more information, see the section Robust Estimates of Scale. This option does not apply if you use a WEIGHT statement.
- ROUND=units
-
specifies the units to use to round the analysis variables prior to computing statistics. If you specify one unit, PROC UNIVARIATE uses this unit to round all analysis variables. If you specify multiple units, a VAR statement is required, and each unit rounds the values of the corresponding analysis variable. If ROUND=0, no rounding occurs. This option reduces the number of unique variable values, thereby reducing memory requirements for the procedure. For example, to make the rounding unit 1 for the first analysis variable and 0.5 for the second analysis variable, submit the following statements:
proc univariate round=1 0.5; var Yieldstrength tenstren; run;When a variable value is midway between the two nearest rounded points, the value is rounded to the nearest even multiple of the roundoff value. For example, with a roundoff value of 1, the variable values of –2.5, –2.2, and –1.5 are rounded to –2; the values of –0.5, 0.2, and 0.5 are rounded to 0; and the values of 0.6, 1.2, and 1.4 are rounded to 1.
- SUMMARYCONTENTS='string'
-
specifies the table of contents entry to use for grouping the summary statistics. You can specify SUMMARYCONTENTS='' to suppress the grouping entry.
-
TRIMMED=values <(options )>
TRIM=values <(options )> -
requests a table of trimmed means, where value specifies the number or the proportion of observations that PROC UNIVARIATE trims. If value is the number n of trimmed observations, n must be between 0 and half the number of nonmissing observations. If value is a proportion p between 0 and ½, the number of observations that PROC UNIVARIATE trims is the smallest integer that is greater than or equal to
, where n is the number of observations. To include confidence limits for the mean and the Student’s t test in the table, you must use the default value of VARDEF=, which is DF. For more information about computing trimmed means, see the section Trimmed Means. This option does not apply if you use a WEIGHT statement.
You can specify one or both of the following options:
- VARDEF=divisor
-
specifies the divisor to use in the calculation of variances and standard deviation. The following table shows the possible values for divisor and associated divisors, where n is the number of observations and
is the weight for the ith observation.
Divisor Description Formula Notes DF Degrees of freedom When you use the WEIGHT statement and VARDEF=DF, the variance is an estimate of where the variance of the ith observation is
. This yields an estimate of the variance of an observation with unit weight.
N Number of observations n WDF Sum of weights minus one WEIGHT | WGT Sum of weights When you use the WEIGHT statement and VARDEF=WGT, the computed variance is asymptotically (for large n) an estimate of where
is the average weight. This yields an asymptotic estimate of the variance of an observation with average weight.
The procedure computes the variance as
where CSS is the corrected sums of squares and equals
. When you weight the analysis variables,
, where
is the weighted mean.
By default, VARDEF=DF, which computes the standard error of the mean, confidence limits, and Student’s t test.
-
WINSORIZED=values <(options )>
WINSOR=values <(options )> -
requests of a table of Winsorized means, where value is the number or the proportion of observations that PROC UNIVARIATE uses to compute the Winsorized mean. If the value is the number n of Winsorized observations, n must be between 0 and half the number of nonmissing observations. If value is a proportion p between 0 and ½, the number of observations that PROC UNIVARIATE uses is equal to the smallest integer that is greater than or equal to
, where n is the number of observations. To include confidence limits for the mean and the Student t test in the table, you must use the default value of the VARDEF= option, which is DF. For more information about computing Winsorized means, see the section Winsorized Means. This option does not apply if you use a WEIGHT statement.
You can specify one or both of the following options: