BART Procedure

OUTPUT Statement

  • OUTPUT OUT=libref.data-table<SAMPLEPRED><ALPHA=number> <COPYVARS=(variables)><keyword <=name>>…<keyword <=name>>;

The OUTPUT statement creates a data table that contains observationwise statistics that PROC BART computes after fitting the model. To avoid data duplication when you have large data tables, the variables in the input data table are not included in the output data table unless you specify them in the COPYVARS= option.

The computation of the output statistics is based on the final fitted estimates. If an error occurs in fitting the model, then the output data table is not created.

You must specify the following option:

OUT=libref.data-table

names the output data table for PROC BART to use. You must specify this option before any other options. libref.data-table is a two-level name, where

libref

refers to a collection of information that is defined in the LIBNAME statement and includes the library, which includes a path to where the data table is to be stored. For more information about libref, see the section Using SAS Viya Workbench.

data-table

specifies the name of the output data table.

You can also specify the following syntax elements:

ALPHA=number

specifies the level for the construction of equal-tail credible limits. The level is 1 – number. The value of number must be between 0 and 1, exclusive.

By default, ALPHA=0.05.

COPYVAR=variable
COPYVARS=(variables)

transfers one or more variables from the input data table to the output data table.

INTO<(cutpoint)>=<name>

names the variable that contains the level of the response into which an observation is classified for binary response models. The default name is _INTO_. If the predicted probability of an observation equals or exceeds the cutpoint, the observation is classified as an event; otherwise it is classified as a nonevent. You can specify the cutpoint value as a number between 0 and 1, exclusive. The default value is 0.5.

LCL<=name>

computes the equal-tail lower credible limit. The default name is Lcl. You can set the credible limit by specifying the ALPHA= option.

PRED<=name>
PREDICTED<=name>
P<=name>

computes predicted values for the response variable. For observations in which the response variable is missing, the predicted values are computed even though these observations do not affect the model fit. The default name is Pred.

RESIDUAL<=name>
RESID<=name>
R<=name>

computes the raw residual, , where is the estimate of the predicted mean. The default name is Residual.

ROLE<=name>

specifies the numeric variable that indicates the role that each observation plays in fitting the model. The default name is _ROLE_. Table 3 shows how this variable is interpreted for each observation.

Table 3: Role Interpretation

Value Observation Role
0 Not used
1 Training
3 Testing


If you do not partition the input data by specifying a PARTITION statement, then the role variable value is 1 for observations that are used to fit the model and 0 for observations that are not used.

SAMPLEPRED

produces predicted values and raw residuals for each posterior sample saved for prediction if you specify the PRED and RESIDUAL keywords.

UCL<=name>

computes the equal-tail upper credible limit. The default name is Ucl. You can set the credible level by specifying the ALPHA= option.

Last updated: May 14, 2026