BART Procedure

Overview: BART Procedure

The BART procedure fits Bayesian additive regression trees (BART) models in SAS Viya Workbench for response variables that follow a normal or binary distribution.

A BART model uses a sum-of-trees ensemble to approximate the conditional mean of a response variable (also called a target variable), given a vector of input variables (also called predictor variables). Training a BART model requires you to specify a sum-of-trees ensemble, a likelihood function, and a prior distribution on the model parameters. The BART procedure uses a Bayesian backfitting Markov chain Monte Carlo (MCMC) algorithm to generate posterior samples of the sum-of-trees ensemble and to fit the model. Predictions from the BART model are obtained by averaging predictions from the posterior samples. You can assess uncertainty in the BART model predictions by examining the variability of the posterior samples. PROC BART supports models of normally distributed response variables and probit models of binary response variables.

The sum-of-trees ensemble that is used in a BART model consists of multiple binary decision trees. A decision tree is a type of predictive model that has been developed independently in the statistics community and the machine learning community. The predictor variables that a tree model uses can be categorical or continuous. The set of all possible combinations of the predictor variables is called the predictor space. A single decision tree partitions the predictor space into a set of nonoverlapping segments. In the terminology of the tree metaphor, segments in the partition correspond to the terminal nodes, or leaves, of the tree. The partition is described by a series of splitting rules that are used to route an observation, starting from the root node of the tree, to one of the terminal nodes. The data in a leaf are used to estimate parameters that enable you to predict the value of the response.

In the terminology of machine learning, BART models are predictive models. A predictive model defines a relationship between the input variables and a target variable. The purpose of a predictive model is to predict a target value from the inputs. You create a predictive model by using training data in which the target values are known. You can then apply the model to observations in which the target is unknown. If the predictions fit the new data well, the model is said to generalize well. Good generalization is the primary goal of predictive modeling. BART models use a regularization prior for the sum-of-trees ensemble in order to limit overfitting of the training data and obtain good generalization. The regularization prior controls the complexity of the trees in the ensemble and limits the contribution of an individual tree to the ensemble prediction.

For more information about BART models, see the section Details: BART Procedure.

Last updated: May 14, 2026