SUPERLEARNER Procedure
PROC SUPERLEARNER Statement
The PROC SUPERLEARNER statement invokes the procedure. Table 1 summarizes the available options in this statement.
Table 1: PROC SUPERLEARNER Statement Options
| Option | Description |
|---|
|
DATA= | Specifies the input data table |
|
METHOD= | Specifies the method to use for fitting the meta-learner model |
|
RESTORE= | Specifies the input item store |
|
SEED= | Specifies the seed for the pseudorandom number generator |
You can specify the following options:
-
DATA=libref.data-table
-
names the input data table for PROC SUPERLEARNER to use. The default is the most recently created data table. libref.data-table is a two-level name, where
- libref
refers to a collection of information that is defined in the LIBNAME statement and includes the library, which includes a path to the data. For more information about libref, see the section Using SAS Viya Workbench.
- data-table
specifies the name of the input data table.
-
METHOD=method
-
specifies the meta-learning method to be used to estimate the coefficients of the super learner model. By specifying a meta-learning method, you specify an algorithm that is used to combine the individual base learner models in the library. For more information, see the section Meta-learning Methods. You can specify one of the following methods:
- CCLOGLIK
estimates the coefficients by maximizing the convex-constrained binomial log likelihood. This option is available only if the response variable is binary.
- CCLS
estimates the coefficients by fitting a convex-constrained least squares regression with no intercept. This option is available only if the response variable is continuous.
- CVSELECTOR
estimates the coefficients by assigning a coefficient of 1 to a single base learner model on the basis of cross-validated performance and setting coefficients of all other base learner models to 0.
By default, METHOD=CCLS if the target variable is numeric and METHOD=CCLOGLIK if the target variable is binary.
-
RESTORE=libref.data-table
-
specifies the name of the item store that contains a model that is fitted and stored from a previous analysis. The item store is created by a STORE statement from a previous PROC SUPERLEARNER call.
libref.data-table is a two-level name, where libref refers to the library, and data-table specifies the name of the input data table. For more information about this two-level name, see the section Using SAS Viya Workbench.
You can use the previously fitted model to either score data or compute predictions at fixed values of the predictors. For each task, you use the DATA= option to specify the input data table to be used. To score the input data, you specify the OUTPUT statement to create a data table that contains the observationwise predictions. To compute predictions at fixed values of the predictors, you specify one or more MARGIN statements and specify the MARGINPRED option in the OUTPUT statement. You can also request observationwise predictions from each base learner model that you use to train the super learner model. To do so, specify the LEARNERPRED option in the OUTPUT statement.
Because the previously fitted model is stored and retrieved for the analysis, you cannot include statements and options that are specific to training a model. In particular, the BASELEARNER, CROSSVALIDATION, INPUT, STORE, and TARGET statements are not available when you specify the RESTORE= option.
-
SEED=n
specifies the initial seed to start the pseudorandom number generator for k-fold partitioning and base learner model building, where n is an integer. If you omit the SEED= option or if n is negative or 0, the procedure uses the time of day from the computer’s clock to obtain the initial seed.
Last updated: May 14, 2026