The NNET Procedure

TRAIN Statement

  • TRAIN OUTMODEL=CAS-libref.data-table <options>;

The TRAIN statement causes the NNET procedure to use the training data that are specified in the PROC NNET statement to train a neural network model whose structure is specified in the ARCHITECTURE, INPUT, TARGET, and HIDDEN statements. The goal of training is to determine a set of network weights that best predicts the targets in the training data while still doing a good job of predicting targets of unseen data (that is, generalizing well and not overfitting).

Training starts with a pseudorandomly generated set of initial weights. PROC NNET then computes the objective function for the training partition, and the optimization algorithm adjusts the weights. This process is repeated until any one of the following conditions is met:

  • The objective function that is computed using the training partition stops improving.

  • The objective function that is computed using the validation partition stops improving.

  • The process has been repeated the number of times specified in the MAXITER= and MAXTIME= options in the OPTIMIZATION statement.

When you are training, you must include exactly one TRAIN statement. The TRAIN statement is not allowed when you are doing stand-alone scoring.

You must specify the following option:

OUTMODEL=CAS-libref.data-table

specifies the final model from training. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.

You can use the model data table later to score a different input data table as long as the variable names and types of the variables in the new input data table match those in the training data table.

You can also specify the following options:

DROPOUTHIDDEN=ratio

specifies the dropout ratio of hidden layers. This option is valid only when you specify ALGORITHM=SGD or ADAM in the OPTIMIZATION statement and when all the connections use the linear combination function. The ratio must be between 0 and 1, inclusive.

By default, DROPOUTHIDDEN=0.

DROPOUTINPUT=ratio

specifies the dropout ratio of input layers. This option is valid only when you specify ALGORITHM=SGD or ADAM in the OPTIMIZATION statement and when all the connections use the linear combination function. The ratio must be between 0 and 1, inclusive.

By default, DROPOUTINPUT=0.

NUMTRIES=number

specifies the number of times the network is to be trained using a different starting point. Specifying this option helps ensure that the optimizer finds the table of weights that truly minimizes the objective function and does not return a local minimum. The value of number must be an integer between 1 and 20,000, inclusive. By default, NUMTRIES=1.

Note: When NUMTRIES > 1, the ODS tables "OptIterHistory" and "ConvergenceStatus" are suppressed.

RESUME

trains with the initial weight that is specified in the INMODEL= option in the PROC NNET statement. If you specify the RESUME option, you must also specify the INMODEL= option.

STAGNATION=number

specifies the number of iterations that result in no improvement for the validation subset before early stopping takes effect during training. This option is valid only when the VALIDATION= option or the PARTITION statement is specified.

By default, STAGNATION=3.

VALIDATION=CAS-libref.data-table

specifies a separate data table for validation during training. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.

If you specify both the VALIDATION= option and the PARTITION statement, the PARTITION statement is ignored. The VALIDATION= data table must have the same variables that you specify in the DATA= option in the PROC NNET statement.

VALIDGOAL=number

specifies the number targeted goal of validation error before early stopping takes effect during training. This option is valid only when the VALIDATION= option or the PARTITION statement is specified.

By default, VALIDGOAL=0.

WSEED=random-seed
SEED=random-seed

specifies the seed for generating initial random weights. If you do not specify a seed or you specify a value less than or equal to 0, the seed is generated by reading the time of day from the computer’s clock. This option enables you to reproduce the same sample output.

Last updated: November 11, 2020