The HPNEURAL Procedure
Training
The HPNEURAL procedure can either train a neural network model or use a previously trained model to score a data set. We will first discuss training.
The HPNEURAL procedure does not have many parameters that you must specify. It simply needs to know where the training data is (the DATA= option in the PROC statement), the names and types of the input variables (the INPUT statement), the names and types of the target variables (the TARGET statement), the number of hidden layers, and the number of neurons in each hidden layer (the HIDDEN statement), and, the number of training tries with each try using different randomly generated initial weights (the TRAIN statement).
Optionally, you can also specify where to write the score file that contains targets from the input file and predicted targets from the trained network, where to write the model file that contains the parameters of the trained network (the SCORE statement), and where to write the SAS data step statements that can be used to score new data sets (the CODE statement).
The most important parameters you can specify are the number of hidden layers and number of hidden neurons in each hidden layer in the network. A good strategy is to start with a single hidden layer by specifying a single HIDDEN statement with a small number of hidden neurons, and slowly increase the number until the validation error stops improving.
The next most important parameter you can specify is the number of times the network is to be retrained using different sets of initial weights (the NUMTRIES option in the TRAIN statement). A good strategy is to start with 5 (the default) and increase by 2 until the validation error stops improving.
Finally, unless your training data set is very large, you should set the MAXITER= option in the TRAIN statement to 1,000 or more to prevent the optimization algorithm from stopping prematurely. The value of the MAXITER= option is only a limit. Specifying MAXITER=1000 does not mean that the algorithm will run for 1,000 iterations. Most training runs will use far fewer iterations. If you have a large data set, you can start with MAXITER=1 to see how long a single iteration takes, and then make it larger.
The following small example trains a neural network to predict the type of iris plant, given several measurements, and then scores the same data set that was used for training. The DATA step contains 150 observations derived from the R. A. Fisher (1936) Iris data set:
title 'Fisher (1936) Iris Data';
proc format;
value specname
1='Setosa '
2='Versicolor'
3='Virginica ';
run;
data iris;
input SepalLength SepalWidth PetalLength PetalWidth Species @@;
format Species specname.;
datalines;
50 33 14 02 1 64 28 56 22 3 65 28 46 15 2 67 31 56 24 3
63 28 51 15 3 46 34 14 03 1 69 31 51 23 3 62 22 45 15 2
59 32 48 18 2 46 36 10 02 1 61 30 46 14 2 60 27 51 16 2
65 30 52 20 3 56 25 39 11 2 65 30 55 18 3 58 27 51 19 3
68 32 59 23 3 51 33 17 05 1 57 28 45 13 2 62 34 54 23 3
77 38 67 22 3 63 33 47 16 2 67 33 57 25 3 76 30 66 21 3
49 25 45 17 3 55 35 13 02 1 67 30 52 23 3 70 32 47 14 2
64 32 45 15 2 61 28 40 13 2 48 31 16 02 1 59 30 51 18 3
55 24 38 11 2 63 25 50 19 3 64 32 53 23 3 52 34 14 02 1
49 36 14 01 1 54 30 45 15 2 79 38 64 20 3 44 32 13 02 1
67 33 57 21 3 50 35 16 06 1 58 26 40 12 2 44 30 13 02 1
77 28 67 20 3 63 27 49 18 3 47 32 16 02 1 55 26 44 12 2
50 23 33 10 2 72 32 60 18 3 48 30 14 03 1 51 38 16 02 1
61 30 49 18 3 48 34 19 02 1 50 30 16 02 1 50 32 12 02 1
61 26 56 14 3 64 28 56 21 3 43 30 11 01 1 58 40 12 02 1
51 38 19 04 1 67 31 44 14 2 62 28 48 18 3 49 30 14 02 1
51 35 14 02 1 56 30 45 15 2 58 27 41 10 2 50 34 16 04 1
46 32 14 02 1 60 29 45 15 2 57 26 35 10 2 57 44 15 04 1
50 36 14 02 1 77 30 61 23 3 63 34 56 24 3 58 27 51 19 3
57 29 42 13 2 72 30 58 16 3 54 34 15 04 1 52 41 15 01 1
71 30 59 21 3 64 31 55 18 3 60 30 48 18 3 63 29 56 18 3
49 24 33 10 2 56 27 42 13 2 57 30 42 12 2 55 42 14 02 1
49 31 15 02 1 77 26 69 23 3 60 22 50 15 3 54 39 17 04 1
66 29 46 13 2 52 27 39 14 2 60 34 45 16 2 50 34 15 02 1
44 29 14 02 1 50 20 35 10 2 55 24 37 10 2 58 27 39 12 2
47 32 13 02 1 46 31 15 02 1 69 32 57 23 3 62 29 43 13 2
74 28 61 19 3 59 30 42 15 2 51 34 15 02 1 50 35 13 03 1
56 28 49 20 3 60 22 40 10 2 73 29 63 18 3 67 25 58 18 3
49 31 15 01 1 67 31 47 15 2 63 23 44 13 2 54 37 15 02 1
56 30 41 13 2 63 25 49 15 2 61 28 47 12 2 64 29 43 13 2
51 25 30 11 2 57 28 41 13 2 65 30 58 22 3 69 31 54 21 3
54 39 13 04 1 51 35 14 03 1 72 36 61 25 3 65 32 51 20 3
61 29 47 14 2 56 29 36 13 2 69 31 49 15 2 64 27 53 19 3
68 30 55 21 3 55 25 40 13 2 48 34 16 02 1 48 30 14 01 1
45 23 13 03 1 57 25 50 20 3 57 38 17 03 1 51 38 15 03 1
55 23 40 13 2 66 30 44 14 2 68 28 48 14 2 54 34 17 02 1
51 37 15 04 1 52 35 15 02 1 58 28 51 24 3 67 30 50 17 2
63 33 60 25 3 53 37 15 02 1
;
proc hpneural data=iris;
input SepalLength SepalWidth PetalLength PetalWidth;
target Species / level=nom;
hidden 2;
train outmodel=model_iris maxiter=1000;
score out=scores_iris;
run;
Figure 1 displays the SAS log output, which shows the percentage of validation observations that were misclassified by the trained network. If there had been any interval targets, the log would have shown the absolute average percentage error and the absolute maximum percentage error for each interval target.
Figure 1: SAS Log Output
| NOTE: The HPNEURAL procedure is executing in single-machine mode. |
| NOTE: Reading data... |
| NOTE: 150 usable observations in input data set. |
| NOTE: Training... |
| NOTE: Try 1 complete after 34 iterations. Reason for stopping: Training |
| error=0.000000 |
| NOTE: Try 2 complete after 32 iterations. Reason for stopping: Training |
| error=0.000000 |
| NOTE: Try 3 complete after 40 iterations. Reason for stopping: Training |
| error=0.000000 |
| NOTE: Try 4 complete after 35 iterations. Reason for stopping: Training |
| error=0.000000 |
| NOTE: Try 5 complete after 30 iterations. Reason for stopping: Training |
| error=0.000000 |
| NOTE: Scoring... |
| NOTE: Misclassification Error for target Species: 5.26%; Maximum Error: 14.29% |
| for level VIRGINICA |
| NOTE: Writing Model File... |
| NOTE: There were 150 observations read from the data set WORK.IRIS. |
| NOTE: The data set WORK.MODEL_IRIS has 11 observations and 18 variables. |
| NOTE: The data set WORK.SCORES_IRIS has 150 observations and 5 variables. |
Figure 2 displays the "Model Information," "Performance Information," and "Number of Observations" tables. The HPNEURAL procedure creates a neural network model for the nominal variable Species. Of the 150 observations, 38 are used as a validation subset, which consists of the first observation and every fourth observation thereafter. The other 112 observations make up the training subset.
Figure 2: Model Information, Performance Information, and Number of Observations Tables
| Fisher (1936) Iris Data |
| Performance Information | |
|---|---|
| Execution Mode | Single-Machine |
| Number of Threads | 16 |
| Model Information | |
|---|---|
| Data Source | WORK.IRIS |
| Architecture | MLP |
| Number of Input Variables | 4 |
| Number of Hidden Layers | 1 |
| Number of Hidden Neurons | 2 |
| Number of Target Variables | 1 |
| Number of Weights | 19 |
| Optimization Technique | Limited Memory BFGS |
| Number of Observations Read | 150 |
|---|---|
| Number of Observations Used | 150 |
| Number Used for Training | 112 |
| Number Used for Validation | 38 |
Figure 3 displays the "Misclassification Table." It shows the results of scoring the validation subset by using the neural network model that is trained on the training subset. This example shows two incorrect classifications: two observations whose target value was "Virginica" were incorrectly classified as "Versicolor."
Figure 3: Misclassification Table for Species
| Misclassification Table for Species | |||
|---|---|---|---|
| Class: | VIRGINICA | VERSICOLOR | SETOSA |
| VIRGINICA | 12 | 2 | 0 |
| VERSICOLOR | 0 | 11 | 0 |
| SETOSA | 0 | 0 | 13 |