The GRADBOOST Procedure
Example 11.1 Scoring New Data by Using a Previous Boosting Model
This example illustrates how you can use the OUTMODEL= option to save a model table, and later use the model table to score a data table. It uses the JunkMail data set in the Sashelp library.
The JunkMail data set comes from a study that classifies whether an email is junk email (coded as 1) or not (coded as 0). The data set contains 4,601 observations with 59 variables. The response variable is a binary indicator of whether an email is considered spam or not. There are 57 predictor variables that record the frequencies of some common words and characters and the lengths of uninterrupted sequences of capital letters in emails.
You can load the Sashelp.JunkMail data set into your CAS session by naming your CAS engine libref in the first statement of the following DATA step:
data mycas.junkmail;
set sashelp.junkmail;
run;
These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined libref.
The following statements train a boosting model and score the training data table. The OUTPUT statement scores the training data and saves the results to a new table named fit_at_runtime.
proc gradboost data=mycas.junkmail outmodel=mycas.gradboost_model seed=12345;
input Address Addresses All Bracket Business CS CapAvg CapLong
CapTotal Conference Credit Data Direct Dollar Edu Email
Exclamation Font Free George HP HPL Internet Lab Labs
Mail Make Meeting Money Order Original Our Over PM Paren
Parts People Pound Project RE Receive Remove Semicolon
Table Technology Telnet Will You Your _000 _85 _415 _650
_857 _1999 _3D / level = interval;
target class /level=nominal;
output out=mycas.score_at_runtime;
ods output FitStatistics=fit_at_runtime;
run;
The preceding statements produce the table shown in Output 11.1.1. The table shows the training statistics.
Output 11.1.1: Fit Statistics, Fit at Run Time
| Fit Statistics | |||
|---|---|---|---|
| Number of Trees | Training Average Square Error | Training Misclassification Rate | Training Log Loss |
| 1 | 0.2087 | 0.3940 | 0.6079 |
| 2 | 0.1842 | 0.2334 | 0.5569 |
| 3 | 0.1645 | 0.1098 | 0.5154 |
| 4 | 0.1479 | 0.0926 | 0.4793 |
| 5 | 0.1330 | 0.0880 | 0.4460 |
| 6 | 0.1216 | 0.0830 | 0.4197 |
| 7 | 0.1113 | 0.0815 | 0.3950 |
| 8 | 0.1024 | 0.0767 | 0.3728 |
| 9 | 0.0952 | 0.0769 | 0.3539 |
| 10 | 0.0888 | 0.0739 | 0.3368 |
| . | . | . | . |
| . | . | . | . |
| . | . | . | . |
| 91 | 0.0276 | 0.0367 | 0.1049 |
| 92 | 0.0273 | 0.0356 | 0.1040 |
| 93 | 0.0271 | 0.0361 | 0.1033 |
| 94 | 0.0270 | 0.0350 | 0.1030 |
| 95 | 0.0268 | 0.0343 | 0.1021 |
| 96 | 0.0267 | 0.0341 | 0.1017 |
| 97 | 0.0266 | 0.0346 | 0.1012 |
| 98 | 0.0264 | 0.0343 | 0.1006 |
| 99 | 0.0261 | 0.0322 | 0.0995 |
| 100 | 0.0259 | 0.0319 | 0.0989 |
The following statements use a previously saved model to score new data:
proc gradboost data=mycas.junkmail inmodel=mycas.gradboost_model;
output out=mycas.score_later;
ods output FitStatistics=fit_later;
run;
When you specify the INMODEL= option to use a previously created boosting model, you see the statistics for the scored data if the target exists in the newly scored data table. In this example, the scored data are the same as the training data, so you can see that the statistics in Output 11.1.2 match those previously seen in Output 11.1.1.
Output 11.1.2: Fit Statistics, Fit Later
| Fit Statistics | |||
|---|---|---|---|
| Number of Trees | Average Square Error | Misclassification Rate | Log Loss |
| 1 | 0.2087 | 0.3940 | 0.6079 |
| 2 | 0.1842 | 0.2334 | 0.5569 |
| 3 | 0.1645 | 0.1098 | 0.5154 |
| 4 | 0.1479 | 0.0926 | 0.4793 |
| 5 | 0.1330 | 0.0880 | 0.4460 |
| 6 | 0.1216 | 0.0830 | 0.4197 |
| 7 | 0.1113 | 0.0815 | 0.3950 |
| 8 | 0.1024 | 0.0767 | 0.3728 |
| 9 | 0.0952 | 0.0769 | 0.3539 |
| 10 | 0.0888 | 0.0739 | 0.3368 |
| . | . | . | . |
| . | . | . | . |
| . | . | . | . |
| 91 | 0.0276 | 0.0367 | 0.1049 |
| 92 | 0.0273 | 0.0356 | 0.1040 |
| 93 | 0.0271 | 0.0361 | 0.1033 |
| 94 | 0.0270 | 0.0350 | 0.1030 |
| 95 | 0.0268 | 0.0343 | 0.1021 |
| 96 | 0.0267 | 0.0341 | 0.1017 |
| 97 | 0.0266 | 0.0346 | 0.1012 |
| 98 | 0.0264 | 0.0343 | 0.1006 |
| 99 | 0.0261 | 0.0322 | 0.0995 |
| 100 | 0.0259 | 0.0319 | 0.0989 |
This example demonstrates that the GRADBOOST procedure can score an input data table by using a previously saved boosting model, which was saved using the OUTMODEL= option in a previous procedure run. If you want to properly score a new data table, you must not modify the table mycas.gradboost_model, because doing so could invalidate the constructed boosting model. As with any scoring of new data, the variables that are used in the model creation must be present in order for you to score a new table.