GRADBOOST Procedure
Getting Started: GRADBOOST Procedure
Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.
A common use of gradient boosting models is to predict whether a mortgage applicant will default on a loan. The home equity data table Hmeq, which is in the Sampsio library that SAS provides, contains observations for 5,960 mortgage applicants. A variable named Bad indicates whether the applicant, after being approved for a loan, paid off or defaulted on the loan.
This example uses the Hmeq data table to build a gradient boosting model that is used to score the data and can be used to score data about new loan applicants. Table 1 describes the variables in Hmeq.
Table 1: Variables in the Home Equity (Hmeq) Data Table
| Variable | Role | Level | Description |
|---|---|---|---|
Bad | Response | Binary | 1 = applicant defaulted on the loan or is seriously delinquent |
| 0 = applicant paid off the loan | |||
CLAge | Predictor | Interval | Age of oldest credit line in months |
CLNo | Predictor | Interval | Number of credit lines |
DebtInc | Predictor | Interval | Debt-to-income ratio |
Delinq | Predictor | Interval | Number of delinquent credit lines |
Derog | Predictor | Interval | Number of major derogatory reports |
Job | Predictor | Nominal | Occupational category |
Loan | Predictor | Interval | Requested loan amount |
MortDue | Predictor | Interval | Amount due on mortgage |
nInq | Predictor | Interval | Number of recent credit inquiries |
Reason | Predictor | Binary | DebtCon = debt consolidation |
HomeImp = home improvement | |||
Value | Predictor | Interval | Value of property |
YoJ | Predictor | Interval | Years at present job |
The following statements load the mylib.hmeq data into your CAS session. For this example, the statements assume that your CAS engine libref is named mylib, but you can substitute any appropriately defined CAS engine libref.
data mylib.hmeq;
length Bad Loan MortDue Value 8 Reason Job $7
YoJ Derog Delinq CLAge nInq CLNo DebtInc 8;
set sampsio.hmeq;
run;
proc print data=mylib.hmeq(obs=10); run;
Figure 1 shows the first 10 observations of mylib.hmeq.
Figure 1: Partial Listing of the mylib.hmeq Data
| Obs | Bad | Loan | MortDue | Value | Reason | Job | YoJ | Derog | Delinq | CLAge | nInq | CLNo | DebtInc |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 1 | 1100 | 25860 | 39025 | HomeImp | Other | 10.5 | 0 | 0 | 94.367 | 1 | 9 | . |
| 2 | 1 | 1500 | . | . | . | . | . | . | . | . | . | ||
| 3 | 1 | 1800 | 48649 | 57037 | HomeImp | Other | 5.0 | 3 | 2 | 77.100 | 1 | 17 | . |
| 4 | 1 | 2000 | . | 62250 | HomeImp | Sales | 16.0 | 0 | 0 | 115.800 | 0 | 13 | . |
| 5 | 1 | 2000 | 45000 | 55000 | HomeImp | Other | 3.0 | 0 | 0 | 86.067 | 2 | 25 | . |
| 6 | 1 | 2200 | 24280 | 34687 | HomeImp | Other | . | 0 | 1 | 300.867 | 0 | 8 | . |
| 7 | 1 | 2300 | 28192 | 40150 | HomeImp | Other | 4.5 | 0 | 0 | 54.600 | 1 | 16 | . |
| 8 | 1 | 2400 | 50000 | 73395 | HomeImp | ProfExe | 5.0 | 1 | 0 | . | 1 | 0 | . |
| 9 | 1 | 2400 | . | 17180 | HomeImp | Other | . | 0 | 0 | 14.567 | 3 | 4 | . |
| 10 | 1 | 2500 | 15000 | 20200 | HomeImp | 18.0 | 0 | 0 | 136.067 | 1 | 19 | . |
PROC GRADBOOST treats numeric variables as interval inputs unless you specify otherwise. Character variables are always treated as nominal inputs. The following statements run PROC GRADBOOST and save the model in a table named mylib.savedModel:
proc gradboost data=mylib.hmeq outmodel=mylib.savedModel seed=12345;
input Delinq Derog Job nInq Reason / level = nominal;
input CLAge CLNo DebtInc Loan Mortdue Value YoJ / level = interval;
target Bad / level = nominal;
ods output FitStatistics=fitstats;
run;
No parameters are specified in the PROC GRADBOOST statement; therefore, the procedure uses all default values. For example, the number of trees in the boosting model is 100, and the number of bins for interval input variables is 20.
The INPUT and TARGET statements are required in order to run PROC GRADBOOST. The INPUT statement indicates which variables to use to build the model, and the TARGET statement indicates which variable the procedure predicts.
Figure 2 displays the "Model Information" table. This table shows the values of the training parameters in the first six rows, in addition to some basic information about the trees in the boosting model.
Figure 2: Model Information
| Model Information | |
|---|---|
| Number of Trees | 100 |
| Learning Rate | 0.1 |
| Subsampling Rate | 0.5 |
| Number of Variables Per Split | 12 |
| Number of Bins | 50 |
| Number of Input Variables | 12 |
| Maximum Number of Tree Nodes | 31 |
| Minimum Number of Tree Nodes | 17 |
| Maximum Number of Branches | 2 |
| Minimum Number of Branches | 2 |
| Maximum Depth | 4 |
| Minimum Depth | 4 |
| Maximum Number of Leaves | 16 |
| Minimum Number of Leaves | 9 |
| Maximum Leaf Size | 2603 |
| Minimum Leaf Size | 5 |
| Seed | 12345 |
| Lasso (L1) penalty | 0 |
| Ridge (L2) penalty | 1 |
| Actual Number of Trees | 100 |
| Average Number of Leaves | 14.17 |
Figure 3 displays the "Number of Observations" table, which shows how many observations were read and used. If you specify a PARTITION statement, the "Number of Observations" table also displays the number of observations that were read and used per partition.
Figure 3: Number of Observations
| Training | |
|---|---|
| Number of Observations Read | 5960 |
| Number of Observations Used | 5960 |
Figure 4 displays the estimates of variable importance. The rows in this figure are sorted by the importance measure. A conclusion from fitting the boosting model to these data is that DebtInc is the most important predictor of loan default.
Figure 4: Variable Importance
| Variable Importance | |||
|---|---|---|---|
| Variable | Importance | Std Dev Importance | Relative Importance |
| DebtInc | 27.9044 | 77.6122 | 1.0000 |
| Delinq | 5.7250 | 7.3421 | 0.2052 |
| Value | 5.5142 | 5.9790 | 0.1976 |
| CLAge | 5.3850 | 5.9920 | 0.1930 |
| Derog | 4.0895 | 4.9382 | 0.1466 |
| CLNo | 3.3499 | 3.1560 | 0.1201 |
| Job | 3.0641 | 2.6463 | 0.1098 |
| YoJ | 2.8628 | 2.9581 | 0.1026 |
| Loan | 2.7176 | 2.9825 | 0.0974 |
| MortDue | 2.5979 | 2.7725 | 0.0931 |
| nInq | 1.9070 | 2.4852 | 0.0683 |
| Reason | 0.3963 | 1.0004 | 0.0142 |
Figure 5 shows the first 10 and last 10 observations of the fit statistics. PROC GRADBOOST computes fit statistics on a per-tree basis. As the number of trees increases, the fit statistics usually improve (decrease) at first and then level off and fluctuate within a small range.
Figure 5: Fit Statistics
| Fit Statistics | |||
|---|---|---|---|
| Number of Trees | Training Average Square Error | Training Misclassification Rate | Training Log Loss |
| 1 | 0.1458 | 0.1995 | 0.459 |
| 2 | 0.1352 | 0.1995 | 0.431 |
| 3 | 0.1258 | 0.1995 | 0.407 |
| 4 | 0.1186 | 0.1995 | 0.388 |
| 5 | 0.1125 | 0.1973 | 0.372 |
| 6 | 0.1076 | 0.1827 | 0.360 |
| 7 | 0.1032 | 0.1567 | 0.348 |
| 8 | 0.0995 | 0.1361 | 0.338 |
| 9 | 0.0961 | 0.1250 | 0.329 |
| 10 | 0.0930 | 0.1190 | 0.320 |
| . | . | . | . |
| . | . | . | . |
| . | . | . | . |
| 91 | 0.0481 | 0.0649 | 0.172 |
| 92 | 0.0477 | 0.0646 | 0.170 |
| 93 | 0.0474 | 0.0648 | 0.170 |
| 94 | 0.0472 | 0.0636 | 0.169 |
| 95 | 0.0469 | 0.0644 | 0.168 |
| 96 | 0.0467 | 0.0629 | 0.167 |
| 97 | 0.0463 | 0.0626 | 0.166 |
| 98 | 0.0460 | 0.0616 | 0.165 |
| 99 | 0.0457 | 0.0611 | 0.164 |
| 100 | 0.0456 | 0.0597 | 0.164 |