SVMACHINE Procedure
Getting Started: SVMACHINE Procedure
Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.
This example trains the model by using German credit benchmark data, which are available in the sampsio.dmagecr data set. This data set contains 1,000 observations, each of which contains an applicant’s information, including the applicant’s credit rating (GOOD or BAD). The binary target is named GOOD_BAD. Other input variables are Checking, Duration, History, and so on.
The contents of the sampsio.dmagecr data set are described at http://support.sas.com/documentation/cdl/en/emgs/59885/HTML/default/a001026918.htm.
You can load the sampsio.dmagecr data set into your CAS session by specifying your CAS engine libref in the second statement in the following DATA step:
data mylib.dmagecr;
set sampsio.dmagecr;
run;
These statements assume that your CAS engine libref is named mylib, as in the section Using CAS Sessions and CAS Engine Librefs, but you can substitute any appropriately defined CAS engine libref.
The following statements execute the SVM algorithm on the mylib.dmagecr data table and produce the results shown in Figure 1 through Figure 3.
proc svmachine data=mylib.dmagecr;
input checking history purpose savings employed marital coapp
property other job housing telephon foreign/level=nominal;
input duration amount installp resident existcr depends age/level=interval;
target good_bad;
run;
The first INPUT statement defines the input variables Checking, History, Purpose, Savings, Employed, Marital, Coapp, Property, Other, Job, Housing, Telephon, and Foreign as categorical variables. The second INPUT statement defines the input variables Duration, Amount, Installp, Resident, Existcr, Depends, and Age as continuous variables. The TARGET statement defines good_bad to be the target variable (the variable that is predicted). The "Training Results" table in Figure 1 shows that the inner product of weights is 11.6121718, the bias is –2.1296773, and the number of support vectors is 531, where 481 of those vectors are on the margin. The table also shows that the maximum decision function value (Maximum F) is 2.57131793 and the minimum decision function value (Minimum F) is –4.6513481.
Figure 1: German Credit Data Training Results
| Training Results | |
|---|---|
| Inner Product of Weights | 11.6121718 |
| Bias | -2.1296773 |
| Total Slack (Constraint Violations) | 492.87883 |
| Norm of Longest Vector | 4.17809329 |
| Number of Support Vectors | 531 |
| Number of Support Vectors on Margin | 481 |
| Maximum F | 2.57131793 |
| Minimum F | -4.6513481 |
| Number of Effects | 20 |
| Columns in Data Matrix | 61 |
The "Misclassification Matrix" table in Figure 2 shows that among the total of 1,000 observations, 700 observations are classified as good and 300 observations are classified as bad. The number of correctly predicted good observations is 626, and the number of correctly predicted bad observations is 158. Thus the accuracy is 78.4%, as indicated in the "Fit Statistics" table in Figure 3.
Figure 2: German Credit Misclassification Matrix
| Misclassification Matrix | |||
|---|---|---|---|
| Observed | Training Prediction | ||
| good | bad | Total | |
| good | 626 | 74 | 700 |
| bad | 142 | 158 | 300 |
| Total | 768 | 232 | 1000 |
Figure 3: German Credit Accuracy
| Fit Statistics | |
|---|---|
| Statistic | Training |
| Accuracy | 0.7840 |
| Error | 0.2160 |
| Sensitivity | 0.8943 |
| Specificity | 0.5267 |
A relatively good model means that misclassification is low while both sensitivity and specificity are high. In PROC SVMACHINE, you can adjust training parameters and use different kernels to obtain a better model.