SVMACHINE Procedure
Example 39.2 Large Simulated Data Table
This example uses a large simulated data table to demonstrate how PROC SVMACHINE can handle relatively large data. The following DATA step generates 10 million observations in the CAS table mylib.bigdata:
data mylib.bigdata ;
array x{5} x1-x5;
drop i n;
do n=1 to 10000000;
do i=1 to dim(x);
x{i} = ranbin(10816, 12, 0.6);
x6 = sum(x2-x4) + ranuni(6068);
end;
if x6 > 0.5 then y = 1;
else if x6 < -0.5 then y = 0;
else y = ranbin(6084, 1, 0.4);
output;
end;
run;
The following statements execute the SVM algorithm on the table mylib.bigdata:
proc svmachine data=mylib.bigdata;
input x1-x6 / level=interval;
target y;
run;
The "Misclassification Matrix" table in Output 39.2.1 shows the classification result. The total number of observations in which is 5,631,506, and the total number of observations in which
is 4,368,494.
Output 39.2.1: Misclassification Matrix
| Misclassification Matrix | |||
|---|---|---|---|
| Observed | Training Prediction | ||
| 1 | 0 | Total | |
| 1 | 5164803 | 466703 | 5631506 |
| 0 | 248256 | 4120238 | 4368494 |
| Total | 5413059 | 4586941 | 10000000 |
The "Fit Statistics" table in Output 39.2.2 shows the accuracy (92.85%) and the error (7.15%) of the model.
Output 39.2.2: Fit Statistics
| Fit Statistics | |
|---|---|
| Statistic | Training |
| Accuracy | 0.9285 |
| Error | 0.0715 |
| Sensitivity | 0.9171 |
| Specificity | 0.9432 |