The SEMISUPLEARN Procedure

Getting Started: SEMISUPLEARN Procedure

Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.

This example shows how to use the SEMISUPLEARN procedure to predict the labels for both the unlabeled and labeled data, from observations in a set of unlabeled data observations and a set of labeled data observations. In this case, the data are from the hmeq data set. This data set contains information about mortgage applicants. The example selects 3,000 applicants for the unlabeled data table, which is the data table without the target variable, and 200 applicants for the labeled data table, which is the data table with the target variable. PROC SEMISUPLEARN returns the predicted target variables for the applicants in both the unlabeled data table and the labeled data table. The analysis uses 10 variables: loan, mortdue, value, yoj, derog, delinq, clage, ninq, clno, and debtinc. The remaining variables in the data table are not used.

You can load the hmeq data set into your CAS session by naming your CAS engine libref in the first statement of the following DATA step. These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.


cas mysess;
libname mycas sasioca sessref=mysess;

data hmeq;
set sampsio.hmeq;
if cmiss(of _all_) then delete;
run;

data mycas.unlabel(drop=bad);
set hmeq(obs=3000);
id =_N_;
run;

data mycas.label;
set hmeq(obs=20);

run;

The following statements run PROC SEMISUPLEARN and output the results to ODS tables:

proc SEMISUPLEARN  data= mycas.unlabel
   label               = mycas.label
   gamma               = 1000;
   input             loan mortdue value yoj derog delinq clage ninq clno debtinc;
   output out          = mycas.out copyvar=(id);
   target              bad;
run;

The INPUT statement specifies that the variables loan, mortdue, value, yoj, derog, delinq, clage, ninq, clno, and debtinc are to be used as inputs. The OUTPUT statement requests that the predicted labels for the unlabeled and labeled tables be written to the data table mycas.out. Figure 1 shows the number of unlabeled observations, number of labeled observations, number of levels for target variable, gamma value, maximum number of iterations, kernel and the loss used in the computation.

Figure 1: Model Information

The SEMISUPLEARN Procedure

Model Information
Labeled Observations Used20
Unlabeled Observations Used3000
Maximum Iterations3
Target Number of Levels2
Gamma1000
Kernel FunctionRBF

Loss
6.159900

Output CAS Tables
CAS LibraryNameNumber
of Rows
Number
of Columns
CASUSERHDFS(tiarno)OUT30203


The following statements sort the output of PROC SEMISUPLEARN by id and show the observations from 100 to 109:

data out2; set mycas.out; run;
proc sort data=out2; by id; run;
proc print data=out2(firstobs=100 obs=109);
run;


Figure 2 shows the id variable, the predicted labels for the unlabeled data, and the indicators for the labeled or unlabeled data.

Figure 2: CAS OUTPUT

ObsidI_BAD_WARN_
1008010
1018110
1028210
1038300
1048400
1058500
1068610
1078710
1088800
1098900


Last updated: November 11, 2020