The KPCA Procedure

Example 13.1 Letter Recognition

In this example, the KPCA procedure is used as a preprocessor to the linear discriminant model for letter recognition.

The letter recognition data set that this example uses is from the UCI Machine Learning Repository (Dua and Graff 2019). The data set includes a large number of black-and-white pixel images of rectangular shape, each of them corresponding to one of the 26 capital letters in the English alphabet. These letter pictures have 20 different fonts, and each picture was randomly altered to generate a total of 20,000 unique instances. Each instance was then transformed into 16 statistical attributes (edge counts and statistical moments). All these attributes were further scaled to a range of integer values from 0 to 15 (Frey and Slate 1991). The objective is to classify each pixel image as one of the 26 capital letters. In this example, a training set of size 16,000 is randomly selected, and the remaining 4,000 instances are used as a test set. Because this is a big data set, fast KPCA is a more appropriate method to use for training while still achieving performance comparable to that of the exact method. To apply fast KPCA, 200 centroids are selected. After the model is trained, the projection values of training and scoring data onto the kernel principal components undergo a multilabel linear discriminant analysis (LDA) for classification.

In the following code, the PROC KPCA statement applies exact and fast KPCA to the data table mycas.letter_train. The statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref. Next, the PROC DISCRIM statement applies linear discriminant analysis to the projections onto the principal component directions. Here letter_train and letter_test are the names of the training data set and test data set, respectively.

   data letter;
infile datalines delimiter=',';
input capital $ var1-var16;
datalines;
T,2,8,3,5,1,8,13,0,6,6,10,8,0,8,0,8

   ... more lines ...   


data letter;
  set letter;
  obsid=_n_;
run;

/* Split data set into train (16000) and test(4000) */
proc surveyselect data=letter
  method=srs n=16000
  seed=100 out=letter_train;
run;

proc sql;
create table letter_test as
select * from letter
where obsid not in (select obsid from letter_train);
quit;

/* Load data to CAS server */
data mycas.letter_train; set letter_train; run;
data mycas.letter_test; set letter_test; run;


/* Apply exact KPCA to training data, extract 200 principal components*/
proc kpca data=mycas.letter_train  method=exact;
  input var1-var16;
  kernel  RBF/bw=7.071;
  output out=mycas.score copyvars=(capital obsid) NPC=200;
  savestate rstore=mycas.state;
run;

/* Score test data */
proc astore;
  setoption kpca_npc 200;
  score data=mycas.letter_test rstore=mycas.state
  out=mycas.outscore copyVars=(capital obsid);
quit;

/* Apply linear discriminant analysis and get multilabel classification error */
proc discrim data=mycas.score method=normal pool=yes short
  testdata=mycas.outscore;
  class capital;
  testclass capital;
run;


/* Apply fast KPCA to training data, extract 200 principal components*/
proc kpca data=mycas.letter_train method=approximate;
  input var1-var16;
  lrapproximation  clusmethod=KMPP  maxclus= 200;
  kernel  RBF/bw=7.071;
  output out=mycas.score_fast copyvars=(capital obsid) NPC=200;
  savestate rstore=mycas.state_fast;
run;

/* Fast score test data */
proc astore;
  setoption kpca_npc 200;
  score data=mycas.letter_test rstore=mycas.state_fast
  out=mycas.outscore_fast copyVars=(capital obsid);
quit;

/* Apply linear discriminant analysis and get multilabel classification error */
proc discrim data=mycas.score_fast method=normal pool=yes short
  testdata=mycas.outscore_fast;
  class capital;
  testclass capital;
run;

In this example, the multilabel classification errors that are obtained by fast KPCA and exact KPCA are close: 0.1662 for fast KPCA and 0.1593 for exact KPCA. However, the training time of fast KPCA is only 10.57 seconds, compared to 272.78 seconds for exact KPCA. The example demonstrates the efficiency of fast KPCA in significantly reducing the run time while not compromising much on the quality of the principal components that it generates. Moreover, in exact KPCA, all the training data along with the eigenvector matrix in the state file must be stored for scoring the test data, whereas in fast KPCA only the k-means centroids and eigenvector matrix must be stored in the state file. This greatly reduces the space requirement to score new data.

You can use the following code to also apply PCA (equivalent to KPCA with linear kernel) to the training data and extract all 16 components to use in linear discriminant analysis:

/* Apply PCA to training data, extract 16 principal components*/
proc kpca data=mycas.letter_train  method=exact;
   input var1-var16;
   kernel  linear;
   output out=mycas.score_pca copyvars=(capital obsid) NPC=16;
   savestate rstore=mycas.state_pca;
run;

/* Score test data */
proc astore;
   setoption kpca_npc 16;
   score data=mycas.letter_test rstore=mycas.state_pca
   out=mycas.outscore_pca copyVars=(capital obsid);
quit;

/* Apply linear discriminant analysis and get multilabel classification error */
proc discrim data=mycas.score_pca method=normal pool=yes short
   testdata=mycas.outscore_pca;
   class capital;
   testclass capital;
run;

Using the preceding code, you can see that even when all components are used, PCA achieves a multilabel misclassification rate of 0.3011, which is much higher than that of KPCA.

Last updated: November 11, 2020