GMM Procedure

Example 14.2 Performing Cluster Analysis and Storing the Model

This example uses the Fish data set in the Sashelp library to demonstrate how to use the GMM procedure to perform cluster analysis and how to use the analytic store to save the model and use the saved model for future clustering.

The Fish data set contains 159 observations and seven variables. Among these variables, Species contains different fish species and the other six numeric variables contain statistics of the fish: height, weight, width, and three measures of length.

Because the statistics of fish in different species overlap, the clusters that PROC GMM discovers do not exactly match the fish species in the Species variable. However, PROC GMM provides a good analysis of the inhomogeneity in the data. Besides, this example shows you how to save the Gaussian mixture model in the analytic store so that you can apply this saved model to other data.

The following DATA step divides the mylib.Fish data into training and testing data tables. PROC GMM uses the first 150 observations in the mylib.Fish data for clustering and saves the model, and it uses the saved model and the last 9 observations directly for cluster score prediction.


data mylib.fish_train;
   title "Using PROC GMM for Cluster Prediction";
   set sashelp.fish (obs=150);
run;
data mylib.fish_test;
   set sashelp.fish (firstobs=151);
run;

These statements assume that your CAS engine libref is named mylib, but you can substitute any appropriately defined CAS engine libref.

The following statements run PROC GMM on the training data and save the trained model in the CAS table mylib.astore:

proc gmm
   data=mylib.fish_train
   nThreads=32
   seed=1234567890
   maxClusters=100
   alpha=1
   inference=VB (maxVbIter=1000 covariance=DIAGONAL threshold=0.001)
   clusterSumOut=mylib.clustersum
   clusterCovOut=mylib.clustercov;
   input _NUMERIC_;
   score out=mylib.score copyvars=(_ALL_);
   ods select nObs descStats modelInfo;
   savestate rstore=mylib.astore(replace=yes);
run;

The following statements run the ASTORE procedure with the saved model on the testing data, and output the cluster scores in the mylib.newscore CAS table:


proc astore;
   score data=mylib.fish_test out=mylib.newscore
   copyvars=(_ALL_) rstore=mylib.astore;
run;

The following statements use PROC PRINT to display the predictions about the testing data in the mylib.newscore CAS table, as shown in Output 14.2.1. In this table, you can see that all nine fish that belong to the smelt species are assigned (very confidently) to the correct cluster, just as they are in the Fish training data.

proc print noobs data=mylib.newscore (obs=9);
run;

Output 14.2.1: Cluster Scores on Testing Data

Using PROC GMM for Cluster Prediction

_CLUSTER_1__CLUSTER_2__CLUSTER_3__CLUSTER_4__CLUSTER_5__CLUSTER_6__CLUSTER_7__CLUSTER_8__PREDICTED_CLUSTER_SpeciesWeightLength1Length2Length3HeightWidth
1.1102E-121.1102E-121.1102E-121.1102E-121.000001.1102E-121.1102E-121.1102E-125Smelt8.710.811.312.61.97821.2852
1.1102E-121.1102E-121.1102E-121.1102E-121.000001.1102E-121.1102E-121.1102E-125Smelt10.011.311.813.12.21391.2838
1.1102E-121.1102E-121.1102E-121.1102E-121.000001.1102E-121.1102E-121.1102E-125Smelt9.911.311.813.12.21391.1659
1.1102E-121.1102E-121.1102E-121.1102E-121.000001.1102E-121.1102E-121.1102E-125Smelt9.811.412.013.22.20441.1484
1.1102E-121.1102E-121.1102E-121.1102E-121.000001.1102E-121.1102E-121.1102E-125Smelt12.211.512.213.42.09041.3936
1.1102E-121.1102E-121.1102E-121.1102E-121.000001.1102E-121.1102E-121.1102E-125Smelt13.411.712.413.52.43001.2690
1.1102E-121.1102E-121.1102E-121.1102E-121.000001.1102E-121.1102E-121.1102E-125Smelt12.212.113.013.82.27701.2558
1.1102E-121.1102E-121.1102E-121.1102E-121.000001.1102E-121.1102E-121.21655E-85Smelt19.713.214.315.22.87282.0672
1.1102E-121.1102E-121.1102E-121.1102E-121.000001.1102E-121.1102E-123.32857E-85Smelt19.913.815.016.22.93221.8792


Last updated: September 04, 2026