KCLUS Procedure
Getting Started: KCLUS Procedure
(View the complete code for this example.)
Note: Input data must be in a table that is accessible in your session. You can refer to this table by using a two-level name. The first level is a libref, and the second level is the table name. For more information, see the section Using SAS Viya Workbench in Chapter 2, Shared Concepts.
This example shows how to use the KCLUS procedure to compute clusters of observations in an input data table.
Suppose you want to group the observations in the input table mylib.inpData, in which the variables are raw measures on interval scales.
The following DATA step creates the input data table, mylib.inpData, in your session. This data table contains four variables: the first two variables are the input variables among which x has missing values, the third variable is the frequency variable, and the last variable is an index variable.
data mylib.inpData;
title 'Using PROC KCLUS to Analyze Data';
drop n;
id=1;
do n=1 to 1000;
x=2*rannor(12345)+20;
y=4*rannor(12345)+20;
freq = 1;
id = id + 1;
output;
end;
do n=1 to 1000;
x=3*rannor(12345)+10;
y=5*rannor(12345)+10;
freq=2;
id = id + 1;
output;
end;
do n=1 to 700;
x=10*rannor(12345);
y=10*rannor(12345);
freq=1;
id = id + 1;
output;
end;
do n=1 to 200;
x=.;
y=10*rannor(12345);
freq=1;
id = id + 1;
output;
end;
run;
These statements assume that your libref is named mylib, but you can substitute any appropriately defined libref.
The following statements run PROC KCLUS and output the results to ODS tables:
proc kclus data=mylib.inpData maxclusters=3;
input x y;
freq freq;
run;
Figure 1 shows that the "Number of Observations Used" is less than the "Number of Observations Read". By default, the KCLUS procedure ignores observations that have missing values, and it does not use them in the analysis. The two additional rows, "Sum of Frequencies Read" and "Sum of Frequencies Used" are displayed when the FREQ statement is specified. They provide information about the frequency values that are read and used.
Figure 1: Number of Observations
| Using PROC KCLUS to Analyze Data |
| Number of Observations Read | 2900 |
|---|---|
| Number of Observations Used | 2700 |
| Sum of Frequencies Read | 3900 |
| Sum of Frequencies Used | 3700 |
Figure 2 shows the values of the parameters that are used in clustering. Because the number of clusters is not estimated by default and MAXCLUSTERS=3, three clusters are generated. Figure 2 shows the number of clusters and the default values for other options.
Figure 2: Model Information
| Model Information | |
|---|---|
| Clustering Algorithm | K-means |
| Maximum Iterations | 10 |
| Stop Criterion | Cluster Change |
| Stop Criterion Value | 0 |
| Clusters | 3 |
| Initialization | Forgy |
| Seed | 578406742 |
| Distance for Interval Variables | Euclidean |
| Standardization | None |
| Interval Imputation | None |
For each cluster, Figure 3 shows the number of observations; the maximum, minimum, and average distances from that cluster’s centroid to the observations in that cluster; the sum of squares error; and the standard deviation. Figure 3 also displays information about the nearest cluster to that cluster and the distance between their centroids.
Figure 3: Cluster Summary
| Cluster Summary for Interval Variables | ||||||||
|---|---|---|---|---|---|---|---|---|
| Cluster | Frequency | Distance from Cluster Centroid to Observation | SSE | Standard Deviation | Nearest Cluster | Distance to Nearest Cluster Centroid | ||
| Minimum | Maximum | Average | ||||||
| 1 | 2038 | 0.2289 | 31.7970 | 5.2019 | 78608.4 | 6.2106 | 2 | 13.8882 |
| 2 | 1175 | 0.0202 | 29.2304 | 4.4741 | 31968.1 | 5.2160 | 1 | 13.8882 |
| 3 | 487 | 0.5652 | 30.4447 | 10.7298 | 73156.7 | 12.2564 | 1 | 18.3687 |
Figure 4 shows the sum of squared errors (SSE) for each iteration. If the variables are interval, then the "Iteration History" table displays SSE Change and Stop Criterion columns. The SSE Change column displays the change in within-cluster distances. The Stop Criterion column displays the stopping criterion for each iteration. If the input variables are nominal, then the "Iteration History" table displays Within Distance Change and Stop Criterion columns. If the input variables are both interval and nominal, then the "Iteration History" table also displays Within Distance Change and Stop Criterion columns, but with the distance in the sense of mixed distance with both the interval and nominal parts as detailed in Clustering Both Interval and Nominal Variables.
Figure 4: Iteration History
| Iteration History | |||
|---|---|---|---|
| Iteration Number | SSE | SSE Change | Stop Criterion |
| 0 | 325431 | ||
| 1 | 204706 | -120725 | 8.851852 |
| 2 | 188187 | -16519 | 3.407407 |
| 3 | 185097 | -3089.291959 | 1.629630 |
| 4 | 184358 | -739.528456 | 0.962963 |
| 5 | 184083 | -275.000574 | 0.777778 |
| 6 | 183876 | -206.393291 | 0.740741 |
| 7 | 183770 | -106.591597 | 0.333333 |
| 8 | 183739 | -30.920207 | 0.074074 |
| 9 | 183733 | -5.722173 | 0 |
| 10 | 183733 | 0 | 0 |
Figure 5 and Figure 6 show statistics for each variable in the INPUT statement. Figure 5 shows the variable statistics for all the observations in the input data table, and Figure 6 shows the variable statistics for the observations that belong to a specific cluster.
Figure 5: Descriptive Statistics
| Descriptive Statistics | ||
|---|---|---|
| Variable | Mean | Standard Deviation |
| x | 11.020648 | 8.189686 |
| y | 10.246546 | 9.472050 |
Figure 6: Within-Cluster Statistics
| Within Cluster Statistics | |||
|---|---|---|---|
| Variable | Cluster | Mean | Standard Deviation |
| x | 1 | 9.5626 | 10.5049 |
| 2 | 18.9930 | 6.0294 | |
| 3 | -2.1130 | 9.0921 | |
| y | 1 | 9.3852 | 11.0225 |
| 2 | 19.5808 | 7.9047 | |
| 3 | -4.7954 | 8.3923 | |