KCLUS Procedure

Getting Started: KCLUS Procedure

(View the complete code for this example.)

Note: Input data must be in a table that is accessible in your session. You can refer to this table by using a two-level name. The first level is a libref, and the second level is the table name. For more information, see the section Using SAS Viya Workbench in Chapter 2, Shared Concepts.

This example shows how to use the KCLUS procedure to compute clusters of observations in an input data table.

Suppose you want to group the observations in the input table mylib.inpData, in which the variables are raw measures on interval scales.

The following DATA step creates the input data table, mylib.inpData, in your session. This data table contains four variables: the first two variables are the input variables among which x has missing values, the third variable is the frequency variable, and the last variable is an index variable.

data mylib.inpData;
   title 'Using PROC KCLUS to Analyze Data';
   drop n;
   id=1;
   do n=1 to 1000;
      x=2*rannor(12345)+20;
      y=4*rannor(12345)+20;
      freq = 1;
      id = id + 1;
      output;
   end;
   do n=1 to 1000;
      x=3*rannor(12345)+10;
      y=5*rannor(12345)+10;
      freq=2;
      id = id + 1;
      output;
   end;
   do n=1 to 700;
      x=10*rannor(12345);
      y=10*rannor(12345);
      freq=1;
      id = id + 1;
      output;
   end;
   do n=1 to 200;
      x=.;
      y=10*rannor(12345);
      freq=1;
      id = id + 1;
      output;
   end;
run;

These statements assume that your libref is named mylib, but you can substitute any appropriately defined libref.

The following statements run PROC KCLUS and output the results to ODS tables:

proc kclus data=mylib.inpData maxclusters=3;
   input x y;
   freq freq;
run;

Figure 1 shows that the "Number of Observations Used" is less than the "Number of Observations Read". By default, the KCLUS procedure ignores observations that have missing values, and it does not use them in the analysis. The two additional rows, "Sum of Frequencies Read" and "Sum of Frequencies Used" are displayed when the FREQ statement is specified. They provide information about the frequency values that are read and used.

Figure 1: Number of Observations

Using PROC KCLUS to Analyze Data

The KCLUS Procedure

Number of Observations Read2900
Number of Observations Used2700
Sum of Frequencies Read3900
Sum of Frequencies Used3700


Figure 2 shows the values of the parameters that are used in clustering. Because the number of clusters is not estimated by default and MAXCLUSTERS=3, three clusters are generated. Figure 2 shows the number of clusters and the default values for other options.

Figure 2: Model Information

Model Information
Clustering AlgorithmK-means
Maximum Iterations10
Stop CriterionCluster Change
Stop Criterion Value0
Clusters3
InitializationForgy
Seed578406742
Distance for Interval VariablesEuclidean
StandardizationNone
Interval ImputationNone


For each cluster, Figure 3 shows the number of observations; the maximum, minimum, and average distances from that cluster’s centroid to the observations in that cluster; the sum of squares error; and the standard deviation. Figure 3 also displays information about the nearest cluster to that cluster and the distance between their centroids.

Figure 3: Cluster Summary

Cluster Summary for Interval Variables
ClusterFrequencyDistance from Cluster Centroid
to Observation
SSEStandard
Deviation
Nearest
Cluster
Distance
to
Nearest
Cluster
Centroid
MinimumMaximumAverage
120380.228931.79705.201978608.46.2106213.8882
211750.020229.23044.474131968.15.2160113.8882
34870.565230.444710.729873156.712.2564118.3687


Figure 4 shows the sum of squared errors (SSE) for each iteration. If the variables are interval, then the "Iteration History" table displays SSE Change and Stop Criterion columns. The SSE Change column displays the change in within-cluster distances. The Stop Criterion column displays the stopping criterion for each iteration. If the input variables are nominal, then the "Iteration History" table displays Within Distance Change and Stop Criterion columns. If the input variables are both interval and nominal, then the "Iteration History" table also displays Within Distance Change and Stop Criterion columns, but with the distance in the sense of mixed distance with both the interval and nominal parts as detailed in Clustering Both Interval and Nominal Variables.

Figure 4: Iteration History

Iteration History
Iteration
Number
SSESSE ChangeStop
Criterion
0325431  
1204706-1207258.851852
2188187-165193.407407
3185097-3089.2919591.629630
4184358-739.5284560.962963
5184083-275.0005740.777778
6183876-206.3932910.740741
7183770-106.5915970.333333
8183739-30.9202070.074074
9183733-5.7221730
1018373300


Figure 5 and Figure 6 show statistics for each variable in the INPUT statement. Figure 5 shows the variable statistics for all the observations in the input data table, and Figure 6 shows the variable statistics for the observations that belong to a specific cluster.

Figure 5: Descriptive Statistics

Descriptive Statistics
VariableMeanStandard
Deviation
x11.0206488.189686
y10.2465469.472050


Figure 6: Within-Cluster Statistics

Within Cluster Statistics
VariableClusterMeanStandard
Deviation
x19.562610.5049
 218.99306.0294
 3-2.11309.0921
y19.385211.0225
 219.58087.9047
 3-4.79548.3923


Last updated: May 14, 2026