The HPCANDISC Procedure

Getting Started: HPCANDISC Procedure

(View the complete code for this example.)

The data in this example are measurements of 159 fish caught in Finland’s Lake Laengelmaevesi; this data set is available from the Puranen. For each of the seven species (bream, roach, whitefish, parkki, perch, pike, and smelt), the weight, length, height, and width of each fish are tallied. Three different length measurements are recorded: from the nose of the fish to the beginning of its tail, from the nose to the notch of its tail, and from the nose to the end of its tail. The height and width are recorded as percentages of the third length variable. The fish data set is available from the Sashelp library.

The following step uses PROC HPCANDISC to find the three canonical variables that best separate the species of fish in the Sashelp.Fish data and create the output data set outcan. When the NCAN=3 option is specified, only the first three canonical variables are displayed. The ID statement adds the variable Species from the input data set to the output data set. The ODS EXCLUDE statement excludes the canonical structure tables and most of the canonical coefficient tables in order to obtain a more compact set of results. The TEMPLATE and SGRENDER procedures create a plot of the first two canonical variables. The following statements produce Figure 51.1 through Figure 51.6:

title 'Fish Measurement Data';

proc hpcandisc data=sashelp.fish ncan=3 out=outcan;
   ods exclude tstruc bstruc pstruc tcoef pcoef;
   id Species;
   class Species;
   var Weight Length1 Length2 Length3 Height Width;
run;

proc template;
   define statgraph scatter;
      begingraph;
         entrytitle 'Fish Measurement Data';
         layout overlayequated / equatetype=fit
            xaxisopts=(label='Canonical Variable 1')
            yaxisopts=(label='Canonical Variable 2');
            scatterplot x=Can1 y=Can2 / group=species name='fish';
            layout gridded / autoalign=(topright);
               discretelegend 'fish' / border=false opaque=false;
            endlayout;
         endlayout;
      endgraph;
   end;
run;

proc sgrender data=outcan template=scatter;
run;

PROC HPCANDISC begins by displaying performance information, data access information, and summary information about the variables in the analysis, as shown in Figure 51.1.

The "Performance Information" table shows the procedure executes in single-machine mode; that is, the data reside and the computation is conducted on the machine where the SAS session executes. This run of the HPCANDISC procedure took place on a multicore machine that had four CPUs; one computational thread was spawned per CPU.

The "Data Access Information" table shows that the input data set and the output data set are both accessed with the V9 (base) engine on the client machine where the MVA SAS session executes.

The summary information includes the number of observations, the number of quantitative variables in the analysis (specified using the VAR statement), and the number of class levels in the classification variable (specified using the CLASS statement). The value and frequency of each class level are also displayed.

Figure 51.1: Fish Data: Performance, Data Access, and Summary Information

Fish Measurement Data

The HPCANDISC Procedure

Performance Information
Execution ModeSingle-Machine
Number of Threads4

Data Access Information
DataEngineRolePath
SASHELP.FISHV9InputOn Client
WORK.OUTCANV9OutputOn Client

Total Sample Size158DF Total157
Variables6DF Within Classes151
Class Levels7DF Between Classes6

Number of Observations Read159
Number of Observations Used158

Class Level Information
SpeciesFrequencyWeightProportion
Bream3434.000000.21519
Parkki1111.000000.06962
Perch5656.000000.35443
Pike1717.000000.10759
Roach2020.000000.12658
Smelt1414.000000.08861
Whitefish66.000000.03797


Figure 51.2 displays the "Multivariate Statistics and F Approximations" table. PROC HPCANDISC performs a one-way multivariate analysis of variance (one-way MANOVA) and provides four multivariate tests of the hypothesis that the class mean vectors are equal. These tests indicate that not all the mean vectors are equal (p < 0.0001).

Figure 51.2: Fish Data: MANOVA and Multivariate Tests

Fish Measurement Data

The HPCANDISC Procedure

Multivariate Statistics and F Approximations
S=6 M=-0.5 N=72
StatisticValueF ValueNum DFDen DFPr > F
Wilks' Lambda0.00036390.7136643.89<.0001
Pillai's Trace3.10465126.9936906<.0001
Hotelling-Lawley Trace52.057997209.2436413.64<.0001
Roy's Greatest Root39.134998984.906151<.0001
NOTE: F Statistic for Roy's Greatest Root is an upper bound.


Figure 51.3 displays the "Canonical Correlations" table. The first canonical correlation is the greatest possible multiple correlation with the classes that you can achieve by using a linear combination of the quantitative variables. The first canonical correlation, displayed in the table, is 0.987463. The figure shows a likelihood ratio test of the hypothesis that the current canonical correlation and all smaller ones are zero. The first line is equivalent to Wilks’ lambda multivariate test.

Figure 51.3: Fish Data: Canonical Correlations

Fish Measurement Data

The HPCANDISC Procedure

 Canonical
Correlation
Adjusted
Canonical
Correlation
Approximate
Standard
Error
Squared
Canonical
Correlation
Eigenvalues of Inv(E)*H
= CanRsq/(1-CanRsq)
Test of H0: The canonical correlations in the current row and all that follow are zero
 EigenvalueDifferenceProportionCumulativeLikelihood
Ratio
Approximate
F Value
Num DFDen DFPr > F
10.9874630.9866710.0019890.97508439.135029.38590.75180.75180.0003632590.7136643.89<.0001
20.9523490.9500950.0074250.9069699.74917.37860.18730.93900.0145789646.4625547.58<.0001
30.8386370.8325180.0236780.7033132.37061.70160.04550.98460.1567113423.6116452.79<.0001
40.6330940.6236490.0478210.4008090.66890.53460.01280.99740.5282034712.099362.78<.0001
50.3441570.3341700.0703560.1184440.13440.13430.00261.00000.881527024.8843000.0008
60.005701.0.0798060.0000330.0000 0.00001.00000.999967490.0011510.9442


Figure 51.4 displays the "Raw Canonical Coefficients" table. The first canonical variable, Can1, shows that the linear combination of the centered variables Can1 = –0.0006 Weight – 0.33 Length1 2.49 Length2 + 2.60 Length3 + 1.12 Height – 1.45 Width separates the species most effectively.

Figure 51.4: Fish Data: Raw Canonical Coefficients

Fish Measurement Data

The HPCANDISC Procedure

Raw Canonical Coefficients
VariableCan1Can2Can3
Weight-0.00064851-0.00523-0.00560
Length1-0.32944-0.62660-2.93432
Length2-2.48613-0.690254.04504
Length32.595651.80318-1.13926
Height1.12198-0.714750.28320
Width-1.44639-0.907030.74149


Figure 51.5 displays the "Class Means on Canonical Variables" table. PROC HPCANDISC computes the means of the canonical variables for each class. The first canonical variable is the linear combination of the variables Weight, Length1, Length2, Length3, Height, and Width that provides the greatest difference (in terms of a univariate F test) between the class means. The second canonical variable provides the greatest difference between class means while being uncorrelated with the first canonical variable.

Figure 51.5: Fish Data: Class Means for Canonical Variables

Class Means on Canonical Variables
SpeciesCan1Can2Can3
Bream10.941420.520780.23497
Parkki2.58904-2.54722-0.49326
Perch-4.47181-1.708231.29281
Pike-4.896898.22141-0.16469
Roach-0.358370.08734-1.10056
Smelt-4.09137-2.35806-4.03836
Whitefish-0.39542-0.420721.06459


Figure 51.6 displays a plot of the first two canonical variables, which shows that Can1 discriminates among three groups: (1) bream; (2) whitefish, roach, and parkki; and (3) smelt, pike, and perch. Can2 best discriminates between pike and the other species.

Figure 51.6: Fish Data: Plot of First Two Canonical Variables

 Fish Data: Plot of First Two Canonical Variables