Introduction to Discriminant Procedures

Example: Contrasting Univariate and Multivariate Analyses

(View the complete code for this example.)

Consider an artificial data set with two classes of observations indicated by 'H' and 'O'. The following statements generate and plot the data:

data random;
   drop n;

   Group = 'H';
   do n = 1 to 20;
      x = 4.5 + 2 * normal(57391);
      y = x + .5 + normal(57391);
      output;
   end;

   Group = 'O';
   do n = 1 to 20;
      x = 6.25 + 2 * normal(57391);
      y = x - 1 + normal(57391);
      output;
   end;

run;

proc sgplot noautolegend;
   scatter y=y x=x / markerchar=group group=group;
run;

The plot is shown in Figure 1.

Figure 1: Groups for Contrasting Univariate and Multivariate Analyses

Groups for Contrasting Univariate and Multivariate Analyses


The following statements perform a canonical discriminant analysis and display the results in Figure 2:

proc candisc anova;
   class Group;
   var x y;
run;

Figure 2: Contrasting Univariate and Multivariate Analyses

The CANDISC Procedure

Total Sample Size40DF Total39
Variables2DF Within Classes38
Classes2DF Between Classes1

Number of Observations Read40
Number of Observations Used40

Class Level Information
GroupVariable
Name
FrequencyWeightProportion
HH2020.00000.500000
OO2020.00000.500000

The CANDISC Procedure

Univariate Test Statistics
F Statistics, Num DF=1, Den DF=38
VariableTotal
Standard
Deviation
Pooled
Standard
Deviation
Between
Standard
Deviation
R-SquareR-Square
/ (1-RSq)
F ValuePr > F
x2.17762.14980.68200.05030.05302.010.1641
y2.42152.44860.20470.00370.00370.140.7105

Average R-Square
Unweighted0.0269868
Weighted by Variance0.0245201

Multivariate Statistics and Exact F Statistics
S=1 M=0 N=17.5
StatisticValueF ValueNum DFDen DFPr > F
Wilks' Lambda0.6420370410.312370.0003
Pillai's Trace0.3579629610.312370.0003
Hotelling-Lawley Trace0.5575425210.312370.0003
Roy's Greatest Root0.5575425210.312370.0003

The CANDISC Procedure

 Canonical
Correlation
Adjusted
Canonical
Correlation
Approximate
Standard
Error
Squared
Canonical
Correlation
Eigenvalues of Inv(E)*H
= CanRsq/(1-CanRsq)
Test of H0: The canonical correlations in the current row and all that follow are zero
 EigenvalueDifferenceProportionCumulativeLikelihood
Ratio
Approximate
F Value
Num DFDen DFPr > F
10.5983000.5894670.1028080.3579630.5575 1.00001.00000.6420370410.312370.0003

Note:The F statistic is exact.


The CANDISC Procedure

Total Canonical Structure
VariableCan1
x-0.374883
y0.101206

Between Canonical Structure
VariableCan1
x-1.000000
y1.000000

Pooled Within Canonical Structure
VariableCan1
x-0.308237
y0.081243

The CANDISC Procedure

Total-Sample Standardized
Canonical Coefficients
VariableCan1
x-2.625596855
y2.446680169

Pooled Within-Class Standardized
Canonical Coefficients
VariableCan1
x-2.592150014
y2.474116072

Raw Canonical Coefficients
VariableCan1
x-1.205756217
y1.010412967

Class Means on Canonical
Variables
GroupCan1
H0.7277811475
O-.7277811475


The univariate R squares are very small, 0.0503 for x and 0.0037 for y, and neither variable shows a significant difference between the classes at the 0.10 level.

The multivariate test for differences between the classes is significant at the 0.0003 level. Thus, the multivariate analysis has found a highly significant difference, whereas the univariate analyses failed to achieve even the 0.10 level. The raw canonical coefficients for the first canonical variable, Can1, show that the classes differ most widely on the linear combination -1.205756217 x + 1.010412967 y or approximately y - 1.2 x. The R square between Can1 and the CLASS variable is 0.357963 as given by the squared canonical correlation, which is much higher than either univariate R square.

In this example, the variables are highly correlated within classes. If the within-class correlation were smaller, there would be greater agreement between the univariate and multivariate analyses.

Last updated: October 28, 2020