The PROBIT Procedure

Example 94.4 An Epidemiology Study

The data in this example, which are from an epidemiology study, consist of five variables: the number, r, of individuals surviving after an epidemic, out of n treated, for combinations of medicine dosage (dose), treatment (treat = A, B), and sex (sex = 0(Female), 1(Male)).

To see whether the two treatments have different effects on male and female individual survival rates, the interaction term between the two variables treat and sex is included in the model.

The following invocation of PROC PROBIT fits the binary probit model to the grouped data:

data epidemic;
   input treat$ dose n r sex @@;
   label dose = Dose;
   datalines;
A  2.17 142 142  0   A   .57 132  47  1
A  1.68 128 105  1   A  1.08 126 100  0
A  1.79 125 118  0   B  1.66 117 115  1
B  1.49 127 114  0   B  1.17  51  44  1
B  2.00 127 126  0   B   .80 129 100  1
;

data xval;
   input treat $ dose sex;
   datalines;
B  2.  1
;
proc probit optc lackfit covout data=epidemic
            outest = out1 xdata = xval
            Plots=(predpplot ippplot lpredplot);
   class treat sex;
   model r/n = dose treat sex sex*treat/corrb covb inversecl;
   output out = out2 p =p;
run;

The results of this analysis are shown in the outputs that follow.

Output 94.4.1 displays the table of level information for all classification variables in the CLASS statement.

Output 94.4.1: Class Level Information

The Probit Procedure

Class Level Information
NameLevelsValues
treat2A B
sex20 1


Output 94.4.2 displays the table of parameter information for the effects in the MODEL statement.

Output 94.4.2: Parameter Information

Parameter Information
ParameterEffecttreatsex
InterceptIntercept  
dosedose  
treatAtreatA 
treatBtreatB 
sex0sex 0
sex1sex 1
treatAsex0treat*sexA0
treatAsex1treat*sexA1
treatBsex0treat*sexB0
treatBsex1treat*sexB1


Output 94.4.3 displays background information about the model fit. Included are the name of the input data set, the response variables used, the numbers of observations, events, and trials, the type of distribution, and the final value of the log-likelihood function.

Output 94.4.3: Model Information

The Probit Procedure

Model Information
Data SetWORK.EPIDEMIC
Events Variabler
Trials Variablen
Number of Observations10
Number of Events1011
Number of Trials1204
Name of DistributionNormal
Log Likelihood-387.2467391


Output 94.4.4 displays the table of goodness-of-fit tests requested with the LACKFIT option in the PROC PROBIT statement. Two goodness-of-fit statistics, the Pearson’s chi-square statistic and the likelihood ratio chi-square statistic, are computed. The grouping method for computing these statistics can be specified by the AGGREGATE= option. The details can be found in the AGGREGATE= option, and an example can be found in the second part of this example. By default, the PROBIT procedure uses the covariates in the MODEL statement to do grouping. Observations with the same values of the covariates in the MODEL statement are grouped into cells and the two statistics are computed according to these cells. The total number of cells and the number of levels for the response variable are reported next in the "Response-Covariate Profile."

In this example, neither the Pearson’s chi-square nor the log-likelihood ratio chi-square tests are significant at the 0.1 level, which is the default test level used by the PROBIT procedure. That means that the model, which includes the interaction of treat and sex, is suitable for this epidemiology data set. (Further investigation shows that models without the interaction of treat and sex are not acceptable by either test.)

Output 94.4.4: Goodness-of-Fit Tests and Response-Covariate Profile

Goodness-of-Fit Tests
StatisticValueDFValue/DFPr > ChiSq
Pearson Chi-Square4.931741.23290.2944
L.R. Chi-Square5.707941.42700.2220

Response-Covariate Profile
Response Levels2
Number of Covariate Values10


Output 94.4.5 displays the Type III test results for all effects specified in the MODEL statement, which include the degrees of freedom for the effect, the Wald Chi-Square test statistic, and the p-value.

Output 94.4.5: Type III Tests

Type III Analysis of Effects
EffectDFWald
Chi-Square
Pr > ChiSq
dose142.1691<.0001
treat116.1421<.0001
sex11.77100.1833
treat*sex113.93430.0002


Output 94.4.6 displays the table of parameter estimates for the model. The PROBIT procedure displays information for all the parameters of an effect. Degenerate parameters are indicated by 0 degree of freedom. Confidence intervals are computed for all parameters with nonzero degrees of freedom, including the natural threshold C if the OPTC option is specified in the PROC PROBIT statement. The confidence level can be specified by the ALPHA= option in the MODEL statement. The default confidence level is 95%.

Output 94.4.6: Analysis of Parameter Estimates

Analysis of Maximum Likelihood Parameter Estimates
Parameter DFEstimateStandard
Error
95% Confidence LimitsChi-SquarePr > ChiSq
Intercept  1-0.88710.3632-1.5991-0.17525.960.0146
dose  11.67740.25831.17112.183742.17<.0001
treatA 1-1.25370.2616-1.7664-0.741022.97<.0001
treatB 00.0000.....
sex0 1-0.46330.2289-0.9119-0.01474.100.0429
sex1 00.0000.....
treat*sexA011.28990.34560.61261.967213.930.0002
treat*sexA100.0000.....
treat*sexB000.0000.....
treat*sexB100.0000.....
_C_  10.27350.09460.08810.4589  


From Output 94.4.6, you can see the following results:

  • The variable dose has a significant positive effect on the survival rate.

  • Individuals under treatment A have a lower survival rate.

  • Male individuals have a higher survival rate.

  • Female individuals under treatment A have a higher survival rate.

Output 94.4.7 and Output 94.4.8 display tables of estimated covariance matrix and estimated correlation matrix for estimated parameters with a nonzero degree of freedom, respectively. They are computed by the inverse of the Hessian matrix of the estimated parameters.

Output 94.4.7: Estimated Covariance Matrix

Estimated Covariance Matrix
 InterceptdosetreatAsex0treatAsex0_C_
Intercept0.131944-0.0873530.0535510.030285-0.067056-0.028073
dose-0.0873530.066723-0.047506-0.0340810.0586200.018196
treatA0.053551-0.0475060.0684250.036063-0.075323-0.017084
sex00.030285-0.0340810.0360630.052383-0.063599-0.008088
treatAsex0-0.0670560.058620-0.075323-0.0635990.1194080.019134
_C_-0.0280730.018196-0.017084-0.0080880.0191340.008948


Output 94.4.8: Estimated Correlation Matrix

Estimated Correlation Matrix
 InterceptdosetreatAsex0treatAsex0_C_
Intercept1.000000-0.9309980.5635950.364284-0.534227-0.817027
dose-0.9309981.000000-0.703083-0.5764770.6567440.744699
treatA0.563595-0.7030831.0000000.602359-0.833299-0.690420
sex00.364284-0.5764770.6023591.000000-0.804154-0.373565
treatAsex0-0.5342270.656744-0.833299-0.8041541.0000000.585364
_C_-0.8170270.744699-0.690420-0.3735650.5853641.000000


Output 94.4.9 displays the computed values and fiducial limits for the first single continuous variable dose in the MODEL statement, given the probability levels, without the effect of the natural threshold, and when the option INSERSECL in the MODEL statement is specified. If there is no single continuous variable in the MODEL specification but the INVERSECL option is specified, an error is reported.

Output 94.4.9: Probit Analysis on Dose

The Probit Procedure

Probit Analysis on dose
Probabilitydose95% Fiducial Limits
0.01-0.85801-1.81301-0.33743
0.02-0.69549-1.58167-0.21116
0.03-0.59238-1.43501-0.13093
0.04-0.51482-1.32476-0.07050
0.05-0.45172-1.23513-0.02130
0.06-0.39802-1.158880.02063
0.07-0.35093-1.092060.05742
0.08-0.30877-1.032260.09039
0.09-0.27043-0.977900.12040
0.10-0.23513-0.927880.14805
0.15-0.08900-0.721070.26278
0.200.02714-0.557060.35434
0.250.12678-0.416690.43322
0.300.21625-0.290950.50437
0.350.29917-0.174770.57064
0.400.37785-0.064870.63387
0.450.453970.041040.69546
0.500.528880.144810.75654
0.550.603800.248000.81819
0.600.679920.352130.88157
0.650.758600.458790.94803
0.700.841510.569851.01942
0.750.930990.687701.09847
0.801.030630.815711.18970
0.851.146770.959261.30171
0.901.292901.128671.45386
0.911.328191.167471.49273
0.921.366541.208671.53590
0.931.408701.252841.58450
0.941.455791.300841.64012
0.951.509491.353971.70515
0.961.572581.414431.78353
0.971.650151.486261.88238
0.981.753261.578332.01720
0.991.915771.717762.23537


If the XDATA= option is used to input a data set for the independent variables in the MODEL statement, the PROBIT procedure uses these values for the independent variables other than the single continuous variable. Missing values are not permitted in the XDATA= data set for the independent variables, although the value for the single continuous variable is not used in the computing of the fiducial limits. A suitable valid value should be given. In the data set xval created by the SAS statements, dose = 2. Only one observation from the XDATA= data set is used to produce a probit analysis table for a combination of classification variable levels. If more than one observation is present in the XDATA= data set, only the last observation is used.

See the section XDATA= SAS-data-set for the default values for those effects other than the single continuous variable, for which the fiducial limits are computed.

In this example, there are two classification variables, treat and sex. Fiducial limits for the dose variable are computed for the highest level of the classification variables, treat = B and sex = 1, which is the default specification. Since these are the default values, you would get the same values and fiducial limits if you did not specify the XDATA= option in this example. The confidence level for the fiducial limits can be specified by the ALPHA= option in the MODEL statement. The default level is 95%.

If a LOG10 or LOG option is used in the PROC PROBIT statement, the values and the fiducial limits are computed for both the single continuous variable and its logarithm.

Output 94.4.10 displays the OUTEST= data set. All parameters for an effect are included. The name of a parameter is generated by combining the variable names and levels in the effect. The maximum length of a parameter name is 32.

Output 94.4.10: Outest Data Set for Epidemiology Study

Obs_MODEL__NAME__TYPE__DIST__STATUS__LNLIKE_rInterceptdosetreatAtreatBsex0sex1treatAsex0treatAsex1treatBsex0treatBsex1_C_
1 rPARMSNormal0 Converged-387.247-1.00000-0.887141.67739-1.253670-0.4632901.289910000.27347
2 InterceptCOVNormal0 Converged-387.247-0.887140.13194-0.087350.0535500.030290-0.06706000-0.02807
3 doseCOVNormal0 Converged-387.2471.67739-0.087350.06672-0.047510-0.0340800.058620000.01820
4 treatACOVNormal0 Converged-387.247-1.253670.05355-0.047510.0684300.036060-0.07532000-0.01708
5 treatBCOVNormal0 Converged-387.2470.000000.000000.000000.0000000.0000000.000000000.00000
6 sex0COVNormal0 Converged-387.247-0.463290.03029-0.034080.0360600.052380-0.06360000-0.00809
7 sex1COVNormal0 Converged-387.2470.000000.000000.000000.0000000.0000000.000000000.00000
8 treatAsex0COVNormal0 Converged-387.2471.28991-0.067060.05862-0.075320-0.0636000.119410000.01913
9 treatAsex1COVNormal0 Converged-387.2470.000000.000000.000000.0000000.0000000.000000000.00000
10 treatBsex0COVNormal0 Converged-387.2470.000000.000000.000000.0000000.0000000.000000000.00000
11 treatBsex1COVNormal0 Converged-387.2470.000000.000000.000000.0000000.0000000.000000000.00000
12 _C_COVNormal0 Converged-387.2470.27347-0.028070.01820-0.017080-0.0080900.019130000.00895


The plots in the following three outputs, Output 94.4.11, Output 94.4.12, and Output 94.4.13, are generated by the PLOTS= option. The first plot, specified with the PREDPPLOT option, is the plot of the predicted probability against the first single continuous variable dose in the MODEL statement. You can specify values of other independent variables in the MODEL statement by using an XDATA= data set or by using the default values.

The second plot, specified with the IPPPLOT option, is the inverse of the predicted probability plot with the fiducial limits. It should be pointed out that the fiducial limits are not just the inverse of the confidence limits in the predicted probability plot; see the section Inverse Confidence Limits for the computation of these limits. The third plot, specified with the LPREDPLOT option, is the plot of the linear predictor against the first single continuous variable with the Wald confidence intervals.

Output 94.4.11: Predicted Probability Plot

Predicted Probability Plot


Output 94.4.12: Inverse Predicted Probability Plot

Inverse Predicted Probability Plot


Output 94.4.13: Linear Predictor Plot

Linear Predictor Plot