PHSELECT Procedure
Cox Regression
(View the complete code for this example.)
Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.
The following DATA step creates the data table mylib.getStarted in your CAS session. This data table consists of 100 observations on a failure time variable (Time), an indicator variable (Status) that has two values (0 for censored observations and 1 for event observations), three classification variables (C1–C3), and four continuous variables (X1–X4). This DATA step assumes that your libref is named mylib, but you can substitute any appropriately defined libref.
data mylib.getStarted;
input Time Status C1$ C2 C3$ X1-X4;
datalines;
53 0 Low 1 M 1.11 2.000 3.6128 12.0
12 0 High 1 M 1.40 1.362 3.8388 8.8
11 1 Low 1 F 1.57 1.672 3.8865 7.5
7 1 Medium 0 M 1.04 2.000 3.7324 5.1
2 1 Low 1 M 1.52 2.000 3.8751 9.8
41 0 High 1 M 1.76 1.447 3.7243 12.8
6 1 Critical 1 F 1.36 1.462 3.5441 9.0
6 1 Critical 1 M 1.42 1.690 3.9294 10.4
16 1 Medium 1 M 1.32 0.699 3.6990 8.8
41 1 Medium 1 M 1.00 1.477 3.4771 10.2
2 1 High 0 M 1.30 2.000 3.7243 5.1
58 1 Medium 1 M 1.20 1.580 3.6990 12.1
11 1 High 1 M 1.08 1.903 3.5051 9.6
12 0 Critical 1 F 1.15 1.146 3.6435 11.6
16 0 High 1 F 1.15 0.903 3.8573 13.0
54 1 Medium 1 M 1.26 1.699 3.7243 9.0
51 1 Low 0 M 1.57 1.041 3.4150 7.7
67 1 Medium 1 M 1.32 1.041 3.6435 12.8
1 1 Low 1 M 1.94 1.954 3.9868 12.0
19 0 Medium 1 M 1.32 2.000 3.7709 13.0
1 1 Medium 1 M 2.22 1.954 3.6628 9.4
35 1 Medium 0 M 1.11 1.176 3.6532 7.0
41 1 High 1 M 1.15 1.342 3.5185 5.0
58 1 Medium 1 M 1.20 1.580 3.6990 12.1
11 1 Medium 1 M 1.11 1.279 3.8808 14.0
41 1 Medium 1 M 1.00 1.477 3.4771 10.2
19 1 High 0 M 1.26 1.929 3.7924 7.5
89 1 High 1 M 1.32 1.623 3.6532 14.0
4 0 Medium 1 F 1.95 0.778 4.0453 10.2
6 1 Critical 1 F 1.36 1.462 3.5441 9.0
57 0 Low 1 F 1.26 1.954 3.9685 12.5
5 1 High 1 M 2.24 1.663 4.9542 10.1
17 1 Medium 1 F 1.59 1.613 3.4314 11.2
77 0 Low 1 F 1.08 0.954 3.6812 14.0
66 1 High 1 M 1.45 1.820 3.7853 6.6
16 1 Medium 1 M 1.32 0.699 3.6990 8.8
8 0 Critical 1 M 1.08 1.653 3.8325 9.9
19 1 High 0 M 1.26 1.929 3.7924 7.5
37 1 High 1 F 1.60 1.204 3.9542 11.0
52 1 Medium 1 M 1.00 1.653 3.8573 10.1
13 0 High 0 F 1.66 1.792 3.6435 4.9
3 1 Medium 1 F 1.54 1.935 4.4757 6.7
51 1 Low 0 M 1.57 1.041 3.4150 7.7
2 1 High 0 M 1.30 2.000 3.7243 5.1
25 1 Medium 1 M 1.00 1.644 3.8195 12.4
11 1 Low 1 F 1.57 1.672 3.8865 7.5
19 1 Low 1 M 1.08 2.000 3.9191 14.4
9 1 Low 1 M 1.72 1.740 3.7993 8.2
6 1 Medium 1 M 1.11 1.398 3.5185 9.7
16 1 Medium 1 M 1.34 2.000 3.9345 9.0
12 0 High 1 M 1.40 1.362 3.8388 8.8
17 1 High 1 M 1.23 1.447 3.8808 10.0
17 1 High 1 M 1.23 1.447 3.8808 10.0
41 0 High 1 M 1.76 1.447 3.7243 12.8
88 1 High 1 F 1.18 1.756 3.5563 10.6
16 0 High 1 F 1.15 0.903 3.8573 13.0
4 0 Low 1 F 1.92 1.623 3.9590 10.0
7 1 Low 1 M 1.18 1.519 3.7243 11.4
19 0 Medium 1 M 1.32 1.519 3.8808 10.8
8 0 Critical 1 M 1.08 1.653 3.8325 9.9
57 0 Low 1 F 1.26 1.954 3.9685 12.5
7 1 Medium 1 M 1.98 1.568 3.3617 9.5
19 1 Low 1 M 1.08 2.000 3.9191 14.4
7 0 Low 1 F 1.53 1.881 3.5911 10.2
5 1 High 1 F 1.68 1.732 3.7324 6.5
2 1 Low 0 M 1.75 1.255 3.8062 11.3
1 1 Low 1 M 1.94 1.954 3.9868 12.0
15 1 Low 1 M 1.60 1.431 3.6902 10.6
26 1 Low 1 M 1.23 2.000 3.6021 11.2
92 1 Low 1 M 1.43 1.415 4.0755 11.0
11 1 Medium 1 M 1.30 1.820 3.7993 13.2
3 1 Medium 1 F 1.54 1.935 4.4757 6.7
66 1 High 1 M 1.45 1.820 3.7853 6.6
1 1 Medium 1 M 2.22 1.954 3.6628 9.4
11 1 Medium 1 M 1.30 1.820 3.7993 13.2
14 1 Medium 1 M 1.40 1.255 3.7243 14.6
32 1 High 1 M 1.32 1.634 3.6990 10.6
24 1 High 1 M 1.30 0.477 4.0899 14.6
18 1 Critical 1 F 1.45 0.903 3.5682 7.5
5 1 High 1 F 1.68 1.732 3.7324 6.5
2 1 Low 0 M 1.75 1.255 3.8062 11.3
26 1 Low 1 M 1.23 2.000 3.6021 11.2
11 1 High 1 M 1.23 1.176 3.7709 12.0
28 0 High 1 M 1.23 1.672 3.7482 7.3
19 0 Medium 1 M 1.32 2.000 3.7709 13.0
54 1 Medium 1 M 1.26 1.699 3.7243 9.0
2 1 Low 1 M 1.52 2.000 3.8751 9.8
7 0 Medium 1 M 1.11 1.857 3.7993 12.4
19 0 Medium 1 M 1.32 1.519 3.8808 10.8
6 1 Critical 1 M 1.42 1.690 3.9294 10.4
6 1 Medium 1 M 1.11 1.398 3.5185 9.7
11 1 Medium 1 M 1.11 1.279 3.8808 14.0
6 1 Critical 0 M 2.11 1.362 3.5441 10.2
13 1 Medium 0 M 0.78 1.398 3.5798 5.5
18 1 Critical 1 F 1.45 0.903 3.5682 7.5
7 0 Low 1 F 1.53 1.881 3.5911 10.2
53 0 Low 1 M 1.11 2.000 3.6128 12.0
11 1 High 1 M 1.08 1.903 3.5051 9.6
11 0 High 1 M 1.61 1.845 3.7324 14.0
89 1 High 1 M 1.32 1.623 3.6532 14.0
;
The following statements fit a Cox proportional hazards model to these data by using three classification effects for the variables C1–C3 and four regressor effects for the variables X1–X4. The ITHIST option displays a table that summarizes the steps of the optimization.
proc phselect data=mylib.getStarted ithist;
class C1-C3;
model Time*Status(0) = C1-C3 X1-X4;
run;
The output from this analysis is presented in Figure 1 through Figure 8.
Figure 1 displays the "Model Information" table. The variable Time is the failure time variable. The variable Status is the censoring variable; the value 0 indicates censored observations. The PHSELECT procedure uses a quasi-Newton algorithm to maximize the partial likelihood to estimate the regression coefficients.
Figure 1: Model Information
| Model Information | |
|---|---|
| Data Source | GETSTARTED |
| Response Variable | Time |
| Censoring Variable | Status |
| Censoring Values | 0 |
| Optimization technique | Dual Quasi-Newton |
Figure 2 displays the "Number of Observations" table. All 100 observations in the data table are used in the analysis; of them, 26 are censored and 74 are uncensored.
Figure 2: Number of Observations
| Number of Observations | |||
|---|---|---|---|
| Description | Total | Event | Censored |
| Number of Observations Read | 100 | 74 | 26 |
| Number of Observations Used | 100 | 74 | 26 |
The classification variables C1–C3 are parameterized using the GLM parameterization, which is the default. The variable C1 has four unique formatted levels; each of the two variables C2 and C3 has two levels. The classification levels are displayed in the "Class Level Information" table in Figure 3.
Figure 3: Class Level Information
| Class Level Information | ||
|---|---|---|
| Class | Levels | Values |
| C1 | 4 | Critical High Low Medium |
| C2 | 2 | 0 1 |
| C3 | 2 | F M |
The "Iteration History" table is shown in Figure 4. The quasi-Newton algorithm converged after 12 iterations, not counting the initial setup iteration.
Figure 4: Iteration History
| Iteration History | ||||
|---|---|---|---|---|
| Iteration | Evaluations | Objective Function | Change | Max Gradient |
| 0 | 4 | 272.56284984 | . | 65.17039 |
| 1 | 3 | 267.70594206 | 4.85690778 | 7.86724 |
| 2 | 2 | 261.80868056 | 5.89726150 | 4.723085 |
| 3 | 2 | 256.62059571 | 5.18808485 | 4.650619 |
| 4 | 2 | 255.59346125 | 1.02713446 | 1.0462 |
| 5 | 2 | 255.44142594 | 0.15203531 | 1.018105 |
| 6 | 3 | 255.40692196 | 0.03450398 | 0.353611 |
| 7 | 3 | 255.39461736 | 0.01230460 | 0.415741 |
| 8 | 3 | 255.39362671 | 0.00099065 | 0.074411 |
| 9 | 3 | 255.39339211 | 0.00023460 | 0.062435 |
| 10 | 3 | 255.39336653 | 0.00002558 | 0.004922 |
| 11 | 3 | 255.39336262 | 0.00000391 | 0.002273 |
| 12 | 3 | 255.39336237 | 0.00000026 | 0.000452 |
Figure 5 displays the final convergence status of the quasi-Newton algorithm. The GCONV=1E-8 convergence criterion is satisfied.
Figure 5: Convergence Status
| Convergence criterion (GCONV=1E-8) satisfied. |
Figure 6 displays the "Dimensions" table for this model. This table summarizes some important sizes of various model components. For example, it shows that the design matrix has 12 columns: 4 columns for the effects that are associated with the classification variable C1, 2 columns for each of the classification variables C2 and C3, and 1 column for each of the continuous variables
X1–X4. However, the rank of the crossproducts matrix is only 9. Because the classification variables C1–C3 use GLM parameterization, there is one singularity in the crossproducts matrix of the model for each classification variable. Consequently, only nine parameters enter the optimization.
Figure 6: Dimensions in Cox Regression
| Dimensions | |
|---|---|
| Number of Effects | 7 |
| Max Effect Columns | 4 |
| Columns in Design | 12 |
| Rank of Design | 9 |
The "Fit Statistics" table is shown in Figure 7. The –2 log likelihood at the converged estimates is 510.78672. You can use this value to compare the model to nested model alternatives by means of a likelihood ratio test. To compare models that are not nested, you can use information criteria such as Akaike’s information criterion (AIC), Akaike’s bias-corrected information criterion (AICC), and the Schwarz Bayesian information criterion (BIC). These criteria penalize the –2 log partial likelihood for the number of parameters.
Figure 7: Fit Statistics
| Fit Statistics | |
|---|---|
| -2 Log Likelihood | 510.78672 |
| AIC (smaller is better) | 528.78672 |
| AICC (smaller is better) | 531.59922 |
| SBC (smaller is better) | 549.52331 |
The "Parameter Estimates" table in Figure 8 shows that many parameters have fairly large p-values, indicating that one or more of the model effects might not be necessary.
Figure 8: Parameter Estimates
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | Chi-Square | Pr > ChiSq |
| C1 Critical | 1 | 0.400493 | 0.489562 | 0.6692 | 0.4133 |
| C1 High | 1 | -0.990602 | 0.325145 | 9.2821 | 0.0023 |
| C1 Low | 1 | -0.665495 | 0.339148 | 3.8504 | 0.0497 |
| C1 Medium | 0 | 0 | . | . | . |
| C2 0 | 1 | 0.410704 | 0.398833 | 1.0604 | 0.3031 |
| C2 1 | 0 | 0 | . | . | . |
| C3 F | 1 | -0.172550 | 0.337940 | 0.2607 | 0.6096 |
| C3 M | 0 | 0 | . | . | . |
| X1 | 1 | 1.970488 | 0.524142 | 14.1335 | 0.0002 |
| X2 | 1 | 0.726605 | 0.415141 | 3.0634 | 0.0801 |
| X3 | 1 | 0.725390 | 0.585658 | 1.5341 | 0.2155 |
| X4 | 1 | -0.131226 | 0.057103 | 5.2810 | 0.0216 |
Finally, the procedure displays the table in Figure 9, which shows the amount of time (in seconds) that PROC PHSELECT required to perform different tasks in the analysis.
Figure 9: Procedure Timing
| Task Timing | ||
|---|---|---|
| Task | Seconds | Percent |
| Setup and Parsing | 0.02 | 9.03% |
| Levelization | 0.02 | 8.67% |
| Model Initialization | 0.01 | 2.18% |
| SSCP Computation | 0.00 | 0.79% |
| Model Fitting | 0.19 | 77.72% |
| Cleanup | 0.00 | 0.46% |
| Total | 0.24 | 100.00% |