CAEFFECT Procedure
Example 6.2 Estimation by Regression Adjustment
(View the complete code for this example.)
This example uses the birthwgt data set available in the Sashelp library. The data are information about infant mortality in 2003 and were obtained from the US National Center for Health Statistics. A random sample of 100,000 observations is used in this example. The analysis of these data is purely illustrative, and definitive conclusions should not be drawn from it.
The following DATA step creates the data table mylib.birthwgt. For this example, it is assumed that your libref is named mylib, but you can substitute any appropriately defined libref.
data mylib.birthwgt;
set sashelp.birthwgt;
run;
For this example, the treatment variable Smoking is an indicator of maternal smoking behavior, with values of 'Yes' and 'No'. The outcome variable of interest, Death, is an indicator of infant death within one year of birth, with values of 'Yes' and 'No'. The following variables are assumed to be confounding variables that you can use in the model for the outcome variable, Death:
AgeGrouprepresents maternal ages of less than 20, between 20 and 35, and greater than 35, with values 1, 2, and 3, respectively.Drinkingis an indicator of maternal drinking during pregnancy, with values 'Yes' and 'No'.Marriedis an indicator of marital status, with values 'Yes' and 'No'.SomeCollegeis an indicator of whether the mother has 12 or more years of education, with values 'Yes' and 'No'.
The following statements use the GRADBOOST procedure, which is part of SAS Viya machine learning software, to fit a model for the outcome variable, Death:
proc gradboost data=mylib.birthwgt ntrees=100 seed=8959;
target Death / level=nominal;
input Smoking AgeGroup Married Drinking
SomeCollege /level=nominal;
savestate rstore=mylib.gbOutMod;
run;
The fitted gradient boosting model is saved in the data table gbOutMod, which you can provide as input to PROC CAEFFECT. The following statements use the saved outcome model to estimate two potential outcome means by regression adjustment:
proc caeffect data=mylib.birthwgt;
treatvar Smoking;
outcomevar Death( event='Yes') / type=Categorical;
outcomemodel restore=mylib.gbOutMod predname=P_DeathYes;
pom treatlev='Yes';
pom treatlev='No';
run;
You use the POM statement to specify the potential outcomes of interest for your analysis. For each potential outcome specification, you must use the TREATLEV= option to identify the level of treatment that defines the potential outcome.
In this example, you use the TYPE= option in the OUTVAR statement to indicate that the outcome variable of interest is categorical. For categorical outcomes, PROC CAEFFECT estimates the probability of a designated outcome event level. The potential outcome means that you estimate therefore correspond to the counterfactual probability of that outcome occurring, and all predicted outcome values must be between 0 and 1, inclusive. You specify the designated event level for the outcome variable by using the EVENT= option.
When you provide a previously fit outcome model as input to PROC CAEFFECT, you use the OUTCOMEMODEL statement. You specify the name of the analytic store that contains the saved model by using the RESTORE= option. You must also use the PREDNAME= option to provide the name of the variable that contains the outcome of interest when you score data by using the analytic store. If you do not know the names of the variables that the analytic store creates, you can use the DESCRIBE statement in PROC ASTORE and check the "Output Variables" table that it produces to find the variable names.
The output from this analysis is displayed by default and is presented in Output 6.2.1 through Output 6.2.3.
The "Model Information" table in Output 6.2.1 summarizes important information, including the data source, treatment variable name, estimation method, and type of outcome variable.
Output 6.2.1: Model Information
| Model Information | |
|---|---|
| Data Source | BIRTHWGT |
| Treatment Variable | Smoking |
| Outcome Type | Categorical |
| Estimation Method | Regression Adjustment |
| Outcome Model Store | GBOUTMOD |
Output 6.2.2 displays the "Number of Observations" table. In this analysis, a number of observations are not used, because the treatment variable Smoking has missing values. To indicate that a missing value is a valid value of the treatment variable, you can specify the COUNTMISSING option in the TREATVAR statement.
Output 6.2.2: Number of Observations
| Number of Observations Read | 100000 |
|---|---|
| Number of Observations Used | 94388 |
Output 6.2.3 displays the potential outcome mean estimates that are obtained by using the gradient boosting model for the outcome variable.
Output 6.2.3: Estimates Using a Stored Model
| POM Estimates | |
|---|---|
| Treatment Level | Estimate |
| Yes | 0.00829 |
| No | 0.00640 |
As an alternative to providing a previously fit outcome model as input to PROC CAEFFECT, you can also specify variables in the input data table that contain predicted counterfactual outcome values for each treatment level of interest. You specify the names of these variables by using the PREDOUT= option for each potential outcome specification in a POM statement. To obtain these predicted values, you can use a fitted model for the outcome and score the data with the treatment set to the counterfactual level of interest.
The following statements demonstrate one way to obtain predicted counterfactual outcome values that you can provide as input to PROC CAEFFECT. This example uses PROC ASTORE to score the data by using the previously fit model that is saved in the gbOutMod data table.
The first step is to create a table that renames the observed treatment variable tempSmoking and sets the value of the treatment variable Smoking to the counterfactual level 'Yes' for scoring the data. The following DATA step performs these tasks and creates the gbPredData data table:
data mylib.gbPredData;
set mylib.birthwgt;
tempSmoking = Smoking;
Smoking = 'Yes';
run;
After creating the gbPredData data table, you can score the counterfactual outcomes by using PROC ASTORE, as in the following statements. To avoid data duplication for large data tables, the variables in the input data table are not included in the output data table that PROC ASTORE creates unless you specify them in the COPYVARS= option. To score the counterfactual outcomes for the treatment level 'No', you must include all the predictors that are used in the outcome model in the COPYVARS= option along with the observed treatment values that are stored in the variable tempSmoking.
proc astore;
score data=mylib.gbPredData out=mylib.gbPredData
rstore=mylib.gbOutMod
copyvars=(tempSmoking AgeGroup Married
Drinking SomeCollege Death);
run;
Because PROC ASTORE uses the same variable names for the predicted values to score the next counterfactual values, you must first rename the variable that contains the predicted outcome probability of interest. You must also set the value of the treatment variable Smoking to the level 'No' to score the remaining counterfactual outcomes. These steps are performed using the following statements:
data mylib.gbPredData;
set mylib.gbPredData;
rename P_DeathYes = SmokingPred;
Smoking = 'No';
run;
Next, the data are scored again for the counterfactual treatment level 'No':
proc astore;
score data=mylib.gbPredData out=mylib.gbPredData
rstore=mylib.gbOutMod
copyvars=(tempSmoking SmokingPred Death);
run;
Because there are no additional counterfactual values to score, you can limit the list of variables in the COPYVARS= option to only those that you need for the effect estimation. Although it is not necessary to do this before using PROC CAEFFECT, you can use the following statements to rename the variable that contains the newly predicted counterfactual outcome values, and you can also replace the temporary name that was applied to the observed treatment variable with the original variable name, Smoking:
data mylib.gbPredData;
set mylib.gbPredData;
rename P_DeathYes = NoSmokingPred tempSmoking=Smoking;
run;
The following statements use the previously computed predicted counterfactual outcome values as input to estimate the potential outcome means:
proc caeffect data=mylib.gbPredData;
treatvar Smoking;
outcomevar Death( event='Yes') / type=Categorical;
pom treatlev='Yes' predOut=SmokingPred;
pom treatlev='No' predOut=NoSmokingPred;
run;
Output 6.2.4 displays the potential outcome mean estimates that are obtained by using the previously estimated counterfactual outcome values. These estimates of the potential outcome means are the same as the estimates shown in Output 6.2.3, because the same model for the outcome variable is used to obtain the predictions.
Output 6.2.4: Estimates Using Previously Predicted Values
| POM Estimates | |
|---|---|
| Treatment Level | Estimate |
| Yes | 0.00829 |
| No | 0.00640 |