DEEPCAUSAL Procedure
Example 15.2 Personalized Discount Policy for an Online Media Company
This example illustrates how a music subscription service from an online media company can offer targeted discounts through a personalized pricing plan based on many features that it observes about its customers to encourage them to buy more songs or beco me members. The main goal is to construct a policy that raises demand enough to boost overall revenue despite decreasing the price for some customers.
The data set is provided by the Microsoft research project ALICE and is available at https://msalicedatapublic.z5.web.core.windows.net/datasets/Pricing/pricing_sample.csv. The data set has 10,000 simulated observations that represent customers’ personal characteristics, such as age and log income, and online behavior history, such as previous purchase and previous online time per week. The treatment variable, t, is a binary variable that indicates whether or not a discount is applied. This variable is generated according to the values that the variable price takes in the data set. The value of t is 0 if the value of price is 1, indicating that no discount is given, and the value of t is 1 if the value of price is less than 1, indicating that a discount is applied. The outcome variable, revenue, is calculated by multiplying the number of songs purchased during the discount season by the price paid for the songs. Table 3 shows the names of the variables that are used in the model, their types, and their definitions. The type can be T (treatment), Y (outcome), (a covariate in the propensity score model), and/or
(a covariate in the outcome model).
Table 3: Model Variables
Assuming that you have downloaded the data set pricing_sample in your session that is associated with the mylib libref, the following statements create an ID variable that has a unique value for each observation, the treatment variable (t), and the outcome variable (revenue) in the data table new_pricing_sample:
data mylib.new_pricing_sample;
set mylib.pricing_sample;
id = put(_threadid_,8.) || '_' || Put(_n_,8.); * ID variable;
if price<1 then t=1; else t=0; * treatment variable;
revenue=price*demand; * outcome variable;
run;
The first step in policy evaluation and policy optimization is to estimate the effect of the treatment and to save to a specified output data table the details of the estimation, including ,
,
, the residual, and the influence functions for each unit. You can do this by using the following statements:
/*--- Estimate the treatment effect and save the estimation details ---*/
proc deepcausal data=mylib.new_pricing_sample;
id id;
psmodel t = account_age age avg_hours days_visited friends_count
has_membership is_US songs_purchased income /
dnn=(nodes=(32 32 32 32)
train=(optimizer=(miniBatchSize=500 regL1=0.0001 maxEpochs=32000
algorithm=adam) nthreads=20 seed=12345 recordseed=67890));
model revenue = account_age age avg_hours days_visited friends_count
has_membership is_US songs_purchased income /
dnn=(nodes=(32 32 32 32)
train=(optimizer=(miniBatchSize=500 regL1=0.001 maxEpochs=32000
algorithm=adam) nthreads=20 seed=12345 recordseed=67890));
infer out=mylib.oest outdetails=mylib.odetails;
run;
For this example, the same covariates are used in both the propensity score mode and the outcome model. The model estimation details are saved in the output data table odetails.
The estimation results are shown in Output 15.2.1. The estimate of the average treatment effect, ATE, is negative and statistically significant, suggesting that the discount, on average, causes revenue that is generated by the whole population to decrease. However, the effect of the discount on the revenue among the customers who received a discount, which is measured by the parameter ATT (average treatment effect on the treated), is considerably different from the ATE estimate. It might suggest that identifying the characteristics of customers to whom the discount matters would help construct the optimum policy so that it uses the fewest resources and earns the most profit.
Output 15.2.1: Parameters of Interest
| Population Parameter Estimates | |||||||
|---|---|---|---|---|---|---|---|
| Parameter | Potential Outcome | Estimate | Standard Error | 95% Confidence Limits | Z | Pr > |Z| | |
| Average | 0 | 14.323260 | 0.075167 | 14.175936 | 14.470584 | 190.55 | <.0001 |
| Average | 1 | 13.577143 | 0.057725 | 13.464004 | 13.690282 | 235.20 | <.0001 |
| ATE | -0.746117 | 0.024569 | -0.794272 | -0.697962 | -30.37 | <.0001 | |
| Subpopulation Parameter Estimates | ||||||||
|---|---|---|---|---|---|---|---|---|
| Parameter | Potential Outcome | Treatment Level | Estimate | Standard Error | 95% Confidence Limits | Z | Pr > |Z| | |
| Conditional Average | 0 | 0 | 18.353066 | 0.230375 | 17.901539 | 18.804594 | 79.67 | <.0001 |
| Conditional Average | 0 | 1 | 11.225712 | 0.135104 | 10.960913 | 11.490510 | 83.09 | <.0001 |
| Conditional Average | 1 | 0 | 16.576263 | 0.204324 | 16.175796 | 16.976730 | 81.13 | <.0001 |
| Conditional Average | 1 | 1 | 11.271842 | 0.120335 | 11.035990 | 11.507693 | 93.67 | <.0001 |
| ATT | 1 | 0.046130 | 0.027921 | -0.008594 | 0.100854 | 1.65 | 0.0985 | |
| ATU | 0 | -1.776803 | 0.036937 | -1.849198 | -1.704409 | -48.10 | <.0001 | |
| Composition Effect | 1 | -5.304422 | 0.305935 | -5.904044 | -4.704799 | -17.34 | <.0001 | |
| Observed Difference | -7.081225 | 0.330048 | -7.728108 | -6.434342 | -21.46 | <.0001 | ||
In policy optimization, the goal is to maximize the expected utility of a policy. For the definition of the expected utility function and how it is estimated, as well as details such as the definition of a policy rule, see the section Policy Evaluation and Comparison. In this example, because a positive treatment effect is preferred, anyone whose value is positive should be given the treatment (that is, a discount). This function is sometimes referred to as the individual treatment effect (ITE).
The variable _beta_ in the data table odetails that is obtained during the previous estimation contains the estimated values for this function for each individual. You can construct an optimal policy by offering a discount to individuals whose _beta_ value is positive. For comparison, the following DATA steps create the optimal policy, s2, along with the other policies: s0, offer n o one a discount; and s1, offer everyone a discount.
data mylib.discountPolicy;
set mylib.odetails;
s0 = 0; * treat none;
s1 = 1; * treat all;
s2 = _beta_>=0; * treat only individuals whose ITEs are positive;
keep id t account_age age avg_hours days_visited
friends_count has_membership is_US songs_purchased
income revenue s0 s1 s2;
run;
You can do policy evaluation by specifying the POLICY= option in the INFER statement and do policy comparison by specifying the POLICYCOMPARISON= option in the same statement. The following statements estimate the model and do policy evaluation and comparison:
/*--- Policy evaluation and comparison ---*/
proc deepcausal data=mylib.discountPolicy;
id id;
psmodel t = account_age age avg_hours days_visited friends_count
has_membership is_US songs_purchased income /
dnn=(nodes=(32 32 32 32)
train=(optimizer=(miniBatchSize=500 regL1=0.0001
maxEpochs=32000 algorithm=adam) nthreads=20
seed=12345 recordseed=67890));
model revenue = account_age age avg_hours days_visited friends_count
has_membership is_US songs_purchased income /
dnn=(nodes=(32 32 32 32)
train=(optimizer=(miniBatchSize=500 regL1=0.001
maxEpochs=32000 algorithm=adam) nthreads=20
seed=12345 recordseed=67890));
infer policy=(t s0 s1 s2) policyComparison=(base=(t)
compare=(s0 s1 s2))
out=mylib.oest2 outdetails=mylib.odetails2;
run;
The option POLICY=(t s0 s1 s2) requests that the policy evaluation be done for the initial discount policy (t), the policy of giving no one a discount (s0), the policy of giving everyone a discount (s1), and the personalized discount policy of giving a discount only to those whose _beta_ value is positive (s2). The suboption COMPARE=(s0 s1 s2) in the POLICYCOMPARISON= option indicates that the policies s0, s1, and s2 are to be compared to the base policy, t, which is specified by the suboption BASE=(t), also in the POLICYCOMPARISON= option.
The policy evaluation results are shown in Output 15.2.2.
Output 15.2.2: Policy Evaluation
| Policy Evaluation | ||||||
|---|---|---|---|---|---|---|
| Policy | Estimate | Standard Error | 95% Confidence Limits | Z | Pr > |Z| | |
| t | 14.383749 | 0.068443 | 14.249603 | 14.517895 | 210.16 | <.0001 |
| s0 | 14.309670 | 0.074717 | 14.163227 | 14.456112 | 191.52 | <.0001 |
| s1 | 13.592299 | 0.058081 | 13.478463 | 13.706135 | 234.02 | <.0001 |
| s2 | 14.835283 | 0.069198 | 14.699658 | 14.970909 | 214.39 | <.0001 |
The optimal policy, s2, indeed has the highest estimate. When the discount is given to customers whose ITE is positive, the average outcome, which in this example is the revenue for each person, is higher than the average outcome that is obtained by the base policy. This is the largest revenue that is generated, and it is generated by the personalized discount policy.
Output 15.2.3 shows the output for the policy comparison. Negative values indicate that the corresponding discount policy is worse than the base policy, and positive values indicate that the corresponding discount policy is better than the base policy. As expected, s2, giving the discount to the right person, produces the best results.
Output 15.2.3: Policy Comparison
| Policy Comparison | |||||||
|---|---|---|---|---|---|---|---|
| Comparison Policy | Base Policy | Estimate | Standard Error | 95% Confidence Limits | Z | Pr > |Z| | |
| s0 | t | -0.074079 | 0.017349 | -0.108082 | -0.040077 | -4.27 | <.0001 |
| s1 | t | -0.791450 | 0.014168 | -0.819218 | -0.763681 | -55.86 | <.0001 |
| s2 | t | 0.451535 | 0.015455 | 0.421243 | 0.481826 | 29.22 | <.0001 |