DEEPCAUSAL Procedure

Example 15.2 Personalized Discount Policy for an Online Media Company

This example illustrates how a music subscription service from an online media company can offer targeted discounts through a personalized pricing plan based on many features that it observes about its customers to encourage them to buy more songs or beco me members. The main goal is to construct a policy that raises demand enough to boost overall revenue despite decreasing the price for some customers.

The data set is provided by the Microsoft research project ALICE and is available at https://msalicedatapublic.z5.web.core.windows.net/datasets/Pricing/pricing_sample.csv. The data set has 10,000 simulated observations that represent customers’ personal characteristics, such as age and log income, and online behavior history, such as previous purchase and previous online time per week. The treatment variable, t, is a binary variable that indicates whether or not a discount is applied. This variable is generated according to the values that the variable price takes in the data set. The value of t is 0 if the value of price is 1, indicating that no discount is given, and the value of t is 1 if the value of price is less than 1, indicating that a discount is applied. The outcome variable, revenue, is calculated by multiplying the number of songs purchased during the discount season by the price paid for the songs. Table 3 shows the names of the variables that are used in the model, their types, and their definitions. The type can be T (treatment), Y (outcome), bold x Superscript left-parenthesis 1 right-parenthesis (a covariate in the propensity score model), and/or bold x Superscript left-parenthesis 2 right-parenthesis (a covariate in the outcome model).

Table 3: Model Variables

Name Type Details
account_age bold x Superscript left-parenthesis 1 right-parenthesis, bold x Superscript left-parenthesis 2 right-parenthesis User’s account age
age bold x Superscript left-parenthesis 1 right-parenthesis, bold x Superscript left-parenthesis 2 right-parenthesis User’s age
avg_hours bold x Superscript left-parenthesis 1 right-parenthesis, bold x Superscript left-parenthesis 2 right-parenthesis Average number of hours user was online per week in the past
days_visited bold x Superscript left-parenthesis 1 right-parenthesis, bold x Superscript left-parenthesis 2 right-parenthesis Average number of days user visited website per week in the past
friend_count bold x Superscript left-parenthesis 1 right-parenthesis, bold x Superscript left-parenthesis 2 right-parenthesis Number of friends user connected to in account
has_membership bold x Superscript left-parenthesis 1 right-parenthesis, bold x Superscript left-parenthesis 2 right-parenthesis Whether user has membership
is_US bold x Superscript left-parenthesis 1 right-parenthesis, bold x Superscript left-parenthesis 2 right-parenthesis Whether user accesses website from US
songs_purchased bold x Superscript left-parenthesis 1 right-parenthesis, bold x Superscript left-parenthesis 2 right-parenthesis Average number of songs user purchased per week in the past
income bold x Superscript left-parenthesis 1 right-parenthesis, bold x Superscript left-parenthesis 2 right-parenthesis User’s income
t T Whether a discount is applied
revenue Y Number of songs purchased during discount season times price paid


Assuming that you have downloaded the data set pricing_sample in your session that is associated with the mylib libref, the following statements create an ID variable that has a unique value for each observation, the treatment variable (t), and the outcome variable (revenue) in the data table new_pricing_sample:

data mylib.new_pricing_sample;
   set mylib.pricing_sample;
       id = put(_threadid_,8.) || '_' || Put(_n_,8.); * ID variable;
       if price<1 then t=1; else t=0; * treatment variable;
       revenue=price*demand; * outcome variable;
run;

The first step in policy evaluation and policy optimization is to estimate the effect of the treatment and to save to a specified output data table the details of the estimation, including alpha left-parenthesis dot right-parenthesis, beta left-parenthesis dot right-parenthesis, p left-parenthesis dot right-parenthesis, the residual, and the influence functions for each unit. You can do this by using the following statements:

 /*--- Estimate the treatment effect and save the estimation details ---*/
 proc deepcausal data=mylib.new_pricing_sample;
    id id;
    psmodel t = account_age age avg_hours days_visited friends_count
                has_membership is_US songs_purchased income /
                dnn=(nodes=(32 32 32 32)
                train=(optimizer=(miniBatchSize=500 regL1=0.0001 maxEpochs=32000
                algorithm=adam) nthreads=20 seed=12345 recordseed=67890));
    model revenue = account_age age avg_hours days_visited friends_count
                    has_membership is_US songs_purchased income /
                    dnn=(nodes=(32 32 32 32)
                    train=(optimizer=(miniBatchSize=500 regL1=0.001 maxEpochs=32000
                    algorithm=adam) nthreads=20 seed=12345 recordseed=67890));
    infer out=mylib.oest outdetails=mylib.odetails;
 run;

For this example, the same covariates are used in both the propensity score mode and the outcome model. The model estimation details are saved in the output data table odetails.

The estimation results are shown in Output 15.2.1. The estimate of the average treatment effect, ATE, is negative and statistically significant, suggesting that the discount, on average, causes revenue that is generated by the whole population to decrease. However, the effect of the discount on the revenue among the customers who received a discount, which is measured by the parameter ATT (average treatment effect on the treated), is considerably different from the ATE estimate. It might suggest that identifying the characteristics of customers to whom the discount matters would help construct the optimum policy so that it uses the fewest resources and earns the most profit.

Output 15.2.1: Parameters of Interest

The DEEPCAUSAL Procedure

Population Parameter Estimates
ParameterPotential
Outcome
EstimateStandard
Error
95% Confidence LimitsZPr > |Z|
Average014.3232600.07516714.17593614.470584190.55<.0001
Average113.5771430.05772513.46400413.690282235.20<.0001
ATE -0.7461170.024569-0.794272-0.697962-30.37<.0001

Subpopulation Parameter Estimates
ParameterPotential
Outcome
Treatment
Level
EstimateStandard
Error
95% Confidence LimitsZPr > |Z|
Conditional Average0018.3530660.23037517.90153918.80459479.67<.0001
Conditional Average0111.2257120.13510410.96091311.49051083.09<.0001
Conditional Average1016.5762630.20432416.17579616.97673081.13<.0001
Conditional Average1111.2718420.12033511.03599011.50769393.67<.0001
ATT 10.0461300.027921-0.0085940.1008541.650.0985
ATU 0-1.7768030.036937-1.849198-1.704409-48.10<.0001
Composition Effect1 -5.3044220.305935-5.904044-4.704799-17.34<.0001
Observed Difference  -7.0812250.330048-7.728108-6.434342-21.46<.0001


In policy optimization, the goal is to maximize the expected utility of a policy. For the definition of the expected utility function and how it is estimated, as well as details such as the definition of a policy rule, see the section Policy Evaluation and Comparison. In this example, because a positive treatment effect is preferred, anyone whose beta left-parenthesis bold x Subscript i Baseline right-parenthesis value is positive should be given the treatment (that is, a discount). This function is sometimes referred to as the individual treatment effect (ITE).

The variable _beta_ in the data table odetails that is obtained during the previous estimation contains the estimated values for this function for each individual. You can construct an optimal policy by offering a discount to individuals whose _beta_ value is positive. For comparison, the following DATA steps create the optimal policy, s2, along with the other policies: s0, offer n o one a discount; and s1, offer everyone a discount.

 data mylib.discountPolicy;
    set mylib.odetails;
    s0 = 0; * treat none;
    s1 = 1; * treat all;
    s2 = _beta_>=0; * treat only individuals whose ITEs are positive;
    keep id t account_age age avg_hours days_visited
         friends_count has_membership is_US songs_purchased
         income revenue s0 s1 s2;
 run;

You can do policy evaluation by specifying the POLICY= option in the INFER statement and do policy comparison by specifying the POLICYCOMPARISON= option in the same statement. The following statements estimate the model and do policy evaluation and comparison:

 /*--- Policy evaluation and comparison ---*/
 proc deepcausal data=mylib.discountPolicy;
    id id;
    psmodel t = account_age age avg_hours days_visited friends_count
                has_membership is_US songs_purchased income /
                dnn=(nodes=(32 32 32 32)
                train=(optimizer=(miniBatchSize=500 regL1=0.0001
                maxEpochs=32000 algorithm=adam) nthreads=20
                seed=12345 recordseed=67890));
    model revenue = account_age age avg_hours days_visited friends_count
                    has_membership is_US songs_purchased income /
                    dnn=(nodes=(32 32 32 32)
                    train=(optimizer=(miniBatchSize=500 regL1=0.001
                    maxEpochs=32000 algorithm=adam) nthreads=20
                    seed=12345 recordseed=67890));
    infer policy=(t s0 s1 s2) policyComparison=(base=(t)
          compare=(s0 s1 s2))
          out=mylib.oest2 outdetails=mylib.odetails2;
 run;

The option POLICY=(t s0 s1 s2) requests that the policy evaluation be done for the initial discount policy (t), the policy of giving no one a discount (s0), the policy of giving everyone a discount (s1), and the personalized discount policy of giving a discount only to those whose _beta_ value is positive (s2). The suboption COMPARE=(s0 s1 s2) in the POLICYCOMPARISON= option indicates that the policies s0, s1, and s2 are to be compared to the base policy, t, which is specified by the suboption BASE=(t), also in the POLICYCOMPARISON= option.

The policy evaluation results are shown in Output 15.2.2.

Output 15.2.2: Policy Evaluation

The DEEPCAUSAL Procedure

Policy Evaluation
PolicyEstimateStandard
Error
95% Confidence LimitsZPr > |Z|
t14.3837490.06844314.24960314.517895210.16<.0001
s014.3096700.07471714.16322714.456112191.52<.0001
s113.5922990.05808113.47846313.706135234.02<.0001
s214.8352830.06919814.69965814.970909214.39<.0001


The optimal policy, s2, indeed has the highest estimate. When the discount is given to customers whose ITE is positive, the average outcome, which in this example is the revenue for each person, is higher than the average outcome that is obtained by the base policy. This is the largest revenue that is generated, and it is generated by the personalized discount policy.

Output 15.2.3 shows the output for the policy comparison. Negative values indicate that the corresponding discount policy is worse than the base policy, and positive values indicate that the corresponding discount policy is better than the base policy. As expected, s2, giving the discount to the right person, produces the best results.

Output 15.2.3: Policy Comparison

Policy Comparison
Comparison
Policy
Base
Policy
EstimateStandard
Error
95% Confidence LimitsZPr > |Z|
s0t-0.0740790.017349-0.108082-0.040077-4.27<.0001
s1t-0.7914500.014168-0.819218-0.763681-55.86<.0001
s2t0.4515350.0154550.4212430.48182629.22<.0001


Last updated: November 24, 2025