CCDM Procedure
Example 7.3 Scenario Analysis with Rich Regression Effects and BY Groups
This example illustrates scenario analysis when frequency and severity models use regression models that contain classification and interaction effects. It also illustrates how you can analyze scenarios for multiple groups of observations in one PROC CCDM step without having to simulate counts externally.
The example in the section Scenario Analysis encodes the discrete-valued, nominal (nonordinal) variables Gender, CarType, and Education as numerical variables that have an implied order. For example, a high school diploma is assigned a smaller number than an advanced degree. This method of forcing an order on otherwise nonordinal (categorical) variables is not natural and might lead to biased estimates. A more accurate approach is to treat such variables as classification variables that enter the statistical analysis or model not through their values but through their levels. For example, when you specify Education as a classification variable, the modeling process creates different parameters for the Education = 'High School' and Education = 'Advanced Degree' levels and estimates a regression coefficient for each. When you specify such variables in the CLASS statement of PROC CNTSELECT and PROC SEVSELECT, those procedures perform the appropriate levelization for you, which is the process of finding and transforming levels into regression parameters. For more information, see the description of the CLASS statement in Chapter 26, SEVSELECT Procedure.
In addition to specifying nominal variables as classification (CLASS) variables, you can include interaction effects in severity and frequency models. For example, you might want to evaluate how the distribution of losses that are incurred by a policyholder who has a college degree and drives an SUV differs from that of a policyholder who has an advanced degree and drives a sedan. You can do this by including an interaction between CarType and Education in your severity model. Similarly, if you want to evaluate how the number of losses that a policyholder incurs per year varies by the number of annual miles for different types of cars, you can include an interaction between CarType and AnnualMiles in your frequency model. Analyzing such a rich set of regression effects can help you make more accurate predictions about the frequency and severity distributions of losses. PROC CCDM is designed to use such rich models to simulate a more accurate distribution of the aggregate loss.
To start the scenario analysis process, let the following programming statements simulate the frequency and severity events that are affected by the regression effects:
/* Define a macro to simulate regressor data */
%macro generateRegvars;
age = MAX(int(rand('NORMAL', 35, 15)),16)/50;
if (rand('UNIFORM') < 0.5) then do;
genderNum = 1; * female;
gender = 'F';
end;
else do;
genderNum = 2; * male;
gender = 'M';
end;
if (rand('UNIFORM') < 0.7) then do;
carTypeNum = 1; * sedan;
carType = 'Sedan';
end;
else do;
carTypeNum = 2; * SUV;
carType = 'SUV';
end;
educationLevel = rand('UNIFORM');
if (educationLevel < 0.5) then do;
educationNum = 1; *high school graduate;
education = 'High School';
end;
else if (educationLevel < 0.85) then do;
educationNum = 2; *college graduate;
education = 'College';
end;
else do;
educationNum = 3; *advanced degree;
education = 'Advanced Degree';
end;
annualmiles = MAX(1000, int(rand('NORMAL', 12000, 5000)))/5000;
carSafety = rand('UNIFORM'); /* scaled to be between 0 & 1 */
income = MAX(15000,int(rand('NORMAL', educationNum*30000, 50000)))/100000;
%mend;
/* Simulate loss severity data */
%let npolicy=1000;
data losses(keep=region policyholderId age gender carType
annualmiles education carSafety
income lossAmount noloss);
call streaminit(12345);
array cx{6} age genderF milesSUV milesSedan educationAdv educationColl;
array cbeta{7} _TEMPORARY_ (1 0.75 -1 -1.2 -0.6 0.5 0.75);
array czx{4} age carTypeSUV educationAdv educationColl;
array czbeta{5} _TEMPORARY_ (-0.75 -1 -1.2 0.6 0.7);
array sx{8} _temporary_;
array sbeta{9} _TEMPORARY_ (5 0.5 1.15 -0.75 -0.3 0.4 0.7 -0.5 -0.3);
length gender $1 carType $8 education $16;
alpha = 0.9; theta = 1/alpha;
do iregion=1 to 2;
if (iregion = 1) then do;
region = 'East';
sigma = 0.5;
end;
else do;
region = 'West';
/* slightly change parameter estimates for this group */
do i=1 to dim(cbeta);
cbeta(i) = cbeta(i) + 0.1;
end;
do i=1 to dim(czbeta);
czbeta(i) = czbeta(i) + 0.1;
end;
do i=1 to dim(sbeta);
sbeta(i) = sbeta(i) + 0.05;
end;
sigma = 0.3;
alpha = 1.6; theta = 1/alpha;
end;
do policyholderId=1 to &npolicy;
/* simulate policyholder and vehicle attributes */
%generateRegvars;
do i=1 to dim(sx);
sx(i) = 0;
end;
if (carTypeNum = 2) then do;
sx(1) = 1;
carTypeSUV = 1;
milesSUV = annualmiles;
milesSedan = 0;
end;
else do;
carTypeSUV = 0;
milesSUV = 0;
milesSedan = annualmiles;
end;
if (genderNum = 1) then do;
sx(2) = 1;
genderF = 1;
end;
else genderF = 0;
sx(3) = carSafety;
sx(4) = income;
if (educationNum = 3 and carTypeNum = 2) then sx(5) = 1;
if (educationNum = 2 and carTypeNum = 2) then sx(6) = 1;
if (educationNum = 3 and carTypeNum = 1) then sx(7) = 1;
if (educationNum = 2 and carTypeNum = 1) then sx(8) = 1;
if (educationNum = 3) then do;
educationAdv = 1;
educationColl = 0;
end;
else do;
educationAdv = 0;
if (educationNum = 2) then educationColl = 1;
else educationColl = 0;
end;
/* simulate number of losses incurred by this policyholder */
czxbeta = czbeta(1);
do i=1 to dim(czx);
czxbeta = czxbeta + czx(i) * czbeta(i+1);
end;
phi = exp(czxbeta)/(1+exp(czxbeta));
if (rand('UNIFORM') < phi) then do;
* zero process is selected with probability phi;
numloss = 0;
end;
else do;
* negbin process is selected with probability (1-phi);
cxbeta = cbeta(1);
do i=1 to dim(cx);
cxbeta = cxbeta + cx(i) * cbeta(i+1);
end;
Mu = exp(cxbeta);
numloss = rand('POISSON',Mu);
end;
/* simulate severity of each loss */
if (numloss > 0) then do;
noloss = 0;
Mu = sbeta(1);
do i=1 to dim(sx);
Mu = Mu + sx(i) * sbeta(i+1);
end;
do iloss=1 to numloss;
lossAmount = exp(Mu) * rand('LOGNORMAL')**Sigma;
output;
end;
end;
else do;
noloss = 1;
lossAmount = .;
output;
end;
end;
end;
run;
/* Load the loss severity data */
data mylib.losses;
set losses;
run;
/* Simulate loss counts data */
proc sort data=losses out=lossesByPolicy;
by region policyholderId;
run;
data losscounts(drop=policyholderId carSafety lossAmount noloss);
set lossesByPolicy;
by region policyholderId;
retain numloss zeroloss 0;
if (noloss eq 1) then
zeroloss = zeroloss + 1;
else
numloss = numloss + 1;
if (last.policyholderId) then do;
output;
numloss = 0;
end;
run;
/* Load the loss counts data */
data mylib.losscounts;
set losscounts;
run;
The following programming statements use the data tables that the preceding statements create to fit the severity and count models that contain a certain set of regression effects:
proc sevselect data=mylib.losses(where=(lossAmount ne .))
covout outstore=mylib.sevstore print=all;
by region;
loss lossAmount;
class carType gender education;
scalemodel carType gender carSafety income education*carType
income*gender carSafety*income;
selection method=backward(select=sbc) details=summary;
dist logn burr;
run;
proc cntselect data=mylib.losscounts store=mylib.cstore;
by region;
class gender carType education;
model numloss = age income gender carType*annualmiles
education / dist=poisson;
zeromodel numloss ~ age income carType education;
selection method=backward(select=sbc) details=summary;
run;
Note the following points about these statements:
The data tables are organized in groups of observations that represent data from two regions, East and West. You can analyze both groups at the same time by specifying the BY statement with
Regionas the BY variable. The observations in the input data table do not need to be sorted by the BY variables.Both severity and count models use three CLASS variables. The severity model includes three interaction effects (
Education*CarType,Income*Gender, andCarSafety*Income) and four main effects.The count model is a zero-inflated model, which is essentially a mixture of two models: a model to estimate the occurrence of zero loss events and a model to estimate nonzero counts. The model for zero counts, which is specified in the ZEROMODEL statement, is a regression model that has four main effects and uses the default logistic link function. The model for nonzero counts, which is specified in the MODEL statement, is a Poisson model that has one interaction effect (
CarType*AnnualMiles) and four main effects.Both the PROC SEVSELECT and PROC CNTSELECT steps use the SELECTION statement to perform automatic effect selection. In particular, they use the backward selection method to choose a subset of regression effects that maximize the Schwarz Bayesian information criterion (SBC). PROC SEVSELECT conducts the effect selection process independently for each distribution.
The PROC SEVSELECT step uses the OUTSTORE= option to store the parameter estimates in an item store. When your scale regression model contains classification or interaction effects, you must store the parameter estimates in an item store instead of storing them in an OUTEST= data table because PROC CCDM cannot obtain the necessary information about classification or interaction effects from an OUTEST= data table.
For severity models, you can inspect the "All Fit Statistics" table for each BY group to decide which severity distribution has the best fit for the data in that BY group. When you specify the SELECTION statement, each row in this table shows the fit statistics of the scale regression model that contains only the final selected set of regression effects for the candidate distribution of that row. The tables in Output 7.3.1 indicate that the model that uses the lognormal distribution is the best according to the majority of fit statistics for both regions, so you can just specify the LOGN distribution in the SEVERITYMODEL statement of the subsequent PROC CCDM step. In cases where different distributions are deemed best for different BY groups, you have two alternatives. The first alternative is to run one PROC CCDM step with the SEVERITYMODEL statement that specifies the best distributions from all BY groups and to later process the results of only the best severity distribution in each BY group. The second alternative is to run separate PROC CCDM steps for each BY group by specifying the best severity distribution for that BY group in the SEVERITYMODEL statement. The latter alternative requires more coding effort but it is more efficient than the first alternative, which runs the analysis for all best severity distributions for all BY groups.
Also note that in some cases, you might see that the likelihood-based fit statistics (–2 log likelihood, AIC, AICC, and BIC) mark one distribution as the best distribution and the EDF-based statistics (KS, AD, and CvM) mark another distribution as the best distribution. In such cases, it is recommended that you conduct aggregate loss simulation by using both severity distributions and compare the summary statistics and percentiles of resulting aggregate loss distributions.
Output 7.3.1: Comparison of Final Models of Severity Distributions
| All Fit Statistics | |||||||
|---|---|---|---|---|---|---|---|
| Distribution | -2 Log Likelihood | AIC | AICC | SBC | KS | AD | CvM |
| Logn | 8715* | 8735* | 8735* | 8782* | 2.68389* | 37.94117 | 2.64223* |
| Burr | 8721 | 8743 | 8743 | 8794 | 2.76081 | 33.56827* | 2.92479 |
| Asterisk (*) denotes the best model in the column. | |||||||
| All Fit Statistics | |||||||
|---|---|---|---|---|---|---|---|
| Distribution | -2 Log Likelihood | AIC | AICC | SBC | KS | AD | CvM |
| Logn | 12989* | 13009* | 13009* | 13060* | 6.34841 | 305.19555 | 15.48743* |
| Burr | 13004 | 13026 | 13026 | 13082 | 6.31408* | 233.46457* | 15.52778 |
| Asterisk (*) denotes the best model in the column. | |||||||
If you want to examine the summary of the effect selection process and the final selected set of regression effects for a severity distribution, you can examine the "Selection Summary" and "Selected Effects" tables as shown in Output 7.3.2. These tables indicate that the best scale regression model for the lognormal distribution in the region=East group is obtained after removing the carSafety*income and income*gender effects.
Output 7.3.2: Summary of the Effect Selection Process for LOGN Severity Model for Region=East
| Selection Summary | |||
|---|---|---|---|
| Step | Effect Removed | Number Effects In | SBC |
| 0 | 8 | 8793.6358 | |
| 1 | carSafety*income | 7 | 8787.0918 |
| 2 | income*gender | 6 | 8781.8787* |
| 3 | income | 5 | 8807.5961 |
| 4 | carType*education | 4 | 8921.5365 |
| * Optimal Value Of Criterion | |||
| Selected Effects: | Intercept carType gender carSafety income carType*education |
|---|
You can examine a similar summary of the selection process that the CNTSELECT procedure follows to select the best effect subset for the count distribution, as shown in Output 7.3.3. The best model for the zero-inflated Poisson (ZIP) distribution in the region=East group is obtained after removing the education, age, and income effects from the model for the zero counts and removing the income effect from the model for nonzero counts. So the final model contains only the carType effect in the zero-model and age, income, gender, carType*annualmiles, and education effects in the nonzero-model.
Output 7.3.3: Summary of the Effect Selection Process for ZIP Count Model for Region=East
| Selection Summary | |||
|---|---|---|---|
| Step | Effect Removed | Number Effects In | SBC |
| 0 | 11 | 2099.9659 | |
| 1 | Inf_education | 10 | 2092.2022 |
| 2 | Inf_age | 9 | 2085.4951 |
| 3 | Inf_income | 8 | 2079.0596 |
| 4 | income | 7 | 2074.1768* |
| 5 | Inf_carType | 6 | 2074.7483 |
| 6 | age | 5 | 2101.6437 |
| 7 | education | 4 | 2162.7788 |
| * Optimal Value Of Criterion | |||
PROC SEVSELECT conducts a separate effect selection analysis for each severity distribution within each BY group, so scale regression models for different severity distributions in the same BY group can contain different final sets of selected effects. Similarly, scale regression models for the same severity distribution in different BY groups can contain different final sets of selected effects. PROC CNTSELECT also conducts a separate effect selection analysis for each BY group, so the final sets of selected effects in the count regression model can be different for different BY groups.
After the CNTSELECT and SEVSELECT procedures have selected the best frequency and severity models, it is time to estimate the distribution of the aggregate loss by using the CCDM procedure. Given that the models depend on regressors, first you need to create the scenario data table, which must contain the final set of regressors that are used in both the severity model and the frequency model. Even if your models contain classification or interaction effects, your scenario data table needs to contain only the columns for individual variables of the effects. PROC CCDM internally performs levelization of each observation, which is the process of expanding the variable values to match them with the parameters of each effect. For example, you do not need to create multiple columns for the parameters of the carType*education effect. You just need to create two columns for carType and education. PROC CCDM internally creates the six columns that correspond to the six parameters that are associated with the carType*education effect. However, you must make sure that the values of each classification variable, such as carType, in the scenario data table belong to the same set of values that appear in the input data tables of PROC SEVSELECT and PROC CNTSELECT.
A typical scenario for an insurance application usually consists of a large number of policyholders. However, for illustration purposes, this example uses a small scenario of only a few policyholders per region that is generated by the following DATA step and shown in Output 7.3.4:
/* Create scenario for internal count simulation */
data scenario(keep=region policyholderId age gender
carType annualmiles education carSafety income);
call streaminit(456);
length region $4 gender $1 carType $8 education $16;
do iregion=1 to 2;
if (iregion = 1) then do;
region = 'East';
npolicy = 3;
end;
else do;
region = 'West';
npolicy = 5;
end;
do policyholderId = 1 to npolicy;
%generateRegvars;
output;
end;
end;
run;
/* Load the scenario data */
data mylib.scenario;
set scenario;
run;
Output 7.3.4: Data Table mylib.Scenario for BY-Group Processing
| Obs | region | gender | carType | education | age | annualmiles | carSafety | income |
|---|---|---|---|---|---|---|---|---|
| 1 | East | F | SUV | High School | 1.16 | 2.1540 | 0.29288 | 0.26090 |
| 2 | East | F | Sedan | High School | 0.86 | 2.3978 | 0.69844 | 0.15000 |
| 3 | East | F | Sedan | Advanced Degree | 0.78 | 1.9926 | 0.59421 | 0.58808 |
| 4 | West | M | Sedan | College | 0.82 | 1.8550 | 0.66849 | 0.15000 |
| 5 | West | M | SUV | College | 0.40 | 3.6240 | 0.23194 | 1.25274 |
| 6 | West | M | Sedan | High School | 0.62 | 3.6162 | 0.86477 | 0.42597 |
| 7 | West | F | Sedan | College | 0.32 | 3.4598 | 0.66294 | 0.36132 |
| 8 | West | M | Sedan | Advanced Degree | 0.90 | 3.2580 | 0.37172 | 0.15000 |
The following PROC CCDM step simulates the aggregate losses for the scenario in the data table mylib.Scenario by each region:
proc ccdm data=mylib.scenario nreplicates=10000 seed=123 print=all
severitystore=mylib.sevstore countstore=mylib.cstore
nperturb=30;
by region;
severitymodel logn;
outsum out=mylib.agglossStats mean stddev skewness kurtosis
pctlpts=(90 97.5 99.5);
run;
The BY statement must contain the same set of variables and in the same order as the variables in the BY statements that you specify in PROC SEVSELECT and PROC CNTSELECT. The SEVERITYSTORE= and COUNTSTORE= options specify the item stores that contain the effect information and parameter estimates of the severity and count models, respectively, for both BY groups. The count item store that PROC CNTSELECT creates always contains the covariance estimates of the count model parameters, and the COVOUT option in the PROC SEVSELECT step ensures that the severity item store includes the covariance estimates. Availability of these covariance estimates makes the perturbation analysis that the NPERTURB= option requests more accurate.
One of the main points that this example illustrates is that irrespective of the complexity of the regression effects in your frequency or severity models, you do not have to do any extra work to convey the effect information to PROC CCDM, except to specify the item stores. The frequency and severity item stores contain all the necessary information that PROC CCDM needs for computing the values of the regressor-dependent parameters of the underlying frequency and severity distributions before making random draws from those distributions. When you use the BY and SELECTION statements in PROC CNTSELECT and PROC SEVSELECT, the respective item stores contain the information about the final selected regression effects for each distribution in each BY group. PROC CCDM uses this information to process only the regression parameters that correspond to the selected regression effects for each distribution in each BY group.
Output 7.3.5 and Output 7.3.6 show the summary of the perturbation analysis for the two regions.
Output 7.3.5: Aggregate Loss Simulation Results for Region=East
| Sample Percentile Perturbation Analysis | ||
|---|---|---|
| Percentile | Estimate | Standard Error |
| 1 | 0 | 0 |
| 5 | 0 | 0 |
| 25 | 1.32068 | 7.11212 |
| 50 | 578.05674 | 67.55693 |
| 75 | 2330.8 | 232.67455 |
| 90 | 3688.4 | 331.99961 |
| 95 | 4495.2 | 384.34629 |
| 97.5 | 5215.7 | 442.70674 |
| 99 | 6096.0 | 486.57812 |
| 99.5 | 6708.0 | 543.85538 |
| Number of Perturbed Samples = 30 | ||
| Size of Each Sample = 10000 | ||
Output 7.3.6: Aggregate Loss Simulation Results for Region=West
| Sample Percentile Perturbation Analysis | ||
|---|---|---|
| Percentile | Estimate | Standard Error |
| 1 | 22.54771 | 33.67374 |
| 5 | 172.06866 | 63.67050 |
| 25 | 478.18324 | 68.06828 |
| 50 | 715.15933 | 69.06287 |
| 75 | 982.39124 | 72.51530 |
| 90 | 1254.1 | 77.78615 |
| 95 | 1434.9 | 82.18084 |
| 97.5 | 1597.6 | 85.80075 |
| 99 | 1793.8 | 88.37708 |
| 99.5 | 1945.0 | 92.41512 |
| Number of Perturbed Samples = 30 | ||
| Size of Each Sample = 10000 | ||