CCDM Procedure

Example 7.3 Scenario Analysis with Rich Regression Effects and BY Groups

This example illustrates scenario analysis when frequency and severity models use regression models that contain classification and interaction effects. It also illustrates how you can analyze scenarios for multiple groups of observations in one PROC CCDM step without having to simulate counts externally.

The example in the section Scenario Analysis encodes the discrete-valued, nominal (nonordinal) variables Gender, CarType, and Education as numerical variables that have an implied order. For example, a high school diploma is assigned a smaller number than an advanced degree. This method of forcing an order on otherwise nonordinal (categorical) variables is not natural and might lead to biased estimates. A more accurate approach is to treat such variables as classification variables that enter the statistical analysis or model not through their values but through their levels. For example, when you specify Education as a classification variable, the modeling process creates different parameters for the Education = 'High School' and Education = 'Advanced Degree' levels and estimates a regression coefficient for each. When you specify such variables in the CLASS statement of PROC CNTSELECT and PROC SEVSELECT, those procedures perform the appropriate levelization for you, which is the process of finding and transforming levels into regression parameters. For more information, see the description of the CLASS statement in Chapter 26, SEVSELECT Procedure.

In addition to specifying nominal variables as classification (CLASS) variables, you can include interaction effects in severity and frequency models. For example, you might want to evaluate how the distribution of losses that are incurred by a policyholder who has a college degree and drives an SUV differs from that of a policyholder who has an advanced degree and drives a sedan. You can do this by including an interaction between CarType and Education in your severity model. Similarly, if you want to evaluate how the number of losses that a policyholder incurs per year varies by the number of annual miles for different types of cars, you can include an interaction between CarType and AnnualMiles in your frequency model. Analyzing such a rich set of regression effects can help you make more accurate predictions about the frequency and severity distributions of losses. PROC CCDM is designed to use such rich models to simulate a more accurate distribution of the aggregate loss.

To start the scenario analysis process, let the following programming statements simulate the frequency and severity events that are affected by the regression effects:

/* Define a macro to simulate regressor data */
%macro generateRegvars;
   age = MAX(int(rand('NORMAL', 35, 15)),16)/50;

   if (rand('UNIFORM') < 0.5) then do;
      genderNum = 1; * female;
      gender = 'F';
   end;
   else do;
      genderNum = 2; * male;
      gender = 'M';
   end;

   if (rand('UNIFORM') < 0.7) then do;
      carTypeNum = 1; * sedan;
      carType = 'Sedan';
   end;
   else do;
      carTypeNum = 2; * SUV;
      carType = 'SUV';
   end;

   educationLevel = rand('UNIFORM');
   if (educationLevel < 0.5) then do;
      educationNum = 1; *high school graduate;
      education = 'High School';
   end;
   else if (educationLevel < 0.85) then do;
      educationNum = 2; *college graduate;
      education = 'College';
   end;
   else do;
      educationNum = 3; *advanced degree;
      education = 'Advanced Degree';
   end;

   annualmiles = MAX(1000, int(rand('NORMAL', 12000, 5000)))/5000;

   carSafety = rand('UNIFORM'); /* scaled to be between 0 & 1 */

   income = MAX(15000,int(rand('NORMAL', educationNum*30000, 50000)))/100000;
%mend;

/* Simulate loss severity data */
%let npolicy=1000;

data losses(keep=region policyholderId age gender carType
                 annualmiles education carSafety
                 income lossAmount noloss);
   call streaminit(12345);
   array cx{6} age genderF milesSUV milesSedan educationAdv educationColl;
   array cbeta{7} _TEMPORARY_ (1 0.75 -1 -1.2 -0.6 0.5 0.75);

   array czx{4} age carTypeSUV educationAdv educationColl;
   array czbeta{5} _TEMPORARY_ (-0.75 -1 -1.2 0.6 0.7);

   array sx{8} _temporary_;
   array sbeta{9} _TEMPORARY_ (5 0.5 1.15 -0.75 -0.3 0.4 0.7 -0.5 -0.3);

   length gender $1 carType $8 education $16;

   alpha = 0.9; theta = 1/alpha;

   do iregion=1 to 2;
      if (iregion = 1) then do;
         region = 'East';
         sigma = 0.5;
      end;
      else do;
         region = 'West';
         /* slightly change parameter estimates for this group */
         do i=1 to dim(cbeta);
           cbeta(i) = cbeta(i) + 0.1;
         end;
         do i=1 to dim(czbeta);
           czbeta(i) = czbeta(i) + 0.1;
         end;
         do i=1 to dim(sbeta);
           sbeta(i) = sbeta(i) + 0.05;
         end;
         sigma = 0.3;
         alpha = 1.6; theta = 1/alpha;
      end;

      do policyholderId=1 to &npolicy;
         /* simulate policyholder and vehicle attributes */
         %generateRegvars;

         do i=1 to dim(sx);
            sx(i) = 0;
         end;
         if (carTypeNum = 2) then do;
            sx(1) = 1;
            carTypeSUV = 1;
            milesSUV = annualmiles;
            milesSedan = 0;
         end;
         else do;
            carTypeSUV = 0;
            milesSUV = 0;
            milesSedan = annualmiles;
         end;

         if (genderNum = 1) then do;
            sx(2) = 1;
            genderF = 1;
         end;
         else genderF = 0;

         sx(3) = carSafety;
         sx(4) = income;
         if (educationNum = 3 and carTypeNum = 2) then sx(5) = 1;
         if (educationNum = 2 and carTypeNum = 2) then sx(6) = 1;
         if (educationNum = 3 and carTypeNum = 1) then sx(7) = 1;
         if (educationNum = 2 and carTypeNum = 1) then sx(8) = 1;

         if (educationNum = 3) then do;
            educationAdv = 1;
            educationColl = 0;
         end;
         else do;
            educationAdv = 0;
            if (educationNum = 2) then educationColl = 1;
            else educationColl = 0;
         end;

         /* simulate number of losses incurred by this policyholder */
         czxbeta = czbeta(1);
         do i=1 to dim(czx);
            czxbeta = czxbeta + czx(i) * czbeta(i+1);
         end;
         phi = exp(czxbeta)/(1+exp(czxbeta));

         if (rand('UNIFORM') < phi) then do;
            * zero process is selected with probability phi;
            numloss = 0;
         end;
         else do;
            * negbin process is selected with probability (1-phi);
            cxbeta = cbeta(1);
            do i=1 to dim(cx);
              cxbeta = cxbeta + cx(i) * cbeta(i+1);
            end;
            Mu = exp(cxbeta);
            numloss = rand('POISSON',Mu);
         end;

         /* simulate severity of each loss */
         if (numloss > 0) then do;
            noloss = 0;
            Mu = sbeta(1);
            do i=1 to dim(sx);
               Mu = Mu + sx(i) * sbeta(i+1);
            end;
            do iloss=1 to numloss;
               lossAmount = exp(Mu) * rand('LOGNORMAL')**Sigma;
               output;
            end;
         end;
         else do;
            noloss = 1;
            lossAmount = .;
            output;
         end;
      end;
   end;
run;

/* Load the loss severity data */
data mylib.losses;
  set losses;
run;

/* Simulate loss counts data */
proc sort data=losses out=lossesByPolicy;
   by region policyholderId;
run;

data losscounts(drop=policyholderId carSafety lossAmount noloss);
   set lossesByPolicy;
   by region policyholderId;
   retain numloss zeroloss 0;
   if (noloss eq 1) then
      zeroloss = zeroloss + 1;
   else
      numloss = numloss + 1;
   if (last.policyholderId) then do;
      output;
      numloss = 0;
   end;
run;

/* Load the loss counts data */
data mylib.losscounts;
  set losscounts;
run;

The following programming statements use the data tables that the preceding statements create to fit the severity and count models that contain a certain set of regression effects:

proc sevselect data=mylib.losses(where=(lossAmount ne .))
        covout outstore=mylib.sevstore print=all;
    by region;
    loss lossAmount;
    class carType gender education;
    scalemodel carType gender carSafety income education*carType
               income*gender carSafety*income;
    selection method=backward(select=sbc) details=summary;
    dist logn burr;
run;

proc cntselect data=mylib.losscounts store=mylib.cstore;
   by region;
   class gender carType education;
   model numloss = age income gender carType*annualmiles
                   education / dist=poisson;
   zeromodel numloss ~ age income carType education;
   selection method=backward(select=sbc) details=summary;
run;

Note the following points about these statements:

  • The data tables are organized in groups of observations that represent data from two regions, East and West. You can analyze both groups at the same time by specifying the BY statement with Region as the BY variable. The observations in the input data table do not need to be sorted by the BY variables.

  • Both severity and count models use three CLASS variables. The severity model includes three interaction effects (Education*CarType, Income*Gender, and CarSafety*Income) and four main effects.

  • The count model is a zero-inflated model, which is essentially a mixture of two models: a model to estimate the occurrence of zero loss events and a model to estimate nonzero counts. The model for zero counts, which is specified in the ZEROMODEL statement, is a regression model that has four main effects and uses the default logistic link function. The model for nonzero counts, which is specified in the MODEL statement, is a Poisson model that has one interaction effect (CarType*AnnualMiles) and four main effects.

  • Both the PROC SEVSELECT and PROC CNTSELECT steps use the SELECTION statement to perform automatic effect selection. In particular, they use the backward selection method to choose a subset of regression effects that maximize the Schwarz Bayesian information criterion (SBC). PROC SEVSELECT conducts the effect selection process independently for each distribution.

  • The PROC SEVSELECT step uses the OUTSTORE= option to store the parameter estimates in an item store. When your scale regression model contains classification or interaction effects, you must store the parameter estimates in an item store instead of storing them in an OUTEST= data table because PROC CCDM cannot obtain the necessary information about classification or interaction effects from an OUTEST= data table.

For severity models, you can inspect the "All Fit Statistics" table for each BY group to decide which severity distribution has the best fit for the data in that BY group. When you specify the SELECTION statement, each row in this table shows the fit statistics of the scale regression model that contains only the final selected set of regression effects for the candidate distribution of that row. The tables in Output 7.3.1 indicate that the model that uses the lognormal distribution is the best according to the majority of fit statistics for both regions, so you can just specify the LOGN distribution in the SEVERITYMODEL statement of the subsequent PROC CCDM step. In cases where different distributions are deemed best for different BY groups, you have two alternatives. The first alternative is to run one PROC CCDM step with the SEVERITYMODEL statement that specifies the best distributions from all BY groups and to later process the results of only the best severity distribution in each BY group. The second alternative is to run separate PROC CCDM steps for each BY group by specifying the best severity distribution for that BY group in the SEVERITYMODEL statement. The latter alternative requires more coding effort but it is more efficient than the first alternative, which runs the analysis for all best severity distributions for all BY groups.

Also note that in some cases, you might see that the likelihood-based fit statistics (–2 log likelihood, AIC, AICC, and BIC) mark one distribution as the best distribution and the EDF-based statistics (KS, AD, and CvM) mark another distribution as the best distribution. In such cases, it is recommended that you conduct aggregate loss simulation by using both severity distributions and compare the summary statistics and percentiles of resulting aggregate loss distributions.

Output 7.3.1: Comparison of Final Models of Severity Distributions

The SEVSELECT Procedure
 
Burr Distribution
Selected Model

All Fit Statistics
Distribution-2 Log
Likelihood
AICAICCSBCKSADCvM
Logn8715*8735*8735*8782*2.68389*37.941172.64223*
Burr87218743874387942.7608133.56827*2.92479
Asterisk (*) denotes the best model in the column.

All Fit Statistics
Distribution-2 Log
Likelihood
AICAICCSBCKSADCvM
Logn12989*13009*13009*13060*6.34841305.1955515.48743*
Burr130041302613026130826.31408*233.46457*15.52778
Asterisk (*) denotes the best model in the column.


If you want to examine the summary of the effect selection process and the final selected set of regression effects for a severity distribution, you can examine the "Selection Summary" and "Selected Effects" tables as shown in Output 7.3.2. These tables indicate that the best scale regression model for the lognormal distribution in the region=East group is obtained after removing the carSafety*income and income*gender effects.

Output 7.3.2: Summary of the Effect Selection Process for LOGN Severity Model for Region=East

The SEVSELECT Procedure
 
Logn Distribution
Selection Details

Selection Summary
StepEffect
Removed
Number
Effects In
SBC
0 88793.6358
1carSafety*income78787.0918
2income*gender68781.8787*
3income58807.5961
4carType*education48921.5365
* Optimal Value Of Criterion

Selected Effects:Intercept carType gender carSafety income carType*education


You can examine a similar summary of the selection process that the CNTSELECT procedure follows to select the best effect subset for the count distribution, as shown in Output 7.3.3. The best model for the zero-inflated Poisson (ZIP) distribution in the region=East group is obtained after removing the education, age, and income effects from the model for the zero counts and removing the income effect from the model for nonzero counts. So the final model contains only the carType effect in the zero-model and age, income, gender, carType*annualmiles, and education effects in the nonzero-model.

Output 7.3.3: Summary of the Effect Selection Process for ZIP Count Model for Region=East

The CNTSELECT Procedure

Selection Summary
StepEffect
Removed
Number
Effects In
SBC
0 112099.9659
1Inf_education102092.2022
2Inf_age92085.4951
3Inf_income82079.0596
4income72074.1768*
5Inf_carType62074.7483
6age52101.6437
7education42162.7788
* Optimal Value Of Criterion


PROC SEVSELECT conducts a separate effect selection analysis for each severity distribution within each BY group, so scale regression models for different severity distributions in the same BY group can contain different final sets of selected effects. Similarly, scale regression models for the same severity distribution in different BY groups can contain different final sets of selected effects. PROC CNTSELECT also conducts a separate effect selection analysis for each BY group, so the final sets of selected effects in the count regression model can be different for different BY groups.

After the CNTSELECT and SEVSELECT procedures have selected the best frequency and severity models, it is time to estimate the distribution of the aggregate loss by using the CCDM procedure. Given that the models depend on regressors, first you need to create the scenario data table, which must contain the final set of regressors that are used in both the severity model and the frequency model. Even if your models contain classification or interaction effects, your scenario data table needs to contain only the columns for individual variables of the effects. PROC CCDM internally performs levelization of each observation, which is the process of expanding the variable values to match them with the parameters of each effect. For example, you do not need to create multiple columns for the parameters of the carType*education effect. You just need to create two columns for carType and education. PROC CCDM internally creates the six columns that correspond to the six parameters that are associated with the carType*education effect. However, you must make sure that the values of each classification variable, such as carType, in the scenario data table belong to the same set of values that appear in the input data tables of PROC SEVSELECT and PROC CNTSELECT.

A typical scenario for an insurance application usually consists of a large number of policyholders. However, for illustration purposes, this example uses a small scenario of only a few policyholders per region that is generated by the following DATA step and shown in Output 7.3.4:

/* Create scenario for internal count simulation */
data scenario(keep=region policyholderId age gender
                   carType annualmiles education carSafety income);
   call streaminit(456);

   length region $4 gender $1 carType $8 education $16;

   do iregion=1 to 2;
      if (iregion = 1) then do;
         region = 'East';
         npolicy = 3;
      end;
      else do;
         region = 'West';
         npolicy = 5;
      end;
      do policyholderId = 1 to npolicy;
         %generateRegvars;
         output;
      end;
    end;
run;

/* Load the scenario data */
data mylib.scenario;
  set scenario;
run;

Output 7.3.4: Data Table mylib.Scenario for BY-Group Processing

ObsregiongendercarTypeeducationageannualmilescarSafetyincome
1EastFSUVHigh School1.162.15400.292880.26090
2EastFSedanHigh School0.862.39780.698440.15000
3EastFSedanAdvanced Degree0.781.99260.594210.58808
4WestMSedanCollege0.821.85500.668490.15000
5WestMSUVCollege0.403.62400.231941.25274
6WestMSedanHigh School0.623.61620.864770.42597
7WestFSedanCollege0.323.45980.662940.36132
8WestMSedanAdvanced Degree0.903.25800.371720.15000


The following PROC CCDM step simulates the aggregate losses for the scenario in the data table mylib.Scenario by each region:

proc ccdm data=mylib.scenario nreplicates=10000 seed=123 print=all
          severitystore=mylib.sevstore countstore=mylib.cstore
          nperturb=30;
   by region;
   severitymodel logn;
   outsum out=mylib.agglossStats mean stddev skewness kurtosis
          pctlpts=(90 97.5 99.5);
run;

The BY statement must contain the same set of variables and in the same order as the variables in the BY statements that you specify in PROC SEVSELECT and PROC CNTSELECT. The SEVERITYSTORE= and COUNTSTORE= options specify the item stores that contain the effect information and parameter estimates of the severity and count models, respectively, for both BY groups. The count item store that PROC CNTSELECT creates always contains the covariance estimates of the count model parameters, and the COVOUT option in the PROC SEVSELECT step ensures that the severity item store includes the covariance estimates. Availability of these covariance estimates makes the perturbation analysis that the NPERTURB= option requests more accurate.

One of the main points that this example illustrates is that irrespective of the complexity of the regression effects in your frequency or severity models, you do not have to do any extra work to convey the effect information to PROC CCDM, except to specify the item stores. The frequency and severity item stores contain all the necessary information that PROC CCDM needs for computing the values of the regressor-dependent parameters of the underlying frequency and severity distributions before making random draws from those distributions. When you use the BY and SELECTION statements in PROC CNTSELECT and PROC SEVSELECT, the respective item stores contain the information about the final selected regression effects for each distribution in each BY group. PROC CCDM uses this information to process only the regression parameters that correspond to the selected regression effects for each distribution in each BY group.

Output 7.3.5 and Output 7.3.6 show the summary of the perturbation analysis for the two regions.

Output 7.3.5: Aggregate Loss Simulation Results for Region=East

The CCDM Procedure
 
Severity Model: Logn
Count Model: ZIP

Sample Percentile Perturbation Analysis
PercentileEstimateStandard
Error
100
500
251.320687.11212
50578.0567467.55693
752330.8232.67455
903688.4331.99961
954495.2384.34629
97.55215.7442.70674
996096.0486.57812
99.56708.0543.85538
Number of Perturbed Samples = 30
Size of Each Sample = 10000


Output 7.3.6: Aggregate Loss Simulation Results for Region=West

Sample Percentile Perturbation Analysis
PercentileEstimateStandard
Error
122.5477133.67374
5172.0686663.67050
25478.1832468.06828
50715.1593369.06287
75982.3912472.51530
901254.177.78615
951434.982.18084
97.51597.685.80075
991793.888.37708
99.51945.092.41512
Number of Perturbed Samples = 30
Size of Each Sample = 10000


Last updated: July 09, 2026