The HPCDM Procedure

Example 16.3 Scenario Analysis with Rich Regression Effects and BY Groups

(View the complete code for this example.)

This example illustrates scenario analysis when frequency and severity models use regression models that contain classification and interaction effects. It also illustrates how you can analyze scenarios for multiple groups of observations in one PROC HPCDM step without your having to simulate counts externally.

The example in the section Scenario Analysis encodes the discrete-valued, nominal (nonordinal) variables Gender, CarType, and Education as numerical variables with an implied order. For example, a high school diploma is assigned a smaller number than an advanced degree. This method of forcing an order on otherwise nonordinal (categorical) variables is not natural and might lead to biased estimates. A more accurate approach is to treat such variables as classification variables that enter the statistical analysis or model not through their values but through their levels. For example, when you specify Education as a classification variable, the modeling process creates different parameters for the Education = 'High School' and Education = 'Advanced Degree' levels and estimates a regression coefficient for each. When you specify such variables in the CLASS statement of PROC COUNTREG and PROC SEVERITY, those procedures perform the appropriate levelization for you, which is the process of finding and transforming levels into regression parameters. For more information, see the description of the CLASS statement in Chapter 28: The SEVERITY Procedure.

In addition to specifying nominal variables as classification (CLASS) variables, you can include interaction effects in severity and frequency models. For example, you might want to evaluate how the distribution of losses that are incurred by a policyholder with a college degree who drives an SUV differs from that of a policyholder with an advanced degree who drives a sedan. You can do this by including an interaction between CarType and Education in your severity model. Similarly, if you want to evaluate how the number of losses that a policyholder incurs per year varies by the number of annual miles for different types of cars, you can include an interaction between CarType and AnnualMiles in your frequency model. Analyzing such a rich set of regression effects can help you make more accurate predictions about the frequency and severity distributions of losses. PROC HPCDM is designed to use such rich models to simulate a more accurate distribution of the aggregate loss.

As an example of the process, first, let the following programming statements fit the severity and count models that contain a certain set of regression effects:

proc severity data=losses(where=(not(missing(lossAmount))))
        covout outstore=work.sevstore print=all plots=none;
    by region;
    loss lossAmount;
    class carType gender education;
    scalemodel carType gender carSafety income education*carType
               income*gender carSafety*income;
    dist logn burr;
run;

proc countreg data=losscounts covout;
   by region;
   class gender carType education;
   model numloss = age income gender carType*annualmiles education / dist=negbin;
   zeromodel numloss ~ age income carType education;
   store cstore;
run;

Note the following points about these statements:

  • You can find the code that prepares the Work.Losses and Work.LossCounts data sets in the PROC HPCDM sample program hcdmex03.sas. The data sets are organized in groups of observations that represent data from two regions, East and West. You can analyze both groups at once by specifying the BY statement with Region as the BY variable.

  • Both severity and count models use three CLASS variables. The severity model includes three interaction effects (Education*CarType, Income*Gender, and CarSafety*Income) and four main effects. PROC SEVERITY uses the same set of regression effects in the scale regression model of each of the two distributions that you specify in the DIST statement, which are LOGN and BURR in this example.

  • The count model is a mixture of two models: a model to estimate the occurrence of zero loss events and a model to estimate nonzero counts. The zero model is a regression model with four main effects and the default logistic link function. The model for nonzero counts is a negative binomial model with one interaction effect (CarType*AnnualMiles) and four main effects.

The "Parameter Estimates" table of the lognormal severity model in Output 16.3.1 for the Region='East' BY group shows that Income*Gender and CarSafety*Income effects are not statistically significant. The "Parameter Estimates" table in Output 16.3.2 shows that those two effects are not statistically significant for the Burr severity model also.

Output 16.3.1: Parameter Estimates for LOGN Severity Model for Region=East

Parameter Estimates
ParameterDFEstimateStandard
Error
t ValueApprox
Pr > |t|
Mu14.982530.02861174.16<.0001
Sigma10.488940.0053591.41<.0001
carType SUV10.517720.0364814.19<.0001
carType Sedan00...
gender F11.166900.0308237.86<.0001
gender M00...
carSafety1-0.715170.04599-15.55<.0001
income1-0.285280.03652-7.81<.0001
carType*education SUV Advanced Degree10.445990.062457.14<.0001
carType*education SUV College10.678520.0441615.36<.0001
carType*education SUV High School00...
carType*education Sedan Advanced Degree1-0.496800.02689-18.47<.0001
carType*education Sedan College1-0.263100.01849-14.23<.0001
carType*education Sedan High School00...
income*gender F10.009880.040100.250.8054
income*gender M00...
carSafety*income1-0.093900.06166-1.520.1278


Output 16.3.2: Parameter Estimates for BURR Severity Model for Region=East

Parameter Estimates
ParameterDFEstimateStandard
Error
t ValueApprox
Pr > |t|
Theta1145.637095.7437125.36<.0001
Alpha10.997830.0647015.42<.0001
Gamma13.587430.0936238.32<.0001
carType SUV10.516480.0370113.96<.0001
carType Sedan00...
gender F11.166640.0308337.84<.0001
gender M00...
carSafety1-0.716360.04590-15.61<.0001
income1-0.295220.03639-8.11<.0001
carType*education SUV Advanced Degree10.436960.063856.84<.0001
carType*education SUV College10.680490.0450115.12<.0001
carType*education SUV High School00...
carType*education Sedan Advanced Degree1-0.501600.02672-18.77<.0001
carType*education Sedan College1-0.264830.01840-14.39<.0001
carType*education Sedan High School00...
income*gender F10.012680.039860.320.7504
income*gender M00...
carSafety*income1-0.077130.06162-1.250.2107


The "Parameter Estimates" table of the count model in Output 16.3.3 shows that the income and Inf_income parameters are insignificant. This implies that the income effect is not significant for the main and zero inflation parts of the count model.

The results for the Region='West' BY group are not shown here, but you can execute the sample program hcdmex03.sas to verify that the same parameters are statistically insignificant in severity and count models of that BY group as well. However, in general, you might find that some effects are significant for some BY groups but insignificant for other BY groups. In such cases, for more accurate results, it is recommended that you create a separate data set for each set of similar BY groups and invoke the SEVERITY, COUNTREG, and HPCDM procedures on each data set to separately analyze each set of similar BY groups.

Output 16.3.3: Count Model Parameter Estimates for Region=East

Parameter Estimates
ParameterDFEstimateStandard
Error
t ValueApprox
Pr > |t|
Intercept11.1566260.1306418.85<.0001
age10.7347980.1122996.54<.0001
income1-0.0407440.081573-0.500.6174
gender F1-0.9990940.053170-18.79<.0001
gender M00...
annualmiles*carType SUV1-1.2664520.045996-27.53<.0001
annualmiles*carType Sedan1-0.6322810.027818-22.73<.0001
education Advanced Degree10.4186510.0994144.21<.0001
education College10.7094790.06959610.19<.0001
education High School00...
Inf_Intercept1-0.5012400.353072-1.420.1557
Inf_age1-0.9456570.329950-2.870.0042
Inf_income1-0.1735410.233461-0.740.4573
Inf_carType SUV1-0.6934270.369119-1.880.0603
Inf_carType Sedan00...
Inf_education Advanced Degree10.6686130.2918212.290.0220
Inf_education College10.4742120.2324992.040.0414
Inf_education High School00...
_Alpha10.7908390.1035227.64<.0001


The following modified PROC SEVERITY and PROC COUNTREG steps refit the severity and count models, respectively, after removing the insignificant effects:

/* Re-fit models after removing insignificant effects. */
proc severity data=losses(where=(not(missing(lossAmount))))
        covout outstore=work.sevstore print=all plots=none;
    by region;
    loss lossAmount;
    class carType gender education;
    scalemodel carType gender carSafety income education*carType;
    dist logn burr;
run;

proc countreg data=losscounts covout;
   by region;
   class gender carType education;
   model numloss = age gender carType*annualmiles education / dist=negbin;
   zeromodel numloss ~ age carType education;
   store cstore;
run;

Note that the PROC SEVERITY step uses the OUTSTORE= option to store the parameter estimates in an item store. When your scale regression model contains classification or interaction effects, you must store the parameter estimates in an item store instead of storing them in an OUTEST= data set, because PROC HPCDM cannot obtain the necessary information about classification or interaction effects from an OUTEST= data set.

The "Parameter Estimates" tables in Output 16.3.4 and Output 16.3.5 show that all parameters are now statistically significant, most at the 95% confidence level and a few at the 90% confidence level. If you want every parameter to be significant at the 95% confidence level, then you might want to continue the process by removing the carType effect with a p-value of 0.0607 from the ZEROMODEL statement and refitting the count model. However, for the purpose of this example, the preceding models are declared to be satisfactory, and the effect selection process stops here.

You need to follow this process of model inspection and effect selection before you use the severity and count models with the HPCDM procedure. For count models, you can use the automatic effect (variable) selection feature of PROC COUNTREG. For more information, see the description of the SELECT= option in the MODEL statement of Chapter 11: The COUNTREG Procedure. For severity models, you need to perform effect selection manually by inspecting the estimates and refitting the model after removing one or a few insignificant effects at a time until you find the final set of significant effects. Although it is not shown in this example, you can also decide which set of effects is better by comparing the fit statistics of two models; the better model might contain certain effects at lower confidence levels than the usual 95% or 90% confidence levels. In fact, the SELECT=INFO option of PROC COUNTREG uses the AIC or BIC of the entire model to select the set of effects instead of using the p-values of individual parameters. You might also want to use some domain knowledge to retain certain effects in the model even if their confidence level is not very high.

Output 16.3.4: Final LOGN Severity Model Parameter Estimates for Region=East

Parameter Estimates
ParameterDFEstimateStandard
Error
t ValueApprox
Pr > |t|
Mu15.008450.02135234.61<.0001
Sigma10.489080.0053591.43<.0001
carType SUV10.515560.0364214.16<.0001
carType Sedan00...
gender F11.172910.0172667.96<.0001
gender M00...
carSafety1-0.772730.02614-29.56<.0001
income1-0.327020.01962-16.67<.0001
carType*education SUV Advanced Degree10.448700.062237.21<.0001
carType*education SUV College10.683600.0440415.52<.0001
carType*education SUV High School00...
carType*education Sedan Advanced Degree1-0.495720.02688-18.44<.0001
carType*education Sedan College1-0.262340.01848-14.19<.0001
carType*education Sedan High School00...


Output 16.3.5: Final Count Model Parameter Estimates for Region=East

Parameter Estimates
ParameterDFEstimateStandard
Error
t ValueApprox
Pr > |t|
Intercept11.1361750.1247869.10<.0001
age10.7378050.1123396.57<.0001
gender F1-1.0013110.052996-18.89<.0001
gender M00...
annualmiles*carType SUV1-1.2631780.045809-27.57<.0001
annualmiles*carType Sedan1-0.6314190.027728-22.77<.0001
education Advanced Degree10.4003070.0920604.35<.0001
education College10.7034360.06793510.35<.0001
education High School00...
Inf_Intercept1-0.5856620.338796-1.730.0839
Inf_age1-0.9282940.324629-2.860.0042
Inf_carType SUV1-0.6580890.350886-1.880.0607
Inf_carType Sedan00...
Inf_education Advanced Degree10.5885110.2691952.190.0288
Inf_education College10.4466000.2281511.960.0503
Inf_education High School00...
_Alpha10.7850180.1013277.75<.0001


For severity models, you also need to inspect the "All Fit Statistics" table to decide which severity distributions you want to use for aggregate loss modeling. The table in Output 16.3.6 shows that the lognormal distribution is the best according to the majority of fit statistics, so you can choose that. However, in some cases, you might see that the likelihood-based fit statistics (–2 log likelihood, AIC, AICC, BIC) choose one distribution and the EDF-based statistics (KS, AD, CvM) choose another distribution. In such cases, it is recommended that before making your final decision, you conduct aggregate loss simulation by using both severity distributions and compare the summary statistics and percentiles that each severity distribution produces.

Output 16.3.6: Comparison of Severity Distributions for Region=East

All Fit Statistics
Distribution-2 Log
Likelihood
AICAICCBICKSADCvM
Logn45280*45300*45300*45364*10.31771*613.78765 46.37913*
Burr45346 45368 45368 45437 10.90815 519.83495*49.71973 
Note: The asterisk (*) marks the best model according to each column's criterion.


After you have satisfactorily estimated the severity and frequency models, it is time to estimate the distribution of the aggregate loss by using the HPCDM procedure. The scenario data set must contain the final set of regressors that are used in both the severity model and the frequency model. Note that even if your models contain interaction effects, your scenario data set needs to contain only the columns for individual variables of the effects. PROC HPCDM internally performs levelization of each observation, which is the process of expanding the variable values to match them with the parameters of each effect. A typical scenario for an insurance application might consist of a large number of policyholders, but for illustration purposes, this example uses a small scenario of only a few policyholders per region. Output 16.3.7 shows the contents of the Work.Scenario data set, and the following PROC HPCDM step simulates the aggregate losses for that scenario:

proc hpcdm data=scenario nreplicates=10000 seed=123 print=all
           severitystore=work.sevstore countstore=work.cstore
           nperturb=30;
   by region;
   severitymodel logn;
   outsum out=agglossStats mean stddev skewness kurtosis pctlpts=(90 97.5 99.5);
run;

Output 16.3.7: Work.Scenario Data Set for BY-Group Processing

ObsregiongendercarTypeeducationageannualmilescarSafetyincome
1EastFSUVHigh School1.162.15400.292880.26090
2EastFSedanHigh School0.862.39780.698440.15000
3EastFSedanAdvanced Degree0.781.99260.594210.58808
4WestMSedanCollege0.821.85500.668490.15000
5WestMSUVCollege0.403.62400.231941.25274
6WestMSedanHigh School0.623.61620.864770.42597
7WestFSedanCollege0.323.45980.662940.36132
8WestMSedanAdvanced Degree0.903.25800.371720.15000


The SEVERITYSTORE= and COUNTSTORE= options specify the item stores that contain the effect information and parameter estimates of the severity and counts models, respectively, for both BY groups. The COVOUT option in the preceding PROC SEVERITY and PROC COUNTREG steps ensures that the respective item stores include the covariance estimates that are needed for the perturbation analysis that the NPERTURB= option requests.

Output 16.3.8: Aggregate Loss Simulation Results for Region=East

The HPCDM Procedure
Severity Model: Logn
Count Model: ZINB

Sample Percentile Perturbation Analysis
PercentileEstimateStandard
Error
100
500
2500
50151.6205220.57120
75492.0436533.55686
90917.1802951.54978
951233.363.95801
97.51553.578.97273
991981.2111.13102
99.52308.0127.42680
Number of Perturbed Samples = 30
Size of Each Sample = 10000


Output 16.3.9: Aggregate Loss Simulation Results for Region=West

Sample Percentile Perturbation Analysis
PercentileEstimateStandard
Error
100
500
25134.1640516.72670
50417.8949827.34826
75863.1305348.13708
901453.774.88636
951913.2101.60492
97.52368.8140.43218
992979.5190.75595
99.53462.9242.14530
Number of Perturbed Samples = 30
Size of Each Sample = 10000


Output 16.3.8 and Output 16.3.9 show the summary of the perturbation analysis for the two regions. You can deduce that for the collection of three policyholders in the eastern region of the specified scenario, the 97.5th percentile of their collective aggregate loss is 1553.5 79 units, and for the collection of five policyholders in the western region of the specified scenario, the 99.5th percentile of their collective aggregate loss is 3462.9 242.2.

Last updated: November 05, 2018