The HPCDM Procedure
Example 16.3 Scenario Analysis with Rich Regression Effects and BY Groups
(View the complete code for this example.)
This example illustrates scenario analysis when frequency and severity models use regression models that contain classification and interaction effects. It also illustrates how you can analyze scenarios for multiple groups of observations in one PROC HPCDM step without your having to simulate counts externally.
The example in the section Scenario Analysis encodes the discrete-valued, nominal (nonordinal) variables Gender, CarType, and Education as numerical variables with an implied order. For example, a high school diploma is assigned a smaller number than an advanced degree. This method of forcing an order on otherwise nonordinal (categorical) variables is not natural and might lead to biased estimates. A more accurate approach is to treat such variables as classification variables that enter the statistical analysis or model not through their values but through their levels. For example, when you specify Education as a classification variable, the modeling process creates different parameters for the Education = 'High School' and Education = 'Advanced Degree' levels and estimates a regression coefficient for each. When you specify such variables in the CLASS statement of PROC COUNTREG and PROC SEVERITY, those procedures perform the appropriate levelization for you, which is the process of finding and transforming levels into regression parameters. For more information, see the description of the CLASS statement in Chapter 28: The SEVERITY Procedure.
In addition to specifying nominal variables as classification (CLASS) variables, you can include interaction effects in severity and frequency models. For example, you might want to evaluate how the distribution of losses that are incurred by a policyholder with a college degree who drives an SUV differs from that of a policyholder with an advanced degree who drives a sedan. You can do this by including an interaction between CarType and Education in your severity model. Similarly, if you want to evaluate how the number of losses that a policyholder incurs per year varies by the number of annual miles for different types of cars, you can include an interaction between CarType and AnnualMiles in your frequency model. Analyzing such a rich set of regression effects can help you make more accurate predictions about the frequency and severity distributions of losses. PROC HPCDM is designed to use such rich models to simulate a more accurate distribution of the aggregate loss.
As an example of the process, first, let the following programming statements fit the severity and count models that contain a certain set of regression effects:
proc severity data=losses(where=(not(missing(lossAmount))))
covout outstore=work.sevstore print=all plots=none;
by region;
loss lossAmount;
class carType gender education;
scalemodel carType gender carSafety income education*carType
income*gender carSafety*income;
dist logn burr;
run;
proc countreg data=losscounts covout;
by region;
class gender carType education;
model numloss = age income gender carType*annualmiles education / dist=negbin;
zeromodel numloss ~ age income carType education;
store cstore;
run;
Note the following points about these statements:
You can find the code that prepares the
Work.LossesandWork.LossCountsdata sets in the PROC HPCDM sample programhcdmex03.sas. The data sets are organized in groups of observations that represent data from two regions, East and West. You can analyze both groups at once by specifying the BY statement withRegionas the BY variable.Both severity and count models use three CLASS variables. The severity model includes three interaction effects (
Education*CarType,Income*Gender, andCarSafety*Income) and four main effects. PROC SEVERITY uses the same set of regression effects in the scale regression model of each of the two distributions that you specify in the DIST statement, which are LOGN and BURR in this example.The count model is a mixture of two models: a model to estimate the occurrence of zero loss events and a model to estimate nonzero counts. The zero model is a regression model with four main effects and the default logistic link function. The model for nonzero counts is a negative binomial model with one interaction effect (
CarType*AnnualMiles) and four main effects.
The "Parameter Estimates" table of the lognormal severity model in Output 16.3.1 for the Region='East' BY group shows that Income*Gender and CarSafety*Income effects are not statistically significant. The "Parameter Estimates" table in Output 16.3.2 shows that those two effects are not statistically significant for the Burr severity model also.
Output 16.3.1: Parameter Estimates for LOGN Severity Model for Region=East
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Mu | 1 | 4.98253 | 0.02861 | 174.16 | <.0001 |
| Sigma | 1 | 0.48894 | 0.00535 | 91.41 | <.0001 |
| carType SUV | 1 | 0.51772 | 0.03648 | 14.19 | <.0001 |
| carType Sedan | 0 | 0 | . | . | . |
| gender F | 1 | 1.16690 | 0.03082 | 37.86 | <.0001 |
| gender M | 0 | 0 | . | . | . |
| carSafety | 1 | -0.71517 | 0.04599 | -15.55 | <.0001 |
| income | 1 | -0.28528 | 0.03652 | -7.81 | <.0001 |
| carType*education SUV Advanced Degree | 1 | 0.44599 | 0.06245 | 7.14 | <.0001 |
| carType*education SUV College | 1 | 0.67852 | 0.04416 | 15.36 | <.0001 |
| carType*education SUV High School | 0 | 0 | . | . | . |
| carType*education Sedan Advanced Degree | 1 | -0.49680 | 0.02689 | -18.47 | <.0001 |
| carType*education Sedan College | 1 | -0.26310 | 0.01849 | -14.23 | <.0001 |
| carType*education Sedan High School | 0 | 0 | . | . | . |
| income*gender F | 1 | 0.00988 | 0.04010 | 0.25 | 0.8054 |
| income*gender M | 0 | 0 | . | . | . |
| carSafety*income | 1 | -0.09390 | 0.06166 | -1.52 | 0.1278 |
Output 16.3.2: Parameter Estimates for BURR Severity Model for Region=East
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Theta | 1 | 145.63709 | 5.74371 | 25.36 | <.0001 |
| Alpha | 1 | 0.99783 | 0.06470 | 15.42 | <.0001 |
| Gamma | 1 | 3.58743 | 0.09362 | 38.32 | <.0001 |
| carType SUV | 1 | 0.51648 | 0.03701 | 13.96 | <.0001 |
| carType Sedan | 0 | 0 | . | . | . |
| gender F | 1 | 1.16664 | 0.03083 | 37.84 | <.0001 |
| gender M | 0 | 0 | . | . | . |
| carSafety | 1 | -0.71636 | 0.04590 | -15.61 | <.0001 |
| income | 1 | -0.29522 | 0.03639 | -8.11 | <.0001 |
| carType*education SUV Advanced Degree | 1 | 0.43696 | 0.06385 | 6.84 | <.0001 |
| carType*education SUV College | 1 | 0.68049 | 0.04501 | 15.12 | <.0001 |
| carType*education SUV High School | 0 | 0 | . | . | . |
| carType*education Sedan Advanced Degree | 1 | -0.50160 | 0.02672 | -18.77 | <.0001 |
| carType*education Sedan College | 1 | -0.26483 | 0.01840 | -14.39 | <.0001 |
| carType*education Sedan High School | 0 | 0 | . | . | . |
| income*gender F | 1 | 0.01268 | 0.03986 | 0.32 | 0.7504 |
| income*gender M | 0 | 0 | . | . | . |
| carSafety*income | 1 | -0.07713 | 0.06162 | -1.25 | 0.2107 |
The "Parameter Estimates" table of the count model in Output 16.3.3 shows that the income and Inf_income parameters are insignificant. This implies that the income effect is not significant for the main and zero inflation parts of the count model.
The results for the Region='West' BY group are not shown here, but you can execute the sample program hcdmex03.sas to verify that the same parameters are statistically insignificant in severity and count models of that BY group as well. However, in general, you might find that some effects are significant for some BY groups but insignificant for other BY groups. In such cases, for more accurate results, it is recommended that you create a separate data set for each set of similar BY groups and invoke the SEVERITY, COUNTREG, and HPCDM procedures on each data set to separately analyze each set of similar BY groups.
Output 16.3.3: Count Model Parameter Estimates for Region=East
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Intercept | 1 | 1.156626 | 0.130641 | 8.85 | <.0001 |
| age | 1 | 0.734798 | 0.112299 | 6.54 | <.0001 |
| income | 1 | -0.040744 | 0.081573 | -0.50 | 0.6174 |
| gender F | 1 | -0.999094 | 0.053170 | -18.79 | <.0001 |
| gender M | 0 | 0 | . | . | . |
| annualmiles*carType SUV | 1 | -1.266452 | 0.045996 | -27.53 | <.0001 |
| annualmiles*carType Sedan | 1 | -0.632281 | 0.027818 | -22.73 | <.0001 |
| education Advanced Degree | 1 | 0.418651 | 0.099414 | 4.21 | <.0001 |
| education College | 1 | 0.709479 | 0.069596 | 10.19 | <.0001 |
| education High School | 0 | 0 | . | . | . |
| Inf_Intercept | 1 | -0.501240 | 0.353072 | -1.42 | 0.1557 |
| Inf_age | 1 | -0.945657 | 0.329950 | -2.87 | 0.0042 |
| Inf_income | 1 | -0.173541 | 0.233461 | -0.74 | 0.4573 |
| Inf_carType SUV | 1 | -0.693427 | 0.369119 | -1.88 | 0.0603 |
| Inf_carType Sedan | 0 | 0 | . | . | . |
| Inf_education Advanced Degree | 1 | 0.668613 | 0.291821 | 2.29 | 0.0220 |
| Inf_education College | 1 | 0.474212 | 0.232499 | 2.04 | 0.0414 |
| Inf_education High School | 0 | 0 | . | . | . |
| _Alpha | 1 | 0.790839 | 0.103522 | 7.64 | <.0001 |
The following modified PROC SEVERITY and PROC COUNTREG steps refit the severity and count models, respectively, after removing the insignificant effects:
/* Re-fit models after removing insignificant effects. */
proc severity data=losses(where=(not(missing(lossAmount))))
covout outstore=work.sevstore print=all plots=none;
by region;
loss lossAmount;
class carType gender education;
scalemodel carType gender carSafety income education*carType;
dist logn burr;
run;
proc countreg data=losscounts covout;
by region;
class gender carType education;
model numloss = age gender carType*annualmiles education / dist=negbin;
zeromodel numloss ~ age carType education;
store cstore;
run;
Note that the PROC SEVERITY step uses the OUTSTORE= option to store the parameter estimates in an item store. When your scale regression model contains classification or interaction effects, you must store the parameter estimates in an item store instead of storing them in an OUTEST= data set, because PROC HPCDM cannot obtain the necessary information about classification or interaction effects from an OUTEST= data set.
The "Parameter Estimates" tables in Output 16.3.4 and Output 16.3.5 show that all parameters are now statistically significant, most at the 95% confidence level and a few at the 90% confidence level. If you want every parameter to be significant at the 95% confidence level, then you might want to continue the process by removing the carType effect with a p-value of 0.0607 from the ZEROMODEL statement and refitting the count model. However, for the purpose of this example, the preceding models are declared to be satisfactory, and the effect selection process stops here.
You need to follow this process of model inspection and effect selection before you use the severity and count models with the HPCDM procedure. For count models, you can use the automatic effect (variable) selection feature of PROC COUNTREG. For more information, see the description of the SELECT= option in the MODEL statement of Chapter 11: The COUNTREG Procedure. For severity models, you need to perform effect selection manually by inspecting the estimates and refitting the model after removing one or a few insignificant effects at a time until you find the final set of significant effects. Although it is not shown in this example, you can also decide which set of effects is better by comparing the fit statistics of two models; the better model might contain certain effects at lower confidence levels than the usual 95% or 90% confidence levels. In fact, the SELECT=INFO option of PROC COUNTREG uses the AIC or BIC of the entire model to select the set of effects instead of using the p-values of individual parameters. You might also want to use some domain knowledge to retain certain effects in the model even if their confidence level is not very high.
Output 16.3.4: Final LOGN Severity Model Parameter Estimates for Region=East
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Mu | 1 | 5.00845 | 0.02135 | 234.61 | <.0001 |
| Sigma | 1 | 0.48908 | 0.00535 | 91.43 | <.0001 |
| carType SUV | 1 | 0.51556 | 0.03642 | 14.16 | <.0001 |
| carType Sedan | 0 | 0 | . | . | . |
| gender F | 1 | 1.17291 | 0.01726 | 67.96 | <.0001 |
| gender M | 0 | 0 | . | . | . |
| carSafety | 1 | -0.77273 | 0.02614 | -29.56 | <.0001 |
| income | 1 | -0.32702 | 0.01962 | -16.67 | <.0001 |
| carType*education SUV Advanced Degree | 1 | 0.44870 | 0.06223 | 7.21 | <.0001 |
| carType*education SUV College | 1 | 0.68360 | 0.04404 | 15.52 | <.0001 |
| carType*education SUV High School | 0 | 0 | . | . | . |
| carType*education Sedan Advanced Degree | 1 | -0.49572 | 0.02688 | -18.44 | <.0001 |
| carType*education Sedan College | 1 | -0.26234 | 0.01848 | -14.19 | <.0001 |
| carType*education Sedan High School | 0 | 0 | . | . | . |
Output 16.3.5: Final Count Model Parameter Estimates for Region=East
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Intercept | 1 | 1.136175 | 0.124786 | 9.10 | <.0001 |
| age | 1 | 0.737805 | 0.112339 | 6.57 | <.0001 |
| gender F | 1 | -1.001311 | 0.052996 | -18.89 | <.0001 |
| gender M | 0 | 0 | . | . | . |
| annualmiles*carType SUV | 1 | -1.263178 | 0.045809 | -27.57 | <.0001 |
| annualmiles*carType Sedan | 1 | -0.631419 | 0.027728 | -22.77 | <.0001 |
| education Advanced Degree | 1 | 0.400307 | 0.092060 | 4.35 | <.0001 |
| education College | 1 | 0.703436 | 0.067935 | 10.35 | <.0001 |
| education High School | 0 | 0 | . | . | . |
| Inf_Intercept | 1 | -0.585662 | 0.338796 | -1.73 | 0.0839 |
| Inf_age | 1 | -0.928294 | 0.324629 | -2.86 | 0.0042 |
| Inf_carType SUV | 1 | -0.658089 | 0.350886 | -1.88 | 0.0607 |
| Inf_carType Sedan | 0 | 0 | . | . | . |
| Inf_education Advanced Degree | 1 | 0.588511 | 0.269195 | 2.19 | 0.0288 |
| Inf_education College | 1 | 0.446600 | 0.228151 | 1.96 | 0.0503 |
| Inf_education High School | 0 | 0 | . | . | . |
| _Alpha | 1 | 0.785018 | 0.101327 | 7.75 | <.0001 |
For severity models, you also need to inspect the "All Fit Statistics" table to decide which severity distributions you want to use for aggregate loss modeling. The table in Output 16.3.6 shows that the lognormal distribution is the best according to the majority of fit statistics, so you can choose that. However, in some cases, you might see that the likelihood-based fit statistics (–2 log likelihood, AIC, AICC, BIC) choose one distribution and the EDF-based statistics (KS, AD, CvM) choose another distribution. In such cases, it is recommended that before making your final decision, you conduct aggregate loss simulation by using both severity distributions and compare the summary statistics and percentiles that each severity distribution produces.
Output 16.3.6: Comparison of Severity Distributions for Region=East
After you have satisfactorily estimated the severity and frequency models, it is time to estimate the distribution of the aggregate loss by using the HPCDM procedure. The scenario data set must contain the final set of regressors that are used in both the severity model and the frequency model. Note that even if your models contain interaction effects, your scenario data set needs to contain only the columns for individual variables of the effects. PROC HPCDM internally performs levelization of each observation, which is the process of expanding the variable values to match them with the parameters of each effect. A typical scenario for an insurance application might consist of a large number of policyholders, but for illustration purposes, this example uses a small scenario of only a few policyholders per region. Output 16.3.7 shows the contents of the Work.Scenario data set, and the following PROC HPCDM step simulates the aggregate losses for that scenario:
proc hpcdm data=scenario nreplicates=10000 seed=123 print=all
severitystore=work.sevstore countstore=work.cstore
nperturb=30;
by region;
severitymodel logn;
outsum out=agglossStats mean stddev skewness kurtosis pctlpts=(90 97.5 99.5);
run;
Output 16.3.7: Work.Scenario Data Set for BY-Group Processing
| Obs | region | gender | carType | education | age | annualmiles | carSafety | income |
|---|---|---|---|---|---|---|---|---|
| 1 | East | F | SUV | High School | 1.16 | 2.1540 | 0.29288 | 0.26090 |
| 2 | East | F | Sedan | High School | 0.86 | 2.3978 | 0.69844 | 0.15000 |
| 3 | East | F | Sedan | Advanced Degree | 0.78 | 1.9926 | 0.59421 | 0.58808 |
| 4 | West | M | Sedan | College | 0.82 | 1.8550 | 0.66849 | 0.15000 |
| 5 | West | M | SUV | College | 0.40 | 3.6240 | 0.23194 | 1.25274 |
| 6 | West | M | Sedan | High School | 0.62 | 3.6162 | 0.86477 | 0.42597 |
| 7 | West | F | Sedan | College | 0.32 | 3.4598 | 0.66294 | 0.36132 |
| 8 | West | M | Sedan | Advanced Degree | 0.90 | 3.2580 | 0.37172 | 0.15000 |
The SEVERITYSTORE= and COUNTSTORE= options specify the item stores that contain the effect information and parameter estimates of the severity and counts models, respectively, for both BY groups. The COVOUT option in the preceding PROC SEVERITY and PROC COUNTREG steps ensures that the respective item stores include the covariance estimates that are needed for the perturbation analysis that the NPERTURB= option requests.
Output 16.3.8: Aggregate Loss Simulation Results for Region=East
| Sample Percentile Perturbation Analysis | ||
|---|---|---|
| Percentile | Estimate | Standard Error |
| 1 | 0 | 0 |
| 5 | 0 | 0 |
| 25 | 0 | 0 |
| 50 | 151.62052 | 20.57120 |
| 75 | 492.04365 | 33.55686 |
| 90 | 917.18029 | 51.54978 |
| 95 | 1233.3 | 63.95801 |
| 97.5 | 1553.5 | 78.97273 |
| 99 | 1981.2 | 111.13102 |
| 99.5 | 2308.0 | 127.42680 |
| Number of Perturbed Samples = 30 | ||
| Size of Each Sample = 10000 | ||
Output 16.3.9: Aggregate Loss Simulation Results for Region=West
Output 16.3.8 and Output 16.3.9 show the summary of the perturbation analysis for the two regions. You can deduce that for the collection of three policyholders in the eastern region of the specified scenario, the 97.5th percentile of their collective aggregate loss is 1553.5 79 units, and for the collection of five policyholders in the western region of the specified scenario, the 99.5th percentile of their collective aggregate loss is 3462.9 242.2.