SEVSELECT Procedure
Example 26.6 Scale Regression with Rich Regression Effects
This example illustrates the use of regression effects that include CLASS variables and interaction effects.
Consider that you, as an actuary at an automobile insurance company, want to evaluate the effect of certain external factors on the distribution of the severity of the losses that your policyholders incur. Such analysis can help you determine the relative differences in premiums that you should charge to policyholders who have different characteristics. Assume that when you collect and record the information about each claim, you also collect and record some key characteristics of the policyholder and the vehicle that is involved in the claim. This example focuses on the following five factors: type of car, safety rating of the car, gender of the policyholder, education level of the policyholder, and annual household income of the policyholder (which can be thought of as a proxy for the luxury level of the car). Let these regressors be recorded in the variables CarType (1: sedan, 2: sport utility vehicle), CarSafety (scaled to be between 0 and 1, the safest being 1), Gender (1: female, 2: male), Education (1: high school graduate, 2: college graduate, 3: advanced degree holder), and Income (scaled by a factor of 1/100,000), respectively. Let the historical data about the severity of each loss be recorded in the LossAmount variable of the mylib.Losses data table. Let the data table also contain two additional variables, Deductible and Limit, that record the deductible and ground-up loss limit provisions, respectively, of the insurance policy that the policyholder has. The limit on ground-up loss is usually derived from the payment limit that a typical insurance policy states. Deductible serves as the left-truncation variable, and Limit serves as the right-censoring variable.
The following SAS statements simulate an example of the mylib.Losses data table:
proc format casfmtlib='myfmtlib';
value genderFmt 1='Female'
2='Male';
run;
data losses(keep=gender carType education carSafety income
lossAmount deductible limit);
call streaminit(12345);
array sx{8} _temporary_;
array sbeta{9} _TEMPORARY_ (5 0.6 0.4 -0.75 -0.3 0.4 0.7 -0.5 -0.3);
length carType $8 education $16;
format gender genderFmt.;
sigma = 0.5;
do lossEventId=1 to 6000;
/* Simulate policyholder and vehicle attributes */
do i=1 to dim(sx);
sx(i) = 0;
end;
if (rand('UNIFORM') < 0.5) then do;
gender = 1; * female;
sx(2) = 1;
end;
else do;
gender = 2; * male;
end;
if (rand('UNIFORM') < 0.7) then do;
carInt = 1;
carType = 'Sedan';
end;
else do;
carInt = 2;
carType = 'SUV';
sx(1) = 1;
end;
educationLevel = rand('UNIFORM');
if (educationLevel < 0.5) then do;
eduInt = 1;
education = 'High School';
end;
else if (educationLevel < 0.85) then do;
eduInt = 2;
education = 'College';
if (carInt=1) then
sx(8) = 1;
else
sx(6) = 1;
end;
else do;
eduInt = 3;
education = 'AdvancedDegree';
if (carInt=1) then
sx(7) = 1;
else
sx(5) = 1;
end;
carSafety = rand('UNIFORM'); /* scaled to be between 0 & 1 */
sx(3) = carSafety;
income = MAX(15000,int(rand('NORMAL', eduInt*30000, 50000)))/100000;
sx(4) = income;
/* Simulate lognormal severity */
Mu = sbeta(1);
do i=1 to dim(sx);
Mu = Mu + sx(i) * sbeta(i+1);
end;
lossAmount = exp(Mu) * rand('LOGNORMAL')**Sigma;
loglossAmount = log(lossAmount);
deductible = lossAmount * rand('UNIFORM');
if (rand('UNIFORM') < 0.25) then
limit = lossAmount;
output;
end;
run;
data mylib.losses;
set losses;
run;
The variables CarType, Education, and Gender each contain a known, finite set of discrete values. By specifying such variables as classification variables, you can separately identify the effect of each level of the variable on the severity distribution. For example, you might be interested in finding out how the magnitude of loss for a sport utility vehicle (SUV) differs from that for a sedan. This is an example of a main effect. You might also want to evaluate how the distribution of losses that are incurred by a policyholder with a college degree who drives a SUV differs from that of a policyholder with an advanced degree who drives a sedan. This is an example of an interaction effect. You can include various such types of effects in the scale regression model. For more information about the effect types, see the section Specification and Parameterization of Model Effects in Chapter 4, Shared Concepts.
Analyzing such a rich set of regression effects can help you make more accurate predictions about the losses that a new applicant with certain characteristics might incur when he or she requests insurance for a specific vehicle, which can further help you with ratemaking decisions.
The following PROC SEVSELECT step fits the scale regression model with a lognormal distribution to data in the mylib.Losses data table, and stores the model and parameter estimate information in the mylib.Est data table on the CAS server:
/* Fit scale regression model with different types of regression effects */
proc sevselect data=mylib.losses outest=mylib.est print=all;
loss lossAmount / lt=deductible rc=limit;
class carType gender education;
scalemodel carType gender carSafety income education*carType
income*gender carSafety*income;
dist logn;
run;
The SCALEMODEL statement in the preceding PROC SEVSELECT step includes two main effects (carType and gender), two singleton continuous effects (carSafety and income), one interaction effect (education*carType), one continuous-by-class effect (income*gender), and one polynomial continuous effect (carSafety*income).
When you specify a CLASS statement, it is recommended that you observe the "Class Level Information" table. For this example, the table is shown in Output 26.6.1. Note that if you specify BY-group processing, then the class level information might change from one BY group to the next, potentially resulting in a different parameterization for each BY group.
Output 26.6.1: Class Level Information Table
| Class Level Information | ||
|---|---|---|
| Class | Levels | Values |
| carType | 2 | SUV Sedan |
| gender | 2 | Female Male |
| education | 3 | AdvancedDegree College High School |
The regression modeling results for the lognormal distribution are shown in Output 26.6.2. The "Initial Parameter Values and Bounds" table is important especially because the preceding PROC SEVSELECT step uses the default GLM parameterization, which is a singular parameterization—that is, it results in some redundant parameters. As shown in the table, the redundant parameters correspond to the last level of each classification variable; this correspondence is a defining characteristic of a GLM parameterization. An alternative would be to use the reference parameterization by specifying the PARAM=REFERENCE option in the CLASS statement, which does not generate redundant parameters for effects that contain CLASS variables and enables you to specify a reference level for each CLASS variable.
Output 26.6.2: Initial Values for the Scale Regression Model with Class and Interaction Effects
| Initial Parameter Values and Bounds | |||
|---|---|---|---|
| Parameter | Initial Value | Lower Bound | Upper Bound |
| Mu | 4.85200 | -709.78271 | 709.78271 |
| Sigma | 0.52443 | 1.05367E-8 | Infty |
| carType SUV | 0.54686 | -709.78271 | 709.78271 |
| carType Sedan | Redundant | -709.78271 | 709.78271 |
| gender Female | 0.34893 | -709.78271 | 709.78271 |
| gender Male | Redundant | -709.78271 | 709.78271 |
| carSafety | -0.63504 | -709.78271 | 709.78271 |
| income | -0.24031 | -709.78271 | 709.78271 |
| carType SUV * education AdvancedDegree | 0.32719 | -709.78271 | 709.78271 |
| carType SUV * education College | 0.68899 | -709.78271 | 709.78271 |
| carType SUV * education High School | Redundant | -709.78271 | 709.78271 |
| carType Sedan * education AdvancedDegree | -0.44650 | -709.78271 | 709.78271 |
| carType Sedan * education College | -0.26834 | -709.78271 | 709.78271 |
| carType Sedan * education High School | Redundant | -709.78271 | 709.78271 |
| income * gender Female | 0.00843 | -709.78271 | 709.78271 |
| income * gender Male | Redundant | -709.78271 | 709.78271 |
| carSafety * income | -0.04744 | -709.78271 | 709.78271 |
The convergence and optimization summary information in Output 26.6.3 indicates that the scale regression model for the lognormal distribution has converged with the default optimization technique in five iterations.
Output 26.6.3: Optimization Summary for the Scale Regression Model with Class and Interaction Effects
| Convergence Status |
|---|
| Convergence criterion (GCONV=1E-8) satisfied. |
| Optimization Summary | |
|---|---|
| Optimization Technique | Trust Region |
| Iterations | 5 |
| Function Calls | 14 |
| Log Likelihood | -19716.25785 |
The "Parameter Estimates" table in Output 26.6.4 shows the distribution parameter estimates and estimates for various regression effects. You can use the estimates for effects that contain CLASS variables to infer the relative influence of various CLASS variable levels. For example, on average, the magnitude of losses that are incurred by the female drivers is times greater than that of male drivers, and an SUV driver with an advanced degree incurs a loss that is on average
times greater than the loss that a college-educated sedan driver incurs. Neither the continuous-by-class effect
income*gender nor the polynomial continuous effect carSafety*income is significant in this example.
Output 26.6.4: Parameter Estimates for the Scale Regression with Class and Interaction Effects
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Mu | 1 | 5.12332 | 0.03792 | 135.11 | <.0001 |
| Sigma | 1 | 0.56941 | 0.00744 | 76.56 | <.0001 |
| carType SUV | 1 | 0.64686 | 0.02989 | 21.64 | <.0001 |
| carType Sedan | 0 | 0 | . | . | . |
| gender Female | 1 | 0.41714 | 0.03213 | 12.98 | <.0001 |
| gender Male | 0 | 0 | . | . | . |
| carSafety | 1 | -0.88078 | 0.05532 | -15.92 | <.0001 |
| income | 1 | -0.38045 | 0.05128 | -7.42 | <.0001 |
| carType SUV * education AdvancedDegree | 1 | 0.42454 | 0.05107 | 8.31 | <.0001 |
| carType SUV * education College | 1 | 0.73228 | 0.03738 | 19.59 | <.0001 |
| carType SUV * education High School | 0 | 0 | . | . | . |
| carType Sedan * education AdvancedDegree | 1 | -0.56685 | 0.03525 | -16.08 | <.0001 |
| carType Sedan * education College | 1 | -0.35208 | 0.02618 | -13.45 | <.0001 |
| carType Sedan * education High School | 0 | 0 | . | . | . |
| income * gender Female | 1 | 0.01288 | 0.04394 | 0.29 | 0.7694 |
| income * gender Male | 0 | 0 | . | . | . |
| carSafety * income | 1 | 0.06499 | 0.07547 | 0.86 | 0.3892 |
If you want to update the model when new claims data arrive, then you can potentially speed up the estimation process by specifying the OUTEST= data table that is created by the preceding PROC SEVSELECT step as an INEST= data table in a new PROC SEVSELECT step. To illustrate, the following PROC SEVSELECT step refits the model on the same input data as the preceding PROC SEVSELECT, but it uses the mylib.Est data table that is created by that step as an INEST= data table:
/* Refit scale regression model on new data */
proc sevselect data=mylib.losses inest=mylib.est print=all;
loss lossAmount / lt=deductible rc=limit;
class carType gender education;
scalemodel carType gender carSafety income education*carType
income*gender carSafety*income;
dist logn;
run;
Because the INEST= data table is used to initialize the distribution and regression parameters, the optimization occurs in very few iterations, as shown in Output 26.6.5.
Output 26.6.5: Optimization Summary for Refitting the Scale Regression Model with the INEST= Option
| Convergence Status |
|---|
| Convergence criterion (ABSGCONV=0.00001) satisfied. |
| Optimization Summary | |
|---|---|
| Optimization Technique | Trust Region |
| Iterations | 0 |
| Function Calls | 4 |
| Log Likelihood | -19716.25785 |