The ECM Procedure

Getting Started: ECM Procedure

This section outlines the use of the ECM procedure to estimate the distribution of the total loss across two dependent lines of business. The example is intended as a gentle introduction to the key functionality of PROC ECM.

The losses that a particular line of business (LoB) incurs often have a compound probability distribution that cannot be expressed in a closed form, so you need to approximate the distribution by simulating a large, empirical sample from that distribution. A common way to do that is to fit probability distribution models for the frequency and severity of losses (which can usually be expressed in parametric forms) and use those models to simulate a large sample of the aggregate loss. You would typically use the CNTSELECT procedure to estimate the frequency model, the SEVSELECT procedure to estimate the severity model, and the CCDM procedure to combine those frequency and severity models to simulate an aggregate loss sample from the compound distribution. However, for the purpose of this example, let the following two DATA steps perform the aggregate loss simulations by assuming that the losses in one LoB follow a negative binomial frequency distribution and a gamma severity distribution, whereas the losses in the second LoB follow a Poisson frequency distribution and a lognormal severity distribution. This example assumes that the CAS engine libref is named mycas, as defined in the section Using CAS Sessions and CAS Engine Librefs, but you can substitute any appropriately named CAS engine libref.

/* Simulate a sample from the compound distribution of
   negative binomial frequency and gamma severity. */
data mycas.loss1sample(keep=loss1) / sessref=mysess;
   call streaminit(12);
   theta = 3.0; p = 0.5; * negative binomial parameters;
   gTheta = 1000; gAlpha = 2.5; * gamma parameters;
   do n = 1 to 10000;
      count = rand('NEGB',p,theta);
      loss1 = 0;
      do c=1 to count;
         loss1 = loss1 + rand('Gamma', gAlpha, gTheta);
      end;
      output;
   end;
run;

/* Simulate a sample from the compound distribution of
   Poisson frequency and lognormal severity. */
data mycas.loss2sample(keep=loss2) / sessref=mysess;
   call streaminit(135);
   lambda = 2; * Poisson parameters;
   mu = log(2500); sigma = 0.85; * lognormal parameters;
   do n = 1 to 10000;
      count = rand('POISSON',lambda);
      loss2 = 0;
      do c=1 to count;
         loss2 = loss2 + exp(mu) * rand('LOGNORMAL')**sigma;
      end;
      output;
   end;
run;

The preceding two DATA steps produce two large sample tables, mycas.Loss1Sample and mycas.Loss2Sample, for the two LoBs. Each data table represents an empirical sample for the aggregate loss distribution of the respective LoB. The use of the SESSREF= option ensures that each DATA step runs on the CAS server in parallel by using multiple threads of execution on all the worker nodes in order to create large samples quickly. When each of these DATA steps is run on a CAS session that has three worker nodes, such that each worker node uses 32 threads of execution, the sample size of each LoB is 960,000, because the loop inside the DATA step runs 10,000 times independently in each thread on each worker node.

The losses in the two LoBs are usually correlated, and you can estimate their dependency structure by fitting a copula model. In the context of copula modeling, loss in each LoB constitutes a marginal random variable, referred to simply as a marginal. For this example, assume that the dependency is captured by a Gaussian copula with a correlation matrix that is recorded in the data table mycas.CorrTab:

data mycas.corrTab;
   loss1 = 1.0; loss2 = 0.35; output;
   loss1 = 0.35; loss2 = 1.0; output;
run;

To simulate the total loss across both LoBs that takes the dependency structure into account, you need to first simulate the probabilities of a loss value that each LoB incurs in the same time period. The following PROC CCOPULA step simulates 500,000 such loss probability combinations and stores them in the data table mycas.LossProb:

proc ccopula;
   var loss1 loss2;
   define cop normal (corr=mycas.corrtab);
   simulate cop / ndraws=500000 seed=234
      outuniform=mycas.lossprob;
run;

Now, you need to perform the following steps to estimate the distribution of the total loss:

  1. For each observation in the copula simulation table, you need to convert the loss probability of each LoB to its loss estimate by computing the quantile function for that LoB’s probability distribution. When the cumulative distribution function (CDF) or the quantile function of a marginal distribution is not available in a closed form, you need to estimate the quantile by the corresponding percentile from an empirical sample of that marginal’s distribution. For this example, you need to use the empirical samples for the two LoBs that are available in the data tables mycas.Loss1Sample and mycas.Loss2Sample. After computing each LoB’s loss estimate as a percentile, you need to add those loss estimates to estimate the total loss for the current observation of the copula simulation table.

  2. You need to perform the preceding step for all observations of the copula simulation table. This results in a sample of the total loss, which represents an empirical estimate of the probability distribution of the total loss. You need to analyze this sample to compute the estimates of various quantities of interest, especially the risk measures such as the value-at-risk (VaR) and tail value-at-risk (TVaR).

The ECM procedure is created precisely to help you implement the preceding steps by submitting a few simple statements and executing them very efficiently by utilizing all the computational resources of a CAS server. The following PROC ECM step implements the preceding steps by using the data table mycas.LossProb (which contains the copula simulations) and the data tables mycas.Loss1Sample and mycas.Loss2Sample (which contain the empirical samples of the two LoBs):

proc ecm data=mycas.lossprob seed=57;
   marginal loss1: data=mycas.loss1sample;
   marginal loss2: data=mycas.loss2sample;
   outsum pctlpts=default tvarpts=90 95 97.5 to 99.5 by 1;
   output out=mycas.tlossSample var=tloss;
run;

Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 4, Shared Concepts.

The DATA= data table must contain the loss probabilities. Use of the OUTUNIFORM= option in the preceding PROC CCOPULA step ensures that the data table mycas.LossProb contains copula simulations on a uniform probability scale.

The SEED= option in the PROC ECM statement ensures that results are similar across multiple runs of the procedure for identical inputs and identical session configuration. This is required because, by default, the algorithm that PROC ECM uses samples the body region of the distribution to reduce the data movement among the worker nodes and to speed up the computations.

The two MARGINAL statements specify the two marginal variables, one for each LoB. The name of the marginal variable, which you must specify first in each MARGINAL statement followed by a colon (:), must match the name of the corresponding variable in the copula simulation table. The DATA= option in each MARGINAL statement specifies the data table that contains the empirical sample of that marginal.

The OUTSUM statement specifies the statistics that you want to estimate for the distribution of the total loss. The PCTLPTS=DEFAULT option requests that the percentiles be computed for the default percentile levels, which are {1, 5, 10, 25, 50, 75, 90, 95, 99, 99.5}. Each percentile is an estimate of the value-at-risk (VaR). The TVARPTS= option requests that the estimates of the tail value-at-risk (TVaR) be computed for 90th, 95th, 97.5th, 98.5th, and 99.5th percentile levels.

The OUTPUT statement requests that the total loss sample be written in the tloss variable in the output data table mycas.TlossSample. You can use this output table especially to compute a risk measure that is different from the two popular risk measures, VaR and TVaR, which PROC ECM computes for you.

Figure 1 shows some of the results that the preceding ECM step prepares. The "Summary Statistics" table displays various summary statistics. The "Percentiles" table displays the percentiles. The Pth percentile is essentially an estimate of the VaR at level alpha equals upper P slash 100, because for a random loss variable L, the VaR at level alpha is defined as the smallest number l such that the probability that the loss exceeds l is no larger than 1 minus alpha. Formally,

StartLayout 1st Row 1st Column VaR Subscript alpha Baseline left-parenthesis upper L right-parenthesis 2nd Column equals inf left-brace l element-of double-struck upper R colon probability left-bracket upper L greater-than l right-bracket less-than-or-equal-to 1 minus alpha right-brace 2nd Row 1st Column Blank 2nd Column equals inf left-brace l element-of double-struck upper R colon upper F Subscript upper L Baseline left-parenthesis l right-parenthesis greater-than-or-equal-to alpha right-brace 3rd Row 1st Column Blank 2nd Column equals upper F Subscript upper L Superscript negative 1 Baseline left-parenthesis alpha right-parenthesis EndLayout

where upper F Subscript upper L and upper F Subscript upper L Superscript negative 1 denote the CDF and the quantile function of L, respectively. A percentile is essentially an empirical estimate of the quantile function, and hence an estimate of the VaR.

Figure 1: Summary Statistics and Percentile Estimates for the Total Loss

The ECM Procedure

Summary Statistics
Sample Size500000
Minimum0
Maximum161698.7
Mean14653.4
Standard Deviation11330.6
Variance128383524
Skewness1.39642
Kurtosis3.33577
Median12307.6
Interquartile Range14061.3

Percentiles
PercentileValue
10
5972.38250
102549.5
256310.9
5012307.6
7520372.2
9029598.5
9536244.7
97.542721.0
98.547460.3
9951283.2
99.557915.0
Percentile Method = 5


The "Tail Value-at-Risk Estimates" table in Figure 2 displays the VaR and TVaR estimates for the percentile levels that are specified in the TVARPTS= option. The Pr > VaR column indicates the probability that a loss exceeds the VaR value. For a level alpha, the TVaR Subscript alpha is defined as the expected value of the loss given that the loss exceeds VaR Subscript alpha. Formally, TVaR Subscript alpha Baseline equals upper E left-bracket upper L vertical-bar upper L greater-than VaR Subscript alpha Baseline right-bracket. Because of this definition, TVaR is also sometimes called a conditional tail expectation (CTE). Computationally, it is the average of the loss values that exceed the Pth percentile. As the "Tail Value-at-Risk Estimates" table shows, the TVaR value is always greater than the VaR value, which is expected because it is the mean of the losses that exceed the VaR value. In fact, TVaR is a better risk measure than the VaR, because TVaR accounts for the shape of the tail of the distribution. To capture the tail more accurately, you need to simulate not only large samples of the loss probabilities (copula simulation) but also large samples of each of the marginals. PROC ECM is specifically designed to handle very large copula simulation and marginal samples. It uses all the computational resources of a CAS server, and it uses a novel parallel and distributed algorithm for approximating the empirical distribution function (EDF) for computing the percentiles of the marginals. Together, these uses enable PROC ECM to perform the necessary computations significantly faster.

Figure 2: Tail Value-at-Risk (TVaR) Estimates for the Total Loss

Tail Value-at-Risk Estimates
PercentilePr > VaRValue at Risk (VaR)Tail Value at Risk
900.129598.539108.7
950.0536244.745682.5
97.50.02542721.052247.9
98.50.01547460.357151.1
99.50.00557915.067992.2


Last updated: May 01, 2023