ECM Procedure

PROC ECM Statement

  • PROC ECM options;

The PROC ECM statement invokes the procedure. You can specify the following options, which are listed in alphabetical order.

DATA=libref.data-table
COPULASIM=libref.data-table

names the input data table for PROC ECM to use. The default is the most recently created data table. libref.data-table is a two-level name, where

libref

refers to a collection of information that is defined in the LIBNAME statement and includes the library, which includes a path to the data, and a session identifier, which defaults to the active session but which can be explicitly defined in the LIBNAME statement. For more information about libref, see the section Using CAS Sessions and CAS Engine Librefs.

data-table

specifies the name of the input data table.

The DATA= data table specifies a copula sample on a uniform scale that is generated by using the OUTUNIFORM= option in the SIMULATE statement in the CCOPULA procedure.

Each observation in this data table is expected to contain the probability of the loss for each of the marginal variables. That probability is determined by accounting for the dependence structure among the marginal variables.

DRAWID=sample-identifier
SAMPLEID=sample-identifier

identifies the sample to use when the marginal data tables contain multiple samples of the marginal variable, where sample-identifier is an integer. A marginal data table can contain multiple samples when it is a result of the perturbation analysis that the CCDM procedure conducts. Each sample is identified by the value of the _DRAWID_ variable. If you specify a sample-identifier value of k, then PROC ECM uses the observations in the marginal data tables whose _DRAWID_=k.

The sample-identifier applies to all the marginal data tables that you specify in one or more MARGINAL statements. You can override the sample-identifier for a specific marginal table by specifying the DRAWID= option in the MARGINAL statement that you specify for the respective marginal variable.

If you omit this option or the DRAWID= option in the MARGINAL statement, then PROC ECM uses a default value of 0 for the marginal tables that contain the _DRAWID_ variable.

EDFACCURACY=number
ACC=number

specifies the accuracy to be achieved when estimating the empirical distribution function (EDF) in the body region of the distribution, where the body region is defined as the region that contains values whose EDF estimate is less than the value that you specify in the TAILSTARTEDF= option. The number you specify represents the desired proportion of observations that each bin in the body region should contain. The number must be between 0 and 1.

PROC ECM uses number for estimating the EDF of each marginal variable. If you specify percentiles in the OUTSUM statement, PROC ECM also uses number when estimating the EDF of the total loss.

By default, EDFACCURACY=1.0E–5.

EXACTFINALCOUNT

computes the empirical distribution function (EDF) by making an additional pass over the data to count the final set of bins when the data are sampled and shuffled by the binning algorithm. This results in better EDF estimates, but at the cost of an additional data pass.

If you omit this option, then the final bin counts are approximated by inflating the counts by the sampling fraction. When the sampling fraction is less than 1, this count approximation happens only for the bins in the body region of the distribution.

If you specify a value of 1 for the SAMPLEFRACTION= option or if you specify the NOSHUFFLE option, then this option is not relevant and PROC ECM ignores it.

MAXITER=number

specifies the maximum number of iterations for the equal-proportions binning algorithm.

By default, MAXITER=50.

NOPRINT

suppresses all displayed output. If you specify this option, then PROC ECM ignores any value that you specify for the PRINT= option.

NOSHUFFLE

suppresses the use of shuffling in the EDF estimation process.

If you omit this option, then PROC ECM shuffles each marginal variable’s data to different worker nodes so that it can perform in parallel the binning for different ranges of the marginal data. However, this additional parallelism comes at the cost of moving data across workers. When the majority of the data are in the body region of the distribution, you can reduce the amount of data that are being moved by using the SAMPLEFRACTION= option. However, if a marginal variable’s data are very large or the SAMPLEFRACTION= option value is closer to 1, then the amount of data movement can make the algorithm slow.

If you specify this option, the binning algorithm does not shuffle the marginal data. Instead, it estimates the boundaries for a global set of bins in each of the regions (body and tail) of the distribution. This estimation increases the number of parameters that the algorithm needs to optimize while reducing the parallelism in the binning algorithm. Also, the algorithm needs to share the bin boundaries with all the worker nodes after every iteration; so if the values of the EDFACCURACY= and TAILEDFACCURACY= options are very small (resulting in a large number of bins), you might not save much in data movement cost. The cost of shuffling the data, which the algorithm incurs only once, might be lower than the cost of sharing the large number of bin boundaries with all worker nodes after each iteration.

If you specify this option, the SAMPLEFRACTION= and SEED= options are irrelevant and PROC ECM ignores them.

PCTLDEF=1 | 2 | 3 | 4 | 5

specifies the method of computing the percentiles. This option applies to the computation of percentiles of the total loss sample. It also applies to each marginal variable for which you do not specify a percentile method in the corresponding MARGINAL statement.

You can specify the following values:

1

uses the weighted average.

2

uses the value closest to the sample size times the percentile.

3

uses the empirical distribution function.

4

uses the weighted average (identical to PCTLDEF=1).

5

uses the empirical distribution function with averaging.

For more information about each method, see the section Percentile Computation Methods. By default, PCTLDEF=5.

PRINT <(global-display-option)> =display-option
PRINT <(global-display-option)> =(display-option …display-option)

specifies the desired displayed output. If you specify more than one display-option, then separate them with spaces and enclose them in parentheses. For more information about the displayed output, see the section Displayed Output.

You can specify the following global-display-option:

ONLY

displays only the output that is requested by the display-options.

You can specify the following display-options:

ALL

displays all the output.

DATASUMMARY
DSUM

displays a summary of the sample that is analyzed for each variable, marginal variable, or total loss variable for which PROC ECM prepares an empirical distribution function (EDF) estimate.

EDFINIT
EINIT

displays the summary of the first initialization stage of the EDF estimation process.

EDFOPTDETAILS
EOPT

displays the details for the data ranges for which the EDF estimation process runs an iterative binning algorithm.

EDFSUMMARY
ESUM

displays the summary of the data ranges that the EDF estimation process analyzes.

NONE

displays no output. If you specify this option, then it overrides all other display options. The default displayed output is also suppressed.

PERCENTILES
PCTL

displays the percentiles of the total loss sample. This includes all the predefined percentiles and percentiles that you request in the OUTSUM statement.

SUMMARYSTATISTICS
SUMSTAT

displays the summary statistics of the total loss sample.

TVAR

displays the tail value-at-risk (TVaR) estimates of the total loss sample.

Not specifying the PRINT= option or the ONLY global-display-option is equivalent to specifying the PRINT=(SUMMARYSTATISTICS) option if you do not specify any percentiles in the OUTSUM statement, or is equivalent to specifying the PRINT=(PERCENTILES SUMMARYSTATISTICS) option if you specify percentiles in the OUTSUM statement.

SAMPLEFRACTION=number
BODYSAMPLEFRACTION=number

specifies the fraction of observations to sample from the body region of the distribution during the EDF estimation process, where number must be between 0 and 1.

Sampling helps reduce the amount of data that need to be communicated across the worker nodes of the CAS server. It can reduce the accuracy of the EDF and percentile estimates, but if your marginal samples have large numbers of observations, then the loss in accuracy is not as significant as the reduction in communication cost that the sampling provides.

By default, SAMPLEFRACTION=0.5.

SEED=number

specifies an integer to use as the seed in generating the pseudorandom numbers that are used for sampling the body region of the distribution.

If you omit this option or if you specify a number that is negative or 0, then PROC ECM uses as the seed a number that depends on the time of day from the computer’s clock.

TAILEDFACCURACY=number
TAILACC=number

specifies the accuracy to be achieved when estimating the empirical distribution function (EDF) in the tail region of the distribution, where number (which must be between 0 and 1) is the desired fraction of observations that each bin in the tail region should contain. The tail region is defined as the region that contains values whose EDF estimate is larger than or equal to the value that you specify in the TAILSTARTEDF= option.

PROC ECM uses number when it estimates the EDF of each marginal variable. If you specify percentiles in the OUTSUM statement, PROC ECM also uses number when it estimates the EDF of the total loss.

By default, TAILEDFACCURACY=1.0E–6.

TAILSTARTEDF=number
TAILSTART=number

specifies the empirical distribution function (EDF) value that marks the beginning of the tail region of the distribution, where number must be between 0 and 1. The first stage of the EDF estimation algorithm finds an approximate value y Subscript t whose EDF is close to number. The part of the sample whose values are less than y Subscript t defines the body region of the distribution, and the remaining values define the tail region of the distribution.

By default, TAILSTARTEDF=0.8.

TOLERANCE=epsilon
TOL=epsilon

specifies the tolerance value that determines when the equal-proportion binning algorithm has converged, where epsilon must be a number between 0 and 1. The algorithm converges when the difference in the fractions of observations in the largest and smallest bins is less than or equal to delta epsilon, where delta is the accuracy of the body or tail region that is being analyzed.

By default, TOLERANCE=0.5.

VARDEF=divisor

specifies the divisor to use in calculating the variance, standard deviation, kurtosis, and skewness of the total loss sample. You can specify one of the following values for the divisor, where N is the sample size:

DF

sets the divisor for variance to upper N minus 1. This also changes the definitions of skewness and kurtosis.

N

sets the divisor to N.

By default, VARDEF=DF.

Last updated: July 09, 2026