CCOPULA Procedure
FIT Statement
FIT type <NAME=name><INIT=(parameter-value-options>) / options;
The FIT statement estimates the parameters for a specified copula type.
You must specify a type:
- type
-
specifies the type of the copula to be estimated, which is one of the following:
- CLAYTON
fits the Clayton copula.
- FRANK
fits the Frank copula.
- GUMBEL
fits the Gumbel copula.
- NORMAL
fits the normal copula.
- T
fits the t copula.
You can also specify the following options:
- INIT=(parameter-value-options)
-
provides the initial values for the numerical optimization.
[parameter-value-options] specify the input parameters that are used to initialize the specified copula. These options must be appropriate for the type of copula that you specify. You can specify the following options:
- CORR=libref.SAS-data-table
-
specifies the data table that contains the correlations matrix to use for elliptical copulas. libref.data-table is a two-level name, where libref refers to the library, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.
All columns of the matrix must be numeric. The order of the column names in the matrix must match their order in the VAR statement. The order of the rows must be such that row n contains the correlations between the nth variable in the VAR statement and the other variables. You can use this option for t copulas.
- DF=value
specifies the degrees of freedom. You can use this option for t copulas.
- KENDALL=libref.SAS-data-table
-
specifies the data table that contains the correlations matrix defined in Kendall’s tau. libref.data-table is a two-level name, where libref refers to the library, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.
All columns of the matrix must be numeric. The order of the column names in the matrix must match their order in the VAR statement. The order of the rows must be such that row n contains the correlations between the nth variable in the VAR statement and the other variables. You can use this option for t copulas.
- SPEARMAN=libref.SAS-data-table
-
specifies the data table that contains the correlations matrix defined in Spearman’s rho. libref.data-table is a two-level name, where libref refers to the library, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.
All columns of the matrix must be numeric. The order of the column names in the matrix must match their order in the VAR statement. The order of the rows must be such that row n contains the correlations between the nth variable in the VAR statement and the other variables. You can use this option for t copulas.
- THETA=value
specifies the parameter value for Archimedean copulas.
For Archimedean copulas, the default initial values of the parameter are computed using the calibration method. The default initial value of the degrees-of-freedom parameter in the t copula is set to 2.0. The following statement shows an initialization for Student’s t copula, where the Kendall’s tau correlations matrix is stored in the
mylib.corrmatdata table and theis set to 2.5:
fit t init=(df=2.5 kendall=mylib.corrmat); - NAME=name
specifies an identifier for the fit.
You can specify the following options after a slash (/):
- MARGINALS=EMPIRICAL<(approx-options)> | UNIFORM
-
specifies the marginal distribution of the individual variables. You can specify the following values:
- EMPIRICAL<(approx-options)>
-
uses an approximation of the marginal empirical cumulative distribution function (CDF) to transform the data and uses the transformed data to fit the copula. You can specify the following approx-options:
- ALG=BIN | SORT
-
specifies the algorithm to use during the approximation of the marginal distribution. You can specify the following values:
- BIN
uses bins to approximate the marginal distributions.
- SORT
stores and sorts each variable’s observations in the input data table.
By default, ALG=BIN.
- INTERP=LINEAR | SPLINE | STEP
-
specifies how to interpolate values within intervals in the distributions’ approximations. You can specify the following values:
- LINEAR
specifies that values within an interval be linearly interpolated between the endpoints.
- SPLINE
specifies that values within an interval be interpolated using a cubic spline function that is constrained to be monotonically increasing.
- STEP
specifies that values within an interval take the value of the left endpoint.
By default, INTERP=LINEAR.
- MAXITERS=integer
specifies the maximum number of iterations to perform when you are generating the distribution approximation function. By default, MAXITERS=10. For more information, see the section Adaptive Approximation of Marginal Distributions.
- REFINERES=integer
specifies the maximum number of intervals to add to the distribution function during each iteration of the empirical approximation. By default, REFINERES=1000.
- UNIFORM
uses the input data without transformation to fit the copula.
- METHOD=MLE | CAL
-
specifies the method to use to estimate parameters. You can specify the following values:
- CAL
specifies the calibration method that uses the correlations matrix (only Kendall’s tau is implemented in this procedure).
- MLE
specifies canonical maximum likelihood estimation (CMLE) or maximum likelihood estimation (MLE).
For the t copula, if METHOD=CAL, then the correlations matrix is estimated using the calibration method with Kendall’s tau and the degrees of freedom are estimated by the MLE. For the normal copula, only METHOD=MLE is supported and METHOD=CAL is ignored. By default for all copula types, METHOD=MLE.
- OUTPSEUDO=libref.data-table
-
specifies the output data table for saving the pseudo-samples with uniform marginal distributions. libref.data-table is a two-level name, where libref refers to the library, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.
The pseudo-samples are obtained by transforming the individual variables of the original data by using the empirical CDFs. The data table is not created if you omit this option.
- PLOTS<(global-plot-options)> <=(specific-plot-options)>
-
controls the plots of the input data table that are produced by the CCOPULA procedure. By default, PROC CCOPULA produces a scatter plot matrix for variables (that is, it displays a symmetric matrix plot that contains the variables specified in the VAR statement).
You can specify the following global-plot-options:
- KENDALL
creates Kendall plots. These plots compare observed quantiles of pairs of variables to quantiles of independent variables. If you specify this option together with the UNPACK option, PROC CCOPULA displays a Kendall plot of each applicable pair of distinct variables that you specify in the VAR statement. If you specify this option without the UNPACK option, PROC CCOPULA displays a scatter plot matrix; the lower triangular section of the matrix shows regular scatter plots between distinct pairs of variables that are specified in the VAR statement, and the upper triangular section of the matrix shows Kendall plots of corresponding pairs of variables.
- NSAMPLES=n
specifies the number points to plot when you create Kendall plots. For each pair of variables, n samples are randomly selected from the DATA= data table. The same samples are plotted for all the Kendall plots. By default, NSAMPLES=1000.
- NVAR=ALL | n
specifies the maximum number of variables that you specify in the VAR statement to display in the matrix plot. NVAR=ALL uses all variables that you specify in the VAR statement. By default, NVAR=5.
- RESOLUTION=n
specifies the number of bins to use for the X and Y axes when you create plots. By default, RESOLUTION=50.
- SCATTER
creates a scatter plot matrix of variables that you specify in the VAR statement. The maximum number of variables that you can display in the scatter plot matrix is 10.
- TAIL | CHI
creates tail dependence plots (chi plots). If you specify this option together with the UNPACK option, PROC CCOPULA displays a chi plot of each applicable pair of distinct variables that you specify in the VAR statement. If you specify this option without the UNPACK option, PROC CCOPULA displays a scatter plot matrix; the lower triangular section of the matrix shows regular scatter plots between distinct pairs of variables that are specified in the VAR statement, and the upper triangular section of the matrix shows chi plots of corresponding pairs of variables.
- UNPACK | UNPACKPANEL
creates scatter plots of pairs of variables. If you specify this option, PROC CCOPULA displays a scatter plot of each applicable pair of distinct variables that you specify in the VAR statement.
You can also specify the following specific-plot-options:
- DATATYPE=ORIGINAL | UNIFORM | BOTH
specifies the data type to be plotted. DATATYPE=ORIGINAL presents the data in their original marginal distribution. DATATYPE=UNIFORM shows the transformed data with a uniform marginal distribution. DATATYPE=BOTH plots both the original and uniform data types. When MARGINALS=UNIFORM, the transformation is omitted and the DATATYPE= option is ignored.
- NONE
suppresses all plots.
- STORE=SAS-item-store
specifies the item store that preserves the properties of the model and the fit results.