CSPATIALREG Procedure

PROC CSPATIALREG Statement

  • PROC CSPATIALREG options;

You can specify the following options in the PROC CSPATIALREG statement.

Data Table Options

You must specify the following option:

DATA=libref.data-table

names the input data table for PROC CSPATIALREG to use. libref.data-table is a two-level name, where

libref

refers to a collection of information that is defined in the LIBNAME statement and includes the library, which includes a path to the data, and a session identifier, which defaults to the active session but which can be explicitly defined in the LIBNAME statement. For more information about libref, see the section Using CAS Sessions and CAS Engine Librefs.

data-table

specifies the name of the input data table.

This is the primary input data table, which contains dependent variables, explanatory variables, and so on.

For all models except a purely linear model, you must also specify the following option:

WMAT=libref.data-table

specifies the secondary spatial weights data table, which you can use to construct the spatial weights matrix bold upper W. libref.data-table is a two-level name, where libref refers to the library, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.

Loosely speaking, the entries of bold upper W, w left-parenthesis bold s Subscript i Baseline comma bold s Subscript j Baseline right-parenthesis, define the amount of influence that a unit bold s Subscript j has over a unit bold s Subscript i. The entries w left-parenthesis bold s Subscript i Baseline comma bold s Subscript j Baseline right-parenthesis must be nonnegative and have zeros on the diagonal; that is, w left-parenthesis bold s Subscript i Baseline comma bold s Subscript j Baseline right-parenthesis greater-than-or-equal-to 0 and w left-parenthesis bold s Subscript i Baseline comma bold s Subscript i Baseline right-parenthesis equals 0, where and n is the total number of spatial units in the data. Any nonzero diagonal elements w left-parenthesis bold s Subscript i Baseline comma bold s Subscript i Baseline right-parenthesis are replaced with 0. The spatial weights matrix can be asymmetric; that is, it is not necessary that w left-parenthesis bold s Subscript i Baseline comma bold s Subscript j Baseline right-parenthesis equals w left-parenthesis bold s Subscript j Baseline comma bold s Subscript i Baseline right-parenthesis. For information about missing spatial weights in bold upper W, see the NONORMALIZE option.

The bold upper W matrix can take two different forms:

  • You can provide a full spatial weights matrix. In this case, the data table that you specify in the WMAT= option has n rows and n plus 1 columns and must have a column for spatial ID variable.

  • You can specify the spatial weights matrix by using a compact form when appropriate. In this form, the number of observations in the data table that you specify in the WMAT= option should match the number of nonzero elements in the spatial weights matrix. Moreover, the number of columns in this data table should be three. The first two columns contain the row and column indices for nonzero entries in the spatial weights matrix, and the third column contains the nonzero entries in the spatial weights matrix. If you use the compact form for the spatial weights matrix, you must include a SPATIALID statement to match observations in the data tables that you specify in the DATA= and WMAT= options. For more information about the SPATIALID statement, see the section SPATIALID Statement. For more information about the compact representation of the spatial weights matrix, see the section Compact Representation of the Spatial Weights Matrix.

For all models except a purely linear model, you can also specify the following option:

NONORMALIZE

suppresses the row standardization of the spatial weights matrix that you specify in the WMAT= option. By default, the spatial weights matrix is row-standardized; that is, the spatial weights matrix has unit row sum. If you specify the NONORMALIZE option, spatial weights are used "as is" except for w left-parenthesis bold s Subscript i Baseline comma bold s Subscript i Baseline right-parenthesis, which is always treated as 0. This implies that an entry w left-parenthesis bold s Subscript i Baseline comma bold s Subscript j Baseline right-parenthesis in the bold upper W matrix cannot be missing for i not-equals j if you specify this option. If you do not specify this option, missing spatial weights are replaced with zeros.

Approximation Control Option

For the SAR, SDM, SEM, SDEM, and CAR models, you can specify the following options:

APPROXIMATION=(approx-options)

specifies options that are related to approximating the Jacobian, as described in the section Approximations to the Jacobian. To invoke approximation, you must specify one or more of the following approx-options:

CHEBYSHEV | TAYLOR

specifies the approximation method. By default, Chebyshev approximation is used. The Taylor approximation is used only if you specify the TAYLOR option.

NMC=number

specifies a positive integer as the number of standard random normal draws for Monte Carlo simulation. By default, NMC=100.

ORDER=number

specifies a positive integer as the order of series in Taylor or Chebyshev approximation. If you specify Taylor approximation, ORDER=50 by default. If you specify Chebyshev approximation, ORDER=5 by default.

SEED=number

specifies an integer seed in the range 1 to 2 Superscript 31 Baseline minus 1 for the random number generator that is used for Monte Carlo simulation. Specifying a seed enables you to reproduce your analysis. By default, SEED=1.

Impact Estimation Control Option

For impact estimation, you can specify the following options:

IMPACT(impactest-option)

specifies options that are related to impact estimation, as described in the section Impact Estimation. To invoke impact estimation, you must specify one or more of the following impactest-options:

NMC=number

specifies a positive integer as the number of iterations for Monte Carlo simulation. By default, NMC=1000.

ORDER=number

specifies a positive integer as the highest order of the Neumann series for approximation. By default, ORDER=50.

SEED=number

specifies an integer seed in the range 1 to 2 Superscript 31 Baseline minus 1 for the random number generator that is used for Monte Carlo simulation. Specifying a seed enables you to reproduce your analysis. By default, SEED=1.

You can use the IMPACT option to summarize the average direct impacts, the average indirect impacts, and the average total impacts of explanatory variables in the model.

Printing Options

You can specify the following options in either the PROC CSPATIALREG statement or the MODEL statement:

CORRB

prints the correlation matrix of the parameter estimates.

COVB

prints the covariance matrix of the parameter estimates.

NOPRINT

suppresses all printed output.

PRINTALL

requests all printing options.

PRINTINTERNALNAMES

prints internal names that are assigned to parameters.

PRINTTIMING

prints a timing report.

Estimation Control Options

You can specify the following options in either the PROC CSPATIALREG statement or the MODEL statement:

COVEST=HESSIAN | OP | QML

specifies the type of covariance matrix for the parameter estimates. You can specify the following types:

HESSIAN

specifies the covariance from the Hessian matrix.

OP

specifies the covariance from the outer product matrix.

QML

specifies the covariance from the outer product and Hessian matrices.

By default, COVEST=HESSIAN. The quasi-maximum-likelihood estimates are computed using COVEST=QML. For all models except the linear and SLX models, only COVEST=HESSIAN is supported.

TYPE=ALL | AUTO | CAR | LINEAR | SAC | SAR | SARMA | SEM | SMA

specifies the type of model to be fit. You can specify the following values:

ALL

fits all models.

AUTO

fits a subset of selected models.

CAR

fits a conditional autoregressive model.

LINEAR

fits a linear model.

SAC

fits a spatial autoregressive confused model.

SAR

fits a spatial autoregressive model.

SARMA

fits a spatial autoregressive moving average model.

SEM

fits a spatial error model.

SMA

fits a spatial moving average model.

By default, TYPE=SAR. For more information about the TYPE=ALL or TYPE=AUTO option, see the section Multiple Model Comparison.

If you specify covariates in the SPATIALEFFECT statement, then their spatial lag is added to the model. In that case, you can specify the following TYPE= option values to get the model that you want:

TYPE=CAR

fits the CAR2 model.

TYPE=LINEAR

fits the SLX model.

TYPE=SAC

fits the SDAC model.

TYPE=SAR

fits the SDM.

TYPE=SARMA

fits the SDARMA model.

TYPE=SEM

fits the SDEM.

TYPE=SMA

fits the SDMA model.

Optimization Control Options

PROC CSPATIALREG uses the nonlinear optimization (NLO) subsystem to perform nonlinear optimization tasks. You can specify the following options in either the PROC CSPATIALREG statement or the MODEL statement:

MAXFUNC=i
MAXFU=i

specifies the maximum number of function calls in the optimization process. By default, MAXFUNC=1000.

The optimization can terminate only after completing a full iteration. Therefore, the number of function calls that are actually used can exceed the number of calls that you specify in this option.

MAXITER=i
MAXIT=i

specifies the maximum number of iterations in the optimization process. By default, MAXITER=200.

MAXTIME=r

specifies an upper limit of r seconds of CPU time for the optimization process. The default value is the largest floating-point double representation available on your computer. The time that you specify in this option is checked only once at the end of each iteration. Therefore, the actual run time can be much longer than r. The actual run time includes the remaining time that is needed to finish the iteration and the time that is needed to generate the output of the results.

METHOD=CONGRA | DBLDOG | NEWRAP | NONE | NRRIDG | QUANEW | TRUREG

specifies the iterative minimization method to use. You can specify the following methods:

CONGRA

specifies the conjugate-gradient method.

DBLDOG

specifies the double-dogleg method.

NEWRAP

specifies the Newton-Raphson method.

NONE

specifies that optimization not be performed.

NRRIDG

specifies the Newton-Raphson ridge method.

QUANEW

specifies the quasi-Newton method.

TRUREG

specifies the trust region method.

By default, METHOD=NEWRAP.

Last updated: November 24, 2025