CSPATIALREG Procedure
PROC CSPATIALREG Statement
PROC CSPATIALREG options;
You can specify the following options in the PROC CSPATIALREG statement.
Data Table Options
You must specify the following option:
- DATA=libref.data-table
-
names the input data table for PROC CSPATIALREG to use. libref.data-table is a two-level name, where
- libref
refers to a collection of information that is defined in the LIBNAME statement and includes the
library, which includes a path to the data, and a session identifier, which defaults to the active session but which can be explicitly defined in the LIBNAME statement. For more information about libref, see the section Using CAS Sessions and CAS Engine Librefs.- data-table
specifies the name of the input data table.
This is the primary input data table, which contains dependent variables, explanatory variables, and so on.
For all models except a purely linear model, you must also specify the following option:
- WMAT=libref.data-table
-
specifies the secondary spatial weights data table, which you can use to construct the spatial weights matrix
. libref.data-table is a two-level name, where libref refers to the library, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.
Loosely speaking, the entries of
,
, define the amount of influence that a unit
has over a unit
. The entries
must be nonnegative and have zeros on the diagonal; that is,
and
, where
and n is the total number of spatial units in the data. Any nonzero diagonal elements
are replaced with 0. The spatial weights matrix can be asymmetric; that is, it is not necessary that
. For information about missing spatial weights in
, see the NONORMALIZE option.
The
matrix can take two different forms:
You can provide a full spatial weights matrix. In this case, the data table that you specify in the WMAT= option has n rows and
columns and must have a column for spatial ID variable.
You can specify the spatial weights matrix by using a compact form when appropriate. In this form, the number of observations in the data table that you specify in the WMAT= option should match the number of nonzero elements in the spatial weights matrix. Moreover, the number of columns in this data table should be three. The first two columns contain the row and column indices for nonzero entries in the spatial weights matrix, and the third column contains the nonzero entries in the spatial weights matrix. If you use the compact form for the spatial weights matrix, you must include a SPATIALID statement to match observations in the data tables that you specify in the DATA= and WMAT= options. For more information about the SPATIALID statement, see the section SPATIALID Statement. For more information about the compact representation of the spatial weights matrix, see the section Compact Representation of the Spatial Weights Matrix.
For all models except a purely linear model, you can also specify the following option:
- NONORMALIZE
suppresses the row standardization of the spatial weights matrix that you specify in the WMAT= option. By default, the spatial weights matrix is row-standardized; that is, the spatial weights matrix has unit row sum. If you specify the NONORMALIZE option, spatial weights are used "as is" except for
, which is always treated as 0. This implies that an entry
in the
matrix cannot be missing for
if you specify this option. If you do not specify this option, missing spatial weights are replaced with zeros.
Approximation Control Option
For the SAR, SDM, SEM, SDEM, and CAR models, you can specify the following options:
- APPROXIMATION=(approx-options)
-
specifies options that are related to approximating the Jacobian, as described in the section Approximations to the Jacobian. To invoke approximation, you must specify one or more of the following approx-options:
- CHEBYSHEV | TAYLOR
specifies the approximation method. By default, Chebyshev approximation is used. The Taylor approximation is used only if you specify the TAYLOR option.
- NMC=number
specifies a positive integer as the number of standard random normal draws for Monte Carlo simulation. By default, NMC=100.
- ORDER=number
specifies a positive integer as the order of series in Taylor or Chebyshev approximation. If you specify Taylor approximation, ORDER=50 by default. If you specify Chebyshev approximation, ORDER=5 by default.
- SEED=number
specifies an integer seed in the range 1 to
for the random number generator that is used for Monte Carlo simulation. Specifying a seed enables you to reproduce your analysis. By default, SEED=1.
Impact Estimation Control Option
For impact estimation, you can specify the following options:
- IMPACT(impactest-option)
-
specifies options that are related to impact estimation, as described in the section Impact Estimation. To invoke impact estimation, you must specify one or more of the following impactest-options:
- NMC=number
specifies a positive integer as the number of iterations for Monte Carlo simulation. By default, NMC=1000.
- ORDER=number
specifies a positive integer as the highest order of the Neumann series for approximation. By default, ORDER=50.
- SEED=number
specifies an integer seed in the range 1 to
for the random number generator that is used for Monte Carlo simulation. Specifying a seed enables you to reproduce your analysis. By default, SEED=1.
You can use the IMPACT option to summarize the average direct impacts, the average indirect impacts, and the average total impacts of explanatory variables in the model.
Printing Options
You can specify the following options in either the PROC CSPATIALREG statement or the MODEL statement:
- CORRB
- COVB
- NOPRINT
- PRINTALL
- PRINTINTERNALNAMES
- PRINTTIMING
Estimation Control Options
You can specify the following options in either the PROC CSPATIALREG statement or the MODEL statement:
- COVEST=HESSIAN | OP | QML
-
specifies the type of covariance matrix for the parameter estimates. You can specify the following types:
- HESSIAN
specifies the covariance from the Hessian matrix.
- OP
specifies the covariance from the outer product matrix.
- QML
specifies the covariance from the outer product and Hessian matrices.
By default, COVEST=HESSIAN. The quasi-maximum-likelihood estimates are computed using COVEST=QML. For all models except the linear and SLX models, only COVEST=HESSIAN is supported.
- TYPE=ALL | AUTO | CAR | LINEAR | SAC | SAR | SARMA | SEM | SMA
-
specifies the type of model to be fit. You can specify the following values:
- ALL
fits all models.
- AUTO
fits a subset of selected models.
- CAR
fits a conditional autoregressive model.
- LINEAR
fits a linear model.
- SAC
fits a spatial autoregressive confused model.
- SAR
fits a spatial autoregressive model.
- SARMA
fits a spatial autoregressive moving average model.
- SEM
fits a spatial error model.
- SMA
fits a spatial moving average model.
By default, TYPE=SAR. For more information about the TYPE=ALL or TYPE=AUTO option, see the section Multiple Model Comparison.
If you specify covariates in the SPATIALEFFECT statement, then their spatial lag is added to the model. In that case, you can specify the following TYPE= option values to get the model that you want:
- TYPE=CAR
fits the CAR2 model.
- TYPE=LINEAR
fits the SLX model.
- TYPE=SAC
fits the SDAC model.
- TYPE=SAR
fits the SDM.
- TYPE=SARMA
fits the SDARMA model.
- TYPE=SEM
fits the SDEM.
- TYPE=SMA
fits the SDMA model.
Optimization Control Options
PROC CSPATIALREG uses the nonlinear optimization (NLO) subsystem to perform nonlinear optimization tasks. You can specify the following options in either the PROC CSPATIALREG statement or the MODEL statement:
-
MAXFUNC=i
MAXFU=i -
specifies the maximum number of function calls in the optimization process. By default, MAXFUNC=1000.
The optimization can terminate only after completing a full iteration. Therefore, the number of function calls that are actually used can exceed the number of calls that you specify in this option.
-
MAXITER=i
MAXIT=i specifies the maximum number of iterations in the optimization process. By default, MAXITER=200.
- MAXTIME=r
specifies an upper limit of r seconds of CPU time for the optimization process. The default value is the largest floating-point double representation available on your computer. The time that you specify in this option is checked only once at the end of each iteration. Therefore, the actual run time can be much longer than r. The actual run time includes the remaining time that is needed to finish the iteration and the time that is needed to generate the output of the results.
- METHOD=CONGRA | DBLDOG | NEWRAP | NONE | NRRIDG | QUANEW | TRUREG
-
specifies the iterative minimization method to use. You can specify the following methods:
- CONGRA
specifies the conjugate-gradient method.
- DBLDOG
specifies the double-dogleg method.
- NEWRAP
specifies the Newton-Raphson method.
- NONE
specifies that optimization not be performed.
- NRRIDG
specifies the Newton-Raphson ridge method.
- QUANEW
specifies the quasi-Newton method.
- TRUREG
specifies the trust region method.
By default, METHOD=NEWRAP.