The SPATIALREG Procedure
Getting Started: SPATIALREG Procedure
(View the complete code for this example.)
The SPATIALREG procedure is similar to other SAS regression model procedures, except that you usually need to provide a secondary data set (in the WMAT= option). The spatial weights matrix defines all pairwise spatial relationships and is the most vital component of a spatial regression model. For more information about how to create spatial weights matrix, see the section Specifying the Spatial Weights Matrix.
The following statements fit a SAR model:
proc spatialreg data=one Wmat=W; model y = x1 x2 / type=SAR; run;
The response variable y is continuous, and the data set W, which you specify in the WMAT= option, contains the spatial relationships among all spatial units in the data. In this case, W is either contiguity or weights. You specify the TYPE=SAR option to request a SAR model.
The following example illustrates PROC SPATIALREG by using a real-world data set. The data set CRIMEOH is taken from Anselin (1988) and can be found in the SAS/ETS Sample Library. This data set contains variables such as INCOME (household income, measured in $1000), HVALUE (housing value by $1000), and CRIME (number of crimes, including residential burglaries and vehicle thefts, measured per 1,000 households) in 49 neighborhoods in Columbus, Ohio. You want to examine how household income and housing value affect the number of crimes in the 49 neighborhoods of interest.
The first 10 observations in the CRIMEOH data set are shown in Figure 32.1.
Figure 32.1: Columbus Crime Data
| Obs | crime | income | hvalue | lat | lon |
|---|---|---|---|---|---|
| 1 | 18.802 | 21.232 | 44.567 | 35.62 | 42.38 |
| 2 | 32.388 | 4.477 | 33.200 | 36.50 | 40.52 |
| 3 | 38.426 | 11.337 | 37.125 | 36.71 | 38.71 |
| 4 | 0.178 | 8.438 | 75.000 | 33.36 | 38.41 |
| 5 | 15.726 | 19.531 | 80.467 | 38.80 | 44.07 |
| 6 | 30.627 | 15.956 | 26.350 | 39.82 | 41.18 |
| 7 | 50.732 | 11.252 | 23.225 | 40.01 | 38.00 |
| 8 | 26.067 | 16.029 | 28.750 | 43.75 | 39.28 |
| 9 | 48.585 | 9.873 | 18.000 | 39.61 | 34.91 |
| 10 | 34.001 | 13.598 | 96.400 | 47.61 | 36.42 |
The following SAS statements fit a linear regression model to the CRIMEOH data set:
proc spatialreg data=crimeoh; model crime = income hvalue / type=LINEAR; run;
The "Model Fit Summary" table, shown in Figure 32.2, lists several fit summary statistics about the model. By default, the SPATIALREG procedure uses the Newton-Raphson optimization technique. The maximum log-likelihood value is shown, in addition to two information measures, Akaike’s information criterion (AIC) and Schwarz’s Bayesian information criterion (SBC). AIC or SBC can be used for model selection. For a set of candidate models, the model with the smallest AIC or SBC is often preferred.
Figure 32.2: Fit Summary Statistics for a Linear Model
The parameter estimates of the model and their standard errors are shown in Figure 32.3. Based on the p-values, both INCOME and HVALUE are significant at the 0.05 level.
Figure 32.3: Parameter Estimates of the Linear Model
The following statements fit a SAR model to the CRIMEOH data set:
proc spatialreg data=crimeoh Wmat=crimeWmat NONORMALIZE; model crime = income hvalue / type=SAR; run;
The NONORMALIZE option requests that the spatial weights matrix that is specified in the CRIMEWMAT data set be used "as is" rather than be row-standardized. The "Model Fit Summary" table, shown in Figure 32.4, lists several fit summary statistics about the SAR model. For this model, the value of AIC is about 374.78—smaller than 382.75, which is the AIC value for the preceding linear model. Based on AIC, the SAR model is preferred.
Figure 32.4: Fit Summary Statistics for a SAR Model
| Model Fit Summary | |
|---|---|
| Dependent Variable | crime |
| Number of Observations | 49 |
| Data Set | WORK.CRIMEOH |
| Spatial Weights | WORK.CRIMEWMAT |
| Model | SAR |
| Log Likelihood | -182.38860 |
| Maximum Absolute Gradient | 2.7871E-7 |
| Number of Iterations | 16 |
| Optimization Method | Newton-Raphson |
| AIC | 374.77720 |
| SBC | 384.23630 |
The parameter estimates of the SAR model and their standard errors are shown in Figure 32.5. According to the p-values, both INCOME and HVALUE are significant at the 0.05 level. In addition, the spatial autoregressive coefficient is estimated to be about 0.431, with a p-value of 0.0005.
Figure 32.5: Parameter Estimates of the SAR Model
The following statements fit an SDM model. Unlike the previous SAR model, SDM accounts for exogenous interaction effects by introducing spatial lags of two explanatory variables—INCOME and HVALUE.
proc spatialreg data=crimeoh Wmat=crimeWmat NONORMALIZE; model crime = income hvalue / type=SAR; spatialeffects income hvalue; run;
The fit summary statistics for the SDM model are shown in Figure 32.6. Parameter estimates are provided in Figure 32.7.
Figure 32.6: Fit Summary Statistics for the SDM Model
| Model Fit Summary | |
|---|---|
| Dependent Variable | crime |
| Number of Observations | 49 |
| Data Set | WORK.CRIMEOH |
| Spatial Weights | WORK.CRIMEWMAT |
| Model | SDM |
| Log Likelihood | -181.39141 |
| Maximum Absolute Gradient | 5.44802E-8 |
| Number of Iterations | 16 |
| Optimization Method | Newton-Raphson |
| AIC | 376.78282 |
| SBC | 390.02556 |
Figure 32.7: Parameter Estimates for the SDM Model
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Intercept | 1 | 42.803457 | 13.924487 | 3.07 | 0.0021 |
| income | 1 | -0.914206 | 0.336439 | -2.72 | 0.0066 |
| hvalue | 1 | -0.293745 | 0.088857 | -3.31 | 0.0009 |
| W_income | 1 | -0.519640 | 0.594772 | -0.87 | 0.3823 |
| W_hvalue | 1 | 0.245716 | 0.176854 | 1.39 | 0.1647 |
| _rho | 1 | 0.426492 | 0.167492 | 2.55 | 0.0109 |
| _sigma2 | 1 | 91.779519 | 18.909222 | 4.85 | <.0001 |
The spatial autoregressive coefficient is estimated to be 0.426 with a p-value of 0.0109 based on an asymptotic t test. This result seems to suggest that there is a significantly positive spatial dependence in the number of crimes.
In the SPATIALREG procedure, the null hypothesis can also be tested against the alternative by using the likelihood ratio (LR) test, Lagrange multiplier (LM) test, and Wald test. For the LR test, the test statistic is equal to , where and are the log likelihoods for the linear regression model and SAR model, respectively. The likelihood ratio test is significant at the 0.05 level, providing strong evidence of spatial dependence in the data.