The MODEL statement specifies the regression model, the error structure that is assumed for the regression residuals, and the estimation technique to be used. The response variable (response) on the left side of the equal sign is regressed on the independent variables (effects), which are listed after the equal sign. You can specify any number of MODEL statements. For each MODEL statement, you can specify only one response.
You can label models. Model labels are used in the printed output to identify the results for different models. If you do not specify a label, the model is referred to by numerical order wherever necessary. You can label models by prefixing the MODEL statement by a label followed by a colon as follows:
label: MODEL …;
The MODEL statement supports many options, some more specific than others. Table 3 summarizes the options available in the MODEL statement. These are subsequently discussed in detail in the order in which they are presented in the table.
You can specify the following options after a slash (/).
Estimation Technique Options
These options specify the assumed error structure and estimation method. You can specify more than one option, in which case the analysis is repeated for each. The default is FIXONE (one-way fixed effects).
requests Amemiya-MaCurdy estimation for a model that has correlated individual (cross-sectional) effects. You specify the correlated effects by using the CORRELATED statement.
BTWNG
estimates a between-groups model.
BTWNT
estimates a between-time-periods model.
DYNDIFF
estimates a dynamic panel by using the generalized method of moments (GMM) on equations that are formed by first differencing.
DYNSYS
estimates a dynamic panel by using GMM on the system that combines first-differenced equations and level equations.
FDONE
estimates a one-way model by using first-differenced methods.
FIXONE
estimates a one-way fixed-effects model that corresponds to cross-sectional effects only.
FIXONETIME
estimates a one-way fixed-effects model that corresponds to time effects only.
FIXTWO
estimates a two-way fixed-effects model.
HTAYLOR
requests Hausman-Taylor estimation for a model that has correlated individual (cross-sectional) effects. You specify the correlated effects by using the CORRELATED statement.
IVBTWNG
requests instrumental variables regression for between-groups estimation for a model that has endogenous effects. You specify the endogenous effects by using the ENDOGENOUS statement, and you specify external instruments by using the INSTRUMENTS statement. You can choose to use two-stage least squares or GMM to estimate the model.
IVFIXONE
requests instrumental variables regression for one-way fixed-effects estimation for a model that has endogenous effects. You specify the endogenous effects by using the ENDOGENOUS statement, and you specify external instruments by using the INSTRUMENTS statement. You can choose to use two-stage least squares or GMM to estimate the model.
IVFIXONETIME
requests instrumental variables regression for one-way time fixed-effects estimation for a model that has endogenous effects. You specify the endogenous effects by using the ENDOGENOUS statement, and you specify external instruments by using the INSTRUMENTS statement. You can choose to use two-stage least squares or GMM to estimate the model.
IVFIXTWO
requests instrumental variables regression for two-way fixed-effects estimation for a model that has endogenous effects. You specify the endogenous effects by using the ENDOGENOUS statement, and you specify external instruments by using the INSTRUMENTS statement. You can choose to use two-stage least squares or GMM to estimate the model.
IVG2SLS
requests instrumental variables regression for generalized two-stage least squares estimation for a model that has endogenous effects. You specify the endogenous effects by using the ENDOGENOUS statement, and you specify external instruments by using the INSTRUMENTS statement. You can choose to use two-stage least squares or GMM to estimate the model.
IVPOOLED
requests instrumental variables regression for pooled regression. You specify the endogenous effects by using the ENDOGENOUS statement, and you specify external instruments by using the INSTRUMENTS statement. You can choose to use two-stage least squares or GMM to estimate the model.
IVRANONE
requests instrumental variables regression for one-way random-effects estimation for a model that has endogenous effects. You specify the endogenous effects by using the ENDOGENOUS statement, and you specify external instruments by using the INSTRUMENTS statement. You can choose to use two-stage least squares or GMM to estimate the model.
IVRANTWO
requests instrumental variables regression for two-way random-effects estimation for a model that has endogenous effects. You specify the endogenous effects by using the ENDOGENOUS statement, and you specify external instruments by using the INSTRUMENTS statement. You can choose to use two-stage least squares or GMM to estimate the model.
POOLED
estimates a pooled (OLS) model.
RANONE
estimates a one-way random-effects model.
RANTWO
estimates a two-way random-effects model.
Estimation Control Options
These options define parameters that control the estimation and can be specific to the chosen technique (for example, how to estimate variance components in a random-effects model).
BIASCORRECTED
computes the bias-corrected covariance matrix of the two-step generalized method of moments (GMM) estimator. This statement is valid only when you perform either dynamic panel estimation or instrumental variables regression.
CLUSTER
specifies the cluster correction for the covariance matrix. You can specify this option when you specify HCCME=0, 1, 2, or 3.
GMM=ONESTEP | TWOSTEP
specifies the number of GMM stages to use in the estimation. This statement is valid only when you perform either dynamic panel estimation or instrumental variables regression.
You can specify the following values:
ONESTEP
uses one-step GMM, which is computationally simple but dependent on model assumptions.
TWOSTEP
uses two-step GMM, which is computationally intensive but more robust to violations of model assumptions.
By default, GMM=ONESTEP.
The following statements perform two-step GMM in a dynamic panel model, which is requested by the DYNDIFF option in the MODEL statement. You can use similar statements for a dynamic panel estimation that uses a full system of difference and level equations by specifying the DYNSYS option instead of the DYNDIFF option in the MODEL statement. The endogenous variable is Sales, and the GMM-style instruments is the predetermined variable Price. Both sets of instruments are included only in the difference equations.
The following statements perform two-step GMM in an instrumental variables model. The endogenous variables are X2 and Z2, and the external instruments are G1 and G2.
proc cpaneldata=mylib.a;
id firm year;
model Y = X1 X2 Z1 Z2 / ivfixone GMM = TWOSTEP;
endogenous X2 Z2;
instruments G1 G2;run;
The IVFIXONE option in the MODEL statement requests IV one-way fixed-effects estimation. You can request other types of IV estimation by specifying IVBTWNG, IVFIXONETIME, IVFIXTWO, IVG2SLS, IVPOOLED, IVRANONE, or IVRANTWO instead of IVFIXONE in the MODEL statement.
HAC <(options)>
specifies the heteroscedasticity- and autocorrelation-consistent (HAC) covariance matrix estimator. This option is not available for between-groups or between-time-periods models, and cannot be combined with the HCCME= option.
You can specify the following options within parentheses and separated by spaces:
ADJUSTDF
makes a small-sample adjustment to the degrees of freedom in the covariance calculation.
BANDWIDTH=number | method
specifies the fixed bandwidth value or bandwidth selection method to be used in the kernel function. You can specify either a fixed value (number) or one of the methods listed after number.
number
specifies a fixed value of the bandwidth parameter.
ANDREWS | AN
specifies the Andrews (1991) bandwidth selection method.
NEWEYWEST<(C=number)> | NW <(C=number)>
specifies the bandwidth selection method of Newey and West (1994) You can also specify C=number to calculate the lag selection parameter; by default, C=12.
SAMPLESIZE<(options)> | SS<(options)>
calculates the bandwidth according to the following equation based on the sample size,
where b is the bandwidth parameter; T is the sample size; and , r, and c are values that you specify using the following options within parentheses and separated by spaces:
CONSTANT=number
specifies the constant c in the equation. By default, CONSTANT=0.5.
GAMMA=number
specifies the coefficient in the equation. By default, GAMMA=0.75.
INTSS
specifies that the bandwidth parameter be an integer; that is, , where denotes the largest integer less than or equal to x.
RATE=number
specifies the growth rate r in the equation. By default, RATE=0.3333.
By default, BANDWIDTH=ANDREWS.
KERNEL=BARTLETT | PARZEN | QS | TH | TRUNCATED
specifies the type of kernel function. You can specify the following values:
BARTLETT
specifies the Bartlett kernel function.
PARZEN
specifies the Parzen kernel function.
QS
specifies the quadratic spectral kernel function.
TH
specifies the Tukey-Hanning kernel function.
TRUNCATED
specifies the truncated kernel function.
By default, KERNEL=QS.
KERNELLB=number
specifies the lower bound of the kernel weight value. Any kernel weight less than number is regarded as 0, which accelerates the calculation in large samples, especially for the quadratic spectral kernel function. By default, KERNELLB=0.
PREWHITENING
requests prewhitening in the covariance calculation.
The following statements perform one-way fixed-effects estimation by using the Newey-West bandwidth and a Bartlett kernel based on a prewhitened gradient vector:
proc cpaneldata=mylib.a;
id firm year;
model Y = X1 X2 Z1 Z2 / ivfixone gmm=twostep
hac(bandwidth=neweywest(c=12) kernel=bartlett prewhitening);run;
HCCME=0 | 1 | 2 | 3
specifies the type of adjustment to the HCCME covariance matrix. You can specify the values 0 through 3, which indicate the type of covariance adjustment.
NOINT
suppresses the intercept parameter from the model.
ROBUST
computes the robust covariance matrix. This option can be applied to dynamic panel models, instrumental variables regression models, and static panel models. For dynamic panel models and instrumental variables regression models, the option provides heteroscedasticity-corrected standard errors according to the GMM estimation method. The following statements perform IV fixed-effects regression with robust standard errors:
proc cpaneldata=mylib.a;
id firm year;
model Y = X1 X2 Z1 Z2 / ivfixone gmm=twostep robust;
endogenous X2 Z2;
instruments G1 G2;run;
For static panel models, this option produces cross-sectional clustered standard errors when HCCME=0. The following statements perform one-way fixed-effects regression with clustered standard errors:
proc cpaneldata=mylib.a;
id firm year;
model Y = X1 X2 Z1 Z2 / fixone robust;run;
VCOMP=FB | NL | SA | WH | WK
specifies the type of variance component estimate to use. You can specify the following values:
FB
uses the Fuller-Battese method.
NL
uses the Nerlove method.
SA
uses the Swamy-Arora method.
WH
uses the Wallace-Hussain method.
WK
uses the Wansbeek-Kapteyn method.
By default, VCOMP=SA.
Dynamic Panel Estimation Options
These options are specific to dynamic panel estimation, which you obtain by specifying the DYNDIFF or DYNSYS option in the MODEL statement.
ARTESTS=integer
specifies the maximum order of the test for the presence of autoregression (AR) effects. By default, ARTESTS=2.
DLAGS=integer
specifies the number of dependent-variable lags to use as regressors. By default, DLAGS=1.
GINV=G2 | G4
specifies what type of generalized inverse to use. You can specify the following values:
G2
uses the G2 generalized inverse.
G4
uses the G4 generalized inverse.
The difference between G2 and G4 becomes evident when you invert singular matrices. The G2 generalized inverse drops rows and columns from singular matrices to produce a viable inverse. The G4 inverse, on the other hand, is the Moore-Penrose generalized inverse, which averages the variance effects between collinear rows. The G4 inverse is usually more stable, but it is computationally intensive. By default, GINV=G2. If you have trouble reproducing published results, often the solution is to switch to GINV=G4.
MAXBAND=integer
if specified, sets the maximum number of GMM-style instruments per observation, for each variable. Because the number of GMM-style instruments grows quadratically with the number of time periods, this option makes estimation more feasible when you have many time periods.
Printed Output Options
These options alter how results are presented.
CORRB
prints the matrix of estimated correlations between the parameter estimates.
COVB
prints the matrix of estimated
covariances between the parameter estimates.
NOLABEL
suppresses variable labels from the printed output.
PRINTFIXED
estimates and prints the fixed effects in models where they would normally be absorbed within the estimation.