The HPF Procedure
FORECAST Statement
FORECAST variable-list / options;
The FORECAST statement lists the numeric variables in the DATA= data set whose accumulated values represent time series to be modeled and forecast. The options specify which forecast model is to be used or how the forecast model is selected from several possible candidate models.
A data set variable can be specified in only one FORECAST statement. Any number of FORECAST statements can be used. The following options can be used with the FORECAST statement.
- ACCUMULATE=option
specifies how the data set observations are to be accumulated within each time period for the variables listed in the FORECAST statement. If the ACCUMULATE= option is not specified in the FORECAST statement, accumulation is determined by the ACCUMULATE= option of the ID statement. For more information, see the ID statement ACCUMULATE= option.
- ALPHA=number
specifies the significance level to use in computing the confidence limits of the forecast. The ALPHA= value must be between 0 and 1. The default is ALPHA=0.05, which produces 95% confidence intervals.
-
CRITERION=option
SELECT=option specifies the model selection criterion (statistic of fit) to be used to select from several candidate models. This option is often used in conjunction with the HOLDOUT= option. The CRITERION= option can also be specified as SELECT=. The default is CRITERION=RMSE.
- HOLDOUT=n
-
specifies the size of the holdout sample to be used for model selection. The holdout sample is a subset of actual time series that end at the last nonmissing observation. If the ACCUMULATE= option is specified, the holdout sample is based on the accumulated series. If the holdout sample is not specified, the full range of the actual time series is used for model selection.
For each candidate model specified, the holdout sample is excluded from the initial model fit and forecasts are made within the holdout sample time range. Then, for each candidate model specified, the statistic of fit specified by the CRITERION= option is computed by using only the observations in the holdout sample. Finally, the candidate model, which performs best in the holdout sample, based on this statistic, is selected to forecast the actual time series.
The HOLDOUT= option is used only to select the best forecasting model from a list of candidate models. After the best model is selected, the full range of the actual time series is used for subsequent model fitting and forecasting. It is possible that one model will outperform another model in the holdout sample but perform less well when the entire range of the actual series is used.
If the MODEL=BESTALL and HOLDOUT= options are used together, the last one hundred observations are used to determine whether the series is intermittent. If the series is determined not to be intermittent, holdout sample analysis is used to select the smoothing model.
- HOLDOUTPCT=number
specifies the size of the holdout sample as a percentage of the length of the time series. If HOLDOUT=5 and HOLDOUTPCT=10, the size of the holdout sample is
where T is the length of the time series with the beginning and ending missing values removed. The default is 100 (100%).
- INTERMITTENT=number
specifies a number greater than one which is used to determine whether or not a time series is intermittent. If the average demand interval is greater than this number, then the series is assumed to be intermittent. This option is used with the MODEL=BESTALL option. The default is INTERMITTENT=1.25.
- MEDIAN
specifies that the median forecast values are to be estimated. Forecasts can be based on the mean or median. By default the mean value is provided. If no transformation is applied to the actual series by using the TRANSFORM= option, the mean and median forecast values are identical.
- MODEL=model-name
-
specifies the forecasting model to be used to forecast the actual time series. A single model can be specified or a group of candidate models can be specified. If a group of models is specified, the model used to forecast the accumulated time series is selected based on the CRITERION= option and the HOLDOUT= option. The default is MODEL=BEST. The following forecasting models are provided:
- NONE
no forecast. The accumulated time series is appended with missing values in the OUT= data set. This option is particularly useful when the results stored in the OUT= data set are subsequently used in regression or autoregression analysis where forecasts of the independent variables are needed to forecast the dependent variable.
- SIMPLE
simple (single) exponential smoothing
- DOUBLE
double (Brown) exponential smoothing
- LINEAR
linear (Holt) exponential smoothing
- DAMPTREND
damped trend exponential smoothing
- ADDSEASONAL|SEASONAL
additive seasonal exponential smoothing
- MULTSEASONAL
multiplicative seasonal exponential smoothing
- WINTERS
Winters multiplicative method
- ADDWINTERS
Winters additive method
- BEST
best candidate smoothing model (SIMPLE, DOUBLE, LINEAR, DAMPTREND, ADDSEASONAL, WINTERS, ADDWINTERS)
- BESTN
best candidate nonseasonal smoothing model (SIMPLE, DOUBLE, LINEAR, DAMPTREND)
- BESTS
best candidate seasonal smoothing model (ADDSEASONAL, WINTERS, ADDWINTERS)
- IDM|CROSTON
intermittent demand model such as Croston’s method or average demand model. An intermittent time series is one whose values are mostly zero.
- BESTALL
best candidate model (IDM, BEST)
The BEST, BESTN, and BESTS options specify a group of models by which the HOLDOUT= option and CRITERION= option are used to select the model used to forecast the accumulated time series based on holdout sample analysis. Transformed versions of the preceding smoothing models can be specified using the TRANSFORM= option.
The BESTALL option specifies that if the series is intermittent, an intermittent demand model such as Croston’s method or average demand model (MODEL=IDM) is selected; otherwise, the best smoothing model is selected (MODEL=BEST). Intermittency is determined by the INTERMITTENT= option.
Chapter 15, Forecasting Process Details, describes the preceding smoothing models and intermittent models in greater detail.
- NBACKCAST=n
specifies the number of observations used to initialize the backcast states. The default is the entire series.
- REPLACEBACK
specifies that actual values excluded by the BACK= option be replaced with one-step-ahead forecasts in the OUT= data set.
- REPLACEMISSING
specifies that embedded missing actual values be replaced with one-step-ahead forecasts in the OUT= data set.
- SEASONTEST=option
-
specifies the options related to the seasonality test. This option is used with the MODEL=BEST and MODEL=BESTALL options.
The following options are provided:
- SEASONTEST=NONE
no test
- SEASONTEST=(SIGLEVEL=number)
significance probability value
Series with strong seasonality have small test probabilities. SEASONTEST=(SIGLEVEL=0) always implies seasonality. SEASONTEST=(SIGLEVEL=1) always implies no seasonality. The default is SEASONTEST=(SIGLEVEL=0.01).
- SETMISSING=option | number
specifies how to assign missing values (either actual or accumulated) in the accumulated time series for variables that are listed in the FORECAST statement. If the SETMISSING= option is not specified in the FORECAST statement, missing values are set based on the SETMISSING= option of the ID statement. For more information, see the SETMISSING= option in the ID statement.
- TRANSFORM=option
-
specifies the time series transformation to apply to the actual time series. The following transformations are provided:
- NONE
no transformation
- AUTO
automatically choose between NONE and LOG based on model selection criteria
- BOXCOX(n)
Box-Cox transformation with parameter number (n); where n is between –5 and 5
- LOG
logarithmic transformation
- LOGISTIC
logistic transformation
- SQRT
square-root transformation
When the TRANSFORM= option is specified, the time series must be strictly positive. After the time series is transformed, the model parameters are estimated by using the transformed series. The forecasts of the transformed series are then computed, and finally the transformed series forecasts are inverse transformed. The inverse transform produces either mean or median forecasts depending on whether the MEDIAN option is specified.
The TRANSFORM= option is not applicable when MODEL=IDM is specified. By default, TRANSFORM=NONE.
- USE=option
-
specifies which forecast values to append to the actual values in the OUT= and OUTSUM= data sets. You can specify the following options:
- PREDICT
The predicted values are appended to the actual values. This option is the default.
- LOWER
The lower confidence limit values are appended to the actual values.
- UPPER
The upper confidence limit values are appended to the actual values.
Thus, the USE= option enables the OUT= and OUTSUM= data sets to be used for worst-, best-, average-, and median-case decisions. By default, USE=PREDICT.
- ZEROMISS=option
specifies how to interpret beginning and ending zero values (either actual or accumulated) in the accumulated time series for variables that are listed in the FORECAST statement. If the ZEROMISS= option is not specified in the FORECAST statement, missing values are set based on the ZEROMISS= option in the ID statement. For more information, see the ZEROMISS=option in the ID statement.