INVESTIGATION INTO PREDICTING MONTHLY INDEX OF INDUSTRIAL PRODUCTION
Data source: https://fred.stlouisfed.org
Using monthly observations on the United States Index of Industrial Production from January 1997 to May 2020, this study compared the results of various forecasting methods to predict the behavior of the index for June 2020 to May 2021., 12 monthly forecasts. An initial regression of the index in levels on an AR(1) lag indicated a unit root value of .98 with a highly significant t stat on the intercept, indicating a random walk with drift time series. A Dickey Fuller test for nonstationarity yielded a p-value of .14, not much evidence against Ho=unit root.
An initial graph of the index against T rendered a cubic-like shape.

A naive forecast for the twelve month lead periods, using last month’s realized value as the prediction returned a forecast mape of 1.98%.
A cubic model in the Trend ( T T^2 T^3) was employed as the starting point. The RMSE was 3.77594 with an R-Sq of .62. The final 12 predicted monthly values yielded a forecast MAPE of 6.9%. Adding a lagged industrial production term to the trend terms resulted in all four explanatory terms being highly significant. Heteroscedasticity was detected, so hetero-robust standard errors were computed, which rendered the T and Tsq terms insignificant. A regression of industrial production on a cubed trend and AR1 term resulted in an unsatisfactory forecast MAPE of 13.4%. First differences were then employed to attenuate the nonstationarity, and the differenced series was regressed on three lags, indicating significance for the first and third lags. An ARIMA model with first degree differencing and AR lags at p=(1,3) was run, showing a good fit. However the forecast error had a horrendous MAPE of 20.4%. A stepwise regression of the differenced series on 2 lags, also returned large forecast errors. Of note was no indication of seasonality, either monthly or quarterly indicated in either the Arima model or the regression model. This was corroborated by a regression of levels on a lag, a Trend, and monthly dummy variables. An F test on the 11 dummies for joint significance yielded a p-value of .67.
A third order autoregressive model was then fitted, without differencing. The twelve forecast values were generated, employing a macro that iterated through one step ahead forecast regressions using each current predicted value to populate the lagged values. This model generated a RMSE of .99213, and R_Sq of .9728, yielding a much better fit than the cubic trend. However, the forecast errors had a 10.1% MAPE, higher than the cubic trend. The third ar term was statistically insignificant, and, additionally, a Breusch-Pagan test indicated the presence of heteroskedasticity. A macro program for obtaining hetero-robust standard errors showed a 4 times increase in standard errors, driving the ar2 term into statistical insignificance. So a first order AR was employed, which returned an unsatisfactory forecast mape of 10%. However, the AR1 term might be used in other models.
Reducing the lead periods from 12 to 6 greatly improved the forecast perfromance of some of the univariate approaches: the curve fitting exercise (T, T^2, T^3) yielded a 2.7% mape. A 6 month arima (0,1,1) model returned a forecast mape of 1.6%, almost beating a 6 month naive forecast mape of 1.44%.
As well documented, univariate approaches are only useful for short term forecasting.

