ATTACHMENT 7 Forecasting Technical Assessment FINAL.pdf

PDF 1 MB Posted

Attached to
Microsimulation modeling and analytical support services Federal contract opportunity
Solicitation number
12-3198-25-R-0002
Issued by
Department of Agriculture Food and Nutrition Service

About this file

This document is a technical review memorandum assessing the Supplemental Nutrition Assistance Program (SNAP) participation forecasting model, prepared by Mathematica for the Food and Nutrition Service (FNS). The study critically examines the current SNAP forecasting methodology, which has been in use since the early 1990s and relies primarily on unemployment levels as a key predictor of SNAP participation. The researchers identified significant changes in the relationship between unemployment and SNAP participation over time, particularly during economic shifts like the Great Recession, which highlighted limitations in the existing forecasting approach.

The analysis explored multiple alternative modeling techniques, including ARIMA (Autoregressive Integrated Moving Average) models, vector autoregression, and machine learning algorithms. Key findings recommend re-specifying the current model as an ARIMA model using first differences of quarterly SNAP data, which demonstrated substantially improved forecasting accuracy. The best-performing model, ARIMA(4,1,0), reduced the Mean Absolute Percent Error (MAPE) for five-to-eight quarter-ahead forecasts by more than half compared to the current model, potentially improving SNAP participation predictions from 4.9 million to 1.6 million participants. The study provides detailed statistical analysis and recommendations for enhancing the forecasting methodology to better capture complex economic and policy dynamics affecting SNAP participation.

View the file

Other files for this federal contract opportunity

Other files attached to Microsimulation modeling and analytical support services, newest first.
File Type Posted
12-3198-25-R-0002 A0004 Questions and Responses.pdf PDF
12-3198-25-R-0002 A0004.pdf PDF
12-3198-25-R-0002 A0003.pdf PDF
Attachment 9a IDIQ Pricing Schedule Revised.xlsx XLSX spreadsheet
12-3198-25-R-0002 A0002..pdf PDF
Attachment 9b Deliverables Schedule Revised.xlsx XLSX spreadsheet
Amendment 0001_Attachment 9a IDIQ Pricing Schedule Rev. 1 _07-31-2025.xlsx XLSX spreadsheet
Amendment 0001_Attachment 9b Deliverables Schedule_Rev. 1 _07-31-2025.xlsx XLSX spreadsheet
Amendment 0001_RFP 12-3198-25-R-0002__07-31-2025.pdf PDF
ATTACHMENT 1 2020-MATH-SIPP-TWP.pdf PDF
Attachment 9a IDIQ Pricing Schedule.xlsx XLSX spreadsheet
Attachment 10a IDIQ Performance Work statement PWS.pdf PDF
ATTACHMENT 4 FY2023-Characteristics-Report.pdf PDF
ATTACHMENT 8 Rules of Thumb Using 2012 QC Data.pdf PDF
10b Task Order Performance Work Statement PWS.pdf PDF
12-3198-25-R-0002.pdf PDF
ATTACHMENT 2 2020 Programmers Guide-v2.pdf PDF
ATTACHMENT 3 FY 2023 QC Tech Doc.pdf PDF
ATTACHMENT 5 Trends-FY-2020-and-FY-2022.pdf PDF
ATTACHMENT 6 SNAP-Participation-Rates-2022.pdf PDF
Attachment 9b Deliverables Schedule.xlsx XLSX spreadsheet
ATTACHMENT 11 QASP.pdf PDF
Show all 22

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

ATTACHMENT 7

INTENTIONALLY BLANK

Memo

To: INTENTIONALLY BLANK

From: INTENTIONALLY BLANK

Date: 12/6/2019

Subject: Technical Review of the SNAP Participation Forecasting Model

Contract Number: AG-3198-B-16-0001/12-3198-19-F-0014

Task Number: 50787.112

Memo Number: 053

A. Introduction

The Food and Nutrition Service (FNS) forecasts Supplemental Nutrition Assistance Program

(SNAP) participation for budgeting purposes. At the end of the first quarter of each fiscal year

(FY), FNS submits estimates of future SNAP participation to the Office of Management and

Budget (OMB). These estimates are based on data through the end of the previous FY. The FNS budgeting cycle requires that FY SNAP participation forecasts be made two years ahead of the current FY. In addition, FNS requires forecasts of SNAP participation for one year ahead of the current FY to assess the adequacy of the budget for the upcoming year and for the last two quarters within the current FY to determine whether a request for supplemental appropriation is necessary.

The current SNAP forecasting model, which has been used since the early 1990s, was originally developed using data from the 1970s and 1980s. A core component of the current model is the relationship between SNAP participation and unemployment, because unemployment is considered to be a leading indicator of future SNAP participation. As part of Mathematica’s previous review of the SNAP forecasting model, Schochet and Needels (2000) noted that this underlying relationship has changed over time. In particular, although the level of SNAP participation and unemployment tracked closely in the 1980s, the relationship changed in the

1990s. SNAP participation increased sharply in the early 1990s to mid-1990s and dropped substantially in the late 1990s. However, the level of unemployment did not fluctuate as sharply during this time. Moreover, policy changes, such as passage of the Personal Responsibility and

Work Opportunity Reconciliation Act (PRWORA) in 1996, also contributed to changes in SNAP participation.

1100 1st Street, NE, 12th Floor, Washington, DC 20002-4221 • (202) 484-9220 phone (202) 863-1763 fax • mathematica-mpr.com

An Affirmative Action/Equal Opportunity Employer

To:

From:

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

Since Mathematica’s last review of the SNAP forecasting model in 2000, SNAP participation patterns have changed again. After declining in the late 1990s and early 2000s, SNAP participation began steadily increasing during and after the Great Recession in FY 2009 and continued to increase years after the subsequent economic recovery. The apparent disconnection between unemployment levels and SNAP participation again raised questions about the nature of the underlying relationship assumed in the current forecasting model. In addition, policies enacted over the last two decades may have impacted SNAP participation. While implementation of PRWORA was only beginning during the last few years covered by Mathematica’s previous technical assessment, the main components that affected SNAP eligibility (and thus participation) were fully implemented as of FY 2003. Other relevant legislation, such as the

Farm Security and Rural Investment Act of 2002, the Food, Conservation, and Energy Act of

2008 (the Farm Bill), and the American Recovery and Reinvestment Act of 2009 (ARRA), as well as State broad-based categorical eligibility (BBCE) policies, took effect. Taken together, the changes that have occurred over the last 20 years suggest that a new review of the adequacy of the forecasting model is appropriate.

In the remainder of this memorandum, we provide a thorough review of the current FNS forecasting model, assess its recent performance, and provide recommendations on ways to improve the model for future forecasts. To accomplish this, we focus primarily on forecasting performance over the past three decades (1990s, 2000s, and 2010s). Accordingly, our analysis also addresses the following research questions posed by FNS:

• Is the current time series model, a first-order autoregressive (AR(1)) model, appropriate for use with the underlying data, which includes historical SNAP participation levels and projected unemployment levels? Would another time series model, such as the autoregressive integrated moving average (ARIMA) model, be better-suited to the data?

• What time series model would best fit the data? Should the data be first differenced to ensure that the data are stationary (that is, that the series stays near its historical mean over time)? Should the model employ a distributed lag for the independent variable (the unemployment level)? If so, what should be the length of the lag?

• Are there other data or variables that should be included in the model to improve its predictive abilities?

• How has the SNAP forecasting model performed in the past—both in the short term and the long term? How does that compare to other econometric models that are used for similar purposes in the public and private sector?

• What sort of precision is the model expected to have in the short term (12 to 24 months) and the long term (24 months to 10 years)?

We address these research questions throughout the memo. The rest of this memo is divided into five sections. First, we provide an overview of the main features and assumptions of the current

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

SNAP forecasting model. Second, we discuss the data and methods that we used to assess the performance of the alternative forecasting models that we examined. Third, we discuss the past performance of the current SNAP forecasting model. Fourth, we assess the performance of the alternative forecasting models. Finally, we present our conclusions.

B. Overview of current SNAP forecasting model

FNS obtains forecasts of SNAP participation in two stages. In the first stage, FNS estimates a regression model by using quarterly data from FY 1977 through the quarter prior to the current

FY. SNAP participation levels are regressed on the unemployment level, a one-quarter lag of the unemployment level, and a limited set of policy and seasonal indicator variables. Underlying economic conditions are assumed to affect SNAP participation through unemployment. In particular, when the economy is strong, more households are expected to have incomes above the threshold for SNAP eligibility. Conversely, when economic conditions weaken, more households are expected to meet SNAP income eligibility requirements. In the current SNAP model, the contemporaneous and one-quarter lag of unemployment serves as a proxy for underlying economic conditions. The directional expectation of the model is that economic downturns will increase unemployment and, operating through the one-quarter lag and the contemporaneous level of unemployment, will increase future SNAP participation. Economic growth is expected to lower unemployment and decrease demand for SNAP benefits.

The error terms related to SNAP participation are assumed to be correlated over time following a first-order autoregressive (AR(1)) process. The errors represent the effects of omitted explanatory variables that are likely to be correlated over time; the AR(1) model represents a way to capture this correlation. This specification assumes that the correlation between successive error terms is high but decreases rapidly as the successive error terms become further apart in time. The assumption regarding the behavior of the error term is an issue that we examine in further depth in this memo.

The current model can be expressed by the following two equations:

(1) 𝑆𝑁𝐴𝑃𝑡 = 𝛽0 + 𝛽1𝑈𝑁𝐸𝑀𝑃𝑡 + 𝛽2𝑈𝑁𝐸𝑀𝑃𝑡−1 + 𝛽3𝐸𝑃𝑅𝑡 + 𝛽4𝑂𝐵𝑅𝐴𝑡 + 𝛽5𝑄𝑇𝑅1𝑡

+ 𝛽6 𝑄𝑇𝑅2𝑡 + 𝛽7𝑄𝑇𝑅3𝑡 + 𝑣𝑡 ,

(2) 𝑣𝑡 = 𝜌𝑣𝑡−1 + 𝑒𝑟𝑟𝑜𝑟𝑡 , where:

𝑆𝑁𝐴𝑃𝑡 is the average SNAP participation per month in quarter t;

𝑈𝑁𝐸𝑀𝑃𝑡 is the U.S. civilian unemployment level (not seasonally adjusted) in quarter t;

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

𝑈𝑁𝐸𝑀𝑃𝑡−1 is the U.S. civilian unemployment level (not seasonally adjusted) in quarter t-1;

𝐸𝑃𝑅𝑡 is the indicator variable for the elimination of the purchase requirement in 1977;1

𝑂𝐵𝑅𝐴𝑡 is the indicator variable for the Omnibus Budget Reconciliation Act of 1981;2

𝑄𝑇𝑅1𝑡, 𝑄𝑇𝑅2𝑡, 𝑄𝑇𝑅3𝑡 are the quarterly indicator variables (1/0) to control for seasonality;

𝛽𝑗 (𝑗 = 0 𝑡𝑜 7) are the parameters to be estimated;

𝑣𝑡 is the error term in quarter t that is assumed to follow an AR (1) process;

𝜌 is the autoregressive parameter to be estimated;

𝑒𝑟𝑟𝑜𝑟𝑡 is the component of the error term in quarter t that is independently and identically distributed over time.

In the second stage of the forecasting procedure, FNS produces forecasts of SNAP participation levels by using the parameter estimates from the regression model specified in Equations (1) and

(2). The model generates quarterly forecasts that are then aggregated and submitted to OMB for budgeting purposes. The model is also used to generate 95 percent upper and lower bound confidence intervals around the forecasted SNAP participation levels.

In addition to the formal modeling procedure, FNS occasionally makes “out-of-sample” adjustments to the forecasts to account for legislative changes or recent economic developments that might impact SNAP participation but cannot be included in the regression model. These adjustments are typically made when FNS believes that SNAP participation will grow faster than changes in unemployment might suggest or when participation might not decline as quickly as unemployment improves. Out-of-sample adjustments can encompass a number of modifications, including using longer lags of unemployment and using growth trends from previous years that are believed to be more analogous to current circumstances. These out-of-sample adjustments are not accounted for in the confidence intervals around the SNAP forecasts. All out-of-sample adjustments must be approved by OMB.

1 The purchase requirement was eliminated in 1977 but was not fully implemented until 1979. Consistent with

Mathematica’s previous report, we set this variable to 0.4 in the second quarter of FY 1979, 0.625 in the third quarter of FY 1979, 0.75 in the fourth quarter of FY 1979, and 0.875 in the first quarter of FY 1980. The variable is set to one for quarters starting in the second quarter of FY 1980 and zero for quarters before the second quarter of FY 1979.

2 This variable is set to one for quarters starting in FY 1982 and to zero for all earlier quarters.

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

C. Data and methods to assess and compare model performance

In this section, we briefly describe the process for creating the historical SNAP series and structuring the data for analysis. We also elaborate on the metrics used to compare and assess both out-of-sample and in-sample performance.

To assess the accuracy of the current and alternative forecasting models, we used published records of monthly SNAP participation that are available online from FNS. We examined aggregated counts of SNAP participants rather than households to match the units that FNS uses in its forecasts. In addition, we created a quarterly data series of SNAP participation from monthly data to be consistent with the level of time aggregation that FNS uses for its forecasts.

To aggregate the monthly series to the quarterly level, we calculated an unweighted average of participation for the three months within each FY quarter. The full series consists of four quarters each from FY 1977 to FY 2018, for a total of 168 quarters of data.

We obtained monthly historical values of non-seasonally adjusted unemployment levels from the

Bureau of Labor Statistics website. We aggregated this series to the quarterly level in the same manner as for SNAP participation. For our forecasting simulations, we used the actual values of unemployment for the time periods covered by our analysis because we did not have access to forecasted values of unemployment. However, this approach underscores an important limitation of our subsequent assessment of the SNAP model. Specifically, when FNS updates its model, it uses forecasted values of unemployment provided by OMB. Using forecasted values of unemployment rather than actual values increases the uncertainty in the resulting forecasts.

Therefore, the performance of the models in this memo that use actual unemployment values may overstate the accuracy and precision of predictions compared to models that use forecasted values. Mathematica’s previous assessment of the SNAP forecasting model also used actual unemployment values.

We simulated the FNS forecasting procedure for quarters starting in FY 1990 through FY 2018.

To obtain parameter estimates to use for forecasting, we estimated (trained) models by using quarterly data between FY 1977 and FY 1990. We then used these parameter estimates to generate forecasts incrementally at one and two quarters ahead, one through four quarters ahead

(one year), and five through eight quarters ahead (two years). As we incrementally generated forecasts for the out-of-sample time periods (FY 1990 to FY 2018), the successive estimation sample included additional quarterly data. For example, to simulate one and two quarter-ahead forecasts for the last two quarters of FY 1990, the estimation (training) sample used data between

FY 1977 and the second quarter (Q2) of FY 1990. Similarly, to generate the eight quarter-ahead forecasts from FY 2010 Q4, we estimated (trained) models on the data up through FY 2010 Q4

(that is, using all data available up through this quarter) and produced forecasts for all four quarters of FY 2011 and FY 2012. By using this approach, we obtained forecasts for each quarter between FY 1990 and FY 2018, which were forecasted from one quarter ahead to eight quarters ahead.

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

To estimate the models, we used data starting in FY 1977. We decided to use data starting from the earliest time period rather than limiting to more recent data for several reasons. First and most importantly, our initial results suggested that the current FNS forecasting model produced better forecasts when we set the estimation sample to begin in FY 1977 as opposed to starting in more recent periods. This could be because a longer series of training data provided more information for estimating the underlying processes governing SNAP participation over time.

Second, using longer time periods to develop forecasting models can help smooth out the influence of more unusual periods in which relationships between key variables may have differed. Accordingly, we opted against truncating the starting point for the estimation series.

1. Assessing model performance

In assessing and comparing the performance of different models, we focused primarily on out-of-sample measures of forecasting performance. We also examined in-sample measures of model fit for models estimated on the entire data series (FY 1977 through FY 2018). However, because the goal was to accurately predict SNAP participation in the future, we considered out-of-sample forecasting performance to be the most important criterion for assessing alternative model specifications. The emphasis on forecasting performance also helps guard against overfitting a statistical model to a specific sample. For data that are not structured as time series, randomly splitting the sample data into separate estimation (training) and validation data sets is a common way to avoid overfitting. With time series data, however, randomly splitting data in this way is not an option, so focusing on performance in the forecasting period (which does not influence the estimation model) helps guard against overfitting.

a. Out-of-sample measures

As noted above, in order to replicate the FNS forecasting procedure, we examined the accuracy of three forecast time horizons: (1) one and two quarter-ahead forecasts; (2) one through four quarter-ahead forecasts (one year); and (3) five through eight quarter-ahead forecasts (two years). In general, forecasts with longer time horizons (that is, those that are projected further into the future) have higher uncertainty, which can affect both variance and accuracy. For all three forecasting time horizons that we examined, we used the same set of measures to assess accuracy.

To standardize the formulas below, we let 𝐹 = {𝑓1, 𝑓2, … , 𝑓𝑁} represent a set of N forecasts from a particular model specification. The subscript, for example, can represent the forecasts made one and two quarters ahead for quarters in the 1990s. This can be extended for the other two forecast horizons (two to four quarters and five to eight quarters out). We also define the corresponding actual level of SNAP participation for a given forecast as {𝑆𝑁𝐴𝑃1, 𝑆𝑁𝐴𝑃2, … , 𝑆𝑁𝐴𝑃𝑁}, where the subscript, n, indexes the forecast for the time point for which actual participation is observed. The set of measures we used were as follows:

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

• Absolute forecasting error (AFE). The AFE measures the average difference between the forecasted and actual SNAP participation levels for a set of forecasted values. It is positive when the forecasts on average overpredict participation and negative when the forecasts on average underpredict participation. The AFE can be expressed as follows:

𝑁

(3) 𝐴𝐹𝐸(𝐹) =

𝑁

∑ 𝑓𝑛 − 𝑆𝑁𝐴𝑃𝑛

𝑛=1

• Mean percent error (MPE). The MPE measures the average percentage difference between the forecasted and actual SNAP participation levels for a set of forecasted values. It is positive when the forecasts on average overpredict participation and negative when the forecasts on average underpredict participation. While the AFE measures prediction error as a count of SNAP participants, the MPE represents the error as a percentage of actual participation for a given time period. The relative values of these measures may be very different when comparing across time periods with vastly different levels of participation. For the current FNS model, the MPE is similar in the 1990s and in the 2010s, but the AFE is much larger in the 2010s due to higher caseloads. The MPE can be expressed as follows:

(4) 𝑀𝑃𝐸(𝐹) = 100 ×

𝑓𝑛 − 𝑆𝑁𝐴𝑃𝑛

𝑁

𝑆𝑁𝐴𝑃𝑛

• Mean absolute deviation (MAD). The MAD measures the average absolute difference between the forecasted and actual SNAP participation levels for a set of forecasted values. In contrast to the AFE, taking the absolute values of the difference between the predicted and observed values ensures that quarters in which SNAP participation was severely overpredicted cannot cancel out quarters that severely underpredicted participation. This is important when we look at summaries over entire decades. The

MAD can be expressed as follows:

(5) 𝑀𝐴𝐷(𝐹) =

𝑁

∑ |𝑓𝑛 − 𝑆𝑁𝐴𝑃𝑛|

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

• Mean absolute percent error (MAPE). The MAPE is similar to the MPE except that it is an average of the absolute value of percentage differences, rather than of the signed differences. As with the MAD and AFE, the MAPE will be more meaningful than the

MPE over longer forecasting horizons. The MAPE can be expressed as follows:

(6) 𝑀𝐴𝑃𝐸(𝐹) = 100 ×

|𝑓𝑛 − 𝑆𝑁𝐴𝑃𝑛|

𝑁

𝑆𝑁𝐴𝑃𝑛

• Root mean squared error (RMSE). The RMSE is similar to the MAD except that it is an average over squared error. This measure weights up large error predictions, so it is less favorable to models that generally predict well but are highly inaccurate in some quarters. The RMSE is a standard measure for evaluating linear regression models because these models are estimated by minimizing the RMSE. The RMSE can be expressed as follows:

(7) 𝑅𝑀𝑆𝐸(𝐹) =

√∑(𝑓𝑛 − 𝑆𝑁𝐴𝑃𝑛)2

b. In-sample measures

In-sample measures are often used to assess model fit within the estimation sample and do not necessarily correlate with the accuracy of forecasts made for quarters that are out-of-sample.

However, all else equal, it is usually advisable to adopt models that fit the estimation data as well as possible. The in-sample measures we considered included (1) Akaike’s Information Criterion

(AIC); (2) the Bayesian Information Criterion (BIC); and (3) the level of statistical significance of the coefficients on the predictor variables and autocorrelations of the error terms. The significance levels of explanatory variables and autocorrelations test whether the individual model elements significantly improve the model fit. The AIC and the BIC are measures that balance between the model fit and model complexity (the number of model parameters). A low

AIC or BIC means that the model strikes a balance between fitting the data well while minimizing unnecessary complexity (thereby, maximizing parsimony). The BIC has a stricter penalty for model complexity than the AIC.

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

2. Comparing models

Our goal was to assess whether alternative model specifications could produce better forecasts of SNAP participation levels than the current model used by FNS. To this end, a secondary goal was to focus on models that are relatively easy to understand, are transparent, and can be replicated. In particular, we were primarily interested in models that improved prediction accuracy during the current decade (FY 2010 through FY 2018) because the underlying processes that determine SNAP caseload are likely to be most similar in the past few years.

However, we were also interested in the performance of models for earlier decades as well

(1990s and 2000s).

a. Considerations for comparing model performance

Overfitting. Overfitting occurs when a model fits a particular data set so closely that it does not generalize beyond those data. There are at least two common causes of overfitting that we attempted to minimize. First, in the forecasting context, it is possible that positive model performance in a few quarters or years is due to the ability of the model to fit certain stretches of time particularly well, even though that performance does not generalize outside of the specific time period. To mitigate this risk, we looked for models that outperformed the current model during the 2010s, but also outperformed in at least one of the other two decades we examined

(either the 1990s or the 2000s). In addition, the more explanatory variables that the model includes, the greater the risk of overfitting. To guard against this risk, we primarily examined models that added a single predictor to the current model and only explored combinations of predictors when there was a compelling theoretical or statistical reason to do so (for example, when two predictors each improved the model on their own, or when multiple predictors were related). Overfitting is a particular concern in this context due to the relatively small sample size of the time series, which means that the degrees of freedom available (observations minus variables) are limited. This also relates to our preference for simpler versus more complex models.

Meaningful improvement. It is possible that a model may perform well in some quarters by chance. To mitigate this possibility, we focused on models that demonstrated a certain magnitude of improvement above a predefined threshold. We are unaware of any specific formal rule regarding what constitutes meaningful improvement. Therefore, as a general guideline, we looked for at least a 0.05 percentage point reduction in the MAPE for the one and two quarter-ahead predictions when averaged across a decade.

Inconsistent improvement across quarters. It is unlikely that a model will improve predictions in every quarter, or even in most quarters. When assessing whether an alternative model performs well relative to the current model, we began by looking at the decade-level prediction summaries, which averages the one and two quarters ahead, one through four quarters ahead, and five through eight quarters ahead forecasts pertaining to the specific decade to which these forecasts apply. For example, a forecast made two years ahead in FY 1999 for FY 2001 would

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019 pertain to the decade summary for the 2000s. If these summaries suggested that the model improved forecasting performance relative to the current model, based on the guidelines that we described above, then we considered these models to be strong contenders for consideration by

FNS. We included detailed summaries of annual predictive performance in Appendix B for the strongest models.

3. Description of the quarterly SNAP participation series

Here we briefly describe the historical SNAP participation time series. Over the past decade, SNAP caseloads have risen substantially (Figure 1). At the previous peak in FY 1994 that was noted by Schochet and Needles (2000), there were about 28 million SNAP participants. At the most recent peak in FY 2013, the caseload reached 48 million participants. Starting in FY 1994, SNAP participation had declined, reaching a low in FY 2000, before levels began rising again.

Since its peak in FY 2013, SNAP participation levels have dropped in more recent years, although the overall level remains high.

Figure 1. Total quarterly SNAP participation in thousands: FY 1977 to FY 2018

D. Performance of current SNAP model

In this section, we briefly review the performance of the current SNAP forecasting model. Figure

2 displays two quarter-ahead, one year-ahead, and two year-ahead forecasts for SNAP participation. The dotted lines show actual SNAP participation levels during each quarter while the shaded bands indicate the upper and lower 95 percent confidence intervals for the forecasts.

The x-axis indicates the FY and quarter for which the forecast was being made. For example, in the two quarter-ahead plot, the forecast with an x-value of FY 2011 Q2 was forecasted with data up to FY 2010 Q4. The y-axis shows the total number of SNAP participants in thousands.

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

Figure 2. Predictive performance of the current forecasting model for forecasts made two quarters, one year, and two years ahead

From the plot, it is clear that the confidence intervals become wider for the forecasts made further out in time. This widening of the confidence intervals occurs because longer-term predictions are based off of the shorter-term predictions, and thus carry forward the uncertainty of the earlier predictions as well as their own uncertainty. Therefore, the precision and accuracy of the model decreases as the length of the forecasting interval increases.

As expected, the prediction error is larger for the five through eight quarter-ahead forecasts than for the one through four quarter-ahead forecasts. The forecasting error is largest during periods where SNAP caseloads are steeply increasing or decreasing. For example, during the surge in

SNAP participation following the Great Recession in 2009, the average forecasting error rose to

-2.59 million participants, which translates into an underestimation of actual participation by 6.7 percent for forecasts made in the first two quarters of FY 2010 and for the second two quarters of

FY 2010. This discrepancy in forecasting performance was also noted by FNS during discussions of recent model performance. Similarly, the average forecasting error for the one year-ahead forecasts peaked at -3.79 million SNAP participants (9.26 percent) for predictions made in FY

2009 for FY 2010, and -9.50 million participants (21.21 percent) for predictions made in FY

2009 for FY 2011. More recently, however, it appears that the forecasted values of SNAP participation are more in line with actual values as the gap between the predicted and actual values has narrowed over the past several quarters. Table 1 provides a summary of the forecasting performance of the current model by decade (Tables A.1 through A.3 in Appendix A show detailed summaries of the current model’s forecasting performance in each year over the past two decades, for predictions made one and two quarters-ahead, one through four quarters-ahead, and five through eight quarters-ahead.)

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

Table 1. Predictive performance by decade for the current model

Decade

1–2 quarters ahead 1–4 quarters ahead 5–8 quarters ahead

MAPE

MAD

(1,000s)

(1,000s)

1990 – 1999 3.16 737 5.10 1,189 12.63 2,946

2000 – 2009 2.91 715 3.98 982 8.77 2,168

2010 – 2018 2.91 1,280 4.55 2,014 10.82 4,857

Note: The number of SNAP participants is expressed in thousands, where 1,000 SNAP participants translates to 1 million. MAD = mean absolute deviation; MAPE = mean absolute percent error.

A similar pattern of underprediction was also evident by looking at the performance of the model in the early 1990s, when a recession impacted participation and may not have been adequately captured by the unemployment variables that are proxies for underlying economic conditions. As was discussed by Schochet and Needels (2000), in the 1980s and the early 1990s, SNAP participation closely tracked with the unemployment level. This can be shown visually based on the series summarized in Figure 3. Starting in the late 1990s, however, this relationship started to change. Following the recession in the early 1990s, the unemployment level recovered after it peaked in the second quarter of FY 1992, while SNAP participation continued to grow until its peak in the second quarter of FY 1994. This observed lag grew during the Great Recession in

2009. Although unemployment peaked in the second quarter of FY 2010, SNAP participation did not reach its peak until the second quarter of FY 2013

Figure 3. SNAP participation and unemployment: FY 1977 to FY 2018

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

FNS is one of many federal agencies to produce forecasts for budgeting and planning purposes.

Given the diversity of programs, data, and forecasting objectives, it is difficult to conduct a formal assessment of how FNS’ current SNAP forecasting procedure compares with other forecasting approaches. This is particularly true because published assessments of forecasting performance across federal agencies and the private sector are often unavailable. However, our informal assessment suggests that SNAP forecasting errors produced by the current approach are broadly comparable to those produced by FNS for estimating participation in the Special

Supplemental Nutrition Program for Women, Infants, and Children (WIC). To produce WIC eligibility and participation forecasts, FNS first uses publicly available data, the Current

Population Survey, to identify the WIC-eligible population. In the second step, FNS projects forward a participation rate for the eligible population and various subgroups. By using this approach, the forecasting errors for WIC participation between 1996 and 2001 ranged from -8 percent (underprediction) to +4 percent (overprediction) (National Research Council 2003). Over the same time period, the one through four quarter-ahead forecasts for SNAP participation ranged between 0.4 percent and 9.6 percent, which suggests that the magnitude of errors produced by these two sets of forecasts are comparable.

More formal macroeconomic models are often used to predict monetary quantities such as future levels of federal debt. Within the federal government, the Congressional Budget Office and

OMB produce separate one year-ahead debt forecasts (Martinez 2015). In terms of performance, the forecast errors for these models are often under 2 percent in absolute value, which is considered small (Ericsson 2017). However, as Ericsson (2017) notes, sometimes these errors can be much larger in magnitude. In fact, forecast errors for U.S. gross federal debt underpredicted actual debt in the periods leading up to the Great Recession and subsequently overpredicted actual debt during and after the recession ended (Ericsson 2017). In this sense, even forecasting models whose specifications and restrictions are derived from more established macroeconomic theory can be subject to substantial errors. Moreover, given the macroeconomic theory behind these models, there can be quite a bit of persistence in the forecast errors moving forward (Ericsson 2017). In assessing the current SNAP model, our primary goal is to identify alternative models that demonstrate improved out-of-sample forecasting performance. Unlike the more structural models used for macroeconomic forecasting, there is very little economic theory that can be used to choose among competing models. Accordingly, although we examined several theoretically plausible enhancements to the current model, we relied primarily on empirical results to judge performance.

E. Comparing performance of alternative models

In this section, we provide a review of alternative SNAP forecasting models that we considered.

This section is broadly organized by the three main approaches that we examined: (1) modifications to the current AR model; (2) ARIMA modeling; and (3) vector autoregression

(VAR).

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

1. Modifications using the current AR model

Here we describe our investigation of whether improvements to the predictive power of the current SNAP forecasting model can be made within the existing AR framework. We explored several modifications to the current model. In particular, we examined (a) incorporating time trends as predictors; (b) developing alternative specifications of the relationship between SNAP participation and unemployment; (c) including additional exogenous time series and policy changes as predictors; (d) developing alternative specifications of the AR error structure; (e) specifying alternative forms of the response variable; and (f) using alternative methods to estimate coefficients used for forecasting.

A first step in our analysis was to examine modifications of the current model that could be estimated using data already available to FNS. While we focused primarily on predictor variables that are forecasted by other federal agencies (although we used actual values in our assessment), we also examined a few predictors that are not forecasted. We focused on the measures of forecasting accuracy for the decade beginning in 2010 (FY 2010 to FY 2018), but also present results for the 1990s and 2000s to examine how the alternative models would have performed compared to the current model during these earlier periods. As discussed above, we focused on models that showed meaningful improvement in multiple decades to reduce the probability that a model showed improvements by chance. We present more detailed results only for selected models that demonstrated improvement in order to keep the presentation manageable. In general, we summarize our findings for each category and present the best models within a category when applicable.

a. Time trend predictors

As discussed, SNAP participation increased substantially over the last decade and these levels have been underpredicted by the current model. To account for this increase over time, we fit models by incorporating linear and quadratic trend terms. The linear time predictor is coded as 1 for the first quarter in the data, 2 for the second quarter, 3 for the third, and so on. The quadratic time predictor is the square of the linear time predictor. Figure 4 gives a graphical summary of the predictive performance of these models. It shows the predicted values for the current model, the model with a linear time trend added, and the model with a quadratic time trend (solid lines) compared to the actual SNAP participation levels (dotted line) for one year-ahead and two year-ahead forecasts (left and right panels, respectively).

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

Figure 4. Predictive performance of models incorporating time trends

Table 2 compares the MAD of SNAP participation and MAPE for the decades beginning in

2010, 2000, and 1990, respectively. The table shows the results for each decade and compares measures of prediction error for forecasts made one and two quarters-ahead, one through four quarters-ahead, and five through eight quarters-ahead. When presenting summaries over entire decades (40 quarters), we opted for measures of absolute prediction error rather than signed prediction error because a small (signed) average forecasting error can be misleading if a model severely overpredicts in some parts of the decade and severely underpredicts in others (that is, balances out on average).

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

Table 2. Predictive performance by decade for models incorporating time trends

Model

(1,000s)

2010–2018

Current model 2.91 1,280 4.55 2,014 10.82 4,857

Linear trend 2.17 930 3.33 1,433 7.99 3,490

Quadratic trend 2.27 972 3.60 1,547 9.03 3,915

2000–2009

Current model 2.91 715 3.98 982 8.77 2,168

Linear trend 2.81 683 3.77 918 8.42 2,052

Quadratic trend 2.75 660 3.62 863 8.01 1,903

1990–1999

Current model 3.16 737 5.10 1,189 12.63 2,946

Linear trend 2.97 678 4.87 1,112 12.85 2,939

Quadratic trend 2.85 658 4.79 1,099 13.98 3,150

On average, the models with a linear trend and quadratic trend outperformed the current model for all forecasting horizons in the 2000s and 2010s, as well as for the one and two quarter-ahead and one through four quarter-ahead forecasts in the 1990s. In the 2010s, the model with the linear trend outperformed the current model, reducing the MAPE from 2.91 to 2.17 (reducing the

MAD from 1.2 million participants to 0.93 million participants) for one to two quarter-ahead forecasts. For the one through four quarter-ahead forecasts, the linear trend model reduced the

MAPE from 4.55 to 3.33 (reducing the MAD from 2.0 million participants to 1.4 million participants). For the five through eight quarter-ahead forecasts, the MAPE dropped from 10.82 for the current model to 7.99 (reducing the MAD from 4.9 million participants to 3.5 million participants) for the linear trend model.

In addition to improving predictive accuracy, the AIC for both the linear and quadratic trend models was smaller than the current model. The time predictor was statistically significant in the linear trend model, while the quadratic time trend was statistically significant in the quadratic trend model at conventional levels of significance (p < 0.05). A more detailed summary of the predictive performance of the linear trend model is provided in Appendix B, Tables B.1 to B.3

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

b. Specifying different lags of unemployment

The current forecasting model predicts SNAP participation during a quarter by using the

(forecasted) unemployment level from that same quarter as well as the (forecasted) unemployment level from the previous quarter. As noted above, the relationship between unemployment and SNAP participation has changed over time, which suggests a need to examine the impact of different unemployment lags on forecasting performance.

To help identify possible alternative lag lengths for unemployment, we examined the correlations between the unemployment and SNAP participation series (both contemporaneous and at lags up to twenty quarters). The strength of the correlations varied somewhat by decade. During the

1990s, for example, the correlations peak at three quarters with a value of 0.73. In the 2000s, the correlation between SNAP participation and unemployment peaks at zero quarters with a value of 0.70, which suggests a more contemporaneous relationship. In the 2010s, the correlation peaks at seven quarters, with a value of 0.63. Based on these correlations, we examined models ranging from a one-quarter lag of unemployment (current model) through 12 lagged quarters of unemployment.

During the 1990s and the 2000s, the model with five lags of unemployment performed best with respect to predictive accuracy. During the 2010s, the model with five lagged quarters continued to perform well. The model with five lagged quarters of unemployment matches our intuition with respect to seasonal patterns in the time series models because the fourth lag corresponds to a one-year lag and the fifth lag corresponds to the previous quarter in the previous year. In addition to the model using five lags of unemployment, we also fit a model that included lags at one, four, and five quarters. This more parsimonious model performed better than the model that included all five lags. Similarly, we estimated a model with lags at one, four, five, and eight quarters, which also performed well. The in-sample summaries also support using lags at one, four, five, and eight quarters. These lags were marginally significant when added to the model.

Table 3 summarizes the predictive performance of the models with Lags 1 through 5; Lags 1, 4, and 5; and Lags 1 through 8 for each of the three decades. To varying degrees, these models improved prediction compared to the current model. However, the improvements are not as substantial as those from the time trend models described in the previous section. Figure 5 shows a graphical summary of the predictions from these alternative lagged unemployment models for one year-ahead and two year-ahead forecasts.

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

Table 3. Predictive performance by decade of models with additional lags of unemployment

(1,000s)

2010–2018

Current model 2.91 1,280 4.55 2,014 10.82 4,857

Five lags 2.41 1,075 3.80 1,701 9.51 4,292

Lags 1, 4, 5 2.37 1,058 3.73 1,673 9.36 4,224

Lags 1, 4, 5, 8 2.20 984 3.51 1,573 9.03 4,070

2000–2009

Current model 2.91 715 3.98 982 8.77 2,168

Five lags 2.72 677 3.69 927 8.05 2,031

Lags 1, 4, 5 2.70 671 3.66 917 7.98 2,015

Lags 1, 4, 5, 8 2.67 665 3.64 915 8.00 2,024

1990–1999

Current model 3.16 737 5.10 1,189 12.63 2,946

Five lags 2.79 646 4.47 1,035 11.23 2,598

Lags 1, 4, 5 2.73 631 4.40 1,018 11.10 2,564

Lags 1, 4, 5, 8 2.99 697 4.80 1,120 12.03 2,816

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

Figure 5. Predictive performance of models incorporating additional unemployment lags

c. Adding additional predictors

We also examined whether including new predictor variables in the current model could improve forecasting performance. The types of variables we examined correspond to four general categories: (i) policies affecting SNAP that were implemented after FY 1999 (the last time period Mathematica examined in the last technical assessment); (ii) recessions; (iii) poverty; and

(iv) changes in SNAP benefits. In general, adding these predictors did not substantially improve forecasting performance compared to the current model. Below, we provide an overview of the predictors and models we considered.

i. Policy variables

In the roughly 20 years since Mathematica last reviewed the SNAP forecasting model, several legislative and policy changes have been implemented that likely affected SNAP participation.

As discussed above, one such policy was PRWORA.3 Although PRWORA passed in 1996, many of its provisions were not fully implemented until later. For example, provisions with restrictions for able-bodied adults without dependents (ABAWDs) did not begin until February 1997;

provisions concerning restrictions on immigrants did not begin until September 1997; and new shelter caps froze deduction amounts at FY 2001 levels and beyond. The updated data from FY

3 PRWORA placed many restrictions on SNAP eligibility and participation. Specifically, it disqualified from SNAP eligibility many legally resident noncitizens; expanded work requirements and set time limits on SNAP benefits for adults age 18 to 49 without disabilities in childless households; changed deduction amounts and the maximum benefit calculation; and changed the structure of cash welfare from an entitlement to temporary assistance designed to move families into the workforce.

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

1996 through FY 2018 allow us to assess whether accounting for the phased-in implementation of key PRWORA provisions improves more recent forecasts of SNAP participation.

While PRWORA is hypothesized to have decreased SNAP participation, other legislative changes in the post-1999 period were expected to increase participation. Three such pieces of legislation were the Farm Security and Rural Investment Act of 2002 (2002 Farm Bill), the 2008

Farm Bill, and the 2009 ARRA. The 2002 Farm Bill restored SNAP eligibility to resident noncitizens with five years of legal residency in the U.S. The 2008 Farm Bill modified several

SNAP eligibility requirements in ways that could have encouraged greater participation.4

Similarly, the ARRA SNAP provisions, which took effect in April 2009, increased SNAP benefits by 13.6 percent and temporarily allowed states to suspend time-limited benefits for

ABAWDs. When the ARRA provision expired on October 31, 2013, maximum benefits returned to 100 percent of the Thrifty Food Plan in the preceding June. Although the SNAP provisions of

ARRA were intended to be temporary, they are hypothesized to have increased SNAP participation between FY 2009 and FY 2013.

To account for these policy changes, we augmented the current model with indicator variables that corresponded to the dates in which the provisions of the legislation would have been active.

For PRWORA, we coded all quarters in FY 1996 onward with a value of one and set all earlier quarters to zero. We coded quarters impacted by the Farm Bill and ARRA in a similar way.

Specifically, for ARRA, we set the value of quarters between FY 2009 Q2 and FY 2014 Q1 to one and all other quarters to zero. This “pulse” specification for ARRA was designed to capture the temporary nature of the SNAP provisions.

Overall, our results indicated that accounting for these legislative and policy changes did not improve forecasting performance. For the models we examined, the predictive accuracy was similar to but slightly worse than the current model without the policy indicators included. In addition, the in-sample performance of these models suggested that they did not fit the data well.

We also examined the impact of BBCE policies that are intended to increase access to SNAP and to streamline the application process. Between FY 2001 and FY 2018, the number of states

(including the District of Columbia, Guam, and the Virgin Islands) implementing BBCE policies increased from 3 states to 43 states (Government Accountability Office, 2012; U.S. Department of Agriculture, Forthcoming). Mathematica retrospectively produces estimates of the annual percentage of SNAP participants that were eligible solely due to BBCE. These estimates are available from calendar year 2007 to 2017. To examine whether BBCE might have improved

4 Specifically, it increased the minimum SNAP benefit for one- and two-person households, setting it at 8 percent of the maximum benefit for a one-person household. It also increased the standard deduction, eliminated the cap on dependent care deductions, excluded most education and retirement accounts from countable resources considered in eligibility determinations, and indexed resource limits to inflation (adjusted to the nearest $250 increment).

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019 more recent SNAP forecasts, we included the annual percentages of participants who were solely eligible due to BBCE.

A few caveats to our approach should be noted. First, the percentage of participants eligible solely through BBCE policies represents only a subset of all SNAP participants. Second, retrospective estimates do not cover most of the quarters that we included for model estimation.

As a result, we set the value of BBCE to zero for all quarters prior to FY 2007 Q1. Consequently, we were unable to retrospectively assess forecasting performance in earlier time periods. Third, the annual estimates do not temporally align with the quarterly structure of the SNAP participation data. This makes the BBCE series appear artificially smoother than the actual

SNAP participation numbers. In addition, the retrospective participation estimates are based on calendar years and thus do not align perfectly with the fiscal year. Finally, we used the 2017 value of BBCE for 2018 because estimates for 2018 were not available at the time of our analysis. Despite these limitations, we wanted to assess the plausibility of using the BBCE percentage estimates as a proxy for BBCE policies in the hopes that it might enhance forecasting for future periods (even if we cannot assess how forecasting performance would have changed if these data were available in previous decades).

Figure 6 compares the BBCE series with SNAP participation levels between FY 2000 and FY

2018. Although the two series are on different scales, the general pattern suggests that the percentage of participants eligible through BBCE appears to move in tandem with overall SNAP participation. Table 4 compares the forecasting performance of the model that includes BBCE to that of the current model. Including BBCE as an explanatory variable improved forecasts over the three forecast horizons that we considered. For the five through eight quarter-ahead forecasts, for example, the MAPE improved (decreased) from 10.82 (current model) to 9.53 (with BBCE).

These preliminary results suggest that BBCE policies may help with forecasting, but we also advise caution. It is uncertain whether the positive predictive performance will continue into the near future. In addition, because we only have a single decade of data on BBCE, it would be very easy for the model to capture spurious relationships between SNAP participation and BBCE during this period. The BBCE variable is also not statistically significant based on our assessment of in-sample measures. This is likely due to the fact that there are limited data to estimate the BBCE coefficient with precision.

Date:

Page:

INTENTIONALLY BLANK

INTENTIONALLY BLANK

12/6/2019

Figure 6. SNAP participation and percentage of participants eligible solely through BBCE

Table 4. Predictive performance of models with percentage eligible solely through BBCE in the 2010s

MAPE…

This is the start of the file's text. The full file is on GovTribe.

File details come from the government source that posted it. Updated .