ATTACHMENT 1 2020-MATH-SIPP-TWP.pdf
PDF 995 KB Posted
- Attached to
- Microsimulation modeling and analytical support services Federal contract opportunity
- Solicitation number
- 12-3198-25-R-0002
About this file
This document is a technical working paper detailing the creation of the 2020 MATH SIPP+ Microsimulation Model and Database, prepared for the U.S. Department of Agriculture's Food and Nutrition Service. The report comprehensively describes the methodology for developing a sophisticated microsimulation model that analyzes the Supplemental Nutrition Assistance Program (SNAP), Temporary Assistance for Needy Families (TANF), and Supplemental Security Income (SSI) programs.
The model uses Survey of Income and Program Participation (SIPP) data and Current Population Survey (CPS) data to simulate program eligibility, participation, and benefits. Key processes include creating model databases, developing state weights, assigning undocumented status to non-citizens, and simulating program participation through complex statistical methods. The document provides detailed technical explanations of how the model handles various demographic and economic factors, such as income limits, asset tests, household composition, and policy changes, to provide policymakers with robust analytical tools for understanding social assistance programs.
View the file
Other files for this federal contract opportunity
Show all 22
On GovTribe
Work with this file on GovTribe
- Download the original file
- Contacts named in this file
- Similar government files
- Ask GovTribe AI about this file
Text version
ATTACHMENT 1
February 28, 2025
Technical Working Paper:
Creation of the 2020 MATH SIPP+ Microsimulation Model and Database
FINAL REPORT
Nondiscrimination Statement ii
In accordance with Federal civil rights law and U.S. Department of Agriculture (USDA) civil rights regulations and policies, the USDA, its Agencies, offices, and employees, and institutions participating in or administering USDA programs are prohibited from discriminating based on race, color, national origin, religion, sex, disability, age, marital status, family/parental status, income derived from a public assistance program, political beliefs, or reprisal or retaliation for prior civil rights activity, in any program or activity conducted or funded by USDA (not all bases apply to all programs). Remedies and complaint filing deadlines vary by program or incident.
Persons with disabilities who require alternative means of communication for program information (e.g., Braille, large print, audiotape, American Sign Language, etc.) should contact the responsible Agency or USDA's TARGET Center at (202) 720-2600 (voice and TTY) or contact USDA through the Federal Relay Service at (800) 877-8339. Additionally, program information may be made available in languages other than English.
To file a program discrimination complaint, complete the USDA Program Discrimination Complaint Form, AD-3027, found online at How to File a Program Discrimination Complaint and at any USDA office or write a letter addressed to USDA and provide in the letter all of the information requested in the form. To request a copy of the complaint form, call (866) 632-9992. Submit your completed form or letter to USDA by: (1) mail: U.S. Department of Agriculture, Office of the Assistant Secretary for Civil Rights, 1400 Independence Avenue, SW, Washington, D.C. 20250-9410; (2) fax: (202) 690-7442; or (3) email: program.intake@usda.gov.
USDA is an equal opportunity provider, employer, and lender.
https://www.usda.gov/oascr/how-to-file-a-program-discrimination-complaint mailto:program.intake@usda.gov
Technical Working Paper:
Creation of the 2020 MATH SIPP+ Microsimulation Model and Database
FINAL REPORT
February 28, 2025
Submitted to:
U.S. Department of Agriculture Food and Nutrition Service 1320 Braddock Place Alexandria, VA 22314 Project Officer:
Contract Number:
Submitted by:
INTENTIONALLY BLANK
iv
Acknowledgments
INTENTIONALLY BLANK
v
Contents I. Introduction
A. Processing steps overview
B. Changes
II. Data Sources for the Model
A. Survey of income and program participation
B. Current population survey
C. Administrative data
III. Creating the Model Database
A. Recode SIPP variables
B. Edit or impute household-level data
C. Extract data for December 2019
D. Convert SIPP data into MATH database
IV. Creation of State Weights
V. Assignment of Undocumented Status
A. Development of the imputation methodology
B. Imputation methodology
C. Assignment of refugees and asylees
VI. Simulating the Supplemental Security Income Program
A. Create potential SSI units
B. Simulate SSI eligibility
C. Select SSI participants
D. SSI participation calibration results
VII. Simulating TANF
A. Create potential TANF units
B. Simulate TANF eligibility
C. Simulate TANF participation
D. TANF simulation results
Contents vi
VIII. Simulating SNAP
A. Determine disability status of individuals
B. Create potential SNAP units
C. Simulate SNAP eligibility and benefits
D. Select program participants
E. SNAP calibration and simulation results
IX. Simulating the Effects of Changes in SNAP Policy
A. Simulate baseline policies
B. Simulate policy changes
C. Calculate change in SNAP caseload and benefits
References
Mathematica® Inc. vii
Tables II.1 SIPP sample sizes and weighted counts
II.2 Comparison of administrative data and reported participation in SIPP, December 2019
IV.1 Population controls for MATH SIPP+ State weights
V.1 Simulated citizenship status by State
V.2 Comparison of reported and simulated citizenship status
V.3 Probability that a newly arrived noncitizen is a refugee or asylee, by year of U.S entry
VI.1 State SSI monthly supplements for individuals and couples living independently, FY 2020
VI.2a December 2019 State SSI control and simulated participant totals, age 0–17
VI.2b December 2019 State SSI control and simulated participant totals, age 18–64
VI.2c December 2019 State SSI control and simulated participant totals, age 65 or older
VI.3 December 2019 average SSI benefits in administrative data versus 2020 MATH SIPP+ model
VII.1 State TANF asset limits, July 2020
VII.2 State TANF income tests and thresholds, July 2020
VII.3 Earnings and dependent care deduction policies, July 2020
VII.4 Minnesota Family Investment Program benefits, FY 2020
VII.5 State minimum and maximum TANF monthly benefits, FY 2020
VII.6 National and State model TANF control and calibration totals
VII.7 Comparison of mean values for simulated and administrative TANF participants in 2020 MATH SIPP+ model
VIII.1 Target percentage of adults age 18–49 without disabilities in childless SNAP units to impute as potentially eligible, by SNAP participation in past year, FY 2020
VIII.2 State BBCE income, asset, and unit composition requirements, FY 2020 (December 2019)
VIII.3 SNAP maximum allowable gross and net monthly income eligibility standards, FY 2020
VIII.4 SNAP standard deductions and maximum excess shelter expense deductions, FY 2020
VIII.5 Standard medical deduction demonstration, FY 2020 (December 2019)
VIII.6a SUAs that vary by SNAP unit size, FY 2020 (December 2019)
VIII.6b SUAs that do not vary by SNAP unit size, FY 2020 (December 2019)
VIII.7 States conferring nominal energy assistance benefits and requirements for receipt, FY 2020 viii
Tables
VIII.8 State policies for counting vehicle assets, FY 2020
VIII.9 Maximum and minimum monthly SNAP benefits, FY 2020
VIII.10a Comparison of control totals with simulated results, national totals
VIII.10b Comparison of control totals with simulated results, State totals
IX.1 Coefficients for equation estimating the probability of participation for newly eligible units under policy change
IX.2 Coefficients for equation estimating the probability of participation for SNAP-eligible units with a change in SNAP benefit amount under policy change
I. Introduction The Supplemental Nutrition Assistance Program (SNAP) is the largest domestic food and nutrition assistance program administered by the U.S. Department of Agriculture’s (USDA) Food and Nutrition Service (FNS), providing millions of Americans with the means to purchase food for a nutritious diet.
During fiscal year (FY) 2020, SNAP served 39.9 million people in an average month and paid a total annual amount of $74.2 billion in benefits (USDA 2025).
Policymakers and administrators want to understand the potential effects of proposed changes in eligibility and benefit determination rules on the SNAP caseload and costs. For instance, they are interested in knowing how a change in the maximum benefit affects the number eligible for SNAP and the amount of total benefits. They are also interested in knowing the characteristics of SNAP participants and nonparticipants to assess whether benefits are effectively reaching subgroups of interest, such as households with members who are elderly, have a disability, are children, or have earned income.
One way to inform policymakers is to use a microsimulation model, which is composed of an underlying database, a set of parameters, and simulation techniques. The database is constructed from a nationally representative sample of households. The set of parameters and simulation techniques apply the rules of a government program to each household to determine its eligibility for, participation in, and benefit amount from that program. By changing the parameters and simulation techniques, an analyst can evaluate whether a change to program rules will have a relatively small or large effect on SNAP caseloads and costs.
FNS uses two microsimulation models to measure the effects of policy reforms to SNAP. This report documents the process of creating the 2020 MATH SIPP+ model. For information about the initial development of the MATH SIPP+ models and the history of the Microanalysis of Transfers to Households (MATH) models, see Smith and Wang (2012). For information about FNS’s other microsimulation model, the QC Minimodel, see Leftin et al. (2024).
The MATH SIPP+ models, first developed in 2006, use Survey of Income and Program Participation (SIPP) data as the underlying database and Current Population Survey Annual Social and Economic Supplement (CPS ASEC) data to contribute additional timely economic and demographic information. The SIPP has advantages over the CPS ASEC data for determining eligibility and benefits for SNAP and other low-income programs. Unlike the CPS ASEC, the SIPP contains extensive monthly information about two major determinants of program eligibility: assets and expenses, the latter of which is also used for benefit determination. The SIPP also includes information on monthly income used to determine SNAP eligibility, while the CPS ASEC provides only annual income. However, the SIPP sample is only about one-fourth the size of the CPS ASEC sample.
The MATH SIPP+ model relies on the primary strength of the SIPP data—monthly income, asset, and expense information—to determine eligibility. It also uses the demographic strengths of the CPS ASEC through a set of State weights. As explained in more detail in Chapter IV, each household is assigned a
Chapter I Introduction
State weight for each of the 50 States and the District of Columbia.1 Using the CPS ASEC–based State weights allows the model to produce State estimates of SNAP eligibility. National estimates are produced by using the original SIPP weights.
In this introductory chapter, we (1) briefly explain the processing steps and associated report chapters and
(2) identify the major changes from the 2011 MATH SIPP+ model described in Leftin et al. (2014). The model processing steps are described extensively in Molinari et al. (2025).
A. Processing steps overview
The 2020 MATH SIPP+ database was created using the following data sources and procedures:
1. Creating the model database (Chapter III)
• We recoded many 2020 SIPP variables.
• We imputed some household-level SIPP data.
• We extracted SIPP data for December 2019.
• We converted the data into MATH format, which is a hierarchical database of households, families, and individuals.
2. Creation of the MATH SIPP+ model State weights (Chapter IV)
• Using the 2020 and 2021 CPS ASEC, we created an average distribution by State for 33 control or target populations as of December 2019.
• We created national totals using SIPP data as of December 2019 for each control population.
• Using the CPS ASEC–based State distribution and the SIPP-based estimates of the national population, we created control populations by State.
• Using a Poisson regression reweighting technique, we created a set of 51 State weights for each household present in the SIPP for December 2019.
3. Noncitizen status (Chapter V)
• We imputed undocumented noncitizen status using a methodology originally developed by Dr.
Jeffrey Passel of the Pew Research Center.
4. Program simulation (Chapters VI, VII, and VIII)
• We simulated Supplemental Security Income (SSI) and Temporary Assistance for Needy Families (TANF) eligibility, participation, and benefits based on program rules and the most recently available administrative data.
1 We refer to the 50 States and the District of Columbia as the 51 States. Guam and the Virgin Islands, which are included in SNAP, and Puerto Rico and the Mariana Islands, which receive a block grant in lieu of SNAP, are not included in the SIPP data, and thus are not represented in the model.
• We simulated December 2019 SNAP eligibility and benefit rules and selected eligible households to participate based on administrative FY 2020 participation and benefit totals.
B. Changes
Leftin et al. (2014) describe the creation of the 2011 MATH SIPP+ model, which used data from Wave 10 of the 2008 SIPP panel (representing August 2011), the 2011 and 2012 CPS ASEC data files, and FY 2011 SNAP program rules. The 2020 MATH SIPP+ model uses data from the 2020 SIPP, the 2020 and 2021 CPS ASEC data files, and updated administrative data. Below, we describe additional changes that we made to the model.
1. Incorporating changes from the FY 2015 Baseline of the 2011 MATH SIPP+ model
In 2015, we developed a new baseline for the 2011 MATH SIPP+ model that included several model changes. We included most of those changes in the 2020 MATH SIPP+ model, such as the following:
• An improved simulation of Minnesota’s Family Investment Program (MFIP) that includes MFIP units in the SNAP simulation process and simulates “Uncle Harry” units
• A Standard Medical Deduction simulation in States with a demonstration program
2. Updating the model to use the 2020 SIPP
The SIPP underwent a major redesign in 2014, so the 2020 SIPP also is substantially different from the 2008 SIPP panel used in the 2011 MATH SIPP+ model. The new SIPP changed many facets of the survey, from the design and content of the survey instrumentation and variables to the frequency of interviews.
Below, we summarize the major changes relevant to the MATH SIPP+ model and how we updated the model to address these changes:
• Reference period. The 2008 SIPP panel’s reference period was four months, while the 2020 SIPP reference period was one year. We updated the model to consider the new reference period for variables that were reported at the reference period level, such as the amount of rental income received.
• Simulation month. We chose December 2019 as the new base month for the MATH SIPP+ model for the following reasons:
– December was the month for which the SIPP collected detailed asset and investment data.
– SIPP calendar year files were weighted to represent December 31 of each reference year.
– The interview recall window was the shortest in December and therefore likely to be the most accurate.
– The single-month approach was consistent with the 2011 MATH SIPP+ model.
– The sample size was sufficient.
• Variable changes. Numerous SIPP variables used in the MATH SIPP+ model changed between 2008 and 2020. Some changes were relatively minor, such as changes to categorical variable codes. Other changes were substantial and required major model updates, such as changes to income and asset categories, the number of jobs reported, household relationships, and disability and survivor income categories. Some 2008 SIPP panel data, such as the amount of energy assistance received by households, were no longer available in 2020. Chapter III describes how we recoded 2020 variables for use in the new model.
• Household definition. The 2020 SIPP had variables that identified household members for each month of the reference year and variables that identified household members at the interview month.
Some data were only available for interview-month households. We developed new processes to reconcile these differences, as described in Chapter III.
• Topical modules. The 2008 SIPP panel was composed of a core survey questionnaire and many topical modules. The 2020 SIPP had no topical modules, and much of the information in previous topical modules were incorporated into a single survey. We made the necessary updates to use variables only from the main survey data.
• Vehicle values. SNAP eligibility rules use the wholesale fair market value (FMV), or average trade-in value, of vehicles. Because the 2008 SIPP panel reported retail FMV, we converted it to wholesale FMV using a methodology based on data from the National Automobile Dealers Association Consumer Price Guide. However, vehicle values in the 2020 SIPP were reported as wholesale FMV, so we did not need to convert these values in the 2020 MATH SIPP+ model.
3. Updating program rules
• We updated the model to use FY 2020 SSI, TANF, and SNAP program rules. Chapter VI describes how we modeled the SSI rules. Chapter VII describes the TANF rules. Chapter VIII describes the SNAP rules.
4. Updating SSI disability checks
• We updated the method for imputing child disability status for SSI. Please see Chapter VII for more information.
5. SNAP unit formation
• We changed the process for forming SNAP units. The 2020 SIPP expanded information about relationships, so it now includes a relationship matrix that describes the relationships between all household members, instead of simply the relationship to the household reference person. We improved the SNAP unit formation process to use this new relationship matrix. We also removed processes related to how household members shared food because this information is no longer available in the 2020 SIPP. See Chapter VIII for details.
6. Assignment of undocumented status
• We updated the undocumented status assignment process to include the L and H-1B visas. See Chapter V for how we model undocumented status.
7. Assignment of refugee status
• Previously, we imputed refugee status using a person’s year of arrival and data from the Yearbook of Immigration Statistics (DHS 2022). In the 2020 MATH SIPP+ model, we improved the assignment status to include a person’s region of birth. We also changed the population we select refugees from.
In the previous model, we selected refugees from permanent residents because the 2008 SIPP panel had a variable that noted whether a person was a permanent resident during the reference period.
This variable no longer exists in the 2020 SIPP, so we recalculated the refugee assignment rates to select from all noncitizens.
Mathematica® Inc. 6
II. Data Sources for the Model The MATH SIPP+ model is based on several data sources: the 2020 SIPP panel; the 2020 and 2021 CPS ASEC; and administrative data for the SSI, TANF, and SNAP programs. The SIPP provides the model with a sample of households that forms a basis for all calculations of SNAP eligibility; the CPS ASEC data are used to derive household State weights that match State distributions of economic and demographic characteristics in the CPS ASEC; and the administrative data are used to simulate SSI, TANF, and SNAP recipient populations that match national and State administrative totals and subgroup characteristics.
A. Survey of income and program participation
The SIPP provides much of the information necessary to simulate program eligibility, making it an excellent choice for a microsimulation model database. In this section, we describe how the SIPP is administered and the types of data it provides. We also describe its weaknesses and changes to the survey since the previous panel.
1. Description of the SIPP
The SIPP is a nationally representative, longitudinal survey providing detailed monthly information on household composition, income, labor force activity, and participation in various government programs such as SNAP, TANF, SSI, and Medicaid. The interviewed population is based on a multistage stratified sample of the noninstitutionalized resident population in the United States. This includes people living in households as well as in group quarters, such as college dormitories and rooming houses. Inmates or residents of institutions, such as homes for elderly individuals, and people living abroad are not included.
Armed forces personnel are included, except for those living in military barracks (U.S. Census Bureau 2021).
People in households participating in the SIPP are interviewed every year. In each round (wave) of interviews, people age 15 or older are asked a set of core questions about their demographic characteristics, income, program participation, and children. Most interview questions are asked about the preceding calendar year, but some questions are asked as of the time of the interview. For the 2020 MATH SIPP+ model, we used data from the 2020 SIPP, which consisted of households interviewed in February through June 2020 about their characteristics in each month of calendar year 2019.
The 2020 SIPP interviewed 21,989 households, or 53,332 people. Weighted, this represents an estimate of 131,669,576 households and 311,273,856 individuals in December 2019 (Table II.1). The weighted totals are less than U.S. population counts because they exclude those living in territories and in institutions.
In 2023, the Census Bureau released a series of patch files that corrected some 2020 SIPP variables, including many variables related to income. We applied the changes in these patch files to variables that are used in the MATH SIPP+ model.
2. Challenges
Focusing on one specific month, December 2019, creates some challenges. First, some data, including householder status, rent and mortgage expenses, and financial assets, are reported as of the interview month. For a small number of households, the household members in December 2019 were not the same
Chapter II Data Sources for the Model
Mathematica® Inc. 7 as household members interviewed in 2020. For households in December 2019 that contained several household reference people or no household reference person, we assigned a single household member as the household reference person by using age and person number variables. For data only available for interview-month households, we edited month-level data for household members by using the assigned householders’ values. Please see Chapter III for more information on this process.
A second challenge with the SIPP is that, like most household surveys, it misreports the number of people participating in government programs (Table II.2). FNS reported 37.2 million SNAP participants in December 2019, while the SIPP estimates that 28.7 million people, 23 percent fewer, received SNAP benefits that month. The Administration for Children and Families (ACF) reported 8.1 million TANF participants in December 2019, while the SIPP estimates that 8.8 million participants, 10 percent more, received TANF benefits that month. To address these differences, we simulated SSI, TANF, and SNAP eligibility according to Federal and State policies and selected participants based on administrative data totals.
A final challenge with the 2020 SIPP is the data collection complication caused by the COVID-19 pandemic (Census Bureau 2021). Due to the pandemic, the 2020 SIPP used only telephone interviews, starting on March 19, 2020, through the end of the interview period. This caused disruptions to the interview process that resulted in a response rate of 45 percent, lower than previous years. Due to this low response rate, nonresponse bias is likely.
B. Current population survey
The CPS ASEC provides estimates of the demographic characteristics of households by State. It is a nationally representative survey of approximately 91,500 households, quadruple the size of the 2020 SIPP panel (Census Bureau 2020). Because the CPS ASEC is representative at the State level, we use it to reweight the SIPP to match the distribution of demographic characteristics of the U.S. population by State.
The CPS is a nationally representative monthly survey of households sponsored jointly by the Census Bureau and the Bureau of Labor Statistics (BLS). Each household is interviewed once per month for four consecutive months in one year and again during the corresponding time period in the following year.
The interviewed population is based on a multistage stratified sample of the noninstitutionalized resident population of the United States. As in the SIPP, this includes people living in households and in group quarters, such as college dormitories and rooming houses, but does not include residents in institutions, such as homes for elderly individuals. Also like the SIPP, individuals living in military barracks are excluded from the survey.
Every month, the CPS ASEC asks a set of basic questions about household composition, demographic characteristics, and labor force participation. It also provides additional detailed data on migration, work experience, household income, noncash benefits, and participation in various government programs, such as TANF, SSI, General Assistance (GA), and SNAP, by adding a set of supplemental questions on a specific topic each month. These supplemental questions make the CPS ASEC an excellent data source for providing distributions by State for many household characteristics.
Chapter II Data Sources for the Model
Mathematica® Inc. 8
C. Administrative data
We used three sources of administrative data for the 2020 MATH SIPP+ model: (1) FY 2020 data from the Social Security Administration (SSA), (2) the FY 2020 TANF data file from ACF, and (3) the FY 2020 pre-pandemic SNAP QC database. The data from SSA provided the number of SSI recipients by age group and State, which we used as a control for the SSI simulation. The TANF data file from ACF contained detailed demographic, economic, and TANF eligibility information for a nationally representative sample. These data were well suited for use as TANF control totals by characteristic and State. The pre-pandemic SNAP QC database was an edited version of the raw data file generated by the SNAP Quality Control System.
This database contained detailed demographic, economic, and SNAP eligibility information for a nationally representative sample of the monthly SNAP caseload. These data were well suited for providing the control totals of SNAP units by State and household characteristic. The FY 2020 pre-pandemic SNAP QC database contained around 18,000 SNAP households sampled between October 2019 and February 2020, before COVID-19 interrupted data collection.
Table II.1. SIPP sample sizes and weighted counts
Unweighted
(2020 SIPP)
Unweighted (December 2019)
Weighted (using household weight)
Households 21,989 22,057 131,669,576 Individuals 53,332 51,473 311,273,856
Source: 2020 Survey of Income and Program Participation User’s Guide and tabulations of the 2020 SIPP.
Note: When tabulating the number of households and individuals read and written into the model development programs, the
MATH SIPP+ model uses the household weight. The unweighted sample size for December 2019 does not include people in households where all members have a weight of zero. There are more households in December 2019 compared to the interview month because some people living in the same household in the interview month may have lived in different households in December.
Table II.2. Comparison of administrative data and reported participation in SIPP, December
Weighted individuals (000s) Administrative data
SNAP (excludes Guam and the Virgin Islands) 37,180
SSI 8,077
TANF (excludes, Guam, the Virgin Islands, and Puerto Rico) 1,965 SIPP data
SNAP 28,727
SSI 8,888
TANF 1,991
Percentage difference
SNAP -22.7
SSI 10.0
TANF 1.3
Sources: Administrative: SNAP Program Operations Data, SSA, ACF, Office of Family Assistance. SIPP Data: 2020 SIPP.
Mathematica® Inc. 9
III. Creating the Model Database The 2020 SIPP provides most of the information needed to simulate SSI, TANF, and SNAP. In this chapter, we describe in general terms how we compiled the information needed to create the MATH SIPP+ model.
A. Recode SIPP variables
We began by assessing differences between 2008 SIPP panel and 2020 SIPP data for all household-, family-, and person-level variables used in the model. For cases where the changes were minimal, we edited the 2020 SIPP variables so that they represented the same information as the 2008 SIPP panel variables used in the previous model version. We edited variables in two ways: (1) renaming 2020 SIPP variables to 2008 variable names and (2) constructing new variables that mimicked 2008 SIPP panel variable codes, universes, or levels. This approach enabled us to minimize the amount of model updates needed. We updated the model to use the new 2020 variables directly for cases where the changes were substantial and where the 2020 SIPP data provided additional information that could improve the model.
B. Edit or impute household-level data
The 2020 SIPP includes two variables that identify households. One variable identifies monthly households over the reference period; the other variable identifies households during the interview month, which occurs three to six months after the reference period. Household-level variables (including variables identifying the household reference person, the value of the cars in the household, the amount of rent or mortgages paid, and whether a household received energy assistance) are consistent with interview-month households. However, they are not always consistent with monthly households during the reference period. For example, if the household reference people of two different households in the interview month lived in the same household in December 2019, then the December 2019 household would have had two designated household reference people, each with their own reported rent expenses or car values. Because our simulation month is December 2019, we needed to edit or impute some household-level variables so that they were consistent with the simulation month.
We imputed household reference person status for December 2019 by first checking whether a household has a single designated householder, using the variable that identifies household reference people in the interview month. In most cases, the household only has a single designated household reference person.
In the few cases where a household had multiple designated household reference people or no household reference person, we designated a household member as the reference person by using their marital status, parental status, amount of earnings, and listed order within the household. After imputing the household reference person, we edited the values of household-level variables for other household members, such as rent or mortgage expenses, to be that of the household reference person. Finally, we removed households where no member had a positive weight.
Mathematica® Inc. 10
Chapter III Creating the Model Database
C. Extract data for December 2019
Next, we extracted these variables and other 2020 SIPP variables relevant to file development for December 2019. These data comprise the bulk of the data elements in the MATH database and include variables related to household composition, family composition, earned and unearned income, assets, and participation in various government programs.
D. Convert SIPP data into MATH database
Finally, we formatted the raw SIPP data into the MATH database, which consists of two files: (1) the data file (MATHPC.BIN) and (2) the header file (MATHPC.HDR). The data file is a hierarchical database of household, family, and person records. The header file is a text file that describes the contents, organization, and data types in the data file. The header file also includes information that needs to be readily available, such as poverty guidelines, year and month of the data, and the model’s version number.
Mathematica® Inc. 11
Target i s h i y=2020 i
IV. Creation of State Weights The 2020 MATH SIPP+ model has two versions: (1) a national model and (2) a State model. The national model uses the household weights from the 2020 SIPP. Because the original SIPP weights were not designed to be representative at the State level, we used the CPS ASEC to create a new set of 51 State household weights for the State model. These weights allow us to use every SIPP household when simulating a particular State’s program rules, regardless of which State the household resides in. This chapter describes the development of the weights used in the State version of the MATH SIPP+ model.
We created State weights for the MATH SIPP+ model by using a Poisson regression algorithm developed by Schirm and Zaslavsky (1997). Each household in the MATH database is given a set of 51 weights, where each weight estimates the number of households that the sample household represents in a particular State. The first State weight estimates the number of households that the sample household represents in Alabama; the 51st State weight estimates the number of represented households in Wyoming. In the 2020 MATH SIPP+ model, the sum of a household’s set of 51 State weights equals the original household weights from the 2020 SIPP.
The first step in creating the State weights is to obtain 33 population control totals, which include demographic, educational, and income characteristics of the population in each State (Table IV.1). For the 2020 MATH SIPP+ model, we used the 2020 and 2021 CPS ASEC and the 2020 SIPP to derive the control totals, or targets, according to the following formula:
where
Dec 2019 i,State
CPS Targeti,State
CPS Targeti,Nation x SIPP National TotalDec 2019 , CPS Target = 0.5 x ∑2021 CPS Estimatey .
In other words, the State targets are constructed to estimate the averaged State distributions in the 2020 and 2021 CPS ASEC, while preserving the national totals in the SIPP for December 2019. We use the CPS ASEC to derive control totals because it contains detailed income as well as demographic information. We combined the 2020 CPS ASEC (representing information from calendar year 2019) and the 2021 CPS ASEC (representing information from calendar year 2020) data files to approximate the U.S. population as of December 2019, because a weighted average of these two time periods coincides with December 2019 and because the use of two cross-sections of the CPS ASEC increases the sample size.
When we use a particular State weight, the sum of all the values for a population target variable will equal the population target for that State. For instance, summing the number of children in the household younger than age 5 over all MATH SIPP+ households while using the Alabama State weight yields the target number of children younger than age 5 for Alabama.
Mathematically, the formula to produce State weights can be described as follows:
β' x +δ wh,s = e , Chapter IV Creation of State Weights
Mathematica® Inc. 12 s where wh,s is the State weight, or expected number of households of type h in State s. A household type is, practically speaking, unique in the database because no two households are exactly alike. Therefore, each household in the database represents its own type. wh,s is the weight that will be given to household h when deriving estimates for State s. xh is a column vector of 33 control variables, or household characteristics for household h. βs is a vector of 33 unknown parameters to be estimated for each State s. δh is an unknown parameter to be estimated for each household h. β' x reflects the prevalence in State s relative to other States of households with the same vector of observed characteristics as household h. δh reflects the national prevalence of households with the same vector of observed characteristics as household h. The βs and δh parameters are estimated using a maximum likelihood method and satisfy the two first order conditions (constraints) of maximum likelihood estimation:
Constraint 1: ∑wh ,s = wh , s where wh is the original national SIPP weight of household h, and
Constraint 2: ∑wh ,s xh ,i = xs ,i h for each s and i , where x s ,i is the control total for control variable i in State s. According to the first constraint, reweighting does not change the total weight given to a household across all States, ensuring that the household contributes the same to a national estimate after reweighting as it does before reweighting. The second constraint stipulates that all control totals are satisfied for every State.
Although the MATH SIPP+ model offers greater precision in estimating eligibility by State, it also introduces measurement biases. The State weight assigned to every household in this database estimates the number of households that the sample household represents in a particular State. Because the State weight is an estimate, it introduces bias in the measure of the number of households that the sample household represents. However, because the measurement bias added to each household is equally likely to be an underestimate or overestimate, the overall effect of these biases is likely to be negligible. In addition, despite the measurement biases, the method used in producing the State weights has been thoroughly tested and shown to produce reliable and robust estimates (Schirm and Zaslavsky 1997).
Chapter IV Creation of State Weights
Mathematica® Inc. 13
Table IV.1. Population controls for MATH SIPP+ State weights
# Category (number of) 1 People who are Black 2 People of Hispanic ethnicity 3 Children younger than age 5 4 Children age 5–17 5 Elderly people (age 60 or older) 6 People with a disability 7 People with a high school diploma or higher 8 People with a bachelor’s degree or higher 9 People with household income at or below 50 percent of the poverty guidelines 10 People with household income 50–99 percent of the poverty guidelines 11 People with household income 100–149 percent of the poverty guidelines 12 People with household income 150–174 percent of the poverty guidelines 13 People with household income 175–200 percent of the poverty guidelines 14 People with household income 201–300 percent of the poverty guidelines 15 Married people 16 People who have a pension 17 People with household income at or below 100 percent of the poverty guidelines and a household size of one 18 People with household income at or below 100 percent of the poverty guidelines and a household size of two 19 People with household income at or below 100 percent of the poverty guidelines and a household size of three or four 20 People living in households with income at or below 100 percent of the poverty guidelines and both a child and someone with earnings within the household 21 Children with household income at or below 100 percent of the poverty guidelines 22 Elderly people with household income at or below 100 percent of the poverty guidelines 23 Earners with household income at or below 100 percent of the poverty guidelines 24 Children with household income at or below 50 percent of the poverty guidelines 25 People living in a household size of one 26 People living in a household size of two 27 Earners 28 People who receive interest or dividends 29 People who are unemployed 30 People who rent housing 31 Noncitizens 32 Number of people 33 Number of households
Source: 2020 MATH SIPP+ database.
Mathematica® Inc. 14
V. Assignment of Undocumented Status Undocumented noncitizens are ineligible for SNAP benefits, so it is important to identify them when simulating SNAP eligibility. According to Passel and Cohn (2018), the undocumented immigrant population as of 2016 was 10.7 million individuals. The SIPP does not ask noncitizens whether they are legally in the United States, so we imputed undocumented immigrant status to a portion of the foreign-born population in the MATH SIPP+ model. In this chapter, we describe that imputation methodology. We first discuss how the method was developed. We then describe the imputation process and its effect on the citizenship status of the foreign-born population in the MATH SIPP+ model.
A. Development of the imputation methodology
The imputation methodology used to assign undocumented immigration status was originally developed by Dr. Passel to assign immigration status and measure the size and characteristics of legal and undocumented immigrant populations by using the CPS (Passel and Clark 1998). Since 2015, Dr. Passel’s estimates have used the American Community Survey (ACS) instead of the CPS (Passel et al. 2018).
However, due to differences between the ACS and SIPP data, we cannot directly apply the ACS-based methodology to the 2020 MATH SIPP+ model but instead must use a method more suited to SIPP data.
The RAND Corporation, with the help of Dr. Passel and support from the Office of the Assistant Secretary for Planning and Evaluation (ASPE), developed a method to assign undocumented status in the SIPP data.
We adapted their method and programming code for use in the MATH SIPP+ model. Our method first identified all foreign-born SIPP survey members as either “potentially legal” or “potentially undocumented,” and then randomly assigned legal or undocumented status up to the point where the total number of legal immigrants in SIPP matched the number estimated by Dr. Passel. The remaining potentially legal or potentially undocumented individuals were then assigned to be undocumented immigrants.
B. Imputation methodology
We took the following steps to impute the number of undocumented immigrants in the sample:
1. Determine each sample member’s self-reported citizenship status.
We divided all SIPP individuals into four self-reported citizenship categories: (1) native, born in the United States; (2) native, born abroad of U.S. citizen parents; (3) naturalized, including those who report naturalization through their own or a spouse’s military service or by adoption; and (4) noncitizen.
2. Verify and correct self-reported status.
We verified self-reported status by using the reported age in December 2019 and a variable for the reporting year of entry. When the year of entry variable represented a range of years (for example, 2017 to 2018), we randomly assigned a year of immigration from within the range. We also determined the year of immigration of the individual’s spouse or parents if present in the SIPP. We then performed the following checks:
Chapter V Assignment of Undocumented Status
Mathematica® Inc. 15
• If the individual reported being naturalized but spent less than five years in the United States, then the individual was reassigned as a noncitizen.
• If the individual reported being naturalized, Hispanic, and from the Americas and was in the country for more than three years, then the individual was reassigned as a noncitizen. SIPP data does not identify smaller geographies within the Americas.
• If the individual reported immigrating before 1980, then the individual was reassigned as naturalized.
• If the individual reported their citizenship status as being a native and born abroad of U.S. citizens, but did not have U.S. citizen parents or a U.S. citizen spouse in the SIPP data, then the individual was reassigned as a noncitizen.
3. Assign certain legal statuses.
We classified foreign-born individuals as legal temporary migrants (LTM) if they were identified as temporarily residing in the United States (less than five years) and worked in the following occupations:2
• Diplomat
• Student
• Visiting professor or graduate assistant
• In the medical services field working as a medical scientist, or as a therapist, medical student, or speech pathologist
• Nurse
• Engineer, technician, or computer operator working for an international organization
• Religious worker
• Athlete or entertainer
• High school exchange student
• Au pair
• Intracompany transfers
• High-tech guest workers
We classified the remaining foreign-born individuals as legal permanent residents (LPRs) if they had the following characteristics:
• Individual married to a native spouse
• Individual or spouse now or ever in the U.S. armed forces
• Individual or spouse receiving SNAP, SSI, TANF or Medicaid
• Government worker or program eligibility interviewer
2 For some occupations, such as doctor, nurse, and engineer, the individual also had to work in an industry that commonly employs LTMs.
Mathematica® Inc. 16
• Medical worker with a professional degree such as a physician, dentist, registered nurse, pharmacist, or therapist
• Teacher
• Lawyer or paralegal
• Police officer or firefighter
• Postal worker
• Accountant
• Inspector
We classified all individuals age 65 or older to be legal and all foreign-born individuals with U.S. citizen spouses to be citizens. We also reclassified non-native individuals younger than age 65 who were not working but had a working noncitizen spouse to be noncitizens.
4. Derive target levels of the legal foreign-born population.
We determined the target size of the SIPP legal foreign-born adult population by State using independently derived estimates of the legal foreign-born population provided by Dr. Passel for March 2020. Following Dr. Passel’s methodology, we combined all States except California, Florida, Illinois, New Jersey, New York, and Texas. We subtracted LPRs, LTMs, and naturalized citizens in the MATH SIPP+ from Dr. Passel’s estimates to derive targets for the next step. National model targets were based on SIPP household weights and State model targets were based on State household weights.
5. Assign legal and undocumented citizenship status to adults.
To assign legal and undocumented citizenship status to individuals in the MATH SIPP+ model, we first divided the adult foreign-born population into two categories. Individuals who reported being naturalized were categorized as potentially documented and noncitizens who were not previously assigned to a legal group were categorized as potentially undocumented. We then assigned a randomly drawn value from a normal distribution to each individual. The distribution of the random value was determined by the following criteria:
• If an individual was in an occupation that requires a professional degree, such as architects and lawyers, and so had a low probability of being undocumented, then the random value was drawn from a normal distribution with mean of 0.1.
• If an individual’s occupation implied a high probability of being undocumented, such as manual labor, then the random value was drawn from a normal distribution with a mean of 0.6.
• If an individual was in neither a low nor a high probability occupation, then we used their age as follows:
– Age 18–39, then the random value was drawn from a uniform distribution between 0 and 1
– Age 40–64, then the random value was drawn from a normal distribution with a mean of 0.12
– Older than age 65, then the random value was drawn from a normal distribution with a mean of 0.02
Mathematica® Inc. 17
We sorted individuals by random value, State group, and potential status, then assigned legal or undocumented status as follows:
• If the MATH SIPP+ model had fewer legal foreign-born individuals than the Passel estimate by State group, then individuals were selected from the potentially undocumented group until the number of legal noncitizens equaled the target amount within an acceptable level of tolerance. All potential noncitizens not selected to be legal were assigned to be undocumented.
• If the MATH SIPP+ model had more legal foreign-born individuals than the Passel estimate by State group, then individuals were selected from the potentially documented group by random value until the number of legal noncitizens equaled the target amount within an acceptable level. The remaining potentially documented were then assigned to be naturalized citizens.
6. Assign legal and undocumented citizenship status to children.
In the last step of the assignment process, we assigned the legal or undocumented status of foreign-born children according to the status of their parents as follows:
• If a foreign-born child was assigned to be undocumented but had a native or naturalized parent, then the child was reclassified as a native or naturalized citizen, respectively.
• If a foreign-born child was reported to be naturalized but did not have a native or naturalized parent (including those who did not have parents in the MATH SIPP+ model), then the child was reclassified as undocumented.
• If a foreign-born child was classified as an LTM but had a parent who was either naturalized or undocumented, then the child was given the status of the mother if she was present in the SIPP, or else the father if he was present.
• If a foreign-born child was reported as native, born abroad of U.S. citizen parents, but did not have citizen parents in the SIPP data, then the child was reclassified as undocumented.
• If a foreign-born child who was classified as a noncitizen:
– Had undocumented parents, then the child was classified as undocumented
– Had a citizen parent, then the child was reclassified as naturalized
– Was born after the year the mother immigrated, then the child was reclassified as a native U.S.
citizen
– Was born before the year the mother immigrated, then the child was assigned the citizenship status of the mother
We summarize the results of the imputation process in the 2020 MATH SIPP+ model in Table V.1 and provide a comparison between reported and simulated citizenship status in Table V.2. Overall, the process assigns around 11.8 million individuals to be undocumented in both the national and State models. It also assigns 14.3 million and 12.0 million individuals to be naturalized in the national and State models, respectively. This difference in the imputation results between the national and State models is due to the difference in the distributions of the foreign-born populations in the two models. The national MATH SIPP+ simulation is based on the SIPP household weight, while the State MATH SIPP+ simulation is based
Mathematica® Inc. 18 on the State household weight, which was created according to the CPS-based distribution of noncitizens.
In both models, the net effect of the imputation process is that more foreign-born individuals are simulated to be categorically ineligible for SNAP.
C. Assignment of refugees and asylees
Noncitizens who are admitted to the United States as refugees or asylees are eligible for SNAP and other needs-based programs but are not identified in SIPP. Therefore, after completing the imputation process, we randomly selected a percentage of noncitizens to be refugees and asylees based on the year of arrival, region of birth, and data from the Yearbook of Immigration Statistics (see Table V.3).
Table V.1. Simulated citizenship status by State
Simulated citizenship status Legal citizens Legal noncitizens
Undocumented noncitizens
Native citizens
Naturalized citizens
Legal permanent residents
Legal temporary migrants
Number in national model (000s)
All States 265,453 14,314 18,711 1,023 11,773 California 27,590 2,945 5,375 219 1,564 Florida 15,627 1,323 2,248 120 683 Illinois 9,883 572 727 25 362 New Jersey 6,423 861 696 12 700 New York 13,465 1,667 2,020 111 814 Texas 22,068 1,404 2,037 74…
This is the start of the file's text. The full file is on GovTribe.
File details come from the government source that posted it. Updated .