PWS Attachment 1 NCEE Guidance for REL Research 2022.docx
DOCX document 5 MB Posted
- Attached to
- PEER REVIEW FOR REGIONAL EDUCATIONAL LABORATORIES Federal contract opportunity
- Solicitation number
- 91990022R0067
About this file
This is a solicitation posted by the U.S. Department of Education seeking proposals for peer review services to support the Regional Educational Laboratories program. The solicitation requests peer review to evaluate proposals, reports, and other products produced by the RELs and published by the Institute of Education Sciences. Peer reviewers will assess REL materials based on criteria outlined in an attachment regarding the relevance of topics addressed, technical writing quality, and rigor of research methodology. The solicitation does not provide pricing or award timing details but rather issues the notice to inform interested parties that the opportunity is now available on the Department of Education's contracting website.
View the file
Other files for this federal contract opportunity
| File | Type | Posted |
|---|---|---|
| Instructions to Offerors Figures E1 E2.pdf | ||
| Past Performance Form.docx | DOCX document | |
| 91990022R0067 SF30 Amendment 0001.pdf | ||
| Subcontracting Plan Review Form.pdf | ||
| Small Business Calculation Tool.xlsx | XLSX spreadsheet | |
| PWS Attachment 4 REL Writers Guide and Style Guide.pdf | ||
| 91990022R0067 FormSF1449.pdf | ||
| PWS Attachment 3 REL Stakeholder Feedback Survey Guidance.docx | DOCX document | |
| PWS Attachment 2 REL Research Proposal Templates.docx | DOCX document |
On GovTribe
Work with this file on GovTribe
- Download the original file
- Contacts named in this file
- Similar government files
- Ask GovTribe AI about this file
Text version
Chapter # Title of Chapter
U. S. DEPARTMENT OF EDUC ATION
July 15, 2022
NCEE Guidance for REL Study Proposals, Reports, and Other Products This document provides guidance for the development and review of proposals, reports, and other products that are produced by the Regional Educational Laboratories (RELs) and published by the Institute of Education Sciences (IES). This guidance is drawn from IES research standards.
VERSION 2.2
BACKGROUND
This document provides guidance for the development and review of proposals, reports, and other products that are produced by the Regional Educational Laboratories (RELs) and published by the Institute of Education Sciences (IES). These criteria are drawn from IES research standards [PL 107-279 Sec 102 (18)—http://ies.ed.gov/pdf/PL107-279.pdf].
CONTENTS
The guidance in this document is arranged in a question and response format. The responses are the criteria. These are the same questions that IES staff and external reviewers use to review proposals, reports, and other products.
The document is organized into sections that correspond to different types of REL products. Sections I and II provide criteria for ensuring the relevance and writing quality of all REL products. Sections III–VI provide criteria for ensuring the products’ technical quality (and utility when applicable), with separate criteria for research studies (Section III), tools (Section IV), applied research methods products (Section V), and literature re views (Section VI). Within each section, most criteria generally apply to both the proposals and the drafts of the products, but specific criteria may be designated as pertaining to just one or the other. The final section of the main body of this document (Section VII) provides reminders about additional content that proposals are required to include. Appendices cover additional guidance regarding how to A) magnitudes, standard errors, statistical significance, and statistical precision, B) address missing data, C) create summary measures, and D) conducting REL Promising Practices Studies.
| SECTION |
| TITLE |
| PAGE |
| I |
| All REL Products: Guidance for Ensuring Relevance |
| 2 |
| II |
| All REL Products: Guidance for Ensuring Writing Quality |
| 4 |
| III |
| Research Studies: Guidance for Ensuring Technical Quality |
A Guidance for all research studies B Additional guidance for reporting and handling missing data in descriptive studies C Additional guidance for promising practices studies D Additional guidance for impact evaluations
| IV |
| Tools: Guidance for Ensuring Technical Quality and Utility |
| 38 |
| V |
| Applied Research Methods Products: Guidance for Ensuring Technical Quality and Utility |
| VI |
| Literature Reviews: Guidance for Ensuring Technical Quality |
A Guidance for all literature reviews B Additional guidance for systematic evidence reviews
| VII |
| Reminders about Additional Content that Proposals Should Cover |
| 52 |
| App. A |
| Guidance on Examining the Magnitudes of Differences Between Groups |
| 55 |
| App. B |
| Guidance on Addressing Missing Data in Descriptive Studies |
| 60 |
App. C
App. D Guidance on Creating Summary Measures Using Factor Analysis and Item Response Theory Guidance for Promising Practices Studies
I. All REL Products: Guidance for Ensuring Relevance Proposals for and drafts of all REL products should meet the criteria in this section to demonstrate the products’ relevance to addressing important needs.
All REL Products: Guidance for Ensuring Relevance
All REL Products: Guidance for Ensuring Relevance 3
Do the proposal and product identify a specific regional need that the product is intended to address? Does addressing this need have the potential to benefit students?
Does existing literature on the issue support the need for the product?
Proposals and products should clearly state a need of a regional stakeholder that the product is supposed to address. The authors should provide sufficient contextual information to make the case that the need is important to the region. When describing this context, the authors should support any factual claims with empirical evidence.
They should also identify the regional stakeholders—usually those participating in a REL partnership—that articulated this need. Products should briefly describe the composition and goals of the stakeholders or partnership to provide a broader audience with the context for why the REL did the study.
The need that authors identify should go beyond stakeholders’ general concern about a problem or general interest in more information on a topic. Rather, the need should include a specific, future action or decision that the product could inform.
The action or decision that the product is meant to inform should have the potential to benefit students. Products may inform decisions focused on intermediate outcomes—such as retention of highly effective teachers—as long as students can ultimately benefit.
In most cases, proposals and products should present a summary of the existing literature as part of the justification for undertaking the study. The literature review should briefly but accurately discuss existing evidence about the product’s specific research questions or objectives—not just about a general problem or issue that motivated those questions. For example, if a study is examining high school characteristics that are related to students’ postsecondary success, the literature review should not only discuss existing concerns about postsecondary success, but should also identify which school characteristics have and have not already been found by prior research to predict postsecondary success. The literature review should allow the audience to see how the proposed product fills a specific gap in the existing body of knowledge, tools, or methods on the particular topic.
Literature reviews should provide a close, critical assessment of the existing evidence. Authors should not only summarize findings but also note key limitations when the basis for those findings is weak. They should place more weight on peer-reviewed research (including journal articles and government reports that have gone through peer review), and cite original sources rather than other writers’ summaries. Literature reviews should also acknowledge, rather than ignore, any inconsistent or contrary findings in past research.
For additional details on criteria relevant to literature reviews and systematic evidence reviews, please see Section VI.
Does the product directly address the need that the authors have identified?
Does the product have a well-defined audience that ideally reaches beyond a single region?
After identifying the actions or decisions that a product is meant to inform (see criterion 1), authors should explain how stakeholders can use the information in the product when undertaking those actions or making those decisions. In particular, proposals and products should state the research questions or objectives such that the answers can directly inform the actions or decisions. Implications sections of reports should discuss the specific ways in which stakeholders could use the study findings. Likewise, the text that accompanies tools should explain how to use the outputs of the tool (for example, the results of a new school climate survey) to make better actions or decisions (for example, to identify expectations for student behavior that staff need to communicate more clearly).
Authors should identify the most important audiences to which the proposed product will be targeted. Although the product will arise out of responding to a regional need, the target audiences should generally not be limited to only stakeholders in the region. Instead, authors should try to describe why the information ought to be important and useful to a more national audience. To engage such an audience, the product should describe the possible uses of the information in circumstances likely to be found in many regions.
II. All REL Products: Guidance for Ensuring Writing Quality REL products should meet the criteria in this section to communicate information clearly and effectively. These criteria pertain mostly to reports and tools rather than proposals. However, criteria 5, 6, 10, and 12 should be addressed in proposals as well.
All REL Products: Guidance for Ensuring Writing Quality
All REL Products: Guidance for Ensuring Writing Quality 5
Do the proposal and product have a precisely defined focus?
Do the proposal and product motivate all research questions or objectives and make them easy to understand?
Is the product well-organized?
Does the product engage the audience through a mix of text and exhibits?
Each product that the authors propose should focus on a single question or objective (or a closely related set of questions or objectives). When authors choose to pursue a line of investigation that covers a range of research questions or objectives, they should consider proposing multiple products to address them. Products with a narrow focus are generally easier to read than a large, comprehensive product that covers a broad range of topics.
The proposal and product should lead the audience to care about the research questions or objectives. As discussed in criterion 3, authors should phrase research questions such that the answers will represent actionable information. In addition, the proposal and early sections of the product should explain the importance of each research question or objective to ensure that the audience will be motivated to read the remainder of the product.
Authors should compose the research questions or objectives such that readers can quickly read and understand them. In particular, the research questions or objectives should be fairly compact. Details related to specific variables, target populations, and timeframes should be spelled out clearly in the narrative that accompanies the research questions or objectives.
Each section of the product should have a clearly defined purpose and a heading that concisely conveys that purpose. Authors should ensure that the information in each section aligns with the purpose of the section. For example, information on the study design and data belongs in a section entitled “What the study examined,” not in a later section on “What the study found.”
Authors should try to use a mix of text and exhibits (tables, figures, and boxes) throughout the entire product to help keep the audience engaged in the content. Long stretches of continuous text may discourage the audience from reading or using the product. Similarly, a long sequence of exhibits may lead the audience to lose focus on the central storyline that the text could have communicated more effectively.
Is the product easy to understand, and are its main messages easy for the reader to ascertain?
Are figures and tables clear and self-explanatory?
To help the audience easily understand the product and its main messages, authors should adhere to the following practices:
· The beginning of the product should summarize the product’s most important findings or messages with statements that can stand alone without additional context.
· Paragraphs should elaborate on a single point and generally begin with that point. Because short paragraphs are easier to grasp, authors should start a new paragraph each time the topic changes.
· Headings and subheadings should convey key messages or findings. For example, a section heading could be, “Academic and graduation outcomes for re-enrollees are mixed,” and a subheading in that section could be, “Most re-enrollees did not earn enough credits to graduate.”
· Descriptions of methods and findings in the main body of the product should be brief and accessible to the intended audience, with more technical details and finely grained findings in appendices.
· Authors should consider using bulleted lists, pullout boxes, and other nontraditional ways to make information accessible.
· The text that refers to a table or figure should focus on key findings or patterns of interest, rather than repeating all of the data from the table or figure.
· Authors should use the plainest language that can accurately convey a point and avoid jargon that is unfamiliar to the intended audience.
· Authors should use acronyms and abbreviations sparingly.
The audience should be able to understand and interpret every figure and table without having to read the text. For figures and tables to be able to stand alone:
· Titles of figures should convey key messages or findings. For example, a figure title could say, “District A students scored highest on reading comprehension,” rather than “Comparison of reading comprehension scores across districts.”
· Authors should clearly label all elements (such as axes, categories, data series, symbols, and column and row headers).
· Notes to the tables and figures should include information that is essential for the reader to understand the sample and data.
In addition, for figures to be as clear as possible:
· Figures should show the most important data needed for readers to grasp the key findings, rather than provide an exhaustive display of all data.
· To allow readers to make precise comparisons, figures should generally direct readers to compare lengths on a common scale (such as the height of bars on a bar chart). Readers are less able to make precise comparisons of areas (for example, portions of pie charts or nonadjacent bars in stacked bar charts). Visual forms that use shading (for example, heat maps) are best for making generic, instead of precise, judgments.
· Figures should be sorted on the most important variable. For example, sorting categories by their outcome (such as most to least) can help readers easily order categories and might be preferable to sorting categories alphabetically.
Proposals should include examples of figures (with hypothetical data) resembling figures that are likely to be appear in the final product.
Are different parts of the product consistent and aligned with each other?
Are the citations fully referenced?
Does the document follow the style and formatting guidance established by NCEE?
Products are easier to read when sections of the product are well aligned. Authors should:
· Discuss topics in the same order across sections.
· Make explicit connections across sections. For example, the text can map each key finding in the findings section to a research question from the introduction. Likewise, a finding that addresses a main research question should generally be bolded as a key finding so that the audience can easily connect research questions with findings.
Likewise, products are easier to read when they refer to key findings, terms, and data points in a consistent manner and use consistent criteria for calling out a finding. Authors should:
· Phrase main findings similarly in the summaries of the product and in the detailed content within the body of the product.
· Define key terms early in the product and use them consistently throughout. Do not introduce multiple terms for the same concept such as “teacher effectiveness,” “teaching quality,” and “teacher performance.”
· Ensure that numbers mentioned in the text appear in some table or figure (unless they are minor points, in which case the authors should make it clear that those points are not shown in any table or figure).
· Provide transparent criteria for classifying differences between groups as large, small, or absent and use those criteria consistently throughout the product (see Appendix A for details).
Citations should follow APA format and must provide adequate information to ensure accessibility to readers and reviewers. (Please see the REL Program Writers Guide and Style Guide for further information on citations.)
The document should:
· Use the product templates provided by NCEE, including the 15-page templates for reports and the 1-page and 4-page templates for summaries.
· Adhere to REL style conventions for capitalization, headings, tables and figures, callouts, lists, notes, references, typography, and other style elements.
· Avoid making sources, figures, or tables the subject of a sentence (for example, “Smith found…,” “Table X shows…”). Instead, citations and callouts can go at the end of the sentence in parentheses, as in: “Enrollment declined dramatically between 2002 and 2012 (table X).”
· Minimize use of the first person, focusing on the study or findings instead of the role or actions of the authors.
· Contain editable versions of tables and figures (not pictures pasted into the file) and alternative text describing figures and equations. These are required for compliance with Section 508 of the Rehabilitation Act, as amended by the Workforce Investment Act of 1998.
(For additional details, see the REL Program Writers Guide and Style Guide.)
III. Research Studies: Guidance for Ensuring Technical Quality Research studies analyze data from a region to provide information that addresses stakeholders’ needs. These studies include descriptive studies and impact evaluations. Research study proposals and reports should meet the criteria in this section to ensure the studies have high technical quality.
A. GUIDANCE FOR ALL RESEARCH STUDIES
The criteria in Section A apply to all research study proposals and reports.
Research Studies: Guidance for Ensuring Technical Quality
Research Studies: Guidance for Ensuring Technical Quality 13
Is the overall study design appropriate given the research questions?
Are the data sources and variables clearly identified and appropriate for the research questions?
Authors should select the most rigorous design that is feasible given time and resource constraints, while also selecting a design that is simple enough to be accessible to the audience.
The proposal and report should:
· Provide details on all data sources.
· Describe which data sources and variables will be used to address specific research questions.
· Identify intended respondents and other sources (existing databases, websites, etc.) explicitly. In proposals, it must be clear that the researchers know the data are available and appropriate for capturing each of the measures identified.
· Justify the choice of variables used in the analyses and years covered, noting how they relate to the research questions.
Do the proposal and report clearly explain how and why samples or subsamples are selected?
Are the samples/ subsamples appropriate for the research questions?
If a sample of respondents cannot be drawn to represent a relevant universe, is this acknowledged and explained?
The sampling universe and the populations of interest must be clearly defined. Any limitations in the representativeness of the sample should be acknowledged.
Ideally, the sample should correspond to the relevant universe that is the focus of the research questions. If the research questions refer to key subgroups, the proposal and product should identify the size of the corresponding subsamples and confirm whether they are adequate to obtain precise estimates. If the unit of sampling is different from the units whose outcomes or status will be examined— for example, if schools will be sampled and outcomes of students in the selected schools will be examined—the proposal must describe how respondents will be selected within each unit and how clustering affects precision. If adjustments are needed to ensure that the results represent a relevant universe, those adjustments should be made (see Tipton and Olsen 2022 for details).
For all studies, the proposal and product must be explicit about what issues can and cannot be addressed with nonrepresentative samples. If the study makes use of a convenience sample to answer some research questions using either qualitative or quantitative methods, the proposal and product must be clear about the limitations of the data. (For example, the study team should note if a dataset includes only those students who participated in a standardized test, while the research questions are focused on a larger population of interest. Also, when a study covers only students who can be followed for a specific number of years authors should clarify that the target population is the set of students who remain in the data during this time. That will help make policy makers aware of the fact that the results may not be relevant for more mobile students). In most cases missing data will need to be addressed (see Appendix B for details).
Similar caveats may be needed if a substantial amount of trimming was needed to make the intervention/treatment and comparison/control groups similar. Also, in general authors should note how their sample might differ from other populations of interest including those in other districts or future cohorts.
If there are substantial missing data and/or trimming, the authors should examine how well their intervention/treatment group, or the full study sample, matches the target population for the study. If it does not match well, the authors could consider selecting a subset of the original target population that is also a population of policy interest, but for which missing data and/or trimming is less problematic.
Authors are also encouraged to minimize missing data for the intervention group or study sample even if data are missing for less than 15 percent of that sample (the cut-point mentioned in Appendix B). In addition, ideally, they should conduct robustness checks to see whether their results depend on whether the intervention group is trimmed and whether the method for adjusting for missing data matters. WWC guidance on missing data provides options that can be used to help bound the potential magnitude of the bias due to missing data.
Are the data collection methods, sources, and instruments clearly described and appropriate for the research questions? Are any data collection instruments that are required included with the proposal?
Does the proposal adequately address protection of confidential data, if applicable, and does the team have a plan to secure consent if necessary?
Are the analysis methods clearly described and appropriate for the research questions? Have the authors avoided the use of any unnecessarily complex methods?
Data collection methods should be selected to elicit the information needed to address each research question.
Generally, instruments should be included in the appendices of the proposal and study product. In the rare cases where draft instruments are not included in the proposal, authors should explain when and how the instruments will be developed and when they will be submitted to IES for review.
In cases where studies are using an existing instrument, authors should provide information on previous studies where the instrument was used as well as information on the validity of the items contained in the instrument.
In all cases, variables should be measured appropriately, and methods for ensuring consistency across multiple data collectors should be discussed.
Reports and searchable databases should not reveal any personally identifiable information about individual students, families, teachers, school staff, or local and state administrators. Information about individual schools or districts that is not publicly available must also not be revealed in IES products. IES has primary responsibility for ensuring these guidelines are observed, but the proposal should identify any instances in which the authors believe the planned study may deal with confidential data and how the authors will keep it secure.
The study proposal should adequately address whether informed consent is necessary (from parents, students, and/or school staff). If consent is necessary, the study proposal should describe a strategy for securing it.
For more information, see the NCEE guidelines for development of Restricted Use Files and Disclosure Analysis Plans.
Authors should use the simplest methods that will fully answer their research questions. In both the proposal and report they should clearly describe their methods; in the report, some of the details of the methods can be confined to an appendix.
The proposal and report should describe any models and procedures for testing statistical significance of differences across groups or observed changes over time.
For studies using qualitative data (e.g., interview write-ups or other documents), the proposal and report should:
· Describe coding protocols (including any plans for ensuring inter-rater reliability).
· Identify how qualitative data are analyzed including specific tools/programs used.
· Explain the method for selecting quotes or vignettes.
· Identify respondents (using pseudonyms as needed) when citing individuals.
· Avoid generalizations and provide information about the extent of the activity, sentiment, or opinion that was described.
If the study creates summary measures, the authors should:
· Clarify what methods are used.
· Provide sufficient information so that the results can be replicated (see Appendix C for details).
If the study will assess alignment between curricula, standards, and/or assessments, authors should:
· Describe the set of dimensions by which alignment will be determined.
· Identify plans for developing the dimensions.
Are the key strengths and especially the key limitations of the data sources and analysis methods described?
Are the report conclusions fully supported by the findings?
In proposals and limitations sections of reports, the authors should describe aspects of the study approach that prevent the study from providing the most accurate or relevant answers to the main research questions. Examples of key limitations include use of a sample that is not representative of the target population, the use of data on past conditions when making decisions about the future, variables that do not capture the intended construct, or inability to measure important factors that might account for differences in outcomes across groups.
In any study that examines associations between policies or practices and outcomes that is not likely to meet the minimum criteria specified for a Promising Practices Study, it is particularly important to caution readers not to interpret such associations as the effects of those policies or practices. Not only should the limitations section include these cautions, but sections on the study’s purpose and implications should make clear why information about these associations can inform decisions or actions even without generating evidence on the effects of those policies or practices. For example, such associations might point to practices that should be assessed more rigorously in the future. Promising Practices Studies that do meet the minimum criteria can go further and discuss implications for policy or practice. See Appendix D for details.
Authors should discuss potential implications of the study findings without speculating beyond the findings presented in the body of the report. Generally, the implications should focus on points that flow directly from the main findings presented in the report summary. Although the implications should make clear how the findings can inform specific stakeholder actions or decisions (see criterion 3), they should not be framed as recommendations for what stakeholders “should” do. Instead, authors can describe how the findings highlight options that stakeholders “might” consider doing (including changes in policies, practices, or future research plans).
Minimum criteria All empirical REL studies, including descriptive studies that use census data, must report statistical significance and standard errors. In addition, the choice of highlighted findings should be based on both the statistical significance and magnitudes of the findings. See appendix A for details.
21.1
Are the standard errors and statistical significance of the findings reported appropriately
B. ADDITIONAL GUIDANCE FOR REPORTING AND HANDLING MISSING DATA IN DESCRIPTIVE STUDIES When some members of the original study sample have missing data, research studies must undertake careful steps to document and address this missing data. The following criteria apply to descriptive studies. A subsequent section (Section C) includes criteria that apply to impact evaluations. Appendix B provides more details.
Missing data can come from a variety of sources, such as survey nonresponse, lack of parental consent for assessments, and unreported information in administrative records. Every effort should be made to minimize the extent of missing data.
When eliminating the occurrence of missing data will not be feasible, authors should plan to report the extent of missing data. Proposals should identify any research question for which it is likely that a low percentage (below 85 percent) of the original study sample will have the data needed to answer that question. The appendix to the report should provide the actual response rate—the actual percentage of the original study sample that has data—for each research question and identify those questions with response rates below 85 percent.
When authors anticipate that the data available to address a question will have a response rate below 85 percent, the proposal should describe plans to conduct a nonresponse bias analysis on that variable. The report should show the results of the nonresponse bias analysis for each research question where the sample with non-missing data is less than 85 percent of the original sample.
The nonresponse bias analysis should examine whether there are systematic differences between sample members with the data needed to answer the research question and the original study sample. These analyses can take a variety of forms but authors should include at least the following two types of analyses:
1. For each research question, assess whether sample members with data and the original study sample differ on other observed characteristics by a substantial magnitude.
2. Assess the most likely reasons for missing data.
Guidelines #23 and #24 below and Appendix B provide more details on these issues.
Does the proposal specify a plan to assess the potential for biased study findings when a substantial amount of the data needed to address the research questions is likely to be missing?
Does the report describe the extent of missing data and carry out the proposed plan for assessing bias?
Does the proposal specify an acceptable method of accounting for missing data? Does the report provide details on how it carried out this method?
When authors anticipate that the size of sample with data needed to address a research question is less than 85 percent of the original sample, they may need to account for missing data in their proposed analysis. Specifically, they must account for differences in observed characteristics between sample members with data sufficient to answer each research question and the original study sample whenever such differences exceed 0.05 standard deviations. Proposals should select one of the following methods of accounting for such differences: multiple imputation, maximum likelihood, or nonresponse weights. All of these methods have the basic aim of adjusting the subsample with data so that it resembles the original study sample on the observed characteristics. If the authors believe that a method other than the ones listed above is appropriate for handling missing data in their study, they should explain their reasoning and provide appropriate references that describe the proposed method.
The appendix of the report must explain in detail the steps taken to implement those methods. For example, reports must specify the models, variables, and distributional assumptions used in these methods, as well as the software used.
When feasible, authors are encouraged to explore whether key study findings are sensitive to using alternative acceptable methods. Implementation of all methods must also account for uncertainty associated with missing data in calculating standard errors and statistical significance.
If no observed characteristics differ by more than 0.05 standard deviations between the subsamples with data sufficient to answer each research question and the original study sample, then reports may conduct analyses on the subsamples without accounting for missing data. In this case, the subsamples may be representative of the original study sample even though the response rate is below 85 percent. See Guideline #24 below and Appendix B for details.
Do the proposal and report carefully describe whether findings are likely to reflect the original study sample or only sample members with data?
Proposals should make an assessment of whether the proposed studies can generate findings that are likely to reflect the original study sample. When the sample available to address a research question is smaller than 85 percent of the original sample, studies can generate findings that are likely to reflect the original study sample only if all of the following conditions are met:
· The study is able to collect data on characteristics that are correlated with the key variables that have missing data.
· These characteristics are measured for all or nearly all members of the original study sample.
· If any such characteristic differs by more than 0.05 standard deviations between sample members with data sufficient to answer a research question and the original study sample, the study uses an acceptable method to account for this difference.
· Sample members with data are unlikely to differ from the original study sample on unobserved factors that are correlated with the key variables.
If the data available to address a research question cannot meet all of the conditions listed above, then authors may have no choice but to generate findings that reflect only sample members with data. In this case, proposals must still provide a clear justification why findings that reflect only sample members with data should be of interest to policymakers or practitioners. In addition, the findings and limitations sections of the report must provide appropriate caveats that restrict the conclusions to sample members with data. In particular, the limitations section should discuss possible ways in which conclusions might differ between the original study sample and sample members with data. See Appendix B for details.
C. ADDITIONAL GUIDANCE FOR PROMISING PRACTICES STUDIES
To provide partners with evidence about potentially promising policies or practices, RELs often conduct studies that examine whether a policy or practice is positively associated with education outcomes. However, some of these studies do not meet What Works Clearinghouse standards with or without reservations, usually because it is not feasible to do so. As a result, study authors must take special care to avoid inferring causation. Nonetheless, ensuring that these studies are as high-quality as possible could enable education leaders to use them in conjunction with other evidence to make better decisions. See Appendix D for more information about the definition of a Promising Practices Study and how to characterize the limitations and implications of their findings in study reports.
There are two sets of criteria for Promising Practices Studies: (1) minimum criteria and (2) other desirable features. A study must meet the minimum criteria to qualify as a Promising Practices Study. When feasible, other desirable study features could further strengthen the designs of such studies. The criteria are described in this section.
· Minimum criteria. Minimum criteria are designed to balance the needs for rigor and feasibility. For example, in an ideal world all Promising Practices studies would adjust for lagged outcomes. In practice, that is often not possible. Therefore, the minimum criteria offer an intermediate solution that considers both the value of rigor and the constraints RELs often face in conducting these kinds of studies.
· Other desirable features beyond the minimum criteria. Other desirable study features will improve the rigor of Promising Practices Studies. REL authors are encouraged to implement these additional features when data and resources allow. For example, when available, RELs should adjust for pre-tests (or other baseline values of outcomes). However, these data are often not available and, for this reason, are not required. The guidance on other desirable features is designed to be flexible. However, if RELs believe that alternative methods not covered in the guidance would be more effective for addressing their research questions, they should propose those methods. In those cases, the authors should cite relevant literature to help REL reviewers understand the benefits of those methods. REL authors are not required to use the latest methods but are encouraged to use rigorous methods when possible.
24.1
Intervention status must be a binary variable that indicates experiencing or being offered specific services or supports. For example: (1) Being assigned a new teacher mentor versus not being given a new teacher mentor, or (2) Attending a school that offers AP courses in mathematics versus attending a school that does not.
Intervention status may take on more than two categories if the research questions focus on pairs of those categories. For example, if intervention status takes on 3 categories, the study can compare category 1 to 2, 2 to 3, and perhaps also 1 to 3. We discourage a regression of the outcome on a continuous variable that describes a practice or policy (or a simple correlation of such a variable with the outcome); a study that does this would not meet the minimum criteria.
Other desirable study features, when feasible Authors should consider both policy relevance and the potential for bias when deciding how to define intervention status. Interventions that policy makers can directly control may be more relevant for policy. However, even if policy-makers can control intervention status, a direct comparison of treated and untreated units may produce biased results because of confounding factors.
For example, it might be easier for policy-makers to control which schools offer an afterschool program than which students choose to participate in that program. At the same time, if policy makers offer an after-school program in almost all low-income schools and in almost no high-income schools, comparing those two sets of schools might not provide a good estimate of the program’s effects. Instead, it might be better in that situation to compare outcomes of participating and non-participating students within each school.
Other examples of interventions to avoid, when possible, are those when the intervention group has a single teacher, school, or district that does not overlap with the comparison group or vice versa. This makes it impossible to distinguish effects of the intervention from effects of being with that single teacher, school, or district (the confounder). When such interventions are evaluated researchers should be clear about the limitations associated with estimating their effects.
Does the study have a clear and evaluable intervention?
24.2
Outcomes must be relevant to stakeholders (the REL’s agency and community partners) and related in some way to student learning or other education outcomes. This can include other outcomes (for example for educators) that are believed to be connected to student outcomes.
Outcomes must be measured in the same way for everyone in both the intervention and comparison groups, or else any measurement differences must be appropriately addressed in the analysis. For example, when the outcomes are measured at substantially different points in time for the intervention and comparison groups, but there is still substantial overlap between the timing, then adjustments should be made for the differences in timing, perhaps using regression or matching. When appropriate, similar adjustments should be made for other differences in measurement between the intervention and comparison groups. In a more extreme case, there is no overlap between the points in time when the outcomes are measured for the intervention and comparison groups. In those cases, an interrupted time series or comparative interrupted time series model may be appropriate. Those models are discussed below in the estimated effects criteria (24.3) .
Other desirable study features, when feasible To the extent possible, authors should strive to ensure that outcomes be measured reliably, have face validity, and not be overaligned with the intervention, as described in WWC guidance and criteria 40 for impact studies.
When focusing on teacher outcomes, there should be some evidence that associated practices can improve student outcomes. In particular, studies should refer to existing literature when possible, and when that is not an option, describe the logic model connecting the non-student outcome to a student outcome.
Are the outcomes well aligned to stakeholder needs, reliable, and valid?
24.3
Each effect must be estimated using the difference between outcomes for an intervention group and a comparison group, with adjustment for covariates when needed (see the section below on evaluating differences between the intervention and comparison groups).
In general, the comparison group will be a separate set of students with outcomes measured at approximately the same point(s) in time as the intervention group’s outcomes. There are at least three exceptions. The comparison group may consist of outcomes for the same set of students during a pre-intervention period when an Interrupted Time Series (ITS) design is used. ITS studies should include at least 4 observations prior to when the intervention started and at least one observation after it started. Alternatively, the comparison group could be for students from earlier cohorts in the same set of schools; this could be done using an ITS design. Finally, a simple pre-post design comparing different cohorts of students is allowed if adjustments are made for changes in the composition of the student body.
Other desirable study features, when feasible Having a comparison group with outcomes measured at the same point(s) in time as the intervention group is preferred.
In addition, while ITS may be used, Comparative ITS is preferred to ITS when possible, with checking for parallel trends. Regardless of the design, matching and regression adjustment is preferred to only matching which is preferred to only doing regression adjustment. Regardless, more time periods are generally preferred to fewer to shed more light on the validity of the design and provide more information about what happened following when the intervention started.
Are the effects estimated using appropriate methods?
Estimated effects must be reported numerically. For example, it would not be acceptable to summarize results using table with + and – signs to identify the direction of an estimated effect and * symbols to identify which results were statistically significant. Rather, authors must provide quantitative estimates of effects. One acceptable way to provide these results would be to report mean outcomes for the intervention group and estimated means for the comparison group. Another option would be to provide estimated adjusted differences.
Other desirable study features, when feasible Authors should report results in units that can be easily interpreted by policy-makers. To help policymakers make sense of the findings, authors should compare results to those found in any previous literature and to effects that would be required to justify a change in policy.
When the outcome is binary, it is often helpful to report results using the fraction or percent change in the outcome associated with being in the intervention group rather than in effect size units or log odds ratios. When the outcome is a continuous test score, then scale score units may be helpful if policy makers are familiar with the scale used. If not, then authors may want to consider standardizing the test (using the standard deviation) or reporting results in percentile units.
24.4
Are results reported with sufficient information and using user-friendly metrics?
Authors must check for differences in the key covariates between the intervention and comparison groups (as discussed in the next row) or adjust for such differences. They must use at least one of the following types of key covariates.
· Type 1: All available variables that are believed to have substantively important associations with intervention status and the outcome, conditional on the other covariates and the intervention status.
· Type 2: A lagged outcome or proxy for the lagged outcome. A proxy for a lagged outcome is a variable that is similar to the outcome and measured before the intervention started. For example, if the outcome is a math test score then a lagged reading test score would be sufficient as a proxy.
Other desirable study features, when feasible Ideally, authors should use both of the covariate types described in the minimum criteria. When possible, the key covariates should include measures at the cluster-level if the intervention is at the cluster-level (i.e., if the intervention varies between clusters but not within cluster). These key covariates should include cluster averages of individual-level key covariates. Also, if subgroup analyses are being run then the cluster-level averages should also be calculated by subgroup.
Authors should also consider using additional covariates including some that capture local area conditions, like school district characteristics, and multiple domains of pre-intervention factors when available, like socio-emotional and academic factors.
24.5
Have the authors identified appropriate key covariates to check and/or adjust for selection bias?
Authors must use sampling, matching, weighting, and/or regression to adjust for the key covariates discussed above. Regression adjustment can be done using ordinary least squares or a multi-level model. When a lagged outcome is used authors can use regression adjustment, a repeated measures ANOVA, simple gain scores, difference-in-differences, or fixed effects to adjust for factors that would affect both the outcome and lagged outcome in the same way.
If the authors believe that the original intervention and comparison groups sampled are “reasonably similar” without any additional adjustments then they must show that similarity using their key covariates. “Reasonably similar” means that the differences between the key covariates’ effect size averages for the intervention and comparison groups are each less than 0.05 standard deviations. The standard deviations used should be pooled standard deviations combining information from the intervention and comparison groups.
Regression adjustment can be done using ordinary least squares or a multi-level model. When a lagged outcome is used authors can use regression adjustment, a repeated measures ANOVA, simple gain scores, difference-in-differences, or fixed effects to adjust for factors that would affect both the outcome and lagged outcome in the same way. The sample used to show similarity should be the analytic sample. Alternatively, the authors can follow WWC guidance regarding baseline equivalence which can differ, for example, when there is substantial missing data.
Other desirable study features, when feasible In most cases authors should consider using both matching (or weighting) and regression adjustment. Also, when doing matching or weighting on multiple covariates they should consider propensity score methods.
Authors should minimize effect size differences in baseline covariates between the intervention and comparison groups, especially for the lagged outcome or proxy, when available, and for variables likely to affect selection into the intervention status and later outcomes. When possible, they should use individual-level pooled standard deviations to calculate effect size differences rather than standard deviations based on aggregate data because the latter standard deviations are much smaller, which will inflate the effect size differences.
When authors find that the intervention and comparison groups are not reasonably similar, they should consider conducting sensitivity tests, for example by varying the set of interactions and higher-order terms used.
24.6
Have the authors used appropriate methods to evaluate differences between the intervention and comparison groups and adjust for those differences if needed?
The outcome must not directly affect intervention status. Thus, for example, authors should not present results showing effects of course taking in grades 9-12 on dropping out of high school since dropping out stops students from taking those courses. Similarly, they should not estimate the effect of receiving test accommodations at any point in high school on grade 11 test scores if those test scores were used to determine which students received those accommodations in grades 11 and/or 12.
Other desirable study features, when feasible No additional guidance.
24.7
Did authors avoid the possibility of reverse causality?
No requirements.
Other desirable study features, when feasible Authors are encouraged to adjust only for pre-intervention covariates that are necessarily exogenous, rather than covariates that might be endogenous. This means that they should avoid adjusting for covariates that capture conditions after the start of an intervention, like attendance or dosage. These types of covariates are often referred to as mediators and might have been affected by the intervention.
24.8
Does the study appropriately handle endogenous covariates?
Promising Practices Studies can draw implications for policy or practice. However, they should still avoid drawing causal conclusions to help ensure that the readers understand the limitations of the research.
24.9
Does the study address causality well in the writing?
D. ADDITIONAL GUIDANCE FOR IMPACT EVALUATIONS
RELs may propose and conduct impact evaluations that use either a randomized controlled trial (RCT) design or a quasi-experimental design (QED).
This is the start of the file's text. The full file is on GovTribe.
File details come from the government source that posted it. Updated .