AttachmentD_WebDASspecs_12_pagenumbers.doc

DOC document 54 KB Posted

Attached to
2011-12 NATIONAL POSTSECONDARY STUDENT AID STUDY (NPSAS:12) Federal contract opportunity
Solicitation number
ED-IES-09-R-0015
Issued by
Department of Education Contracts and Acquisition Management

About this file

Attachment D - Data Analysis Input Specifications(PWS)

View the file

Other files for this federal contract opportunity

Other files attached to 2011-12 NATIONAL POSTSECONDARY STUDENT AID STUDY (NPSAS:12), newest first.
File Type Posted
AttachmentA_PWS 5-28-08.doc DOC document
AttachmentB_AwardFeeExhAB_12_revised 5-28-09.xls XLS spreadsheet
AttachmentF_AdditionalSecurityCompliance_12_pagenumbers.doc DOC document
AttachmentM_SmallBusinessSubcontractingPlan_12_pagenumbers.doc DOC document
RFP - NPSAS.pdf PDF
AttachmentE_Security_Procedures_12_pagenumbers.doc DOC document
Attachment N - Pricing Table.doc DOC document
AttachmentL_ContractorInformationForm_12_pagenumbers.doc DOC document
AttachmentJ_BusinessProposalWorksheet_12_revised_5-28-09.xls XLS spreadsheet
AttachmentG_DataElements_12.xls XLS spreadsheet
AttachmentI_TechnicalProposalWorksheet_12.xls XLS spreadsheet
AttachmentC_Billing_Instructions_12_pagenumbers.doc DOC document
Attachment O - Conflict of Interest Certification Form.doc DOC document
NPSAS Draft SOW 4-15-09.doc DOC document
Show all 14

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

ATTACHMENT D

Data Analysis Input File Specifications

NCES will create DATA ANALYSIS SYSTEMS (DAS) for projects upon receipt of six (6) files with the following specifications. The DAS will create tables and correlation matrices with Taylor standard errors and/or replicate standard errors. To prevent disclosure of individual data, the DAS will compress the data.

The five files are:

(1) INPUT.TXT

General system identification and inputs

(2) DATA.POS

Input data

(3) FMT.TXT

Describes the format of DATA.POS

(4) MASTER.TXT

Detailed information on the variables

(5) ORDER.TXT

Describes the ordering of sections within the DAS display

(6) README.TXT

Readme file for users

DAS systems always are designed for and operate in only one of eight possible modes of calculation, based entirely on the value of Calculation mode as given in the file INPUT.TXT.

INPUT.TXT

The INPUT.TXT file organizes the input for various miscellaneous input items, such as the system title, names of directories, etc. This file is formatted by lines—the parsing program expects certain input items to be in certain line numbers. The lines with # as the first character are comments. Users usually just modify an existing copy, such as the one below which was taken from NPSAS:90 Graduate/First Professionals. Do NOT add additional comment lines or remove existing ones, since then the items will be on the wrong line numbers.

The parsing program that creates the DAS uses INPUT.TXT for many things. The identification of prefix, file names, title, weights, and calculation mode must agree with information included in several other files.

<Next line is the first one in file INPUT.TXT>

# System Name

NPSAS:90 Graduate/First Professionals

# FPrefix = the prefix for this system, exactly three chars

N0G

# Name of MASTER.TXT file is masterg.txt

# Name of FMT.TXT is n0gmap.prn

# Name of the raw data file is n0gdata.raw

# Variable name for strata none

# Variable name for PSU none

# Calculation Mode

# Number of weights

# List of variable names for weights

WTA00

WTB00

WTC00

WTD00

<Above line is the last line in the file INPUT.TXT>

An important feature of this file is Calculation mode, which specifies what kind of system and standard errors are to be used. The possible values for Calculation mode are:

One weight, Taylor method

Handles cases where there is only one weight and standard errors are calculated by the Taylor method. The weight variable, the strata variable, and the PSU variable must NOT be included within MASTER.TXT. They must, however, be included in fmt.txt.

Multiple weights, Taylor method

Handles cases where two or more weights are available for selection by users. Standard errors will be calculated using the Taylor method based on the weight selected by the user. The variable names for the weights must be listed one per line within INPUT.TXT. The strata variable, and the PSU variable must NOT be included within MASTER.TXT, but they must be included in FMT.TXT.

Since the user will be able to select one of several weights for analysis, a description of each weight must be included within MASTER.TXT. Also, in MASTER.TXT, the description windows for variables other than weights must identify which weight users should select for analyses.

One set of BRR weights

One set of Jack knife I weights

One set of Jack knife II weights

Multiple sets of BRR weights

Multiple sets of Jack knife I weights

Multiple sets of Jack knife II weights

All weights must follow the following naming conventions. The names of all weight variables must begin with WT. The first, and in the case of one set of weights only, weight names must start with WTA. The estimation weight gets a suffix of 00—so the first estimation weight must be named WTA00. Each replicate of WTA00 is named beginning with WTA01 through WTAkk, where kk is the number of replicates. In the case of 32 replicate weights, DATA.POS must have 33 weights named WTA00, WTA01, WTA02,…, WTA32.

In cases where there is a second set of weights, the second estimate weight must be named WTB00 and the replicate weights must be named WTB01 through WTBkk. Additional sets of weights follow the alphabetical ordering. For example, four sets of replicate weights with 32 replicates there must be named using 132 variable names as follows: WTA00, WTA01, WTA02, …, WTA32, WTB00, WTB01, WTB02, …, WTB32, WTC00, WTC01, WTC02, …, WTC32, WTD00, WTD01, WTD02, …, WTD32. Replicates are not described in MASTER.TXT, but when multiple sets of replicate weights are specified, each of the WTA00, WTB00, …, WTk00 estimate weights must be included in MASTER.TXT.

The number of replicates will be identified by the parsing program that creates the DAS by the largest value of kk in the WTA00, WTA01, …, WTAkk set of variable names in FMT.TXT.

When Calculation mode is identified as a replicate method (3, 4, 5, 6, 7, or 8) the variable names for strata and PSU are ignored.

DATA.POS

This is the data file. The data file must contain integers greater than or equal to -1. Absolutely no real values or decimal points are allowed. No values less than -1 are allowed. It includes categorical (contiguous values 1, 2, 3, 4,..., and possibly zero) and continuous (zero and positive integer values) with uniform missing values of -1. See MASTER.TXT below for a description of when zeros are legitimate for categorical variables.

For Calculation modes 1 or 2, the variables needed for Taylor standard errors (primary sampling unit or PSU, Stratum) must be included on this file and the file must be sorted by Stratum and PSU. For Calculation modes 3, 4, 5, 6, 7, or 8, the variables for the replicate weights must be included on this file. Some systems may exist with Taylor and replicates—requiring both Stratum, PSU, and replicate weights, with separate INPUT.TXT files for the two Calculation modes.

The maximum record length within DATA.POS is 255. The only formats that are acceptable are F2.0, F4.0, F6.0, and F8.0. Formats should be selected with the smallest field width possible. Examples of formats which are NOT acceptable include F3.0, F10.0, F12.0, and F1.0. All records must be positional fixed format (see FMT.TXT below) and of fixed length (max = 255).

Weights should be rounded to integers.

FMT.TXT

This is a positional file that documents the format of DATA.POS. In column form, the following information must be included (one line for each variable in DATA.POS):

Start position
Width
Description
1
10
Variable name (alphanumeric)
11
5
Record number in file
16
5
Starting position (numeric)
21
5
Number of positions used
26
70
Variable label (table quality characters)

Variable names should be alphanumeric and begin with an alpha character. Underscores should NOT be used in variable names. The Variable name should be left-justified.

Record numbers must be in the range 1,2,3,...,99. Record number should be right-justified.

Starting position should be right-justified.

The only legitimate values for Number of positions used are 2, 4, 6, and 8. Number of positions used should be right-justified.

Variable labels should be ready for tabling--no abbreviations, capitalization for first letter of first word and proper nouns ONLY, no single quotes or double quotes should appear, and all words must be spelled correctly! Try to eliminate NUMBER OF, AMOUNT OF, and RESPONDENT from all labels. The Variable label should be left-justified.

Only the variables listed in FMT.TXT will be included in the final DAS, whether or not additional variables are included on the DATA.POS file. For QC, respondent tracking, or other purposes, other variables may be included on DATA.POS as long as they are NOT on the FMT.TXT file.

MASTER.TXT

This is a text file that documents the data in DATA.POS. It is structured in 4 parts:

Part 1. Label line

Col 1 = \

Col 2-9 = Variable name

Col 11-80 = Variable label

Example:

\VARNAM01 First variable, based on enrollment status

Part 2. Percentages and code labels block

(Consists of an appropriate number of lines, one line for each code value)

Col 1-3 = code value

Col 5-9 = weighted percentage

Col 11-49 = code label

Missing values (-1) are not labeled, but the -1 code and weighted percentage are included.

Continuous variables have their code labels enclosed in braces “{ }”. If the first code label in a percentage and code labels block is enclosed in braces, then the variable is marked as continuous, otherwise the variable is assumed to be categorical. {Zero} and {minimum-maximum, mean/standard deviation values} should be used for the code labels for continuous variables.

For categorical variables, second and later code labels may be enclosed in braces. Code labels not enclosed within braces are assumed to be categories, and the code values must be consecutive integers starting with one (1, 2, 3, 4, 5, ...). The lines with code labels enclosed within braces must all appear after the lines without braces. The lines without braces must be sorted ascending by the code values. In particular, zero is valid as a code value for categorical variables if and only if (1) the associated code label is enclosed in braces, and (2) it is not the first code in the code label block. Braced code values may be used for lumping and filtering only.

Do NOT under any circumstances report the actual weighted frequency for any variable/category with an unweighted frequency of less than thirty (30) observations in MASTER.TXT. Categories such as these may exist in the data file but their frequencies must be set to zero in MASTER.TXT.

Examples:

For a continuous variable:

0 10.0 {Zero} c 80.0 {1-999,999 , 44.3/5.32}

-1 10.0

For a categorical variable:

1 5.2 One

2 15.3 Two

3 59.5 Three

4 0.0 Four (this category had less than 30 unweighted)

0 12.8 {Zero}

-1 5.0

Part 3. Divider line(s)

Contains a \ in column 1 followed by the section label. There may be up to four section labels for each variable; however, only the first section label will be used in the CD version of the DAS.

For the web-DAS, the section labels identify the sets for variables for users based on drilling down a hierarchy. The underscore character “_” is used to identify the breaks in section labels for drill down expansions. The ordering of the section labels in the display is determined by the ordering within the file ORDER.TXT.

For CalcMode=2, 6, 7, and 8 WEIGHTS should have section labels that are “Weight”.

Part 4. Description block

Contains multiple lines that document item wording, recoding, imputation, source, and analytic cautions. When sources are used, they should begin on a separate line with the introductory word "Source: ", and should be placed after any other text in the description block. The description block may refer to Variable names within this DAS, but not to electronic codebook or other variables. Maximum line length is 70.

Cautions

Do NOT include PSU or STRATA into MASTER.TXT when the Calculation mode is 1 or 2. Do NOT include weights in MASTER.TXT when the Calculation mode is 1, 3, 4, or 5.

The variable names and variable labels in FMT.TXT and MASTER.TXT must match exactly. The variable labels should be of print quality—no abbreviations, properly capitalized, and reasonably short.

ORDER.TXT

Describes the ordering of sections within the DAS display by listing all section labels (see Divider line(s) section of MASTER.TXT) in the order they should be displayed. Use one line for each section label. ALL section labels must be included within this file and the ordering must be hierarchical.

README.TXT

This file is used to allow web-DAS users to select the proper DAS for their analysis. It should include general information concerning samples, timing, and policy topics covered by the variables. This file must convey sufficient information to all users to properly select the best DAS system, but it must be reasonably small—users will not read a huge discussion.

The following example is a README.TXT file for NPSAS:90 Graduate/First Professionals DAS:

<Next line is the first one in file README.TXT>

A sample of full- and part-time graduate and first-professional students enrolled during academic year 1989-90 provided these data. Topics include student financial aid (including assistantships), degree programs and major fields of study, employment and expected careers, demographic and family circumstances, plans for further education, goals, and values. These data were collected as part of the National Postsecondary Student Aid Study. Further information is available at http://nces.ed.gov/npsas/.

<above line is last one in file README.TXT>

File details come from the government source that posted it. Updated .