Attachment_7c_Systems_Approach_to_Training_NAVMC_1553.1_3.pdf

PDF 3 MB Posted

Attached to
Communication Training Center Instructional Services Federal contract opportunity
Solicitation number
M67854-15-R-7901
Issued by
United States Marine Corps

About this file

Attachment 7c Systems Approach to Training NAVMC 1553.1_3

View the file

Other files for this federal contract opportunity

Other files attached to Communication Training Center Instructional Services, newest first.
File Type Posted
Attachment_2-COURSE_DESCRIPTION_MatriX_v10_29_May_15_(3).xlsx XLSX spreadsheet
Amendment_1.doc DOC document
Attachment_6-After_Instruction_Report_BinderTemplate.pdf PDF
Q_and_A_spreadsheet3.xlsx XLSX spreadsheet
Attachment_8_Academic_Standard_Operating_Procedures_(ASOP).pdf PDF
A002_Grad_Roster1.pdf PDF
Attachment_3_Draft_Quality_Assurance_Surveillance_Plan.doc DOC document
Exhibit_A_INSTRUCTIONAL_SERVICES_Pricing_Spreadsheet-Section_B-Schedule_of_Supplies_and_Services_for_CLINs_0001-0051.xlsx XLSX spreadsheet
Attachment_6-After_Instruction_Report_BinderTemplate.pdf PDF
Attachment_1_Course_Description_Sheets.pdf PDF
A004_After_Instruction_Report_(AIR)_Binder.pdf PDF
Solicitation_M67854-15-R-7901.doc DOC document
A009_Corrected_Curriculum.pdf PDF
A003_Monthly_Report.pdf PDF
A006_Master_Lesson_File.pdf PDF
Exhibit_B_CURRICULUM_DEVELOPMENT_OTHER__Pricing_Spreadsheet-Section_B-Schedule_of_Supplies_and_Services_for_CLINs_0052-0060.xlsx XLSX spreadsheet
A008_Truancy_Report.pdf PDF
Attachment_5-Privacy_and_Security_Non-Disclosure_Statement.doc DOC document
A005_Weekly_Curriculum_Meetings.pdf PDF
Attachment_7a_Systems_Approach_to_Training_NAVMC_1553.1_1.pdf PDF
A007_Curriculum_Development_Plan_of_Action_and_Milestones_(POAM).pdf PDF
Attachment_12_Staffing_Plan.xlsx XLSX spreadsheet
Attachment_13-Tentative_Class_Schedule.xlsx XLSX spreadsheet
Attachment_9_Roster_Start-Grad.xlsx XLSX spreadsheet
A010_Developed_Curriculums.pdf PDF
Attachment_7b_Systems_Approach_to_Trainiing_NAVMC_1553.1_2.pdf PDF
Attachment_10_Roster_Start-Grad_(Template).xlsx XLSX spreadsheet
Attachment_2-COURSE_DESCRIPTION_MatriX_v10_29_May_15_(3).xlsx XLSX spreadsheet
A001_Start_Roster.pdf PDF
Attachment_4_DD_254.pdf PDF
Attachment_11_Past_Performance_Questionnaire_CTC.docx DOCX document
Show all 31

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

QUANTIFY AND INTERPRET DATA

Quantifying data is systematically assigning numbers to data allowing statistical analysis to be performed so that trends and relationships can be Identified and interpreted. Through quantifying the data, the interpretation of data Is possible.

For test items, item analysis Is used to quantify data so that defective test Items are identified. Another way that data Is quantified and Interpreted is through descriptive statistics. Some of the data may need to be coded prior to performing statistics. In these cases, it is important to understand the scales of measurement: nominal, ordinal, interval, and ratio. The scales of measurement provide an understanding of what statistical procedures can be performed for different types of instruments. Item analysis, descriptive statistics, and assigned numbers allow the evaluator to pinpoint trends. Trends can be defined as a pattern or prevailing theme. These trends can reveal strengths and weaknesses within the Instructional program. Interpreting data also Involves analyzing the test results for validity and reliability. This section will discuss what an evaluator is looking for in the test results to find out If the test is valid and reliable. The use of computer programs can make the process of data interpretation an easier task.

Data must be Interpreted to identify the problems so that recommendations can be made and solutions generated.

Item Analysis

Desaiptive Statistics

Scales of Measurement

Interpreting Quantified Data

Test Reliability and Validity

6-29

NAVMC 1553.1

27 Oct 2010

Figure 6-17. Quantify and Interpret Data.

ITEM ANALYSIS

NAVMC 1553.1

27 Oct 2010

Item analysis provides Information about the reliability and validity of test Items.

Reliability and validity are discussed later in this section. There are two purposes for dOing an item analysis. First, the analysis Identifies defective test items.

Secondly, it indicates areas where learners have not mastered the learning objective(s). Through item analYSiS, trends are identified as far as which test items are problematic. An example of one way to determine item difficulty and item discrimination can be found in Figure 6-18.

1. Item Difficulty. The frequency of students who answered an Item correctly determines the level of difficulty. For example, if 45 of 50 students answer an Item correctly, then the level of difficulty is low (.90) since 90 percent were able to answer correctly. However, If 10 out of 50 students answer correctly, then the level of difficulty Is high (.20). The Individual Response Report in the Student Evaluation module of McnMS provides the number and percentage of students who answered each item correctly. This makes It easy to determine the level of difficulty of an Item through percentages. The difficulty Index Is calculated below.

Pifficulty Index CD). Proportion of students who answered Item correctly.

p = Number of students selecting correct answer Total number of students attempting the test item p= 45 = .90

When Difficulty Index (p level) is less than about .25, the item is conSidered relatively difficult. When Difficulty Index (p level) Is above. 75, the Item is considered relatively easy. Test construction experts try to build tests with an average p level (difficulty) of about .50 for the test.

2. Item Discrimination. A percentage of high-test scorers (U) are compared to a percentage of low-test scorers (L) to determine how both groups of test scorers performed on the same item. To perform item discrimination, a percentage of high-test scorers and low-test scorers must be designated.

(Exampie:-COmpare the top 10% test scorers to the bottom 10% test scorers who answered the test item correctly.) If a high percentage from both groups missed the item, then more extensive evaluation of the test item and/or instructional process Is needed.

Item DiScrimination Index (D). Measure of the extent to which a test item discriminates or differentiates between students who perform well on the overall test and those who do not perform well on the overall test.

(Number who got Item (Number with item D '" correct in upper group) - correct in lower grouP)

Number of students In either group

Some experts Insist that 0 should be at least .30, while others believe that as long as D has a positive value, the item's discrimination ability is adequate.

6-30

There are three types of discrimination indexes:

a, Positive discrimination index - those who did well (U) on the overall test chose the correct answer for a particular item more often than those who did poorly (L) on the overall test.

b. Negative discrimination index - those who did poorly (L) on the overall test chose the correct answer for a particular item more often than those who did well (U) on the overall test.

c. Zero discrimination index - those who did well (U) and those who did poorly (L) on the overall test choose the correct answer for a particular item with equal frequency.

Item U M L Difficulty Discrimination (10 stu) (10 stu) (10 stu) (U + M + l) (U-l)

1 '7 4 3 14 4 2 10 10 9 29 1 3 8 6 4 18 4 4 4 4 6 14 -2 5 6 7 6 19 0 6 8 7 4 19 4 7 3 0 0 3 3 8 10 7 5 22 5 9 1 2 8 11 -7

10 8 5 3 16 5

The table above shows a simple analysis using a percentage of 33 percent to divide a class Into three groups - Upper (U), Middle (M), and Lower (L). For Instance, if you have a class of 30 students, then the students would be divided by test scores into the following groups: 10 (U) students (33 percent), 10 (M) students (33 percent), and 10 (L) students (33 percent).

Using the above table, a measure of item difficulty Is obtained by adding Upper (U) + Middle (M) + Lower (l). The difficulty index for item 2 is found by dividing 29 by 30 equaling .97 (97% of students answered correctly). Either the material is covered extremely well In the class or the question does not have convincing dlstracters. MCTIMS' Individual Response Report provides a look at the distracters and is discussed in the next section. On Item 7, 3 students answered the question correctly. This is an Indicator that the material has not been covered adequately, the test question Is poorly written, or answer Is miskeyed.

A rough index (ratiO) of the discriminative value (Upper test scorers compared to the Lower test scorers) of each item can be provided by subtracting the number of Individuals answering an Item correctly In the Lower (l) group from the number of individuals answering an item correctly in the Upper (U) group (Ex: U-L). Negative numbers indicate that there were more students from the Upper (U) group who missed the question. Positive numbers indicate that more students In the Lower (L) group missed the item. Zero indicates that there was no difference between the Upper (U) group and the Lower (L) group.

6-31

NAVMC l553.l 27 Oct 2010

Figure 6;'18.ltem. ..

Analysis; Number of , Learners Giving Correct Response In Each Criterion Group.

Figure 6-20. Frequency DIStribution.

NAVMC 1553.1

27 Oct 2010

1. Frequency of Collection. Descriptive statistics should be calculated every time a test, questionnaire, survey, etc., is administered. Even if these data are not used Immediately to summarize results in a report or to provide feedback to respondents, these data can be useful for future analysis to Identify trends or relationships, among groups.

2. Types of Descriptive Statistics' This section presents information and detail conceming descriptive statistics.

a. Frequency. Frequendes are determined by counting the number of occurrences. As example in Figure 6-20, the score 75 has a frequency of 3 because It occurs 3 times. Frequency counts are used to describe data (e.g., responses, scores, factors, variables) in raw numbers. Arranging variables Into a' frequency distribution makes the description of the variables easier than It would be if the scores were just listed In order. TO Illustrate, Figure 6-20 presents ten scores on a test and the same ten scores listed in a frequency distribution below.

llHl. Frequency counts are useful·for counting the number qf students who took a particular test, the number of students who passed a particular test, the number of students who selected answer A on item 13 of a test, the number of people who responded to a survey questionnaire, the number of people who rated an instructional program as effective, etc.

FREQUENCY-

Test Scores: 75, 75, 85, 90, 60, 65, 65, 75, 100, 85

Frequency Distribution

~ Fregueng

100 1 .' 90 1 85 2- 75 3 65 2 60 1

Appropriate Scale of Measurement. Frequency counts can be performed on data represented by nominal, ordinal, interval, and ratiO scales (Scales of Measure~ent will be discussed in detail later in this section).

6-32

~ -~ z :J (J

~ z w :J a w 0:

Graphic Representation. Frequency distribution data can be readily Interpreted by the use of graphs.

(1) The simplest graph( known as a frequency polygon( involves representing the frequency count (expressed in raw numbers or by percent) on the V-axis (vertical). Test scores should be divided Into equal intervals and plotted on the X-axis (horizontal). Using the data in Figure 6-20 and grouping the test scores in three intervals, Figure 6-21 displays the frequency distribution In graphic form. A frequency polygon is useful for displaying data within a group or data across groups. An example,of data within a group is student scores on a test. Subsequent class scores can be plotted on the same graph to display data across groups.

FREQUENCY POLYGON

4.5

3.5

2.5

1.5

0.5

0 50 100 150

TEST SCORES (x)

(3) Figure 6-22 presents different frequency distributions in graphic form. A frequency distribution is said to be "normal" when it

. represents a bell-shaped curve. It is Important to graph data to see if it is "normal" before performing any statistical analyses • .A frequency distribution In which scores trail off at either the high end or the low end of the spectrum Is said to be skewed. Where these scores trail off is referred to as the "tail" of the distribution. If the tail of the distribution extends toward the low or negative end of the scale, the distribution is considered negatively skewed; if the tail extends toward the high or positive end of the scale, the distribution is positively skewed.

6-33

NAVMC 1553.1

27 Oct 2010

Figure 6-21. Frequency Polygon.

Figure 6·22. Frequency Distributions.

f

SCORES

NORMAL

FrequencY Distributions-f

NEGATIVELY

SKEWED

NAVMC 1553.1

27 Oct 2010 f

SCORES

POSITIVELY

SKEWED

b. Measures of Central Tendency. While frequency distributions typically represent a breakdown of individual scores or variables among many, it is often useful to characterize a group as a whole. Measures of central tendency are measures of the location of the middle or the center of a distribution. The definition of "middle" or "center" Is p~rposely left somewhat vague so that the term "central tendency" can refer to a wide variety of measures. Three measures of central tendency are the mode, median, and mean. The mean is the most commonly used measure of central tendency. Figure 6-23 provides a description and sample of how to determine each.

-'-~~~1.f,f~-y~i}f:f~j;:Jr~'rit';;ij?- I The mode is the most frequently occurring response or I score.

I Mode Sample Test SCOres: 52, 78, 85, 88, 90, 93, 93, 100

Figure 6·23 •. Measures of ., Mode = 93 Centraftendency. \ , NOTE: More than one mode can exist in a set of data.

=1 =====~I ~~~~~~~~~~~ I . r-Th-e-m-edia-niS the score above and below which 50 ._-

I percent of the,sc_ ores in the sample fall. It Is sometimes referred to as the "breaking point."

1,1 1. Place numbers in order from least to greatest.

2. If number of scores is even, then the median Is the 1 Median central number or midpoint.

_ 3. If number of scores Is odd, then add the two middle

Mean

, scores and divide by two.

I Sample Test SCOres: 52, 78, 85, 88, 90, 93, 93, 100 1 88+90=178/2=89 ! Median = 89' j

I

Mean is the "average" score. I Sample Test Scores: ,52, 78, 85, 88, 90, 93, 93, 100 I 52+78+85+88+90+93+.93+100 = 679/8 = 84.875 Mean = 84.875

6-34

Figure 6-24 provides the scales of measurement (to be discussed next), the data types (I.e., test Items, questionnaires), and how the measures of central tendency can be used for each.

(1) H2ds:. As the most frequently occurring response, mode Is simple to compute. The mode Is not affected by extreme values.

However, it is usually not very descriptive of the data so it is Important that other measures of central tendency are used to describe the data.

A. Mode Is useful for determining what most students score on a given test or test item.

B. Mode is particularly useful for determining what response most students select In a multiple-choice test item, thereby allowing analysis of the Item's ability to clearly discriminate between correct and incorrect responses (a good multiple-choice test Item has a clear "correct" response and several plausible dlstracters).

(2) Median. Median is useful for splitting a group Into halves. The median is the middle of a dlstrlbution; half the scores are above the median and half are below the median. The median is less sensitive to extreme scores than the mean and this makes It a better measure than the mean for highly skewed distributions.

For example, the median Income Is usually more Informative than the mean income.

A. The median is not affected by extreme values and it always exists.

B. Though median is easy to compute, the numbers must be properly ordered to compute the correct '!'edlan.

(3) !:km. Mean is the "average."

A. Mean is calculated to produce an average response per test item across a class or to produce an average response per respondent.

B. Mean is also useful for determining overall attitudes toward a topic when using a Likert rating scale. For example, using a five-response Likert scale, a student rates the overall effectiveness of a course by answering 20 questions concerning course content, Instructor performance, use of media, etc. The value circled for each response can then be summed for a total score. This score is then divided by the number of questions

(20) to come up with the mean. In this case, the mean Is a total rating of course effectiveness.

6-35

27 Oct 2010

Flgpre 6-24. Type of Da~Measured By Central Tendency.

NAVMC 1553.1

27 Oct 2010

I

C. Mean Is generally the preferred measure of central tendency because it is the most ~nsistent or stable measure from sample to sample. The mean is good measure of central tendency for roughly symmetric distributions but can be misleading in skewed distributions since It can· be greatly Influenced by extreme scores.

For example, ten students score the following: 20,86,88,94,92, 90, 40, 88, 76, and 83. Although the mean is 76, it hardly reflects the typical score in the set. Mode or median may be more representative of that group's performance as a whole. When the distribution of scores is widely diSpersed, median is the most appropriate measure of central tendency. For example, If five students achieved test scores of 60, 65, 70, 72, and 74, and three students achieved scores of 90, 95, and 100, the overall class score should be reported as a median score. Since the scores achieved by the second group of students are much higher than those of the first group, calculating a mean score would inflate the value of the scores achieved by the lower scoring group. In this example, the mean score is 78, while the median score Is 73. When a distribution Is extremely skewed, it is recommended that all three measures be reported and the data be interpreted based on the direction and amount of skew.

Measure of Measurement Instrument Type of Data

Central Tendency

Scale Type Measured

Mode Nominal Scale . Student Data Most frequent score , Test Data Most frequent

Questionnaires answer- Interview.

Ordinal Scale Test Data - Useful for splitting

Median Interval Scale groups In halves i.e.

Ratio Scale Mastery and Non-

Masterv Ordinal Scale Test Data Avg. response per test item Questionnaires Avg. response per respondent Interview Overall attitudes toward topic/total

Mean rating of course effectiveness

Interval Scale Test Data Allows comparisons Ratio Scale Questionnaires of Individuals to

'. Interview overall class mean (test scores, , responses to particular items)

6-36

c. Variability, The variability of a set of scores Is the typical degree of spread among the scores. Range,·variance, and standard deviation are used to report variability.

(1) BJ..nu. Range Is the difference between the highest and the lowest, scores in the set. Range Is typically not the best measure of variability because it is dependent upon the spread in a set of scores, which can vary widely. For example, 10 students take a test and score as follows: 100, 92, 94, 94, 96, 100, 90, 93, 97, and 62. The range of scores varies from 100 to 62 so the range is 38 (100-62 = 38). If the lowest score were dropped, the range would be 10 (100- 90 = 10), which more accurately reflects the sample. Range serves as a rough index to variability and can be useful to report when the mean of a set of scores is not really representative due to a wide ranging of scores.

(2) variance' Variance is a more widely accepted measure of variability because it measures the average squared distance of the scores from the mean of the set in which they appear. An example of how to determine variance from a population Is shown in Figure 6-25. The variance (136) is the average of the squared deviation of the scores and is used to calculate standard deviation, which is the most widely accepted measure of variability.

Student Scores (X) lQQ

Number of Scores = 5

X-Mean

100-88 = 12

90-88 = 2 70-88 = -18 80-88 = -8

100-88 = 12

Mean = : X = ~ = 88 Number of Scores 5

Variance = (X-Mean)2 = 680 = 136 Number of Scores 5

Standard Deviation = v'136 = 11.7

6-37

(X-Mean)2

NAVMC 1553.l 27 Oct 2010

Figure 6-25, VarianCe and Standard Deviation of Test SCores. . .. ,' , 27 Oct 2010

(3) Standard DeWatlon. Standard deviation Is the 'square root of the variance for a set of variables. Standard deviation can reflect the amount of variability amo'ng a set of variables, responses, characteristics, scores, etc. In Figure 6-25, the variance score is

136. When the square root of 136 is taken, the standard deviation is 11.7. This means that the average distance of the students' scores from the class mean is 11.7. As another example, the mean score on a test Is 70 with a standard deviation of 10. Thus, the average amount students deviated from the mean score of 70 Is 10 points. If student A scored a 90 on the test, 20 points above the mean score, we interpret this as.a very good score, deviating from the norm twice as much as the average student. This Is often referred to as deviating from the mean by 2 standard deviation (SO) units (z score or standard score). If student B scored a 30 on the test, 40 points below the mean score, we Interpret this as a very bad score, deviating from the norm four times as much as the average student

SCALES OF MEASUREMENT

Scales of measurement specify how the numbers assigned to variables relate to what is being evaluated or measured. It tells whether a number is a label (nominal), a ranking order (ordinal), represented in equal intervals (interval), or describing a relationship between two variables (ratio). The type of measurement scale used affectsthe way data is statistically analyzed. Scales of measurement represent the varying degree of a particular variable. Figure 6-30 provides the types of statistical analysis that can be performed for different Instruments using the scales. Sample questions illustrating the use of the following scales can be found in Section 5604, Design Evaluation Instruments.

1. Nominal Scale. A nominal scale measurement is simply a classification system. For Instance, observation data can be labeled and categorized into mutually exclusive categories. Nominal numbering involves arbitrarily assigning labels to whatever is being measured. Assigning a 1 to a "yes" response and a 0 to a "no" response is an example of nominal numbering; .

so is assigning a 1 to "male" respondents and a 0 to "female" respondents.

Quantification of data by nominal numbering should be done only when an arbitrary number is needed to distinguish between groups, responses, etc.

Characteristics of a nominal scale are listed in Figure 6-26.'

6-38

Characteristics of a Nominal Scale It! Characterized by a lack of degree of magnitude. In other words, assigning a 1 to a variable does not mean that it is of a greater value than a variable assigned a O. Using the example below, answering "yes" is not of greater value than answering "no." The numbers serve only to distinguish among different responses or different characteristics.

YES = 1

NO = 0

o Does not reflect equal intervals between assigned numbers. For example, the numbers distinguIshing the military branches are just data labels.

Air Force = 1 Army = 2 Navy = 3 Marine Corps = 4

It! Does not have a true zero; because a variable is assigned a 0 does not mean that·it lacks the property being measured. Using the example below, assigning the number "0" to those who answered female on a student data sheet does not mean that the participant lacks gender.

MALE = 1

FEMALE = 0

6-39

NAVMC 1553.1

27 Oct 2010

'Flgu~:~·26.Characteristia.

Of'alNOWiiiial:ScaI.~" .,' '.,,'

Figure 6-27.

Characteristics of Ordinal Scale.

Figui'e6-30~' . . . .' ChariCteriftiCsi of Inte'nl'al Scale.

NAVMC 1553.1

27 Oct 2010

2. Ordinal Scale. The ordinal scale permits a "ranking" between values.

Differences cannot be "quantified" between two ordinal values. A Ukert scale is an example of an ordinal scale. For example, rating the effectiveness of Instruction from 1 {Ineffective) to 5 (very effective) permits comparisons to be made regarding the level of effectiveness; a larger number Indicates more of the property being measured. Characteristics of an ordinal scale are listed in Figure 6-27.

Characteristics of an Ordinal Scale Strongly Disagree Disagree Neutral Agree

Strongly ~gree

1 2 3 4 5 It] Degree of magnitude exists in ordinal numbers because each higher rating indicates more of the property being measured. Above, the level of agreement is being measured. A 5 indicates a higher level of agreement than a 2.

It] Equal intervals do not exist between ordinal numbers. For example, a rating of a 4 in the above example is not twice as effective as a rating of

2. Numbers used In an ordinal scale should not be added or multiplied because this can produce misleading results [i.e., two ratings of 2 (disagree) do not equal a single rating-of 4'(agree). A 4 means something totally different than a 2.

iii There is no true zero in an ordinal scale. In the above example, it Is meaningless to assign a 0 to a variable to Indicate a lack of effectiveness because a ratln of 1 indicates "ineffective."

3. Interval Scale. Interval numbering allows comparisons about the extent of differences between variables. For example, on test X, student A scored 20 points higher than student B. An example of an interval numbering system is a response to a question asking the respondent's age, the number of years in grade, etc. Characteristics of an Interval scale are listed In Figure 6-28. This will help the evaluator determine when to quantify data using an Interval scale.

Characteristics of an Interval Scale iii Degree of magnitude exists In Interval numbers because each higher rating indicates more of the property being measured. For example, a score of 95 Is better than a score of 90 on a test.

iii Equal Intervals exist between interval numbers. For example, 30 years In service Is 10 more years than 20 years In service, which is 10 more years than 10 years In service.

iii There is no true zero in an interval scale. Temperature is an example because temperatures can dip below 0 degrees and a temperature of 0 d rees does not indicate an absence of tern rature.

6-40

4. Ratio Scale; A ratio scale has equal Intervals and a meaningful zero point Point values assigned to responses to score a test is an example of a ratio scale. Ratio numbering permits precise relationships among variables to be made. For example, student A received a score of 40 on the test, which is twice as good as student 8's score of 20. Characteristics of a ratio scale are listed in Figure 6-31. This will help the evaluator determine when to quantify data using a ratio scale.

Characteristics of a Ratio Scale o Degree of magnitude exists in a ratio numbering scale. Test scores are an example of a ratio scale illustrating degree of magnitude (e.g., a score of 80 is better than a score of 70).

It! Equal intervals exist on a ratio numbering scale (e.g., a score of 90 is twice as good as a score of 45).

o A true zero exists in a ratio numbering scale (e.g., a score of 0 Indicates no score). A ratio numbering system Is typically used to quantify pass/fall data on a performance checklist, with "pass" quantified by a 1 and "fail" quantified by a O.

6-41

NAVMC 1553.1

27 Oct 2010

Figure 6·31. Characteristics of a Ratio Scale.

GUIDE TO OUANTIFYING DATA TO PERMIT STATISTICAL

ANALYSIS

The following is presented to aid the evaluator In quantifying data and selecting appropriate statistical analyses based on the evaluation instrument being used (see Figure 6-30).

See Figure 6-30 on the next page.

6-42

GUIDELINES FOR QUANTIFYING DATA TO PERMIT STATISTICAL ANALYSIS

Evaluation Scale of Instl"Ument Measurement

Multiple- Nominal choice Test Item

Ratio

True/False Ratio Test Item

Fill-In-the- Ratio blank Short- Answer Test Item

Examples of

Qualifying Data

A = 1, B = 2, C = 3, 0= 4, etc.

Point system:

1 = correct answer

0= Incorrect answer

Point system:

1 = correct answer

0= incorrect answer

Point system:

Points for correct response and partial credit

Statistical Analyses That Can Be Performed

Frequency counts of responses per test item Mode (most frequently selected responses per test Item)

Item Analysis (when used In conjunction with a ratio scale)

Frequency counts for correct/incorrect responses

0 Per test Item

0 Per student Mean (calculated to produce Item difficulty)

Median (score for overall test which splits dass in half)

Item analysis (but cannot determine where problem lie)

Overall test score per student

Variability of overall test scores

Frequency counts for correct/Incorrect responses

0 Per test Item

0 Per student

Mean (calculated to produce Item difficulty) Median (score for overall test which splits class. In half)

Item analysis Overall test score per student Variability of overall test scores

Frequency counts of responses

0 Per test item

0 Per student

Mean score per test item and per student Mode (most frequently scored points per test item)

Median (score for overall test which splits dass in half) Preliminary item analysis Overall test score per student Variability of overall test scores and points scored per test Item

Statistical Analyses

That Cannot Be Performed

Mean {average response per test item or per student)

Median Overall test score per student Varlabllity (range, variance, standard deviation)

Frequency counts for all Incorrect responses (dlstracters)

Mean (average response per test Item or per student) Mode (most frequently selected response per test Item)

Variability of responses per test Item

Mean (average response per test Item or student) Mode (most frequently selected response per test Item)

Figure 6-30. Guidelines for Quantifying Data to Permit Statistical Analysis.

6-43

GUIDELINES FOR QUAt1JTIFYING DATA TO PERMIT STATISTICAL ANALYSIS (cont.)

Evaluation Instrument

FIII-in-the, blank!

Short- Answer Test Item (cant.)

Performance- Based Test Item

Scale of MeaSurement

Ratio (cont.)

Nominal

RatiO

Examples of Qualifying

Data

. Point syste!l1:

l=correct answer O=incorrect answer

Categorize responses and assign a number to each response

·Point system:

l=pass O=fall

Statistical Analyses

That Can Be Performed

Frequency counts of corred/incorrect responses iii Per student 0 Per test Item

Mean (calculated to produce item difficulty)

Median

Preliminary item analysis

Overall test score per student

Frequency counts of ali responses per test Item

Mode (most frequently occurring response)

Item analysis

Frequency counts of pass/fail 0 Per student ' iii . Per test Item

Mean (calculated to produce item difficulty)

Median

Preliminary Item a-na~SiS

. Overall test score per student

Statistical Analyses

That Cannot Be Performed

Frequency counts for all incorrect responses

Mean (average response per test Item or per student)

Mode (most frequent response per test item)

Variability of responses per test Item

Mean (average response per test Item or per student)

Median

Overall test score per student

Variability of responses per

Mean (average response per test Item or per student)

MOde (most frequent response per test item)

Variability of outcomes per test item

Figure 6-30. Guidelines for Quantifying Data to Permit Statistical Analysis (cont.).

6-44

GUIDELINES FOR QUANTIFYING DATA TO PERMIT STATISTICAL ANALYSIS

(cont. )

Evaluation Instrument

Intervlewl Survey Questionnaire

Scale of Measurement

Nominal

Ordinal

Interval

Examples of Qualifying

Data

Categorize responses and assign a number to each response

Ukert scale

Response serves as the code when response is numerical (e.g., age, years in service)

Statistical Analyses

That Can Be Performed

Frequency counts of responses per Item

Mode (most frequently occurring response)

Frequency counts of responses per Item

Mean response per item

Mean response per student (assuming scale Is same throughout survey)

Median (response per item which splits respondent group In half)

Mode (most frequently occurring response per item)

Frequency counts of responses per Item

-I

Mean response per Item

Median (response per Item which splits respondent group In half)

Mode (most frequentfy occurring response)

Statistical Analyses

That C)nnot Be Performed

Frequency counts per student

Mean response per Item and per student

Mean response per student

Figure 6-30. Guidelines for Quantifying Data to Permit Statistical Analvsls (cont.).

6-45 fl9'g"J~"~l. Student Tes.: O'ata::.··

27 Oct 2010

INTERPRETING QUANTIFIED DATA

1. Multiple-Choice Test Item. Both nominal and ratio scales can be used for multiple-chOice test items. Using these scales to analyze multiple-choice test items is explained below.

a. Nominal Scale. Labels are assigned to different responses .. For example, in a 4-cholce Item, answer "a" Is coded as 1, answer "b" as 2, answer "e" as 3, and answer "d" as 4.

Test Item

1.

2.

3.

4.

5.

6.

7.

8.

9.

10.

(1) A nominal scale permits frequency counts, mode, and item analysis of Individual test items to be performed. Figure 6-34 presents data from three students who took the same lO-item test and their responses to each question. Next to each response is the number assigned to categorize the response (an asterisk Indicates an Incorrect response). Nominal numbers can be added across test items to calculate frequency counts (e.g., two out of three students selected response "a" on test Item 1; all three students selected response lib" on test item 2). Mode can be determined for an item by looking for the most frequently occurring response an item (e.g., the mode for test item 1 is "a").

Student / Student Student #1 #2 #3 a 1 *d 4 a 1 b 2 b 2 b 2

*a 1 *d 4 *b 2 c 3 c 3 *a 1 d 4 d 4 d 4 a 1 a 1 a 1 b 2 *d 4 b 2 d 4 *a 1 d 4 c 3 c 3 c 3 d 4 d 4 d 4

(2) Nominal numbers cannot be summed to provide an overall score on the test for each student because a nominal scale only assigns labels to responses and does not reflect degree of magnitude (a higher score does not reflect a better score). In Figure 6-31, It would be incorrect to sum the coded responses to provide an overall score of 2S for student #1, 30 for student #2, and 24 for student #3. In actuality, student #1 performed the best with only 1 incorrect answer, student #3 performed second best with two incorrect answers, and student #2 had four incorrect answers.

6-46

(3) A nominal scale cannot be used to calculate mean, median, or variability (range, variance, and standard deviation) because these data are meaningless in this context. For example, in Figure 6-34, a calculated mean or average response to test item #1 [(1 + 4 + 1) divided by 3 = 2J Is meaningless because it would reflect that the average response to Item #lls "b." It would also be incorrect to calculate a mean by, for example, adding student #1'5 scores for each item and dividing by the number of items (25 divided by 10) to produce a mean response of 2.5. To Interpret this would mean that the average response is halfway between a response of "b" and a response of "c," which Is a meaningless calculation.

b. Ratio Scale. A ratio scale can be used In conjunction with a nominal scale when quantifying responses to multiple-choice test items or it may be used as the only means of quantifying the data.

(1) If an evaluator is solely Interested in how many questions a student answers correctly,· a simple scoring system is needed to count the number of correct and incorrect responses so a total score for the test can be calculated for each student. To do .this, multiple-choice test items can be quantified using a ratio scale (e.g., 1 point is given to each correct answer and a 0 is given to each incorrect answer). This numbering system permits some frequency count data to be gathered (e.g., 22 of 50 students answered test item #1 correctly), but It does not permit frequency counts to be made across responses. This is because every incorrect response Is assigned a 0, making it impossible to discern how many students selected any response other than the correct response. This numbering system permits preliminary Item analysis to be performed (e.g., determining the percentage of students who got the answer right and those who did not), but It does not permit further item analysis to determine the Item difficulty level of each response.

(2) The evaluator can code the data using a ratiO scale by assigning paint values for correct responses and no points for an incorrect response.

This allows the calculation of an overall test score per student by summing the point values for each question. A median (i.e., score for overall test which splits the class In half) can also be calculated, as can the variability of overall test scores.

6-47

27 Oct 2010

(3) A ratio scale can also enable the calculation of mean to produce item difficulty rating. When responses are quantified with either of two numbers (e.g., 0 and 1), the evaluator can sum the responses to get a frequency count. The frequency counts relate to the number of correct

_ and Incorrect answers. The frequency count is then used to calculate item difficulty. Item difficulty Is calculated by dividing the number of students who got the item correct by the total number of students taking the test. Therefore, if 20 students answered a test Item correctly and five answered incorrectly, the item difficulty would be .80.

# of Students Who Answered Correctly = IQ = .SO # of Students Taking the Test 25

(4) Quantifying data using a ratio scale does not, however, permit calculation of a mean response per student or per test item. Variability is not calculated for the same reason.

(5) Mode (I.e., most frequently selected response per test Item) Is not calculated when using a ratiO scale on a mUltiple-choice test Item. This is because test data are coded as incorrect or correct rather than labeling all of the responses as is done with a nominal scale.

2. True/False Test Items. A true/false test item is typically quantified using a ratio scale (1 point for a correct response 0 points for an incorrect response).

This allows frequency counts, mean (calculated to produce item difficulty), median (overall test score that splits the class in half), an overall test score per student, and variability of overall tests scores to be calculated. However, a mean response per test item or per Student and a mode cannot be calculated because the actual response of "true" or "false" Is not quantified; the correctness of the answer is.

3. . Fill-In-the-Blank and Short-Answer Test Items. Fill-in-the-blank and short-answer test items can be quantified using a ratio scale and a nominal scale.

a. One method for quantlfyirig this type of data is to devise a scoring system so that answers are given paints based on the "correctness" of the response. This is typically done by creating an answer key that details the levels of acceptable responses to each question. For instance, a test question may require the student to list, In order, the seven essential qualities of leadership. The answer key may be established so that the student recelves·1 point for each correct quality listed and another 3 points if they are listed in correct order. This creates a scale of measurement that ranks performance on each item by the responseis level of correctness. This Is a good scale of measurement if there is some flexibility in the answers so that partial credit may be given to some Information.

6-48

(1) This type of scoring system permits frequency counts of responses per test item and per student, a mean score per test item, a mode (most frequently scored points) per test item, a median test score that splits the class in half, preliminary item analysis, an overall test score per student, the variability (range, variance, and standard deviation) of overall test scores, and the variability in the point spread among students per their overall test scores and per test Item.

(2) Item difficulty and item dlscriminability may be calculated per test item to determine the percentage of students who answered correctly and the percentage who did not. However, an analysis of responses to determine if students responded incorrectly, but in similar ways, cannot be performed. For Instance, It may be useful to know that students who missed a particular test question all responded with the same "wrong" answer. These data would help determine if the question was worded poorly so that it may be reworded in the future to remove any uncertainty or misinterpretation of Its meaning. This can only be accomplished through use of a nominal scale.

b. Another ratio scale involves establishing a scale of measurement with equal intervals and a true zero. Unlike the previous example where each response is keyed to a point system that mayor may not be the same for each response, this method uses a point system that is the same for all responses. Such a system may be as simple as assigning a 1 to a correct response and a 0 to an Incorrect response. This scale of measurement is only useful if there is a clearly defined correct and incorrect response for the item. This scoring system permits the same statistical analyses to be performed that a ratio scale for a multiple-choice test item permits.

c. Flll-in-the-blank and short-answer test items can also be quantified using a nominal scale, although this can be time consuming. To quantify data using a nominal scale, the responses must first be categorized into same or like responses. This can be difficult If the responses in the group vary greatly. If the responses can be categorized, the data are then quantified by assigning a number to each category through use of a nominal scale. Frequency counts, mode, and item analysis can be calculated. Mean (i.e., average response per test Item or per student), median, an overall test score per student, and variability cannot be calculated.

4. Performance·Based Test Items. Performance-based test items are typically pass/fail items quantified as either a 1 (pass) or a 0 (fail). This scoring system permits the same statistical analyses to be performed that a ratio scale for a multiple-choice test item permits.

6-49

NAVMC 1553.1

Figure 6-32. categqrizirtg.

ReSponses to an Opei1~ Ended Question.

NAVMC 1553.1

27 Oct 2010

5. Interview Data/Survey Ouestionnaires. Interview data and survey questionnaires are structured to collect data through flll-in-the-blank/short answer questions, multiple-choice items, and Likert rating scales.

a. Nominal

(1) FIlI-in-the=Blank/Short-Answer Response. Survey and Interview data of this nature can be difficult to quantify because they require a subjective judgment by the evaluator to categorize responses Into meaningful groups of like responses. Unlike test data, survey and Interview data are not quantified by "points" that can be added up for a total score but, rather, by using numbers to assign labels to responses (nominal scale). The difficulty lies in grouping the responses because an open-ended question can produce a multitude of different responses. For example, Figure 6-32 presents an open-ended question. Just below the question are the categories of responses· Identified during analysis ofthe test. The responses should be categorized into the smallest number of groups possible. In this example, all responses were easily categorized into one of five groups and quantified accordingly. Care should be taken when constructing a survey questionnaire to minimize fili-in-the-blank/short-answer Items so the data can be easily quantified and analyzed (see Section 5604). In this example, the question was better suited to be a multiple-choice item that could have been quantified readily by allowing respondents to select their responses.

CATEGORIZING RESPONSES TO AN OPEN-ENDED QUEmON

How often did you receive hands-on training with the equipment while attending the Radio Repairman Course?

Less than once a week = 1 Once a week = 2 Twice a week = 3 Three times a week = 4 More than three times a week = 5

(2) Multigll-g,oiCJ ResDOnH. Survey and interview data that use a . multiple-choice response format can be quantified like their counterpart knowledge-based test items using a nominal scale to assign labels to responses.

6-50

b. Ordinal. An ordinal scale is used to measure responses gathered using a Likert rating scale. A Likert rating scale Is the primary data collection tool that employs an ordinal scale. Typically, responses to a subject are rated across a continuum using a scale that varies from three to seven possible responses. The bottom of the scale typically represents a low amount of the property being measured while the top of the scale typically represents a high amount of the property being measured.

(1) Unlike knowledge and performance-based test items and other types of survey/intervIew questions, a Likert rating scale Is a measure where the mean (I.e., typical response per Item) per respondent is calculated.

When using a Likert scale, it is appropriate to add the responses and divide by the number of questlons per student to produce a student's overall response or attitude to a subject. For example, a survey evaluating the improvements made to a training program uses a 3-point Likert scale. Respondents answer questions cdncerning the Improvements made with 1 = "not improved," 2 ;:; "improved," and 3 = ligreatly Improved." In this example, it would be appropriate to calculate a mean response to the survey per student. It would be possible for a student's mean response to be 2.5 which could be interpreted as the training program overall Is considered to be .

improved. '

(2) A mean is calculated using a Likert scale only if the same scale is used throughout the survey and the whole survey measures the same topiC.

For example, half of a survey measures the effectiveness of graduate job performance on a 5-polnt Likert scale from "Ineffective" to "very effective." The other half of the survey measures graduate training in terms of effectiveness by using the same 5-polnt scale. It would be inappropriate to calculate an average response per respondent to the overall survey when the survey is measuring two different topics.

c. Interval. Responses to a survey questionnaire or Interv!ew that are numerical in nature (e.g., respondent's age, years in service) are quantified using an interval scale. An Interval scale quantifies the responses by the value of the response. If a respondent answers 23 to a question asking his age, his response is coded as 23. An interval scale permits the following statistics to be performed on a per Item basis only:

frequency counts, mean response, mode (most frequently occurring response), median (the response that splits the respondent pool in half), and variability (range, variance, and standard deviation). Unlike a Likert scale that may be the same scale used throughout a survey, an interval scale Is not usually the same throughout a survey. A survey is usually designed with interval questions to gather primarily demographic data.

Therefore, it is not appropriate to sum responses in an Interval scale to calculate the above descriptive statistics for the overall survey.

6-51

TEST RELIABILITY AND VAUDITY

The reliability and validity of a test provide the foundation for effectlve evaluation of student performance. Both the reliability and validity of a test should be assessed to identify the appropriateness of the test as an accurate measure of instructional effectiveness.

1. Reliability, Reliability refers to the ability of an instrument to measure skills and knowledge consistently. The reliability of a test is determined based on the calculation of a reliability coeffident (r). It is recommended that this coefficient be computed using a computer statistical analysis software package. A reliability coefficient is the correlation, or degree of association, between two sets of scores. Correlation coeffidents range from -1.0 to + 1.0. The closer a coefficient gets to -1.0 or to + 1.0, the stronger the relationship. The sign of the 'coefficient tells whether the relationship Is positive or negative.

Coefficient Strength Direction r = - .85 Strong Negative r = +.B2 Strong Positive r = +.22 Weak Positive

I r = +.03 Very Weak Positive I

I

· r = - .42 - Moderate· Negative I . __ ._. _____ . _______ . _______ ---.J

The different methOds of estimating reliability fall within three categories:

determining the Internal consistency of a test, determining the stability of a test over time, and determining the equivalence of two forms of a test

a. Test-Retest, Test-retest is a method of estimating reliability by giving the test twice and comparing the first set of scores and the second set of scores. For example, suppose a test on Naval correspondence is given to six students on Monday and again on the following Monday without any teaching between these times. - If the test scores do not fluctuate, then it is concluded that the test is reliable. The problem with test-retest reliability is that there Is usually some memory or experience involved the second time the test Is taken. Generally, the longer the Interval between test administration, the lower the correlation.

I " -.. -.' --------.. , .. -...... Fi;;t------·------------- ----S~_.;d ._-- ... -, Administration Administration I Student Score Score I

1 85 87 II

2 93 93 3 78 75 I

4 BO 85 I

I 5 65 61;

'I 6 83 . 80 ! ,. ____ • •• __ ·ri ____ • _____ ·~ __ ..... _________ •• __ • _______ ._. ____ ._, ____ ••• _____ ••• _'

6-52

b. Alternate Forms. If there are two equivalent forms of a test, these forms can be used to obtain an estimate of the reliability of the test. Both forms of the test are administered to the same group of students and the correlation between the two sets of scores is determined. If there is a large difference In a student's score on the two forms of the test that are suppose to measures the same behavior, then it Indicates that the test is unreliable. To use this method of estimating reliability, two equivalent forms of the test must be available and they must be administered under conditions as nearly equivalent as possible.

c. Split-Half Method. If the test in question is designed to measure a single basic concept, then the split-half method can be used to determine reliability.

To find the split-half (or odd-even) reliability, each item is assigned to one half or the other. Then, the total score for each student on each half is determined and the correlation between the two total scores for both halves is computed.

Essentially, one test Is used to make two shorter alternate forms. This method has the advantage that only one test administration is required, so memory or practice effects are not issues. This method underestimates what the actual reliability of the full test would be.

2. Interpreting Reliability

a. Scoring reliability limits test reliability. If tests are unreliably scored, then error is introduced that limits the reliability of the test.

b. The more items Included in a test. the higher the test's reliability, When more items are added to a test, the test is better able to sample the student's knowledge or skill that is being measured.

c. Reliability tends to deqease as tests are too eaSY or too difficult.

Score distributions become similar which makes it tough to know whether the instrument is measuring knowledge and skills consistently. When tests are too difficult, guessing is encouraged which creates a source of error in the test results.

3. Yalldity. The term validity refers to how well an instrument measures what it Is suppose to measure. Validity can be assessed for tests, questionnaires, Interviews, etc. However, validity is most often calculated for tests. Without establishing its validity, a test is of questionable usage since the evaluator does not know for sure whether the test Is measuring the concepts it Is intended to measure. There are several types of validity that can be determined.

a. Content Validity, Content validity assesses the relevance of the test Items to the subject matter being tested. Content validity Is established by examining an instrument to determine whether it provides an adequate representation of the skills and knowledge It Is designed to measure. No statistical test is used to establish content validity. To determine whether a test has content validity, SMEs review the test items and make a judgment regarding the validity of each item. For this approach to be effective, two major assumptions must be met.

First, the SMEs must have the background and expertise to make a judgment regarding the content of the test. Second, the objectives to which the test is compared must be valid.

6-53

b. Criterion-Related validity. Criterion-related validity Is…

This is the start of the file's text. The full file is on GovTribe.

File details come from the government source that posted it. Updated .