DD1423-2__CDRL_A023 Appendix Metrics.docx
DOCX document 55 KB Posted
- Attached to
- Information Technology Capabilities Contract (ITCC) II Federal contract opportunity
- Solicitation number
- FA4600-14-R-0017
About this file
This document contains metrics for monitoring and reporting on various information technology applications and services required under the Information Technology Capabilities Contract (ITCC) II awarded by the Department of the Air Force Air Combat Command. Key applications and services include the Mission Planning and Analysis System, Decision Support System - Classified, Force Survivability Analysis and Management System, Strategic Mission Assurance Data System, Data Federation and Synchronization, SMART.neXt cross-domain messaging, and the Enterprise Database. A variety of system availability, performance, and workload metrics are defined including operational availability, mean time between failures, number of users and transactions by hour, response times, data replication rates and latencies to measure continuity of operations. Reporting is required on a monthly basis with historical data and trend analysis when available. Records of system outages and their resolution must also be maintained for the life of the contract.
Not Listed
View the file
Other files for this federal contract opportunity
Show all 50
Information Technology Capabilities Contract (ITCC) II has more files on GovTribe.
On GovTribe
Work with this file on GovTribe
- Download the original file
- Contacts named in this file
- Similar government files
- Ask GovTribe AI about this file
Text version
Metrics The below table are a set of metrics that shall be delivered as part of the CDRL A023. Unless explicitly stated, metrics will be collected on a monthly basis for each of the systems identified in paragraph 2. Where historical data exists or as history is established, data from the prior 365 days should also be reported to aid in the identification of trends. Archive of the metrics and reports should be maintained for the life of the contract and be available upon request for analysis of longer term trends. Suggested report formats are identified, but this is flexible and can/should be adapted to be consistent with existing tools and processes.
Applications or Services requiring monitoring and reporting
1) Mission Planning and Analysis System (MPAS)
a) Air Vehicle Planning System (APS)
b) Analysis Tools
i) Integrated Weapons of Mass Destruction Toolkit (IWMDT)
ii) Extended Air Defense Simulation (EADSIM)
c) Automated Windows Planning System (AWPS)
d) Dynamic Application and Rapid Targeting System (DARTS)
e) Document Production System (DPS)
f) Decision Support Tool (DST)
g) National Target Base (NTB) and National Desired Ground Zero (NDL) Integrated Development System (NIDS II)
h) Theater Integrated Planning System (TIPS)
2) Decision Support System – Classified (DSS-C) consisting of
a) Force Status & Readiness (FSR) System – Secret
b) FSR - TS
c) Global Information Grid (GiG) Area of Responsibility (AOR) Decision Support System (GADSS)
3) Force Survivability, Analysis, and Management (FSAM) System
4) Strategic Mission Assurance Data System (SMADS).
5) Data Federation and Synchonization (DF&S)
6) SMART.neXt (Cross Domain & Legacy Messaging Data Service)
7) Enterprise Database (ISPAN EDB, C2EDB).
Metrics and Reporting v1.0
| Metric |
| Format |
| Description |
| Calculation |
| Purpose |
Application/Service Availability
| Ao – Operational availability of application or service |
| Value |
| For each of the applications or services listed in paragraph 2, report the end to end operational availability percentage for the reporting period (1 month) and for the service interval (1 year). |
| Ao = |
for all planned and unplanned outages
How available is my application capability end-to-end?
| Mean Time Between Failure (MTBF) |
| Value |
| For each of the applications or services listed in paragraph 2, report the average time duration between failures for the service interval (1 year) up to date for the service interval as of the current reporting interval. |
| MTBF = |
How often does my application capability fail?
| Metric |
| Format |
| Description |
| Calculation |
| Purpose |
Application/Service Workload & Performance
| Average # of users by the hour |
| Graph |
| For each of the applications listed in paragraph 2, report the average number of users by the hour. Every x seconds or minutes sample and store the number of users logged in to each application. Hourly, sum samples and divide by x (number of samples in one hour). Store hourly average. |
| Average # of users hourly = |
How busy is my application?
| Maximum # of users by the hour |
| Graph |
| For each of the applications listed in paragraph 2, every hour report the maximum number of users. Every x seconds or minutes sample and store the number of users logged in to each application. Compare the new sample to the saved sample, if greater discard the stored sample and save the new sample. Hourly, record the saved sample as the hourly maximum and zero the periodic sample. |
| Maximum # of users by the hour = |
When are my peak processing times?
| Application or Service Response Time by the hour |
| Graph |
| For each of the applications listed in paragraph 2, every 5 minutes run a representational set of application actions and record the response times. Hourly, calculate the average response time for each transaction and store. [endnoteRef:1] [1: Alternatively, calculate approximate response times from the application server.] |
Response time =
How does my application perform?
Application/Service Workload & Performance (database)
| # database transactions by the hour |
| Graph |
| For the enterprise databases (ISPAN EDB and C2 EDB) count the number of data requests and responses. Hourly, record the number of requests and reset the volume sum to zero. |
| # database transactions hourly = |
How many transactions do I process? How busy is my database?
| Database transaction volume by the hour |
| Graph |
| For the enterprise databases (ISPAN EDB and C2 EDB) continually sum the volume of data requests and responses. Hourly, record the volume of data and reset the volume sum to zero. |
| Total volume of database transactions hourly = |
How much total data does my database process? Can transaction size be optimized?
| Database transaction average latency by the hour |
| Graph |
| For the enterprise databases (ISPAN EDB and C2 EDB) continually sum the duration of each data request from initiation to completion. Hourly, divide the sum of the durations of data by the number of requests and record the resulting average. Reset the volume sum to zero. [endnoteRef:2] [2: If transaction size varies considerably, consider sub-dividing transactions into small, medium, and large categories to avoid skewing the results.] |
Database transaction average hourly latency =
How long does an average database transaction take for my database?
| Database transaction maximum latency by the hour |
| Graph |
| For the enterprise databases (ISPAN EDB and C2 EDB) transaction compare the duration to the stored maximum. If current duration is greater, replace the stored maximum with the current value. Hourly, record the stored maximum as the hourly maximum and reset the stored maximum to zero.ii |
| Maximum latency = |
maximum database transaction latency for each hourly interval What is the longest time a transaction takes? Can that transaction type be optimized?
| Metric |
| Format |
| Description |
| Calculation |
| Purpose |
Application/Service Workload & Performance (cross domain)
| # cross domain transactions by the hour |
| Graph |
| For the cross domain capabilities (Smart.neXt), continually keep a count of the cross domain transactions. Hourly, record the count and reset the counter to zero. |
| # cross domain transactions hourly = |
How many cross domain transactions do I process? How busy is my cross domain capability?
| Cross domain transaction volume by the hour |
| Graph |
| For the cross domain capabilities, continually sum the volume of cross domain requests. Hourly, record the volume and reset the volume sum to zero. |
| Total volume of cross domain transactions hourly = |
How much total data does my cross domain capability process? Can transaction size be optimized?
| Cross domain transaction average latency |
| Graph |
| For the cross domain capabilities, continually sum the duration of each cross domain request from initiation to completion. Hourly, divide the sum of the durations of data by the number of requests and record the resulting average. Reset the volume sum to zero.ii |
| Cross domain transaction average latency = |
How long does an average cross domain transaction?
| Cross domain transaction maximum latency |
| Graph |
| For each cross domain transaction, compare the duration to the stored maximum. If current duration is greater, replace the stored maximum with the current value. Hourly, record the stored maximum as the hourly maximum and reset the stored maximum to zero.ii |
| Maximum latency = |
maximum cross domain transaction latency for each hourly interval What is the longest time a cross domain transaction takes?
| Metric |
| Report Format |
| Description |
| Calculation |
| Purpose |
Application/Service Outages
| Outage Summary Report |
| List |
| List of all outages, by outage type, that impact a service or capability (application). |
| List of all outages during reporting cycle |
including outage cause and outage resolution For all planned and unplanned outages, when my capability was unavailable, what was the cause and how was it resolved?
| Actual Duration of planned outages |
| Histo-gram |
| Plot of duration for all planned outages. |
| Simple plot, no calculation |
| Are there any planned outages that were more significant? |
| Average duration of planned outages |
| Time value |
| Average time duration of outages where prior notification had occurred and approval is granted (scheduled in advance). |
| Average duration of planned outages = |
When an outage is planned, what is the average time my capability was unavailable?
Maximum duration of longest planned outage
| Time duration of the single outage which was scheduled in advance that occurred during the reporting period and had the longest duration. |
| Maximum duration of longest planned outage = |
When an outage is planned, what outage impacted availability the most?
Total duration of planned outages
| Sum total of time duration for all outages scheduled in advance during the reporting period. |
| Total duration of planned outages = |
How much total time was my application unavailable for scheduled downtimes?
| Deviation from planned outage time |
| Graph |
| For each planned outage, subtract planned outage time from actual outage time and plot. |
| Outage deviation = |
for each planned outage (actual outage time – planned outage time) How close are outage request to actual downtimes?
| Durations of unplanned outages |
| Histo-gram |
| Plot of duration for all unplanned outages. |
| Simple plot, no calculation |
| Are there any unplanned outages that were more significant? |
| Average duration of unplanned outages |
| Time value |
| Average time duration of outages where prior notification did not occur or no approval granted. |
| Average duration of unplanned outages = |
When an outage is unplanned, what is the average time my capability was unavailable?
Maximum duration of longest unplanned outage
| Time duration of the single unplanned outage during the reporting period that had the longest duration. |
| Maximum duration of all unplanned outages = |
When an outage is unplanned, what outage impacted availability the most?
Total duration of all unplanned outages
| Sum total of time duration for all unplanned outages during the reporting period. |
| Total duration of unplanned outages = |
for all unplanned outages
How much total time was my application unavailable for unscheduled downtimes?
Total downtime
| Sum total of time duration for all planned and unplanned outages during the reporting period |
| Total downtime = |
for all planned and unplanned outages
How much total time was my application unavailable?
| Metric |
| Format |
| Description |
| Calculation |
| Purpose |
Continuity of Operations
| Data replication rate |
| Graph |
| For an hour interval, keep a count of the # of replication transactions. At the end of the hour record the count and reset the counter to zero. |
| Data replication rate = |
number of transactions for each hour interval Does replication of my data to a remote continuity site correspond with my data transaction rates?
Average data replication latency
| For each replication transaction record the latency of the replication to the remote operating locations. Hourly, sum the latency and divide the # of replication transactions. Store hourly average. |
| Average data replication latency = |
If I need to use my remote site, how out of date will my data be?
Maximum data replication latency
| For each replication transaction during the reporting period, measure the latency. If latency of current transaction is greater than stored value then store current latency as the current maximum. Hourly store the current maximum value as the hourly maximum and reset current maximum to zero. |
| Maximum data replication latency = |
Is my data ever so out of date that it is not of what is operationally acceptable?
Table of Metrics and Reporting
Glossary
| Term |
| Definition |
| Application |
| Software that provides functions which are required by an IT service. Each application may be part of more than one IT service. An application runs on one or more servers or clients. See also application management; application portfolio. |
| Availability (Ao) |
| (ITIL Service Design) Ability of an IT service or other configuration item to perform its agreed function when required. Availability is determined by reliability, maintainability, serviceability, performance, and security. Availability is usually calculated as a percentage. This calculation is often based on agreed service time and downtime. It is best practice to calculate availability of an IT service using measurements of the business output. |
| COOP (Continuity of Operations) |
| Continuity of Operations (COOP) is a United States federal government initiative, required by U.S. Presidential directive, to ensure that agencies are able to continue performance of essential functions under a broad range of circumstances. |
National Security Presidential Directive-51 (NSPD-51)/Homeland Security Presidential Directive-20 (HSPD-20), National Continuity Policy, specifies certain requirements for continuity plan development, including the requirement that all federal executive branch departments and agencies develop an integrated, overlapping continuity capability. FCD 1 also serves as guidance to state, local, and tribal governments.
| Data latency |
| Measure of time delay that describes how long it takes for a packet of data to move from one designated point to another |
| Escalation |
| (ITIL Service Operation) An activity that obtains additional resources when these are needed to meet service level targets or customer expectations. Escalation may be needed within any IT service management process, but is most commonly associated with incident management, problem management and the management of customer complaints. There are two types of escalation: functional escalation and hierarchic escalation. |
| Closed |
| (ITIL Service Operation) The final status in the lifecycle of an incident, problem, change, etc. When the status is closed, no further action is taken. |
| Failure |
| Any service impacting event that results in the capability (application or service) to be unusable in accomplishing the mission from a user perspective. |
| First-line (1st call) resolution rate |
| Percentage of service desk incident requests that are resolved by 1st level support on the first call. |
| First-line (1st level) support |
| (ITIL Service Operation) The first level in a hierarchy of support groups involved in the resolution of incidents. Each level contains more specialist skills, or has more time or other resources. See also escalation. |
| Incident |
| (ITIL Service Operation) An unplanned interruption to an IT service or reduction in the quality of an IT service. Failure of a configuration item that has not yet affected service is also an incident – for example, failure of one disk from a mirror set. |
| Incident resolution |
| If service impacting an incident is considered resolved if service is restored. If non-service impacting an incident is considered resolved if configuration item causing the incident is repaired or incident is referred to next level of support. |
| IT |
| Information Technology |
| ITSM |
| IT service management |
| Event |
| (ITIL Service Operation) A change of state that has significance for the management of an IT service or other configuration item. The term is also used to mean an alert or notification created by any IT service, configuration item or monitoring tool. Events typically require IT operations personnel to take actions, and often lead to incidents being logged. |
| Metric |
| (ITIL Continual Service Improvement) Something that is measured and reported to help manage a process, IT service, or activity. |
| Mean time between failures (MTBF) |
| (ITIL Service Design) A metric for measuring and reporting reliability. MTBF is the average time that an IT service or other configuration item can perform its agreed function without interruption. This is measured from when the configuration item starts working, until it next fails. |
| Mean time to repair (MTTR) |
| The average time taken to repair an IT service or other configuration item after a failure. MTTR is measured from when the configuration item fails until it is repaired. MTTR does not include the time required to recover or restore. It is sometimes incorrectly used instead of mean time to restore service. |
| Mean time to restore service (MTRS) |
| The average time taken to restore an IT service or other configuration item after a failure. MTRS is measured from when the configuration item fails until it is fully restored and delivering its normal functionality. See also maintainability; mean time to repair. |
| Outage |
| Any planned or unplanned interruption to an IT service or reduction in quality of an IT service. |
| Patch |
| (Merriam-Webster.com) A minor correction or modification in a computer program. To mend, cover, fill up a hole or weak spot in; to apply a patch to (a computer program). |
| Planned outage |
| Any event that impacts an application or service that is scheduled in advance. |
| Problem |
| (ITIL Service Operation) A cause of one or more incidents. The cause is not usually known at the time a problem record is created, and the problem management process is responsible for further investigation. |
| Problem resolution |
| A problem is considered closed if a root cause fix has been implemented that resolves the problem and any associated incidents. |
| Remedy request |
| Reference for IT service request specifically at USSTRATCOM that draws its name from the BMC Remedy ITSM software solution used by J6/ITCC to manage service requests. |
| Reporting interval |
| Time interval between service reporting. Typically 1 month. |
| Resolution |
| (ITIL Service Operation) Action taken to repair the root cause of an incident or problem, or to implement a workaround. In ISO/IEC 20000, resolution processes is the process group that includes incident and problem management. |
| Response time |
| A measure of the time taken to complete an operation or transaction. Used in capacity management as a measure of IT infrastructure performance, and in incident management as a measure of the time taken to answer the phone, or to start diagnosis. |
| Root cause |
| (ITIL Service Operation) The underlying or original cause of an incident or problem. |
| Second-line (2nd level) support |
| (ITIL Service Operation) The second level in a hierarchy of support groups involved in the resolution of incidents and investigation of problems. Each level contains more specialist skills, or has more time or other resources. |
| Service desk |
| (ITIL Service Operation) The single point of contact between the service provider and the users. A typical service desk manages incidents and service requests, and also handles communication with the users. |
| Service interval |
| Defined time frame for calculation and reporting of a service metric. For example Operational Availability Ao might be calculated for the prior year long interval from the current reporting date. |
| Service level agreement (SLA) |
| (ITIL Continual Service Improvement) (ITIL Service Design) An agreement between an IT service provider and a customer. A service level agreement describes the IT service, documents service level targets, and specifies the responsibilities of the IT service provider and the customer. A single agreement may cover multiple IT services or multiple customers. See also operational level agreement. |
| Service request |
| (ITIL Service Operation) A formal request from a user for something to be provided – for example, a request for information or advice; to reset a password; to install a workstation for a new user; to install a new application; to provide hosting for a new capability. Service requests are managed by the request fulfilment process, usually in conjunction with the service desk. Service requests may be linked to a request for change as part of fulfilling the request. |
| Unplanned outage |
| Any event that impacts service or capability (application) that was not scheduled in advance. |
| User facing metric |
| Measurement of an IT event that has more direct correlation to the end user experience. |
File details come from the government source that posted it. Updated .