J-11_-_Critical_Incident_Management_Overview.pdf
PDF 913 KB Posted
- Attached to
- Enterprise IT Shared Services (EITSS) Federal contract opportunity
- Solicitation number
- 693JK419R500005
About this file
This document includes a request for proposals for enterprise IT shared services. The Department of Transportation Office of the Chief Information Officer is seeking a contractor to provide infrastructure and standard operations support for the DOT Common Operating Environment. Proposals will be accepted in person only on January 8, 2019 between 10:00 am and 3:00 pm Eastern Time at DOT headquarters. Offerors should submit twenty thumb drives with their technical and business volumes in a sealed envelope including their company name and point of contact. The North American Industry Classification System code is 541513 with a small business size standard of $27.5 million. The unrestricted solicitation is open to all businesses while the small business set aside requires small businesses to perform 51% of the work.
J-11 - Critical Incident Management Overview
View the file
Other files for this federal contract opportunity
Show all 50
Enterprise IT Shared Services (EITSS) has more files on GovTribe.
On GovTribe
Work with this file on GovTribe
- Download the original file
- Contacts named in this file
- Similar government files
- Ask GovTribe AI about this file
Text version
Critical Incident Management Workflow
Identification & Recording
Resolution & Recovery
Closure Investigation &
Diagnosis
Contact IMC by Phone, E-mail, Web, Walk-up or Event
Create, Categorize & Prioritize
Incident
Sy st em o r
U se r
IM
C C
O E
Ti e
II
o
II
I Escalate to access system &/or Resolve
? Incident Closed
Incident Resolved
Complete Sat Survey
Resolve as
FCR
4b
Problem Mgt Process Change Mgt Process
M o d e
/O A
P
O C
Determine & Broadcast
Critical Incident Degradation or
Outage
Broadcast Critical Service
Restored
Se rv ic e
P ro vi d e r/ V e n d o r
Escalate to access system &/or resolve
Escalate to access system &/or resolve
4c
4d
4a
Triggers for Critical Incident Management
• Reported by Telephone, E-mail, or Walk-up
– Degradation or Outage - Service Desk or IMC receives greater than 5 Calls on a particular issue within 15-30 minutes that they cannot connect or having performance issues with a service or Security / SPAM issue. Service & Users are
IMPACTED.
– Example >5 Calls for VDI Degradation or Outage or SPAM E-mail received
• Reported by Event Management
– Degradation or Outage – IMC detects a Node is Down (Red) from SolarWinds or other Monitoring System, and the Node is not part of a redundant high availability architecture. Service & Users are IMPACTED.
– Example: Network Circuit, Switch, Router, Server or Power is down.
• Outage definition
– Priority 1 - Significant problem affecting multiple users; a critical business function or entire application is inaccessible; and multiple customers’ immediate work flow impacted. All functionality DOWN.
• Degradation definition
– Priority 2 - System operations are severely degraded; potential loss of critical business function is imminent; and can impact one or multiple customers. Partial functionality DEGRADED.
Critical Incident Impact, Urgency, & Priority Determination
• Impact – Reflects impact on number of people or systems affected.
– 1-Extensive/Widespread – Major impact to organization. Multiple users at multiple sites affected. Normal business cannot be conducted.
– 2-Significant/Large – Multiple users affected but only one site or business unit involved. Normal business impeded.
– 3-Moderate/Limited – One customer has lost complete functionality. Limited impact to business but one customer cannot do anything.
– 4-Minor/Localized – One customer is affected & lost some functionality, but can still work.
• Urgency – Indicates speed necessary to resolve incident.
– 1-Critical – Customer needs immediate response as they cannot perform normal duties.
– 2-High – Customer has a deliverable & cannot meet their deadline until issue resolved.
– 3-Medium – Customer can work in a limited way.
– 4-Low – Minor inconvenience to the customer due to the issue.
• Priority
– P1 Critical - Significant problem affecting multiple users; a critical business function or entire application is inaccessible; multiple customers’ immediate work flow impacted.
– P2 High - System operations are severely degraded; potential loss of critical business function is imminent; can impact one or multiple customers.
– P3 Medium - Operational performance of one system that is moderately impaired or may be needed, while most other business operations continue to function.
– P4 Low - Problems typically affect a single user or a routine service request; does not impact core workflow.
Critical Incident Management as Categorized by IMPACT, URGENCY, & PRIORITY in Remedy ITSM
Determines occurrence of DEGRADATION or OUTAGE OF COE IT Technical Services impact on DOT Mission &/or Business Services
Incident Management Center is Responsible & Accountable for End to End Critical Incident Management
Required Critical Incident Management Timeline
Steps Minutes Critical Incident Activities & Tasks
1 0 IMC notified of degradation or outage by Phone, Walk-up, E-mail or Event
2 10 Create Critical Incident using Critical Incident template
3 15 E-mail Internal Informational Broadcast Notification to Federal Managers, Operations Managers, and the Incident Management Center using Internal Broadcast template
4 20 Open IMC Conference Bridge
5 20 Notify Tier III, Modal POC &/or Vendor & coordinate substantiation, next steps & ETR of degradation or outage
6 25 E-mail External Degradation or Outage Broadcast Notification to impacted LIST-MODE NAME-ESCALATIONS- Group using External Broadcast template
7 30 Steps 1-5 MUST be completed in <30 minutes of notification of degradation or outage
8 45 E-mail follow-up Internal Degradation or Outage Broadcast Notification to Federal Managers, Operations Managers, and the Incident Management Center.
9 60 E-mail follow-up External Degradation or Outage Broadcast Notification to impacted LIST-MODE NAME- ESCALATIONS-Group
▪ Initiated & maintained by IMC Critical Incident Analyst until Critical Incident is resolved
▪ Maintained & logged in Remedy Incident Work Detail
▪ End to end timeline recording the outage or degradation from critical Incident identification, recording, broadcast notification, investigation, diagnosis, root cause, known error identification, workaround, resolution, recovery, closure & after action review.
Broadcasts go out hourly after the initial 60 minute mark has passed
File details come from the government source that posted it. Updated .