DOE-HDBK-1208-2012, Accident and Operational Safety Analysis, Volume I, Accident Analysis Techniques (Part 1 of 2, links to all Parts)
Functional areas: Functional Technical Basis, Accident Prevention, Investigation Principles, Investigation Practice
This handbook has been organized along a logical sequence of the application of the DOE "core analytical techniques" for conducting DOE Federal-, or contractor-let Accident Investigation of an OSR in order to prevent accidents. The analysis techniques presented in this Handbook have been developed and informed from academic research and validated through industry application and practice.
Version history and related documents
Document text
Text extracted from the attached file. Refer to the original document for the authoritative version.
Section 1
SENSI
NOT MEAS UREMENT
TIVE
DDOE‐HDBK‐11208‐2012
July 2012
DOEE HAANDBOOKK
Acccideent andd Opperaational
Saafetyy Annalyysis
Volumee I: Acccideent AAnalyysis
Tecchniqques
U.S. Deparrtmennt of Ennergy
Wasshingtoon, D.CC. 205 85
DOE‐HDBK‐1208‐2012
INTRODUCTION - HANDBOOK APPLICATION AND SCOPE
Accident Investigations (AI) and Operational Safety Reviews (OSR) are valuable for evaluating
technical issues, safety management systems and human performance and environmental
conditions to prevent accidents, through a process of continuous organizational learning. This
Handbook brings together the strengths of the experiences gained in conducting Department of
Energy (DOE) accident investigations over the past many years. That experience encourages us
to undertake analyses of lower level events, near misses and, adds insights from High Reliability
Organizations (HRO)/Learning organizations and Human Performance Improvement (HPI).
The recommended techniques apply equally well to DOE Federal-led accident investigations
conducted under DOE Order (O) 225.1B, Accident Investigations, dated March 4, 2011,
contractor-led accident investigations or under DOE O 231.1A, Chg. 1, Environment, Safety and
Health Reporting, dated June 3, 2004, or Operational Safety Reviews as a element of a
“Contractor Assurance Program.” However, the application of the techniques described in this
handbook are not mandatory, except as provided in, or referenced from DOE O 225.1B for
Federally-led investigations.
The application of the techniques described as applied to contractor-led accident investigations
or OSRs are completely non-mandatory and are applied at the discretion of contractor line
managers. Only a select few accidents, events or management concerns may require the level
and depth of analysis described in this Handbook, by the contractor’s line management.
This handbook has been organized along a logical sequence of the application of the DOE “core
analytical techniques” for conducting a DOE Federal-, or contractor-led Accident Investigation
or an OSR in order to prevent accidents. The analysis techniques presented in this Handbook
have been developed and informed from academic research and validated through industry
application and practice.
The techniques are for performance improvement and learning, thus are applicable to both AI
and OSR. This handbook serves two primary purposes: 1) as the training manual for the DOE
Accident investigation course, and the Operational Safety and Accident Analysis course, taught
through the National Training Center (NTC) and, 2) as the technical basis and guide for persons
conducting accident investigations or operational safety analysisi while in the field.
Volume I - Chapter 1; provides the functional technical basis and understanding of accident
prevention and investigation principles and practice.
Volume 1 - Chapter 2; provides the practical application of accident investigation techniques as
applicable to a DOE Federally-led Accident Investigation under DOE O 225.1B. This includes:
the process for organizing an accident investigation, selecting the team, assigning roles,
collecting and recording information and evidence; organizing and analyzing the information,
The term operational safety analysis for the purposes of this Handbook should not be confused with
application of other DOE techniques contained within nuclear safety analysis directives or standards
such as 10 CFR 830 Subpart B, or DOE-STD-3009.
Section 2
i
i
DOE‐HDBK‐1208‐2012
forming Conclusions (CON) and Judgments of Need (JON), and writing the final report. This
chapter serves as a ready easily available reference for Board Chairpersons and members during
an investigation.
Volume II provides the adaptation of the above concepts and processes to an OSR, as an
approach to go deeper within the contractor’s organization and prevent accidents by revealing
organizational weaknesses before they result in an accident.
Simply defined, the process in this Handbook includes:
Determining What Happened;
Determining Why It Happened and,
Developing Conclusions and Judgments of Needs to Prevent Re-Occurrence.
To accomplish this, we use:
Event and Causal Factor Charting and Analysis.
And, apply the core analytical techniques of:
Barrier analysis;
Change analysis,
Root cause analysis, and
Verification analysis.
Each of these analyses includes the integration of tools to analyze, DOE and Contractor
management systems, organizational weaknesses, and human performance. Other specific
analysis, beyond these core analytical techniques may be applied if needed, and are also
discussed in this Handbook.
ii
DOE‐HDBK‐1208‐2012
ACKNOWLEDGEMENTS
This DOE Accident and Operational Safety Analysis Handbook was prepared under the
sponsorship of the DOE Office of Health Safety and Security (HSS), Office of Corporate Safety
Programs, and the Energy Facility Contractors Operating Group (EFCOG), Industrial Hygiene
and Safety Sub-group of the Environmental Health and Safety (ES&H) Working Group.
The preparers would like to gratefully acknowledge the authors whose works are referenced in
this document, and the individuals who provided valuable technical insights and/or specific
reviews of this document in its various stages of development:
Writing Team Co-Chairs:
David Pegram, DOE Office of Health Safety and Security (HSS)
Richard DeBusk, Lawrence Berkley National Laboratory (LBNL)
Writing Team Members:
Marcus Hayes, National Nuclear Security Administration (NNSA)
Jenny Mullins, DOE Oak Ridge Site Office (ORO)
Bill Wells, Lawrence Berkley National Laboratory (LBNL)
Roger Kruse, Los Alamos National Laboratory (LANL)
Rick Hartley, Babcock and Wilcox Technical Services Pantex (BW-PTX)
Jeff Aggas, Savannah River Site (SRS)
Gary Hagan, Oak Ridge Y-12 National Security Complex (Y12)
Advisor:
Earl Carnes, DOE Office of Health Safety and Security (HSS)
Technical Editors:
Susan Keffer, Project Enhancement Corporation
Erick Reynolds, Project Enhancement Corporation
iii
DOE‐HDBK‐1208‐2012
iv
DOE‐HDBK‐1208‐2012
Table of Contents
INTRODUCTION - HANDBOOK APPLICATION AND SCOPE ................................................... i
ACKNOWLEDGEMENTS ........................................................................................................... iii
ACRONYMS ................................................................................................................................ xi
FOREWORD ................................................................................................................................. 1
Section 3
CHAPTER 1. DOE’S ACCIDENT PREVENTION AND INVESTIGATION PROGRAM ............1-1
1. Fundamentals.................................................................................................................. 1-1
1.1 Definition of an Accident................................................................................................1-1
1.2 The Contemporary Understanding of Accident Causation .........................................1-1
1.3 Accident Models – A Basic Understanding..................................................................1-2
1.3.1 Sequence of Events Model..................................................................................................1‐2
1.3.2 Epidemiological or Latent Failure Model ............................................................................1‐3
1.3.3 Systemic Model ...................................................................................................................1‐4
1.4 Cause and Effect Relationships ....................................................................................1-5
1.4.1 Investigations Look Backwards ...........................................................................................1‐5
1.4.2 Cause and Effect are Inferred .............................................................................................1‐6
1.4.3 Establishing a Cause and Effect Relationship ......................................................................1‐6
1.4.4 The Circular Argument for Cause ........................................................................................1‐6
1.4.5 Counterfactuals ...................................................................................................................1‐7
1.5 Human Performance Considerations............................................................................1-8
1.5.1 Bad Apples...........................................................................................................................1‐9
1.5.2 Human Performance Modes – Cognitive Demands ............................................................1‐9
1.5.3 Error Precursors ................................................................................................................1‐11
1.5.4 Optimization......................................................................................................................1‐13
1.5.5 Work Context ....................................................................................................................1‐13
1.5.6 Accountability, Culpability and Just Culture .....................................................................1‐15
1.6 From Latent Conditions to Active Failures.................................................................1-16
1.7 Doing Work Safely - Safety Management Systems ....................................................1-18
1.7.1 The Function of Safety Barriers .........................................................................................1‐20
1.7.2 Categorization of Barriers .................................................................................................1‐22
1.8 Accident Types/ Individual and Systems....................................................................1-25
1.8.1 Individual Accidents ..........................................................................................................1‐25
1.8.2 Preventing Individual Accidents ........................................................................................1‐26
Section 4
1.8.3 System Accident ................................................................................................................1‐27
1.8.4 How System Accidents Occur............................................................................................1‐28
1.8.5 Preventing System Accidents ............................................................................................1‐29
1.9 Diagnosing and Preventing Organizational Drift .......................................................1-30
v
DOE‐HDBK‐1208‐2012
1.9.1 Level I: Employee Level Model for Examining Organizational Drift ‐‐Monitoring
the Gap – “Work‐as‐Planned” vs. “Work‐as‐Done”..........................................................1‐31
1.9.2 Level II: Mid‐Level Model for Examining Organizational Drift – Break‐the‐Chain ...........1‐32
1.9.3 Level III: High Level Model for Examining Organizational Drift ........................................1‐35
1.10 Design of Accident Investigations ..............................................................................1-36
1.10.1 Primary Focus – Determine “What” Happened and “Why” It Happened ........................1‐37
1.10.2 Determine Deeper Organizational Factors .......................................................................1‐38
1.10.3 Extent of Conditions and Cause ........................................................................................1‐39
1.10.4 Latent Organizational Weaknesses ...................................................................................1‐39
1.10.5 Organizational Culture ......................................................................................................1‐41
1.11 Experiential Lessons for Successful Event Analysis ................................................1-45
CHAPTER 2. THE ACCIDENT INVESTIGATION PROCESS ..................................................2-1
2. THE ACCIDENT INVESTIGATION PROCESS ................................................................2-1
2.1 Establishing the Federally Led Accident Investigation Board and Its Authority ......2-1
2.1.1 Accident Investigations’ Appointing Official .......................................................................2‐1
2.1.2 Appointing the Accident Investigation Board .....................................................................2‐3
2.1.3 Briefing the Board ...............................................................................................................2‐5
2.2 Organizing the Accident Investigation..........................................................................2-6
2.2.1 Planning...............................................................................................................................2‐6
2.2.2 Collecting Initial Site Information .......................................................................................2‐6
Section 5
2.2.3 Determining Task Assignments ...........................................................................................2‐6
2.2.4 Preparing a Schedule ..........................................................................................................2‐7
2.2.5 Acquiring Resources ............................................................................................................2‐8
2.2.6 Addressing Potential Conflicts of Interest...........................................................................2‐9
2.2.7 Establishing Information Access and Release Protocols .....................................................2‐9
2.2.8 Controlling the Release of Information to the Public .......................................................2‐10
2.3 Managing the Investigation Process...........................................................................2-11
2.3.1 Taking Control of the Accident Scene ...............................................................................2‐11
2.3.2 Initial Meeting of the Accident Investigation Board .........................................................2‐12
2.3.3 Promoting Teamwork .......................................................................................................2‐13
2.3.4 Managing Evidence, Information Collection .....................................................................2‐15
2.3.5 Coordinating Internal and External Communication ........................................................2‐15
2.3.6 Managing the Analysis ......................................................................................................2‐17
2.3.7 Managing Report Writing..................................................................................................2‐18
2.3.8 Managing Onsite Closeout Activities ................................................................................2‐19
2.3.8.1 Preparing Closeout Briefings....................................................................................2‐19
2.3.8.2 Preparing Investigation Records for Permanent Retention .....................................2‐19
2.3.9 Managing Post‐Investigation Activities .............................................................................2‐21
2.3.9.1 Corrective Action Plans ............................................................................................2‐21
2.3.9.2 Tracking and Verifying Corrective Actions ...............................................................2‐21
2.3.9.3 Establishing Lessons Learned ...................................................................................2‐22
2.4 Controlling the Investigation .......................................................................................2-23
2.4.1 Monitoring Performance and Providing Feedback ...........................................................2‐23
2.4.2 Controlling Cost and Schedule ..........................................................................................2‐23
vi
Section 6
DOE‐HDBK‐1208‐2012
2.4.3 Assuring Quality ................................................................................................................2‐24
2.5 Investigate the Accident to Determine “What” Happened ........................................2-24
2.5.1 Determining Facts .............................................................................................................2‐24
2.5.2 Collect and Catalog Physical Evidence ..............................................................................2‐26
2.5.2.1 Document Physical Evidence ...................................................................................2‐28
2.5.2.2 Sketch and Map Physical Evidence ..........................................................................2‐28
2.5.2.3 Photograph and Video Physical Evidence ................................................................2‐29
2.5.2.4 Inspect Physical Evidence.........................................................................................2‐30
2.5.2.5 Remove Physical Evidence .......................................................................................2‐30
2.5.3 Collect and Catalog Documentary Evidence .....................................................................2‐31
2.5.4 Electronic Files to Organize Evidence and Facilitate the Investigation.............................2‐32
2.5.5 Collecting Human Evidence...............................................................................................2‐34
2.5.6 Locating Witnesses............................................................................................................2‐34
2.5.7 Conducting Interviews ......................................................................................................2‐35
2.5.7.1 Preparing for Interviews ..........................................................................................2‐35
2.5.7.2 Advantages and Disadvantages of Individual vs. Group Interviews ........................2‐36
2.5.7.3 Interviewing Skills ....................................................................................................2‐37
2.5.7.4 Evaluating the Witness’s State of Mind ...................................................................2‐39
2.6 Analyze Accident to Determine “Why” It Happened ..................................................2-40
2.6.1 Fundamentals of Analysis .................................................................................................2‐40
2.6.2 Core Analytical Tools ‐ Determining Cause of the Accident or Event ...............................2‐41
2.6.3 The Backbone of the Investigation – Events and Causal Factors Charting .......................2‐43
2.6.3.1 ECF Charting Symbols...............................................................................................2‐47
2.6.3.2 Events and Causal Factors Charting Process Steps ..................................................2‐47
2.6.3.3 Events and Causal Factors Chart Example ...............................................................2‐58
2.6.4 Barrier Analysis.................................................................................................................. 2‐60
2.6.4.1 Analyzing Barriers ....................................................................................................2‐60
Section 7
2.6.4.2 Examining Organizational Concerns, Management Systems, and Line
Management Oversight...........................................................................................2‐65
2.6.5 Human Performance, Safety Management Systems and Culture Analysis ......................2‐69
2.6.6 Change Analysis.................................................................................................................2‐69
2.6.7 The Importance of Causal Factors.....................................................................................2‐76
2.6.8 Causal Factors ...................................................................................................................2‐77
2.6.8.1 Direct Cause .............................................................................................................2‐78
2.6.9 Contributing Causes ..........................................................................................................2‐79
2.6.10 Root Causes .......................................................................................................................2‐79
2.6.10.1 Root Cause Analysis .................................................................................................2‐80
2.6.11 Compliance/Noncompliance .............................................................................................2‐83
2.6.12 Automated Techniques .....................................................................................................2‐86
2.7 Developing Conclusions and Judgments of Need to “Prevent” Accidents in
the Future ...................................................................................................................... 2-87
2.7.1 Conclusions .......................................................................................................................2‐87
2.7.2 Judgments of Need ...........................................................................................................2‐88
2.7.3 Minority Opinions .............................................................................................................2‐91
2.8 Reporting the Results...................................................................................................2-92
2.8.1 Writing the Report ............................................................................................................2‐92
vii
DOE‐HDBK‐1208‐2012
2.8.2 Report Format and Content ..............................................................................................2‐93
2.8.3 Disclaimer..........................................................................................................................2‐95
2.8.4 Appointing Official’s Statement of Report Acceptance ....................................................2‐95
2.8.5 Acronyms and Initialisms ..................................................................................................2‐96
2.8.6 Prologue ‐ Interpretation of Significance ..........................................................................2‐97
2.8.7 Executive Summary ...........................................................................................................2‐98
Section 8
2.8.8 Introduction ....................................................................................................................2‐100
2.8.9 Facts and Analysis ...........................................................................................................2‐102
2.8.10 Conclusions and Judgments of Need ..............................................................................2‐106
2.8.11 Minority Report...............................................................................................................2‐108
2.8.12 Board Signatures .............................................................................................................2‐108
2.8.13 Board Members, Advisors, Consultants, and Staff .........................................................2‐110
2.8.14 Appendices ......................................................................................................................2‐110
2.9 Performing Verification Analysis, Quality Review and Validation of
Conclusions ................................................................................................................2-111
2.9.1 Structure and Format ......................................................................................................2‐111
2.9.2 Technical and Policy Issues .............................................................................................2‐111
2.9.3 Verification Analysis ........................................................................................................2‐111
2.9.4 Classification and Privacy Review ...................................................................................2‐112
2.9.5 Factual Accuracy Review .................................................................................................2‐112
2.9.6 Review by the Chief Health, Safety and Security Officer ................................................2‐112
2.9.7 Document the Reviews in the Records ...........................................................................2‐112
2.10 Submitting the Report ................................................................................................2-113
Appendix A. Glossary ..................................................................................................... A-1
Appendix B. References ................................................................................................. B-1
Appendix C. Specific Administrative Needs ................................................................. C-1
Appendix D. Forms.......................................................................................................... D-1
Attachment 1. ISM Crosswalk and Safety Culture Lines of Inquiry ....................................1-1
Attachment 2. Bibliography .................................................................................................... 2-1
viii
Section 9
DOE‐HDBK‐1208‐2012
Table of Tables
Table 1‐1: Common Organizational Weaknesses ..............................................................................1‐40
Table 2‐1: DOE Federal Officials and Board Member Responsibilities ................................................2‐1
Table 2‐2: DOE Federal Board Members Must Meet These Criteria ...................................................2‐4
Table 2‐3: These Activities should be Included in an Accident Investigation Schedule.......................2‐7
Table 2‐4: The Chairperson Establishes Protocols for Controlling Information ................................2‐10
Table 2‐5: The Chairperson Should Use These Guidelines in Managing Information Collection
Activities. ...........................................................................................................................2‐17
Table 2‐6: Use Precautions when Handling Potential Blood Borne Pathogens .................................2‐28
Table 2‐7: These Sources are Useful for Locating Witnesses.............................................................2‐34
Table 2‐8: Group and Individual Interviews have Different Advantages ...........................................2‐37
Table 2‐9: Guidelines for Conducting Witness Interviews .................................................................2‐38
Table 2‐10: Benefits of Events and Causal Factors Charting ................................................................2‐46
Table 2‐11: Common Human Error Precursor Matrix ..........................................................................2‐53
Table 2‐12: Sample Barrier Analysis Worksheet ..................................................................................2‐64
Table 2‐13: Typical Questions for Addressing the Seven Guiding Principles of Integrated Safety
Management. ....................................................................................................................2‐67
Table 2‐14: Sample Change Analysis Worksheet .................................................................................2‐74
Table 2‐15: Case Study: Change Analysis Summary.............................................................................2‐75
Table 2‐16: Case Study Introduction ....................................................................................................2‐77
Table 2‐17: Compliance/Noncompliance Root Cause Model Categories ............................................2‐85
Table 2‐18: These Guidelines are Useful for Writing Judgments of Need ...........................................2‐90
Table 2‐19: Case Study: Judgments of Need ........................................................................................2‐90
Table 2‐20: Useful Strategies for Drafting the Investigation Report ...................................................2‐93
Table 2‐21: The Accident Investigation Report Should Include these Items .......................................2‐94
Table 2‐22: Facts Differ from Analysis ...............................................................................................2‐104
Table of Figures
Figure 1‐1: IAEA‐TECDOC‐1329 – Safety Culture in Nuclear Installations.............................................1‐8
Figure 1‐2: Performance Modes..........................................................................................................1‐11
Section 10
Figure 1‐3: Error Precursors ................................................................................................................1‐12
Figure 1‐4: Organizational Causes of Accidents ..................................................................................1‐17
Figure 1‐5: Five Core Functions of DOE’s Integrated Safety Management System ............................1‐20
Figure 1‐6: Barriers and Accident Dynamics – Simplistic Design ........................................................1‐21
Figure 1‐7: Individual Accident ............................................................................................................1‐26
Figure 1‐8: System Accident ................................................................................................................1‐28
Figure 1‐9: How System Accidents Happen ........................................................................................1‐29
Figure 1‐10: Prevent a System Accident................................................................................................1‐30
ix
DOE‐HDBK‐1208‐2012
Figure 1‐11: Level I ‐ “Work‐as‐Done” Varies from “Work‐as‐Planned” at Employee Level ................1‐32
Figure 1‐12: Level II ‐ Physics‐Based Break‐the‐Chain Framework .......................................................1‐35
Figure 1‐13: Level III ‐ High‐Level Model for Examining Organizational Drift .......................................1‐36
Figure 1‐14: Factors Contributing to Organizational Drift ....................................................................1‐37
Figure 1‐15: Assessing Organizational Culture ......................................................................................1‐42
Figure 2‐1: Typical Schedule of Accident Investigation .........................................................................2‐8
Figure 2‐2: Example of Electronic File Records To Keep for the Investigation....................................2‐33
Figure 2‐3: Analysis Process Overview ................................................................................................2‐42
Figure 2‐4: Simplified Events and Causal Factors Chart for the July 1998 Idaho Fatality CO2
Release at the Test Reactor Area ......................................................................................2‐49
Figure 2‐5: Sequence of Events and Actions Flowchart ......................................................................2‐50
Figure 2‐6: Decisions before Actions Flowchart ..................................................................................2‐50
Figure 2‐7: Conditions and Context of Human Performance and Safety Management Systems
Flowchart ..........................................................................................................................2‐51
Figure 2‐8: Context of Decisions Flowchart......................................................................................... 2‐52
Section 11
Figure 2‐9: Racked Out Air Breaker .....................................................................................................2‐59
Figure 2‐10: Excerpt from the Accident ECF Chart ................................................................................2‐60
Figure 2‐11: Summary Results from a Barrier Analysis Reveal the Types of Barriers Involved ............2‐61
Figure 2‐12: The Change Analysis Process ............................................................................................2‐70
Figure 2‐13: Determining Causal Factors ..............................................................................................2‐76
Figure 2‐14: Roll Up Conditions to Determine Causal Factors ..............................................................2‐78
Figure 2‐15: Grouping Root Causes on the Events and Causal Factors Chart .......................................2‐83
Figure 2‐16: Facts, Analyses, and Causal Factors are needed to Support Judgments of Need.............2‐89
x
DOE‐HDBK‐1208‐2012
ACRONYMS
AEC Atomic Energy Commission
AI Accident Investigation
AIB Accident Investigation Board
BAM Barrier Analysis Matrix
BTC Break-the-Chain
CAM Culture Attribute Matrix
CFA Causal Factors Analysis
CFR Code of Federal Regulations
CON Conclusions
CTL Comparative Timeline
DOE Department of Energy
DOE G DOE Guide
DOE M DOE Manual
DOE O DOE Order
DOE P DOE Policy
E1 Electrician 1
E2 Electrician 2
ECAQ Extraneous Conditions Adverse Quality
ECFA Expanded Causal Factors Analysis
ECF Events and Causal Factors
EFCOG Energy Facility Contractors Operating Group
ERDA Energy Research and Development Administration
ES&H Environment, Safety and Health
FOIA Freedom of Information Act
FOM Field Office Manager
HPI Human Performance Improvement
HRO High Reliability Organization
HSS Office of Health, Safety and Security
IAEA International Atomic Energy Agency
INPO Institute of Nuclear Power Operations
ISM Integrated Safety Management
xi
DOE‐HDBK‐1208‐2012
ISMS Integrated Safety Management System
IWD Integrated Work Document
LOW Latent Organizational Weakness Table
LOTO Lockout/Tagout
JON Judgment of Need
MOM Missed Opportunity Matrix
MORT Management Oversight and Risk Tree Analysis
NNSA National Nuclear Security Administration
NTC National Training Center
OPI Office of Primary Interest
ORPS Occurrence Reporting and Processing System
OSHA Occupational Safety and Health Administration
OSR Operational Safety Review
PM Preventive Maintenance
PPE Personal Protection Equipment
SSDC Safety Management System Center
TWIN Task, Work Environment, Individual Capabilities, Human Nature (TWIN)
Analysis Matrix (Human Performance Error Precursors)
WAD “Work-as-Done”
WAP “Work-as-Planned”
xii
DOE‐HDBK‐1208‐2012
FOREWORD
“The … (DOE) has exemplary programs for the control of accidents and fires, signified by
numerous awards. Its work in such areas as reactors, radiation, weapons, and research has
developed new methods of controlling unusual and exotic problems, including safe methods
of utilizing new materials, energy sources, and processes.
Despite past accomplishments, human values and other values stimulate a continual desire to
improve safety performance. Emerging concepts of systems analysis, accident causation,
human factors, error reduction, and measurement of safety performance strongly suggest the
practicality of developing a higher order of control over hazards.
Section 12
Our concern for improved preventive methods, nevertheless, does not stem from any specific,
describable failure of old methods as from a desire for greater success. Many employers
attain a high degree of safety, but they seek further improvement. It is increasingly less
plausible that the leading employers can make further progress by simply doing more, or
better, in present program. Indeed, it seems unlikely that budget stringencies would permit
simple program strengthening. And some scaling down in safety expenditures (in keeping
with other budgets) may be necessary.
Consequently, the development of new and better approaches seems the only course likely to
produce more safety for the same or less money. Further, a properly executed safety system
approach should make a major contribution to the organization's attainment of broader
performance goals.”
These words were written by W. G. Johnson in 1973, in The Management Oversight and Risk
Tree – MORT, a report prepared for the U.S. Atomic Energy Commission (AEC). While written
almost 40 years ago Johnson’s words and the context in which the MORT innovation in accident
prevention and investigation was developed remains as vital today as then. [Johnson, 1973]1
The MORT approach described in the report was converted into the first accident investigation
manual for the Energy Research and Development Administration (ERDA), the successor to
AEC, in 1974. In 1985 the manual was revised. The introduction to that revision explained that:
“In the intervening years since that initial publication, methods and techniques that were
new at that time have been further developed and proven, and Johnson's basic concepts and
principles have been further defined and expanded. Experience in using the manual in
conducting high quality, systematic investigations has identified areas for additional
development and has generated need for yet higher levels of investigative excellence to meet
today's safety and loss control needs.
This revision is intended to meet those needs through incorporating developments and
advances in accident investigation technology that have taken place since Johnson’s first
accident investigation manual was written.”
1
DOE‐HDBK‐1208‐2012
This new DOE Operational Safety and Accident Analysis Techniques Handbook was prepared in
the tradition of Johnson’s original report and its subsequent revisions. It incorporates
“developments and advances in accident investigation technology that have taken place since…”
issuance of the DOE Accident Investigation Workbook, Rev. 2, 1999.
What are those developments that prompted issuance of a new Handbook? One researcher
expresses the current situation thus:
“Accident models provide a conceptualisation of the characteristics of the accident, which
typically show the relation between causes and effects. They explain why accidents occur,
and are used as techniques for: risk assessment during system development, and post hoc
accident analysis to study the causes of the occurrence of an accident.
The increasing complexity in highly technological systems such as aviation, maritime, air
traffic control, telecommunications, nuclear power plants, space missions, chemical and
petroleum industry, and healthcare and patient safety is leading to potentially disastrous
failure modes and new kinds of safety issues. Traditional accident modelling approaches are
not adequate to analyse accidents that occur in modern sociotechnical systems, where
accident causation is not the result of an individual component failure or human error.”
[Qureshi, 2007]2
Section 13
In 1978, sociologist Barry Turner wrote a book called Man-Made Disasters, in which he
examined 85 different accidents and found that in common they had a long incubation period
with warning signs that were not taken seriously. Safety science today views serious accidents
not as the result of individual acts of carelessness or mistakes; rather they result from a
confluence of influences that emerge over time to combine in unexpected combinations enabling
dangerous alignments sometimes catastrophically. [Turner and Pidgeon, 1978]3
The accidents that stimulated the new safety science are now indelibly etched in the history of
safety: Challenger and Columbia, Three Mile Island, Chernobyl, Bophal, Davis Besse, Piper-
Alpha, Texas City, and Deepwater Horizon. The list is long. These accidents have introduced
new concepts and new vocabulary: normal accidents, systems accidents, practical drift, normal
deviance, latent pathogens, the gamblers dilemma, organizational factors, and safety culture. As
explained by Roger Boisjoly in an article after the 1986 Challenger accident: “It is no longer the
individual that is the locus of power and responsibility, but public and private institutions. Thus,
it would seem, it is no longer the character and virtues of individuals that determine the
standards of moral conduct, it is the policies and structures of the institutional settings within
which they live and work.” [Ermann and Lundman, 1986]4
The work of Johnson and the System Safety Development Center at the Idaho National
Engineering Laboratory was among the early contributions to a systems view. The accident at
the Three Mile Island nuclear plant in 1979 prompted new directions in safety and organizational
performance research going beyond human actions and equipment as initiating events to examine
the influence of organizational systems. Charles Perrow’s 1984 book, Normal Accidents,
challenged long held beliefs about safety and accident causation. Publication of his book was
followed by the Bhopal chemical leak (1984), the Chernobyl disaster (1986), and the Challenger
2
DOE‐HDBK‐1208‐2012
explosion (1986); contributed additional urgency for rethinking conventional wisdom about
safety and performance in complex systems. [Perrow, 1984]5
In 1987, the first research paper on what have come to be known as Highly Reliable
Organizations (HRO) was published, The Self-Designing High-Reliability Organization: Aircraft
Carrier Flight Operations at Sea by Gene I. Rochlin, Todd R. La Porte, and Karlene H. Roberts
published in the Autumn 1987 issue of Naval War College Review. HRO concepts were
formally introduced to DOE through the Defense Nuclear Federal Safety Board Tech 35 Safety
Management of Complex, High-Hazard Organizations, December 2004, and subsequently
adopted as design principles in the Department’s “Action Plan - Lessons Learned from the
Columbia Space Shuttle Accident and Davis-Besse Reactor Pressure-Vessel Head Corrosion
Event.” DOE’s adoption of Human Performance Improvement (derived from commercial
nuclear power and aviation successful approaches and socio-technical system research)
reinforced the findings of high reliability research with specific practices and techniques.
[Rochlin, La Porte, Roberts, 1987]6
Section 14
Early HRO studies were expanded to other hazardous domains over a period of some 20 years.
The broad body of research revealed common characteristics among diverse mission high hazard
organizations that are able to accomplish their missions safely over long time periods with few
adverse events. HRO research has been further expanded though the perspective of Resilience
Engineering. This perspective counters the historical deterministic view that safety is an inherent
property of well-designed technology and reveals how technology is nested in complex
interrelationships of social, organizational, and human factors. Viewing safety though the lens of
complexity theory illuminates an understanding that it is the ability of people in organizations to
adapt to the unexpected that produces resilient systems, systems in which safety is continually
created by human expertise and innovation under circumstances not foreseen or foreseeable by
technology designers.
Erik Hollnagel, a pioneer of the Resilience Engineering perspective, has explained that accident
investigation and risk assessment models focus on what goes wrong and the elimination of
"error.” While this principle may work with machines, it does not work with humans.
Variability in human performance is inevitable, even in the same tasks we repeat every day.
According to Hollnagel; our need to identify a cause for any accident has colored all risk
assessment thinking. Only simple technology and simple accidents may be said to be “caused.”
For complex systems and complex accidents we don't "find" causes; we "create" them. This is a
social process which changes over time just as thinking and society change. After the Second
World War and until the late 1970s, most accidents were seen as a result of technical failure.
The Three Mile Island accident saw cause shift from technical to human failure. Finally in the
1980s, with the Challenger disaster, cause was not solely technical or human but organizational.
Hollnagel and other resilience thinking proponents see the challenge not as finding cause. The
challenge is to explain why most of the time we do things right and to use this knowledge to shift
accident investigation and prevention thinking away from cause identification to focus on
understanding and supporting human creativity and learning and performance variability. In
other words, understanding how we succeed gains us more than striving to recreate an
unknowable history and prescribing fixes to only partially understood failures. [Hollnagel,
2006]7
3
DOE‐HDBK‐1208‐2012
It has been suggested that we are living in the fifth age of safety. The first was a technical age,
the second a systems age, and the third a culture age. Metaphorically, the first may be
characterized by engineering, the second by cybernetics and systems thinking, and the third by
psychology and sociology. The fourth age, the “integration age,” builds on the first three ages
not abandoning them but blending them into a trans-disciplinary socio-technical paradigm, thus
prompting more complex perspectives to develop and evolve. The fifth age is an “adaptive age.”
It does not displace the former, but rather transcends the other ages by introducing the notion of
complex adaptive systems in which the roles of expertise, professional practice, and naturalistic
observation attain primacy in resolving the duality of “work-as-imagined” versus “work as
done.” [Borys, Else, Leggett, October 2009]8
Section 15
At present, we see mere glimpses of the implications of the adaptive age on how we think about
“accident investigation.” How we may view accidents though fourth Age lens is somewhat
clearer. Though still myopic, we do have examples of fourth age investigation reports beginning
with the Challenger Accident. Dianne Vaughn wrote, “The Challenger disaster was an accident,
the result of a mistake. What is important to remember from this case is not that individuals in
organizations make mistakes, but that mistakes themselves are socially organized and
systematically produced. Contradicting the rational choice theory behind the hypothesis of
managers as amoral calculators, the tragedy had systemic origins that transcended individuals,
organization, time and geography. Its sources were neither extraordinary nor necessary
peculiar to NASA, as the amoral calculator hypothesis would lead us to believe. Instead, its
origins were in routine and taken for granted aspects of organizational life that created a way of
seeing that was simultaneously a way of not seeing.” [Vaughan, 1996]9
The U.S. Chemical Safety Board enhanced our fourth age vision by several diopters in its report
on the British Petroleum Texas City Refinery accident. Organizational factors, human factors
and safety culture were integrated to suggest new relationships that contributed to the nation’s
most serious refinery accident. Investigations of the Royal Air Force Nimrod and the Buncefield
accidents followed suit. More recent investigations of the 2009 Washington Metro crash and the
Deepwater Horizon catastrophe were similarly inspired by the BP Texas City investigation and
the related HRO framework.
This revision of DOE’s approach to accident investigation and organizational learning is by no
means presented as an exemplar of fifth nor even fourth age safety theory. But it was developed
with awareness of the lessons of recent major accident investigations and what has been learned
in safety science since the early 1990s. Still grounded in the fundamentals of sound engineering
and technical knowledge, this version does follow the fundamental recognition by Bill Johnson
that technical factors alone explain little about accidents. While full understanding of the
technology as designed is necessary, understanding the deterministic behavior of technology
failure offers little to no understanding about the probabilistic, even chaotic interrelationships of
people, organization and social environmental factors.
The Handbook describes the high level process that DOE and DOE contractor organizations
should use to review accidents. The purpose of accident investigation is to learn from experience
in order to better assure future success. As Johnson phrased it: “Reduction of the causes of
failures at any level in the system is not only a contribution to safety, but also a moral obligation
4
DOE‐HDBK‐1208‐2012
to serve associates with the information and methods needed for success.” We seek to develop
an understanding of how the event unfolded and the factors that influenced the event. Classic
investigation tools and enhanced versions of tools are presented that may be of use to
investigators in making sense of the events and factors. Further-more, an example is provided of
how such tools may be used within an HRO framework to explore unexpected occurrences, so
called “information rich, low consequence, no consequence events”, to perform organizational
diagnostics to better understand the “work-as-imagined” versus “work-as-performed” dichotomy
and thus maintain reliable and resilient operations. [Johnson, 1973]1
Section 16
This 2012 version of the Handbook retains much of the content from earlier versions. The most
important contribution of this new version is the reminder that tools are only mechanisms for
collecting and organizing data. More important is the framework; the theory derived from
research and practice, that is used for interpreting the data.
Johnson’s 1973 report contained a scholarly treatment of the science and practice that underlay
the techniques and recommendations presented. The material presented in this 2011 version
rests similarly on extensive science and practice, and the reader is challenged to develop a
sufficient knowledge of both as a precondition to applying the processes and techniques
discussed. Johnson and his colleagues based the safety and accident prevention methodologies
squarely on the understanding of psychology, human factors, sociology, and organizational
theory. Citing from the original AEC report “To say that an operator was inattentive, careless
or impulsive is merely to say he is human” (quoting from Chapanis). “…each error at an
operational level must be viewed as stemming from one or more planning or design errors at
higher levels.”
This new document seeks to go a step beyond earlier versions in DOE’s pursuit of better ways to
understand accidents and to promote the continuous creation of safety in our normal daily work.
Fully grounded in the lessons and good practices of those who preceded us, the contributors to
this document seek as did our predecessors to look toward the future. This Accident and
Operational Safety Analysis Techniques Handbook challenges future investigators to apply
analytical tools and sound technical judgment within a framework of contemporary safety
science and organizational theory.
5
DOE‐HDBK‐1208‐2012
6
DOE‐HDBK‐1208‐2012
CHAPTER 1.
DOE’S ACCIDENT PREVENTION AND INVESTIGATION PROGRAM
1. Fundamentals
This chapter discusses fundamental concepts of accident dynamics, accident prevention, and
accident analysis. The purpose of this chapter is to emphasize that DOE accident investigators
and improvement analysts need to understand the theoretical bases of safety management and
accident analysis, and the practical application of the DOE Integrated Safety Management (ISM)
framework. This provides investigators the framework to get at the relevant facts, surmise the
appropriate causal factors and to understand those organizational factors that leave the
organization vulnerable for future events with potentially worse consequences.
1.1 Definition of an Accident
Accidents are unexpected events or occurrences that result in unwanted or undesirable outcomes.
The unwanted outcomes can include harm or loss to personnel, property, production, or nearly
anything that has some inherent value. These losses increase an organization’s operating cost
through higher production costs, decreased efficiency, and the long-term effects of decreased
employee morale and unfavorable public opinion.
How then may safety be defined? Dr. Karl Weick has noted that safety is a “dynamic non
event.” Dr. James Reason offers that “safety is noted more in its absence than its presence.”
Scholars of safety science and organizational behavior argue, often to the chagrin of designers,
that safety is not an inherent property of well designed systems. To the contrary Prof. Jens
Rasmussen maintains that “the operator’s role is to make up for holes in designers ‘work’.” If
the measurement of safety is that nothing happens, how does the analyst then understand how
systems operate effectively to produce nothing? In other words, since accidents are probabilistic
outcomes, it is the challenge to determine by evidence if the absence of accidents is by good
design or by lucky chance. Yet, this is the job of the accident investigator, safety scientists and
analysts.
Section 17
1.2 The Contemporary Understanding of Accident Causation
The basis for conducting any occurrence investigation is to understand the organizational,
cultural or technical factors that left unattended could result in future accidents or unacceptable
mission interruption or quality concerns. Guiding concepts may be summarized as follows:
Within complex systems human error does not emanate from the individual but is a bi
product or symptom of the ever present latent conditions built into the complexity of
organizational culture and strategic decision-making processes.
The triggering or initiating error that releases the hazard is only the last in a network of
errors that often are only remotely related to the accident. Accident occurrences emerge
1
.2
1
.1
1‐1
u
r
f d
DOE‐HDBK‐1208‐20012
1
.3
fromm the organizzation’s commplexity, takiing many facctors to overrcome systemms’ networkk of
barriiers and alloowing a threaat to initiate the hazard rrelease.
Inveestigations reequire delvinng into the basic organizzational processes: designning,
consstructing, op erating, mai ntaining, commmunicating, selecting, and trainingg, supervisinng,
and managing thhat contain thhe kinds of llatent condittions most likkely to constitute a threaat to
the ssafety of the system.
The inherent natture of organnizational cuulture and strrategic decission-making means latennt
condditions are innevitable. Syystems and oorganizationnal complexiity means noot all problemms
can bbe solved inn one pass. RResources arre always limmited and saffety is only oone of manyy
commpeting priorities. There fore, event i nvestigatorss should targget the latent conditions mmost
in neeed of urgennt attention aand make theem visible too those who mmanage the organizationn so
theyy can be corrected. [Holllnagel, 20044]10 [Dekker,, 2011]11 [Reeiman and OOedewald,
9]122009
1.3 AAccident Models –– A Basic Understanding
An acciddent model iss the frame oof reference, or stereotyppical way of thinking aboout an accident,
that are uused in tryingg to understaand how an accident happpened. Thee frame of reeference is offten
an unspooken, but commmonly heldd understandding, of how accidents occcur. The addvantage is tthat
communiication and uunderstandinng become mmore efficiennt because soome things ( e.g., commoon
terminoloogy, commoon experienc es, commonn points-of-reeference, or typical sequuences) can bbe
taken forr granted. Thhe disadvanttage is that itt favors a sinngle point off view and ddoes not conssider
alternate explanationns (i.e., the hypothesis mm a recognizeed solution, ccausing the uusery odel creates
to discardd or ignore iinformation iinconsistent with the moodel). This iis particularlly important
when adddressing humman componnent because preconceiveed ideas of hhow the acciddent occurreed
can influence the invvestigators’ aassumptions of the peoplles’ roles andd affect the lline of
questioniing. [Hollnaggel, 2004]10
What invvestigators loook for whenn trying to unnderstand annd analyze aan accident ddepends on hhow
it is belieeved an acciddent happens. A model,, whether forrmal or simpply what youu believe, is
extremelyy helpful be cause it brinngs order to aa confusing situation andd suggests wways you cann
explain relationships. But the moodel is also cconstrainingg because it vviews the ac cident in a
particularr way, to thee exclusion oof other viewwpoints. Acccident modeels have evollved over timme
Section 18
4]10and can bbe characteriized by the tthree modelss below. [Hoollnagel, 2000
1.3.1 SSequence oof Events MModel
This is a simple, line ar cause andd effect mod el where
accidentss are seen thee natural cullmination off a series
of eventss or circumsttances, which occur in a specific
and recoggnizable ordder. The moddel is often
representted by a chaiin with a weeak link or a series of
falling doominos. In tthis model, aaccidents aree
preventedd by fixing oor eliminatinng the weak link, by
removingg a domino, or placing a barrier betwween two
1‐2
DOE‐HDBK‐1208‐2012
dominos to interrupt the series of events. The Domino Theory of Accident Causation developed
by H.W. Heinrich in 1931 is an example of a sequence of events model. [Heinrich, 1931]13
The sequential model is not limited to a simple series and may utilize multiple sequences or
hierarchies such as event trees, fault trees, or critical path models. Sequential models are
attractive because they encourage thinking in causal series, which is easier to represent
graphically and easier to understand. In this model, an unexpected event initiates a sequence of
consequences culminating in the unwanted outcome. The unexpected event is typically taken to
be an unsafe act, with human error as the predominant cause.
The sequential model is also limited because it requires strong cause and effect relationships that
typically do not exist outside the technical or mechanistic aspect of the accident. In other words,
true cause and effect relationships can be found when analyzing the equipment failures, but
causal relationships are extremely weak when addressing the human or organizational aspect of
the accident. For example: While it is easy to assert that “time pressure caused workers to take
shortcuts,” it is also apparent that workers do not always take shortcuts when under time
pressure. See Section 1.4, Cause and Effect Relationships.
In response to large scale industrial accidents in the 1970’s and 1980’s, the epidemiological
models were developed that viewed an accident the outcome of a combination of factors, some
active and some latent, that existed together at the time of the accident. [Hollnagel, 2004]10
1.3.2 Epidemiological or Latent Failure Model
This is a complex, linear cause and effect model where
accidents are seen as the result of a combination of
active failures (unsafe acts) and latent conditions
(unsafe conditions). These are often referred to as
epidemiological models, using a medical metaphor
that likens the latent conditions to pathogens in the
human body that lay dormant until triggered by the
unsafe act. In this model, accidents are prevented by
strengthening barriers and defenses. The “Swiss
Cheese” model developed by James Reason is an example of the epidemiological model.
[Reason, 1997]14
This model views the accident to be the result of long standing deficiencies that are triggered by
the active failures. The focus is on the organizational contributions to the failure and views the
human error as an effect, instead of a cause.
The epidemiological models differ from the sequential models on four main points:
Performance Deviation – The concept of unsafe acts shifted from being synonymous with
human error to the notion of deviation from the expected performance.
1‐3
DOE‐HDBK‐1208‐2012
Conditions – The model also considers the contributing factors that could lead to the
performance deviation, which directs analysis upstream from the worker and process
Section 19
deviations.
Barriers – The consideration of barriers or defenses at all stages of the accident
development.
Latent Conditions – The introduction of latent or dormant conditions that are present within
the system well before there is any recognizable accident sequence.
The epidemiological model allows the investigator to think in terms other than causal series,
offers the possibility of seeing some complex interaction, and focuses attention on the
organizational issues. The model is still sequential, however, with a clear trajectory through the
ordered defenses. Because it is linear, it tends to oversimplify the complex interactions between
the multitude of active failures and latent conditions.
The limitation of epidemiological models is that they rely on “failures” up and down the
organizational hierarchy, but does nothing to explain why these conditions or decisions were
seen as normal or rational before the accident. The recently developed systemic models start to
understand accidents as unexpected combinations of normal variability. [Hollnagel, 2004]10
[Dekker, 2006]15
1.3.3 Systemic Model
This is a complex, non-linear model where
both accidents (and success) are seen to
emerge from unexpected combinations of
normal variability in the system. In this
model, accidents are triggered by unexpected
combinations of normal actions, rather than
action failures, which combine, or resonate,
with other normal variability in the process to
produce the necessary and jointly sufficient
conditions for failure to succeed. Because of the complex, non-linear nature of this model, it is
difficult to represent graphically. The Functional Resonance model from Erik Hollnagel uses a
signal metaphor to visualize this model with the undetectable variabilities unexpectedly
resonating to result in a detectable outcome.
The JengaTM game is also an excellent metaphor for describing the
complex, non-linear accident model. Every time a block is pulled
from the stack, it has subtle interactions with the other blocks that
cause them to loosen or tighten in the stack. The missing blocks
represent the sources of variability in the process and are typically
described as organizational weaknesses or latent conditions.
Realistically, these labels are applied retrospectively only after what
was seen as normal before the accident, is seen as having contributed
to the event, but only in combination with other factors. Often, the
1‐4
DOE‐HDBK‐1208‐2012
worker makes an error or takes an action that seems appropriate, but when combined with the
other variables, brings the stack crashing down. The first response is to blame the worker
because his action demonstrably led to the failure, but it must be recognized that without the
other missing blocks, there would have been no consequence.
A major benefit of the systemic model is that it provides a more complete understanding of the
subtle interactions that contributed to the event. Because the model views accidents as resulting
from unexpected combinations of normal variability, it seeks an understanding of how normal
variability combined to create the accident. From this understanding of contributing interactions,
latent conditions or organizational weaknesses can be identified.
1.4 Cause and Effect Relationships
Section 20
Although generally accepted as the overarching purpose of the investigation, the identification of
causes can be problematic. Causal analysis gives the appearance of rigor and the strenuous
application of time-tested methodologies, but the problem is that causality (i.e., a cause-effect
relationship) is often constructed where it does not really exist. To understand how this happens,
we need to take a hard look at how accidents are investigated, how cause – effect relationships
are determined, and the requirements for a true cause - effect relationship.
1.4.1 Investigations Look Backwards
The best metaphor for how accidents are investigated is a simple
maze. If a group of people are asked to solve the maze as quickly
as possible and ask the “winners” how they did it, invariably the
answer will be that they worked it from the Finish to the Start.
Most mazes are designed to be difficult working from the Start to
the Finish, but are simple working from the Finish to the Start.
Like a maze, accident investigations look backwards. What was
uncertain for the people working forward through the maze
becomes clear for the investigator looking backwards.
Because accident investigations look backwards, it is easy to
oversimplify the search for causes. Investigators look backwards
with the undesired outcome (effect) preceded by actions, which is opposite of how the people
experienced it (actions followed by effects). When looking for cause - effect relationships (and
there many actions taking place along the timeline), there are usually one or more actions or
conditions before the effect (accident) that seem to be plausible candidates for the cause(s).
There are some common and mostly unavoidable problems when looking backwards to find
causality. As humans, investigators have a strong tendency to draw conclusions that are not
logically valid and which are based on educated guesses, intuitive judgment, “common sense”, or
other heuristics, instead of valid rules of logic. The use of event timelines, while beneficial in
understanding the event, creates sequential relationships that seem to infer causal relationships.
A quick Primer on cause and effect may help to clarify.
1
.4
1‐5
DOE‐HDBK‐1208‐2012
1.4.2 Cause and Effect are Inferred
Cause and effect relationships are normally inferred from observation, but are generally not
something that can be observed directly.
Normally, the observer repeatedly observes Action A, followed by Effect B and conclude that B
was caused by A. It is the consistent and unwavering repeatability of the cause followed by the
effect that actually establishes a true cause – effect relationship.
For example: Kink a garden hose (action A), water flow stops (effect B), conclusion is kinking
garden hose causes water to stop flowing. This cause and effect relationship is so well
established that the person will immediately look for a kink in the hose if the flow is interrupted,
Accident investigations, however, involve the notion of backward causality, i.e., reasoning
backward from Effect to Action.
The investigator observes Effect B (the bad outcome), assumes that it was caused by something
and then tries to find out which preceding Action was the cause of it. Lacking the certainty of
repeatability (unless the conditions are repeated) and a causal relationship can only be assumed
because it seems plausible. [Hollnagel, 2004]10
1.4.3 Establishing a Cause and Effect Relationship
Section 21
A true cause and effect relationship must meet these requirements:
The cause must precede the effect (in time).
The cause and effect must have a necessary and constant connection between them, such
that the same cause always has the same effect.
This second requirement is the one that invalidates most of the proposed causes identified in
accident investigations. As an example, a cause statement such as “the accident was due to
inadequate supervision” cannot be valid because the inadequate supervision does not cause
accidents all the time. This type of cause statement is generally based on the simple “fact” that
the supervisor failed to prevent the accident. There are generally some examples, such as not
spending enough time observing workers, to support the conclusion, but these examples are
cherry-picked to support the conclusion and are typically value judgments made after the fact.
[Dekker, 2006]15
1.4.4 The Circular Argument for Cause
The example (inadequate supervision) above is what is
generally termed a “circular argument.” The statement is
made that the accident was caused by “inadequate XXX.”
But when challenged as to why it was judged to be
inadequate, the only evidence is that it must be inadequate
because the accident happened. The circular argument is
usually evidenced by the use of negative descriptors such
1‐6
DOE‐HDBK‐1208‐2012
as inadequate, insufficient, less than adequate, poor, etc. The Accident Investigation Board
(AIB) needs to eliminate this type of judgmental language and simply state the facts. For
example, the fact that a supervisor was not present at the time of the accident can be identified as
a contributing factor, although it is obviously clear that accidents do not happen every time a
supervisor is absent.
True cause and effect relationships do exist, but they are almost always limited to the
mechanistic or physics-based aspects of the event. In a complex socio-technical system
involving people, processes and programs, the observed effects are uaually emergent phenomena
due to interactions within the system rather than resultant phenomena due to cause and effect.
With the exception of physical causes, such as a shorted electrical wire as the ignition source for
a fire, causes are not found; they are constructed in the mind of the investigator. Since accidents
do happen, there are obviously many factors that contribute to the undesired outcome and these
factors need to be addressed. Although truly repeatable cause and effect relationships are almost
impossible to find, many factors that seemed to have contributed to the outcome can be
identified. These factors are often identified by missed opportunities and missing barriers which
get miss labeled as causes. Because it is really opinion, sufficient information needs to be
assembled and presented in a form that makes the rationale of that opinion understandable to
others reviewing it.
The investigation should focus on understanding the context of decisions and explaining the
event. In order to understand human performance, do not limit yourself to the quest for causes.
An explanation of why people did what they did provides a much better understanding and with
understanding comes the ability to develop solutions that will improve operations.
1.4.5 Counterfactuals
Section 22
Using the maze metaphor, what was complex, with multiple paths and unknown outcomes for
the workers, becomes simple and obvious for the investigator. The investigator can easily
retrace the workers path through the maze and see where they chose a path that led to the
accident rather than one that avoided the accident. The result is a counterfactual (literally,
counter the facts) statement of what people should or could have done to avoid the accident. The
counterfactual statements are easy to identify because they use common phrases like:
“they could have …”
“they did not …”
“they failed to …”
“if only they had …”
The problem with counterfactuals is that they are a statement of what people did not do and does
not explain why the workers did what they did do. Counterfactuals take place in an alternate
reality that did not happen and basically represent a list of what the investigators wish had
happened instead.
1‐7
DOE‐HDBK‐1208‐2012
1
.5
Discrepancies between a static procedure and actual work practices in a dynamic and ever
changing workplace are common and are not especially unique to the circumstances involved in
the accident. Discrepancies are discovered during the investigation simply because considerable
effort was expended in looking for them, but they could also be found throughout the
organization where an accident has not occurred. This does not mean that counterfactual
statements should be discounted. They can be essential to understanding why the decisions the
worker made and the actions (or no actions) that the worker took were seen as the best way to
proceed. [Dekker, 2006]15
1.5 Human Performance Considerations
In order to understand human performance, do not limit yourself to the quest for causes. The
investigation should focus on understanding the context of decisions and explaining the event.
An explanation of why people did what they did provides a much richer understanding and with
understanding comes the ability to develop solutions that will improve operations.
The safety culture maturity model from the International Atomic Energy Agency (IAEA)
provides the basis for an improved understanding the human performance aspect of the accident
investigation. IAEA TECDOC 1329, Safety Culture in Nuclear Installations: Guidance for Use
in the Enhancement of Safety Culture, was developed for use in IAEA’s Safety Culture Services
to assist their Member States in their efforts to develop a sound safety culture. Although the
emphasis is on the assessment and improvement of a safety culture, the introductory sections,
which lay the groundwork for understanding safety culture maturity, provide a framework to
understand the environment which forms the organization’s human performance.
Organizational Maturity
Rule
Based
Improvement
Based
Goal
Based
People who make Management’s Mistakes are seen as
mistakes are blamed response to process variability with
for their failure to mistakes is more emphasis is on
comply with rules controls, understanding what
procedures, and happened, rather than
training finding someone to
blame
Figure 1-1: IAEA-TECDOC-1329 – Safety Culture in Nuclear Installations
The model (Figure 1-1) defines three levels of safety culture maturity and presents characteristics
for each of the maturity levels based on the underlying beliefs and assumptions. The concept is
illustrated below with the characteristics for how the organization responds to an accident.
Section 23
1‐8
DOE‐HDBK‐1208‐2012
Rule Based –Safety is based on rules and regulations. Workers who make mistakes are
blamed for their failure to comply with the rules.
Goal Based –Safety becomes an organizational goal. Management’s response to mistakes
is to pile on more broadly enforced controls, procedures and training with little or no
performance rationale or basis for the changes.
Improvement Based –The concept of continuous improvement is applied to safety. Almost
all mistakes are viewed in terms of process variability, with the emphasis placed on
understanding what happened rather than finding someone to blame, and a targeted response
to fix the underlying factors.
When an accident occurs that causes harm or has the potential to cause harm, a choice exists: to
vector forward on the maturity model and learn from the accident or vector backwards by
blaming the worker and increasing enforcement. In order to do no harm, accident investigations
need to move from the rule based response, where workers are blamed, to the improvement
based response where mistakes are seen as process variability needing improvement.
1.5.1 Bad Apples
The Bad Apple Theory is based on the belief that the system in which people work is basically
safe and worker errors and mistakes are seen as the cause of the accident. An investigation based
on this belief focuses on the workers’ bad decisions or inappropriate behavior and deviation from
written guidance, with a conclusion that the workers failed to adhere to procedures. Because the
supervisor’s role is seen as enforcing the rules, the investigation will often focus on supervisory
activities and conclude that the supervisor failed to adequately monitor the worker’s performance
and did not correct noncompliant behavior. [Dekker, 2002]16
From the investigation perspective, knowing what the outcome was creates a hindsight bias
which makes it difficult to view the event from the perspective of the worker before the accident.
It is easy to blame the worker and difficult to look for weaknesses within the organization or
system in which they worked. The pressure to find an obvious cause and quickly finish the
investigation can be overpowering.
1.5.2 Human Performance Modes – Cognitive Demands
People are fallible, even the best people make mistakes. This is the first principle of Human
Performance Improvement and accident investigators need to understand the nature of the error
to determine the appropriate response to the error. Jen Rasmussen developed a classification of
the different types of information processing involved in industrial tasks. Usually referred to as
performance modes, these three classifications describe how the worker’s mind is processing
information while performing the task. (Figure 1-2) The three performance modes are:
Skill mode - Actions associated with highly practiced actions in a familiar situation usually
executed from memory. Because the worker is highly familiar with the task, little attention
is required and the worker can perform the task without significant conscious thought. This
1‐9
DOE‐HDBK‐1208‐2012
mode is very reliable, with infrequent errors on the order of 1 in every 10,000 iterations of
the task.
Rule mode - Actions based on selection of written or stored rules derived from one’s
recognition of the situation. The worker is familiar with the task and is taking actions in
response to the changing situation. Errors are more frequent, on the order of 1 in 1,000, and
are due to a misrepresentation of either the situation or the correct response.
Section 24
Knowledge mode - Actions in response to an unfamiliar situation. This could be new task
or a previously familiar task that has changed in an unanticipated manner. Rather than using
known rules, the worker is trying to reason or even guess their way through the situation.
Errors can be as frequent as 1 in 2, literally a coin flip.
The performance modes refer to the amount of conscious control exercised by the individual
doing the task, not the type of work itself. In other words, the skill performance mode does not
imply work by crafts; rule mode does not imply supervision; and the knowledge mode does not
imply work by professionals. This is a scale of the conscious thought required to react properly
to a hazardous condition; from drilled automatic response, to conscious selection and compliance
to proper rules, to needing to recognize there is a hazardous condition. The more unfamiliar the
worker is with the work environment or situation, the more reliance there is on the individual’s
alert awareness, rational reasoning and quick decision-making skills in the face of new hazards.
Knowledge mode would be commonly relied on in typically simple, mundane, low hazard tasks.
All work, whether performed by a carpenter or surgeon, can exist in any of the performance
modes. In fact, the performance mode is always changing, based on the nature of the work at the
time. [Reason and Hobbs, 2003]17
Understanding the performance mode the worker was in when he/she made the error is essential
to developing the response to the accident (Figure 1-2). Errors in the skill mode typically
involve mental slips and lapses in attention or concentration. The error does not involve lack of
knowledge or understanding and, therefore, training can often be inappropriate. The worker is
literally the expert on their job and training is insulting to the worker and causes the organization
to lose credibility. Likewise, changing the procedure or process in response to a single event is
inappropriate. It effectively pushes the worker out of the skill mode into rule-based until the new
process can be assimilated. Because rule mode has a higher error rate, the result is usually an
increase in errors (and accidents) until the workers assimilate the changes and return to skill
mode. Training can be appropriate where the lapse is deemed due to a drift in the skills
competence, out-of-date mindset, or the need for a drilled response without lapses.
Training might be appropriate for errors that occurred in rule mode because the error generally
involved misinterpretation of either the situation or the correct response. In these instances,
understanding requirements and knowing where and under what circumstance those
requirements apply is cognitive in nature and must be learned or acquired in some way.
Procedural changes are appropriate if the instructions were incorrect, unclear or misleading.
1‐10
DOE‐HDBK‐1208‐2012
High
A
tt
e
n
ti
o
n
(
to
ta
sk
)
Inaccurate
Mental Picture
Misinterpretation
Inattention
Low Famil iarity (w/task) High
Low
Figure 1-2: Performance Modes
Training might also be appropriate for errors that occurred in the knowledge mode, if the
workers’ understanding of the system was inadequate. However, the problem might have been
issues like communication and problem-solving during the event, rather than inadequate
knowledge.
1.5.3 Error Precursors
Section 25
“Knowledge and error flow from the same mental sources, only success can tell the one from the
other.” The idea of human error as “cause” in consequential accidents is one that has been
debunked by safety science since the early work by Johnson and the System Safety Development
Center (SSDC) team. As Perrow stated the situation “Formal accident investigations usually
start with an assumption that the operator must have failed, and if this attribution can be made,
that is the end of serious inquiry. Finding that faulty designs were responsible would entail
enormous shutdown and retrofitting costs; finding that management was responsible would
threaten those in charge, but finding that operators were responsible preserves the system, with
some soporific injunctions about better training.” [Mach, 1976]18 [Perrow, 1984]5
In contemporary safety science the concept of error is simply when unintended results occurred
during human performance. Error is viewed as a mismatch between the human condition and
environmental factors operative at a given moment or within a series of actions. Research has
demonstrated that presence of various factors in combination increase the potential for error;
1‐11
DOE‐HDBK‐1208‐2012
these factors may be referred to as error precursors. Anticipation and identification of such
precursors is a distinguishing performance strategy of highly performing individuals and
organizations. The following Task, Work Environment, Individual Capabilities and Human
Nature (TWIN) model is a useful diagnostic tool for investigation (Figure 1-3).
TWIN Analysis Matrix
(Human Performance Error Precursors)
Task Demands Individual Capabilities
Time Pressure (in a hurry) Unfamiliarity with task / First time
High workload (large memory) Lack of knowledge (faulty mental model)
Simultaneous, multiple actions New techniques not used before
Repetitive actions / Monotony Imprecise communication habits
Irreversible actions Lack of proficiency / Inexperience
Interpretation requirements Indistinct problem‐solving skills
Unclear goals, roles, or responsibilities Unsafe attitudes
Lack of or unclear standards Illness or fatigue; general poor health or injury
Work Environment Human Nature
Distractions / Interruptions Stress
Changes / Departure from routine Habit patterns
Confusing displays or controls Assumptions
Work‐arounds Complacency / Overconfidence
Hidden system / equipment response Mind‐set (intentions)
Unexpected equipment conditions Inaccurate risk perception
Lack of alternative indication Mental shortcuts or biases
Personality conflict Limited short‐term memory
Figure 1-3: Error Precursors
1‐12
DOE‐HDBK‐1208‐2012
1.5.4 Optimization
Human performance is often summarized as the individual working within organizational
systems to meet the expectations of leaders. Performance variability is all about meeting
expectations and actions intended to produce a successful outcome.
Section 26
To understand performance variability, an investigator
must understand the nature of humans. Regardless of
the task, whether at work or not, people constantly strive
to optimize their performance by striking a balance
between resources and demands. Both of these vary
over time as people make a trade-off between
thoroughness and efficiency. In simple terms,
thoroughness represents the time and resources
expended in preparation to do the work and efficiency is
the time and resources expended in completing the
work. To do both completely requires more time and
resources than is available and people must choose
between them. The immediate and certain reward for meeting schedule and production
expectations easily overrides the delayed and uncertain consequence of insufficient preparation
and people lean towards efficiency. They are as thorough as they believe is necessary, but
without expending unnecessary effort or wasting time.
The result is a deviation from expectation and the reason is obvious. It saves time and effort
which is then available for more important or pressing activities. How the deviation is judged
afterwards, is a function of the outcome, not the decision. If organizational expectations are met
without incident, the deviations are typically disregarded or may even be condoned and rewarded
as process improvements. If the outcome was an accident, the same actions can be quickly
judged as violations. This is the probabilistic nature of organizational decision-making which is
driven by the perceptions or misperceptions of risks. A deviation or violation is not the end of
the investigation; it is the beginning as the investigator tries to understand what perceptions were
going on in the system that drove the choice to deviate. [Hollnagel, 2009]19
1.5.5 Work Context
Context matters and performance variability is
driven by context. The simple sense – think – act
model illustrates the role of context. Information
comes to the worker, he makes a decision based on
the context, and different actions are possible, based
on the context.
The context of the decision relate to the goals,
knowledge and focus of the worker. Successful
completion of the immediate task is the obvious
goal, but it takes place within the greater work
environment where the need to optimize the use of
1‐13
DOE‐HDBK‐1208‐2012
time and resources is critical. Workers have knowledge, but the application of knowledge is not
always straight forward because it needs to be accurate, complete and available at the time of the
decision. Goals and knowledge combine together to determine the worker’s focus. Because
workers cannot know and see everything all the time, what they are trying to accomplish and
what they know drives where they direct their attention.
All this combines to create decisions that vary based on the influences that are present at the time
of the decision and the basic differences in people. These influences and differences include:
Organization - actions taken to meet management priorities and production expectations.
Knowledge - actions taken by knowledgeable workers with intent to produce a better
outcome.
Social – actions taken to meet co-worker expectations, informal work standards.
Experience – actions based on past experience in an effort to repeat success and avoid
failure.
Inherent variability – actions vary due to individual psychological & physiological
differences.
Section 27
Ingenuity and creativity – adaptability in overcoming constraints and under specification.
The result is variable performance. From the safety perspective, this means that the reason
workers sometimes trigger an accident is because the outcome of their action differs from what
was intended. The actions, however, are taken in response to the variability of the context and
conditions of the work. Conversely, successful performance and process improvement also
arises from this same performance variability. Expressed another way, performance variability is
not aberrant behavior; it is the probabilistic nature of decisions made by each individual in the
organization that can result in both success and failure emerging from same normal work
sequence.
In accident investigations, performance variability needs to be acknowledged as a characteristic
of the work, not as the cause of the accident. Rather than simply judging a decision as wrong in
retrospect, the decision needs to be evaluated in the context in which it was made. In accident
investigation, the context or influences that drive the deviation need to be understood and
addressed as contributing factors. Stopping with worker’s deviation as the cause corrects
nothing. The next worker, working in the same context, will eventually adapt and deviate from
work-as imagined until chance aligns the deviation to other organization system weaknesses for
a new accident.
Performance variability is not limited to just the worker who triggers the accident. People are
involved in all aspects of the work, and the result is variability of all factors associated with the
work. This can include variation in the actions of the co-workers, the expectations of the leaders,
accuracy of the procedures, the effectiveness of the defenses and barriers, or even the basic
1‐14
DOE‐HDBK‐1208‐2012
policies of the organization. This is reflected in the complex, non-linear (non-Newtonian)
accident model where unexpected combinations of normal variability can result in the accident.
1.5.6 Accountability, Culpability and Just Culture
“Name, blame, shame, retrain” is an oft used phrase for older ineffective paradigms of safety
management and accident analysis. Dr. Rosabeth Moss Kanter of Harvard Business School
phased the situation this way: “Accountability is a favorite word to invoke when the lack of it
has become so apparent.” [Kanter, 2009]20
The concepts of accountability, culpability and just culture are inextricably entwined.
Accountability has been defined in various ways but in general with this characterization; “The
expectation that an individual or an organization is answerable for results; to explain actions, or;
the degree to which individuals accept responsibility for the consequences of their actions,
including the rewards or sanctions.” As Dr. Kanter explains “The tools of accountability — data,
details, metrics, measurement, analyses, charts, tests, assessments, performance evaluations —
are neutral. What matters is their interpretation, the manner of their use, and the culture that
surrounds them. In declining organizations, use of these tools signals that people are watched
too closely, not trusted, about to be punished. In successful organizations, they are vital tools
that high achievers use to understand and improve performance regularly and rapidly.”
Section 28
Culpability is about considering if the actions of an individual are blame worthy. The concept of
culpability in safety is based largely on the work of Dr. James Reason as a function of creating a
Just Culture. The purpose is to pursue a humane culture in which learning as individuals and
collectively is valued and human fallibility is recognized as simply part of the human condition.
Being human however is to be distinguished from being a malefactor. He explains; “The term
‘no-blame culture’ flourished in the 1990’s and still endures today. Compared to the largely
punitive cultures that it sought to replace, it was clearly a step in the right direction. It
acknowledged that a large proportion of unsafe acts were ‘honest errors’ (the kinds of slips,
lapses and mistakes that even the best people can make) and were not truly blameworthy, nor
was there much in the way of remedial or preventative benefit to be had by punishing their
perpetrators. But the ‘no-blame’ concept had two serious weaknesses. First, it ignored – or at
least, failed to confront – those individuals who willfully (and often repeatedly) engaged in
dangerous behaviors that most observers would recognize as being likely to increase the risk of a
bad outcome. Second, it did not properly address the crucial business of distinguishing between
culpable and non-culpable unsafe acts.”
“…a safety culture depends critically on first negotiating where the line should be drawn
between unacceptable behaviour and blameless unsafe acts. There will always be a grey
area between these two extremes where the issue has to be decided on a case by case basis.”
“… the large majority of unsafe acts can be reported without fear of sanction. Once this crucial
trust has been established, the organization begins to have a reporting culture, something that
provides the system with an accessible memory, which, in turn, is the essential underpinning to a
learning culture. There will, of course, be setbacks along the way. But engineering a just culture
is the all-important early step; so much else depends upon it.” [GAIN Working Group E, 2004]21
1‐15
DOE‐HDBK‐1208‐2012
1
.6
Along the road to a Just Culture organizations may benefit from explicit “amnesty” programs
designed to persuade people to report their personal mistakes. In complex events, individual
actions are never the sole causes. Thus determination of individual culpability and personnel
actions that might be warranted should be explicitly separated from the accident investigation.
Failure to make such separation may result in reticence or even refusal of individuals involved to
cooperate in the investigation, may skew recollections and testimony, may prevent investigators
from obtaining important information, and may unfairly taint the reputations and credibility of
well intended individuals to whom no blame should be attached.
1.6 From Latent Conditions to Active Failures
An organizational event causal story developed by James Reason starts with the organizational
factors: strategic decisions, generic organizational processes – forecasting, budgeting, allocating
resources, planning, scheduling, communicating, managing, auditing, etc. These processes are
colored and shaped by the corporate culture or the unspoken attitudes and unwritten rules
concerning the way the organization carries out its business. [Reason, 1997]14
Section 29
These factors result in biases in the management decision process that create “latent conditions”
that are always present in complex systems. The quality of both production systems and
protection systems are dependent upon the same underlying organizational decision processes;
hence, latent conditions cannot be eliminated from the management systems, since they are an
inevitable product of the cultural biases in strategic decisions. [Reason, p. 36, 1997]14
Figure 1-4 illustrates an example of latent conditions produced from the pressures of
commitment to a heavy work load as an organizational factor at the base of the pyramid. This
passes into the organization as a local work place factor in the form of stress in the work place.
This is the latent condition that is a precursor or contributing factor to the worker cutting corners
(the active failure of the safety system).
A distinction between active failures and latent conditions rests on two differences. The first
difference is the time taken to have an adverse impact. Active failures usually have immediate
and relatively short-lived effects. Latent conditions can lie dormant, doing no particular harm,
until they interact with local circumstances to defeat the systems’ defenses. The second
difference is the location within the organization of the human instigators. Active failures are
committed by those at the human-system interface, the front-line activities, or the “sharp-end”
personnel. Latent conditions, on the other hand, are spawned in the upper echelons of the
organization and within related manufacturing, contracting, regulatory and governmental
agencies that are not directly interfacing with the system failures.
The consequences of these latent conditions permeate throughout the organization to local
workplaces—control rooms, work areas, maintenance facilities etc. —where they reveal
themselves as workplace factors likely to promote unsafe acts (moving up the pyramid in Figure
1-4). These local workplace factors include undue time pressure, inadequate tools and
equipment, poor human-machine interfaces, insufficient training, under-manning, poor
supervisor-worker ratios, low pay, low morale, low status, macho culture, unworkable or
ambiguous procedures, and poor communications.
1‐16
DOE‐HDBK‐1208‐20012
Within thhe workplacee, these locaal workplace factors can combine with natural huuman
performaance tendenccies such as llimited attenntion, habit ppatterns, assuumptions, coomplacency, or
mental shhortcuts. Thhese combinaations produuce unintentiional errors aand intentionnal violationns —
collectiveely termed ““adaptive actts”—committted by indivviduals and tteams at the “sharp end,”” or
the directt human-system interfacce (active errror).
Large nuumbers of theese adaptive acts will haappen (small red arrows iin Figure 1-44), but very few
will alignn with the hooles in the deefenses (holees are createed by the lateent conditionns deep withhin
the organnization). WWith defense--in-depth prooviding a muulti-barrier ddefense, it takkes multiplee
human peerformance errors to breeach the mul tiple defensees. However, when defeenses have
become ssufficiently fflawed and oorganizational behavior cconsistently drifts from desired behaavior
accidentss can occur. In such eveents causes aare multiple aand only thee most superfficial analysis
would suuggest otherwwise.
FFigure 1-4: Organizaational Cauuses of Acccidents
1‐17
DOE‐HDBK‐1208‐2012
1
.7
Section 30
1.7 Doing Work Safely - Safety Management Systems
Safety Management Systems (SMS) were developed to integrate safety as part of an
organization’s management of mission performance. The benefits of process based management
systems is a well established component of quality performance. As organizations and the
technologies they employ became more complex and diverse, and the rate of change in pace of
societal expectations, technical innovations, and competitiveness increased, the importance of
sound management of functions essential to safe operations became heightened.
A SMS is essentially a quality management approach to controlling risk. It also provides the
organizational framework to support a sound safety culture. Systems can be described in terms
of integrated networks of people and other resources performing activities that accomplish some
mission or goal in a prescribed environment. Management of the system’s activities involves
planning, organizing, directing, and controlling these assets toward the organization’s goals.
Several important characteristics of systems and their underlying process are known as “process
attributes” or “safety attributes” when they are applied to safety related operational and support
processes.
The SMS for DOE is the Integrated Safety Management System (ISMS), defined in Federal
Acquisition Regulation and amplified though DOE directives and guidance. The ISMS is the
overarching safety system used by DOE to ensure safety of the worker, the community and the
environment. The DOE ISMS is characterized by seven principles and five core functions:
Seven Principles
Line management responsibility for safety
Line management is directly responsible for the protection of workers, the public and the
environment.
Clear roles and responsibilities
Clear and unambiguous lines of authority and responsibility for ensuring safety is
established and maintained at all organizational levels and for subcontractors.
Competence commensurate with responsibilities
Personnel are required to have the experience, knowledge, skills and capabilities necessary
to discharge their responsibilities.
Balanced priorities
Managers must allocate resources to address safety, as well as programmatic and operational
considerations. Protection of workers, the public and the environment is a priority whenever
activities are planned and performed.
Identification of safety standards and requirements
Before work is performed, the associated hazards must be evaluated, and an agreed-upon set
of safety standards and requirements must be established to provide adequate assurance that
workers, the public and the environment are protected from adverse consequences.
1‐18
DOE‐HDBK‐1208‐2012
Hazard controls tailored to work being performed
Administrative and engineering controls are tailored to the work being performed to prevent
adverse effects and to mitigate hazards.
Operations authorization
The conditions and requirements to be satisfied before operations are initiated are clearly
established and agreed upon.
Five Core Functions (Figure 1-5)
Define the scope of work
Missions are translated into work, expectations are set, tasks are identified and prioritized
and resources are allocated.
Analyze the hazards
Hazards associated with the work are identified, analyzed and categorized.
Develop and implement hazard controls
Applicable standards, policies, procedures and requirements are identified and agreed upon;
controls to prevent/mitigate hazards are identified; and controls are implemented.
Section 31
Perform work within controls
Readiness is confirmed and work is performed safely.
Provide feedback and continuous improvement
Information on the adequacy of controls is gathered, opportunities for improving the
definition and planning of work are identified, and line and independent oversight is
conducted.
1‐19
DOE‐HDBK‐1208‐20012
Figure 1-5: Five Core Funcctions of DDOE’s Integgrated Safeety Manageement Sysstem
1.7.1 TThe Functioon of Safetty Barrierss
The use oof controls oor barriers too protect the people fromm the hazardss is a core prrincipal of saafety.
Barriers aare employeed to serve twwo purposes; to prevent release of haazardous eneergy and to
mitigate harm in the event hazarddous energy is released. Energy is ddefined broaddly as used hhere,
and incluudes multiplee forms, for example; kinnetic, biologgical, acoustiical, chemic al, electricall,
mechaniccal, potential, electromaggnetic, thermmal, or radiattion.ii
For a detailed disccussion of barrriers refer to “Barriers and Accident Preevention” by EErik Hollnage l,
2004.
1‐20
ii
DOE‐HDBK‐1208‐2012
The dynamics of accidents may be categorized into five basic components, illustrated in Figure
1-6: 1) the threat or triggering action or energy, 2) the prevention barrier between the threat and
the hazard, 3) the hazard or energy potential, 4) the mitigation barrier to mitigate hazardous
consequences towards the target, 5) the targets in the path of the potential hazard consequences.
When these controls or barriers fail, they allow unwanted energy to flow resulting in an accident
or other adverse consequence.
Preventing System Accidents
Initiating
Hazards Targets Source or
Threats
Prevention Mitigation
Undesired
Energy
Flow
Human
Error Workers
Public
Environment
Attack or
Sabotage
Natural
Forces
Equipment
Failure
Barrier (e.g. Barrier (e.g.
spark secondary
inhibitors) containment)
Figure 1-6: Barriers and Accident Dynamics – Simplistic Design
The objective is to contain or isolate hazards though the use of protective barriers. Prevention
barriers are intended to preclude release of hazards by human acts, equipment degradation, or
natural phenomena. Mitigation barriers are used to shield, contain, divert or dissipate the
hazardous energy if it is released thus precluding negative consequences to the employees or the
surrounding communities. Distance from the hazard is a common mitigating barrier.
Barrier analysis is based on the premise that hazards are associated with all accidents. Barriers
are developed and integrated into a system or work process to protect personnel and equipment
from hazards. For an accident to occur the design of technical systems did not provide adequate
barriers, work design did not specify use of appropriate barriers, or barriers failed. Investigators
use barrier analysis to identify hazards associated with an accident and the barriers that
should/could have prevented it. Barrier analysis addresses:
1‐21
DOE‐HDBK‐1208‐2012
Barriers that were in place and how they performed
Barriers that were in place but not used
Barriers that were not in place but were required
The barrier(s) that, if present or strengthened, would prevent the same or a similar accident
from occurring in the future.
Section 32
All barriers are not the same and differ significantly in how well they perform. The following are
some of the general characteristics of barriers that need to be considered when selecting barriers
to control hazards. When evaluating the performance of a barrier after an accident, these
characteristics also suggest how well we would expect the barrier to have performed to control
the hazard.
Effectiveness – how well it meets its intended purpose
Availability – assurance the barrier will function when needed
Assessment – how easy to determine whether barrier will work as intended
Interpretation – extent to which the barrier depends on interpretation by humans to achieve
its purpose
1.7.2 Categorization of Barriers
Barriers may also be categorized according to a hierarchy of cost/reliability and according to
barrier function. The barrier cost/reliability hierarchy includes:
Physical or engineered barriers – These are the structures that are built, or sometimes naturally
exist, to prevent the flow of energy or personnel access to the hazards. These barriers require an
investment to design and build and have a cost to maintain and update. Examples: Personnel
cage around a multi-story ladder, a guard rail on a platform, or a barricade to prevent access.
Administrative or management policy barriers – These include rules, procedures, policies,
training, work plans that describe the requirements to avoid hazards. These barriers require less
capital investment but have a cost in the development, review, updating, training,
communication, and enforcement to assure adequacy and compliance. Examples: Requirement
to use harness and strap ties while climbing a multi-story ladder, a prescriptive process procedure
sequence, or laws against trespassing.
Personal knowledge or skill barriers – These include human performance aspects of:
fundamental lessons-learned, knowledge, common sense, life experiences, and education that
contribute to the individuals’ survival instincts and decision-making ability. These barriers
require little or no investment except in the screening and selection process for qualified
personnel used in a task and providing supervision. Examples: The decision not to climb a
1‐22
DOE‐HDBK‐1208‐2012
ladder with a tool in one hand, the decision not to violate one of the administrative barriers, or
recognizing a dangerous situation.
Another analysis system divides barriers into four categories that reflect the nature of the
barriers’ performance function. These four categories can be useful in the barrier analysis for
characterizing more precisely the purpose of the barrier and its type of weakness. Examples for
each of the four categories are as follows:
Physical– physically prevents an action from being carried out or an event from happening
Containing or protecting - walls, fences, railings, containers, tanks
Restraining or preventing movement - safety belts, harnesses, cages
Separating or protecting – crumple zones, scrubbers, filters
Functional– impedes actions through the use of pre-conditions
Prevent movement/action (hard) – locks, interlocks, equipment alignment
Prevent movement/action (soft) – passwords, entry codes, palm readers
Impede actions – delays, distance (too far for single person to reach)
Dissipate energy/extinguish – air bags, sprinklers
Symbolic– requires an act of interpretation in order to achieve their purpose
Countering/preventing actions – demarcations, signs, labels, warnings
Section 33
Regulating actions – instructions, procedures, dialogues (pre-job brief)
System status indications – signals, warnings, alarms
Permission/authorization – permits, work orders
Incorporeal– requires interpretation of knowledge in order to achieve their purpose
Process – rules, restrictions, guidelines, laws, training
Comply/conform – self-restraint, ethical norms, morals, social or group pressure
Within DOE organizations, there is typically a defense-in-depth policy for reducing the risks of a
system failure or an accident due to the threats. This policy maintains a multiple layered barrier
system between the threats or hazards and the requirement to correct any weaknesses or failures
identified in a single layer. Therefore, an accident involving such a protected system requires
1‐23
DOE‐HDBK‐1208‐2012
either a uniquely improbable simultaneous failure of multiple barriers, or poor barrier concepts
or implementation, or a period of neglect allowing cascading deterioration of the barriers.iii
Defense-in-depth can be comprised of layers of any combination of these types of barriers.
Obviously, it is much more difficult to overcome multiple layers of physical or engineered
barriers. This is the most reliable and most costly defense. Risk management analysis
determines the basis and justification for the level of barrier reliability and investment, based on
the probability and consequence of a hazard release scenario. For low probability, low
consequence events the level of risk often does not justify the investment of physical barriers.
Cost and schedule conscious management may influence selection of non-physical barriers on all
but the most likely and catastrophically hazardous conditions. Such choices place greater
reliance on layers of the less reliable barriers dependent on human behavior. Adding multiple
barrier layers can appear to add more confidence, but multiple layers may also lead to
complacency and diminish the ability to use and maintain the individual barrier layers. Complex
barrier systems and barrier philosophies place heightened importance on the context of
organizational culture and human performance becomes a major concern in the prevention of
accidents as barrier systems become more complex and individual barrier layer functionality
become less apparent.
A cascading effect can occur in aging facilities. Engineered barriers can become out-of-date, fall
into disrepair or wear out; or be removed as part of demolition activity. Management should
transition to reliance on a substitute administrative barrier, but this need may not be recognized.iv
For example, a fire protection system, temporarily or permanently disable, is replaced by a fire
watch until the protection system is restored, replaced, or the fire potential threat is removed.
Administrative barriers may weaken due to inadequate updates to rules, inadequate
communication and training, and inadequate monitoring and enforcement. This results in
managements’ often unintentional reliance on the personal knowledge barriers. Personal
knowledge barriers can be weakened by the inadequate screening for qualifications, inadequate
assignment selections, or inadequate supervision.
Section 34
An alignment of cascading weaknesses in barriers can result in an unqualified worker
unintentionally violating an administrative control and defeating a worn out physical barrier to
initiate an accident. Effective management of any of the barriers would have prevented the
accident by breaking the chain of events. Therefore, investigating a failure of defense-in-depth
requires probing a series of management and individual decisions that form the precursors and
chain of actions that lead to the final triggering action.
iii A common use of “defense-in-depth” is the Lockout-Tagout (LOTO) Procedure. This procedure
administratively requires that a hazardous energy be isolated by a primary physical barrier (e.g., valve
or switch), a secondary physical barrier (a lock) that controls inadvertent defeat of the primary barrier,
and a tertiary administrative barrier (tagging) controls the removal of the physical barriers. It is
understood that omitting any one of these barriers is a violation of the LOTO procedure.
iv An example of a cascading effect, related to LOTO, is the discovery that some old facilities have used
the out-of-date practice of common neutrals in old electrical systems or that facility circuit diagrams
and labeling were not maintained accurately. These latent conditions potentially defeat LOTO
entirely, requiring an additional administrative barrier procedure to do de-energized-circuit verification
prior to accessing old wiring systems. Latent conditions are explained further in section 1.4.
1‐24
http:recognized.iv
DOE‐HDBK‐1208‐2012
1.8 Accident Types/ Individual and Systems
There are two fundamental types of accidents which DOE seeks to avoid; individual and system
accidents. Confusion between individual and system safety has been frequently cited as causal
factors in major accidents.v In the ISMS framework, individual accidents are most often
associated with failures at the level of the five core functions. System accidents involve failures
at the principles level involving decision making, resource allocation and culture factors that may
shift the focus and resources of the organization away from doing work safely to detrimental
focus on cost or schedule.
1.8.1 Individual Accidents
Individual accidents - an accident occurs wherein the worker is not protected from the hazards of
an operation and is injured (e.g., radiation exposure, trips, slips, falls, industrial accident, etc.).
The focus of preventing individual accidents is to protect the worker from hazards inherent in
mission operation (Figure 1-7). The inherent challenges in investigating an individual accident
are due to the source of the human error and the victim or target of the accident can often be the
same individual. This can lead to a limited or contained analysis that fails to consider the larger
organizational or systemic contributors to the accident. These types of accidents involving
individual injuries can overly focus on the mitigating barriers or personnel protection equipment
(PPE) that avoid injuries and not consider the appropriate preventative barriers to prevent the
actual accident.
1
.8
Texas City, Buncefield, Deepwater Horizon
1‐25
v
DOE‐HDBK‐1208‐2012
Preventing Individual Accidents
Initiating
Hazards Targets Source or
Threats
Undesired
Energy
Flow
Section 35
Human
Error
Individual
worker
Equipment
Failure
Prevention Mitigation
Barrier (e.g. Barrier
LOTO policy) (e.g. PPEs)
Figure 1-7: Individual Accident
1.8.2 Preventing Individual Accidents
To prevent recurrence of individual injury accidents, corrective actions from accident
investigations must identify what barriers failed and why [i.e., stop the source and the flow of
energy from the hazards to the target (the worker)]. The mitigating barriers are important to
reducing or eliminating the harm or consequences of the accident, but emphasis must be on
barriers to prevent the accident from occurring. However, it is possible to find conditions where
the threat is deemed acceptable if the consequence can be adequately mitigated.vi
An example of reliance on a mitigating barrier would be in the meat cutting process where chain-mail
gloves protect hands from being cut. The threat or initiating energy is the knife moving towards the
hand or vice-versa. The hazard energy is the cutting action of the blade. Since the glove does not
prevent the knife from impacting the hand, the glove is a mitigation barrier that reduces the hazardous
cutting consequence of the impact. Implementing a prevention barrier would require redesigning the
process to block or eliminate the need for the hand to be in cutting area. The absence of the
prevention barrier is the result of a bias in the organizational decision-making process discussed later
in this handbook.
1‐26
vi
http:mitigated.vi
DOE‐HDBK‐1208‐2012
1.8.3 System Accident
A system accident is an accident wherein the protective and mitigating systems collectively fail
allowing release of the hazard and adversely affecting many people, the community and
potentially the environment. A system accident can be characterized as an "unanticipated
interaction of multiple failures in a complex system. This complexity can either be technological
or organizational, and often is both.” [Perrow, 1984]5
The focus of preventing system accidents is to maintain the physical integrity of operational
barriers such that they prevent threats that may result from human error, malfunctions in
equipment or operational processes, facility malfunctions or from natural disasters or such that
they mitigate the consequences of the event in case prevention fails. (Figure 1-8).
System hazards are typically managed from cradle to grave through risk management. Risk
management processes identify the potential threats, weaknesses, and failures as risks to the
design, construction, operations, maintenance, and disposition of the system. Risk management
establishes and records the risk parameters (or basis) and the investment decisions, the control
systems, and policies to mitigate these risks. Risk management, in a broad organizational sense,
can include financial, political, cultural, and social risks. While not excluding the broader
societal factors, the principal focus of this handbook is on socio-technical systems and related
life-cycle management (design, build, operate, maintain, dispose) system risks.
It is important to recognize the distinction between individual accidents and system accidents as
it affects the way the accident is investigated, in particular the way the barriers are analyzed.
The most likely differentiation of the type of accident investigation is from experience that
individual accidents are likely to be influenced by work practices, plans and oversight, while
system failures will most likely be influenced by risk management process for design,
operations, or maintenance. System accidents require a more in-depth investigation into the
policies and management culture that drives risk management decision-making. Naturally, there
is often an overlap that combines individual work hazards control practices and the system risk
management policies as potential areas of investigation.
Section 36
1‐27
DOE‐HDBK‐1208‐2012
System Accident
An accident wherein the system fails allowing a
threat to release the hazard and as a result
many* people are adversely affected
* Workers, Enterprise, Environment, Country
Focus
Protect the operations
Th e emph asis on th e system acciden t
from the threats in n o way degrades th e importan ce of
in dividual safety, it is a pre-requ isite of
system safety, bu t focu s on in dividu als
safety is n ot en ou gh .
Figure 1-8: System Accident
1.8.4 How System Accidents Occur
In order to prevent system accidents and incidents, it is important to first understand (via a
mental model) how they occur. Figure 1-9 represents a simple schematic of how system
accidents (accidents with large consequences affecting many people) can occur.
As defined in this figure a threat can come from four sources:
Human error such as someone dropping high explosives resulting in detonation.
Failure of a piece of equipment, tooling or facility. For example a piece of tooling with
faulty bolts causes high explosive to drop on the floor resulting in detonation.
From a natural disaster such as an earthquake resulting in falling debris that could
detonation high explosives.
“Other” as of yet undiscovered to accommodate future discoveries.
1‐28
DOE‐HDBK‐1208‐2012
Based on this simplistic system accident scenario it is clear technical system integrity must be
protected from deterioration from physical and human/social factors.
How System Accidents Happen
(Consider all Threats)
UNWANTED ENERGY FLOW
Equip/
tooling /
facilities
Human
Error
Natural
Disasters
Other
Hazard
to
Protect
& to
Minimize
System
Accident to
Avoid
THREATS * HAZARDS CONSEQUENCE OR
SYSTEM ACCIDENT
Unwanted energy flows as a result of the threat to a plant hazard
potentially resulting in a catastrophic consequence.
* Categories of threats adapted from MORT, DOE G 231.1‐2 and TapRoot
Figure 1-9: How System Accidents Happen
1.8.5 Preventing System Accidents
Figure 1-10 provides a simplistic view of how to prevent a system accident. Hazards can be
energy in the form of leaks, projectiles, explosions, venting, radiation, collapses, or other ways
that produce harm to the work force, the surrounding community, or the environment. The idea
is that one wants to isolate these hazards from those things that would threaten to release the
unwanted energy or material, such as human errors, faulty equipment, sabotage, or natural
disasters such as wind and lightning through the use of preventive barriers. If this is done, work
can proceed safely (accidents are avoided).
1‐29
DOE‐HDBK‐1208‐2012
1
.9
DOE takes a system approach (ISMS) to preventing system
accidents. The system is predicated on identifying hazards to
protect, identifying threats to those hazards, implementing controls
(barriers) to protect the hazard from the threats, and reliably
performing work within the established safety envelope.
Equipment, tooling,
facility malfunctions
Human Errors
Consequence
Barrier
(to prevent)
Hazard
Threats
Energy
Flow
Barrier
(to mitigate)
Natural Disasters
Figure 1-10: Prevent a System Accident
Section 37
1.9 Diagnosing and Preventing Organizational Drift
Recognizing the hazards or risks and establishing and maintaining the barriers against accidents
are continuous demands on organizations at all levels. Work, organizations, and human activity
are dynamic, not static. This means conditions are always changing, even if only through aging,
resource turnover, or creeping complacency to routine. Similarly to the Second Law of
Thermodynamics—the idea that everything in the created order tends to dissipate rather than to
coalesce – organizations left untended trend in the direction of disorder. In the safety literature
this phenomena is referred to as organizational drift. Organizational drift, if not halted, will
lead to weakened or missing barriers.
In order to recognize, diagnose and hopefully to prevent organizational drift from established
safety systems (ISMS), models (mental pictures) are needed. Properly built models help
investigators recognize aberrations by providing an accepted reference to compare against (i.e., a
mental picture of how the organization is supposed to work). Models in combination with an
understanding of organizational behavior also allow investigators to extrapolate individual events
to a broader organizational perspective to determine if the problem is pervasive throughout the
organization (deeper organizational issues).
1‐30
DOE‐HDBK‐1208‐2012
Three levels of models are introduced in this section to aid the investigators putting their event
into perspective.
Level I at the employee level,
Level II at the physics level - Break-the-Chain Framework (BTC),
Level III at the organization or system level.
1.9.1 Level I: Employee Level Model for Examining Organizational Drift --
Monitoring the Gap – “Work-as-Planned” vs. “Work-as-Done”
The Employee Level Model provides the most detailed examination of organizational drift by
comparing “work-as-done” on the shop floor with how work was planned by management and
process designers. At this level, the effect of organization drift could result in an undesirable
event because this is where the employees contact the hazards while performing work.
DOE organizations develop policies, procedures, training etc. to provide a management system
envelope of safety within which they want their people to work. This safety envelope is
developed through the ISMS “Define the Scope of Work, Analyze the Hazards, and Develop and
Implement Hazard Controls” and can be referred to as “work-as-planned.” The way work is
actually accomplished under ISMS “Perform Work within Controls,” referred to as “work-as
done”, can be compared to the work-as-planned. Every organization’s goal is to have “work-as
done” to equal work-as-planned (i.e., actual work performed within the established safety
envelope – left side of Figure 1-11).
There will always be a performance gap between “work-as-planned” and “work-as-done” work
performance gap (ΔWg) because of the variability in the execution of every human activity (right
side of Figure 1-11). When the ΔWg becomes a problem because an accident or an information-
rich, high-consequence or reoccurrence event occurs, a systematic investigative process helps to
understand first “what” the variation is and second, determine “why” the variation exists. Figure
1-11 illustrates the comparison of the ideal or desirable organizational work performance goal on
the left side, with the more likely or realistic work performance gap on the right. Recognizing
and reducing the gap is the objective of “Provide Feedback and Continuous Improvement”
activities.
Section 38
Within this handbook, the term “physics of safety” is used to represent the science and
engineering principles and methods used to assure the barriers designed into the systems are
effective against the nature of the threats and hazards. Only with sound “physics of safety” basis
behind the purpose of the barriers can management truly rely on a “work-as-planned” safety
performance envelope. A typical gap analysis must explore weaknesses in the “work-as
planned” and the “work-as-performed.”
Because the “work-as-planned” truly represents the requisite safety/security/quality process that
management wants their employees to follow; the investigative process reduces the gap ΔWg by
1‐31
DOE‐HDBK‐1208‐2012
systematically addressing the broadest picture of what went wrong, and focuses the Judgments of
Need and Corrective Actions to reduce the gap.
Systematically Evaluate
Organizational Goal Organizational Reality
Work-as-Planned
Work-as-Done
Work-as-Planned
Work-as-Done
∆Wg
“What”
“Why” Goal: Align, tighten, and sustain
spectrum of performance to keep
work-as-planned the same as
work-as-done.
Where we want to be Where we probably are
∆wg = gap in “work-as-don e” vs. as “plan n ed”
Figure 1-11: Level I - “Work-as-Done” Varies from “Work-as-Planned” at
Employee Level
1.9.2 Level II: Mid-Level Model for Examining Organizational Drift – Break-the-
Chain
The Mid-Level Model for examining organizational drift focuses on the Break-the-Chain (BTC)
framework. Based on the simplistic representation displayed in Figure 1-12, the BTC framework
provides a broader, more complete model to help organizations avoid the threat potential of
catastrophic events posed by the significant hazards, dynamic tasks, time constraints, and
complex technologies that are integral to ongoing missions. And, when an event does occur, it
also provides a logical and systematic framework to diagnose the event to determine which step
in the process broke down to allow focusing corrective actions in only those areas found
deficient. The BTC model is designed to stop the system accident as shown in Figure 1-10 but
1‐32
DOE‐HDBK‐1208‐2012
can be applied equally to individual accidents. The BTC model is nothing but a logical, physics-
based application of the ISM core functions. The six basic components of the BTC model are:
Step #1 – Focus on the System Accident (Pinnacle/Plateau Event) to Avoid: The first step
focuses on the last link of the chain, the consequences of the system accident that the
organization is trying to prevent. Once the catastrophic consequences have been identified, they
should be listed in priority order. This prioritization is important for four reasons:
It serves as an important reminder to all employees of the potential catastrophic
consequences they must strive to avoid each day.
It pinpoints where defensive barriers are most needed; as one would expect, the probability
of an event and the severity of the consequences will drive the number and type of barriers
selected.
It ensures that the defensive barriers associated with the highest priority consequences will
receive top protection against degradation.
It encourages a constant review of resources against consequences focusing attention on
making sure the most severe consequences are avoided at all times.
Section 39
Prioritization is a critical organizational dynamic. Efforts to protect against catastrophic
consequential events should be the first priority. Focus must be maintained on the priority
system accidents to assure that the needed attention and resources are available to prevent them.
Step #2 – Recognize and Minimize Hazard: Identify and minimize the physical hazard, while
maintaining production. After identifying the hazard, there are two approaches to minimize it.
First, actions are taken to reduce the physical hazard that can be impacted by the threat (for
example minimizing the amount of combustible material in facilities). Second, attempts are
made to reduce the interactive complexity and tight coupling within the operation or, conversely,
to increase the response time of the organization so an event can be recognized and responded to
more quickly. The intent of these two approaches is to remove or reduce the hazard so that the
consequences of an accident are minimized to the extent possible.
Step #3- Recognize Threat Posed by Human Errors, Failed Equipment, Tooling or
Facilities, Mother Nature (i.e., natural disasters) or Other as of yet Unknown Things: A
key component of consequence avoidance is identifying and minimizing all significant knowable
threats that could challenge the hazard (i.e., allow the flow of unwanted energy). Note the use of
the word “all.” The intent is that if not all threats are identified and addressed; the organization
is vulnerable to failure. Organizations should ensure the system event does not occur, not hope it
does not occur (i.e., they prove operations safe). The categories of threats from human error and
failed equipment, tooling or facilities, and natural disasters have been adapted from a
combination of MORT, DOE Guide (G) 231.1-1 and TapRoot® .
Step #4 – Manage Defenses: Based on the threats identified, one must ensure the right barriers
are identified to prevent or to reduce the probability of the flow of energy to the hazard (red,
blue, brown, and purple barriers in Figure 1-12) or if that fails to mitigate the consequences of a
1‐33
DOE‐HDBK‐1208‐2012
system accident (shown by granite encasement around system event box in Figure 1-12). The
type and number of barriers and the level of effort needed to protect them are dictated by level of
consequence and type of hazard associated with the operation. The decrease in the number of
threats or probability of occurrence as a result of the application of various barriers or defenses is
indicated in Figure 1-12 by the reduction in the number of colored arrows that can reach the
hazard.
Step #5 – Foster a Culture of Reliability: Steps 1 through 4 make the operational hazard less
vulnerable to threats. To execute these steps successfully and consistently without observable
signs of degradation or significant events, requires an army of trained and experienced personnel
who conscientiously follow the proven work practices. These workers must maintain their
proficiency through continuous hands-on work and be trained so they can make judgment calls
on the shop floor that will reflect the shared organizational values. They also need to have the
authority to make time-critical decisions when situations require this action. They must be part
of an organization that has a strong culture of reliability.
Section 40
Step #6 – Learn from Small Errors to Prevent Big Ones: Gaps between “work-as-planned”
by the process designer and “work-as-done” by the employees exist in every operation and
reflects the challenges an organization will face sustaining the BTC framework (Figure 1-12).
The fact that these gaps exist should be of no surprise, they exist in every organization. The
problem occurs when the organization is unaware of the gaps or does not know the magnitude or
extent of the gaps across the operation. Because of the importance of DOE sites remaining
within the established safety basis (ISMS), the investigation process as described in this
document places special emphasis on evaluating and closing the gap between “work-as-planned”
and “work-as-done”.
BTC parallels and complements the ISMS functions. The levels of formality or rigor to which
the six process components (or process steps) are applied are proportional to the complexity and
consequences of the operations (e.g., for nuclear operations where the potential consequences are
severe, the full rigor of 10 Code of Federal Regulations (CFR) Part 830, nuclear safety is
employed). Detailed application of this process can be found in Volume II, Chapter 1.
1‐34
DOE‐HDBK‐1208‐2012
Break‐the‐Chain Framework to Prevent System Accidents
Human
Performance
Error
Precursors
Human
Error
Equip/
tooling /
facilities
Natur al
Disasters
Other
Hazard
to
Protect
&
to
Minimize
System
Accident
to Avoid
Step #3
Step #6
Learn from Small Errors
Step #5
Foster a Culture of Reliability
Recognize Step #4 Step #2 Step #1
Threats Manage Defenses Recognize &
Minimize
Focus on the
System Accident
Hazard
Figure 1-12: Level II - Physics-Based Break-the-Chain Framework
1.9.3 Level III: High Level Model for Examining Organizational Drift
The High Level Model for examining organizational drift, shown in Figure 1-13, was adapted
from work by the Institute of Nuclear Power Operations (INPO)vii . It is intended to represent a
systematic view for analysis of both individual and system accidents. The model breaks down
work into four sequences that one typically finds at DOE sites: 1) Organizational Processes &
Values; 2) Job Site Conditions (work-as-planned); 3) Worker Behaviors (“work-as-done”); 4)
Operational Results. An explanation of each category of work can be found in Figure 1-13.
Also shown in Figure 1-13 are the quality assurance checks (green ovals) that take place before
transitioning from one sequence of work to the next sequence of work. These process check
points are additional examples of barriers put in place to ensure readiness to go the next sequence
of work. DOE uses many similar quality assurance readiness steps in both its high hazard
nuclear operations and industrial operations.
vii INPO Human Performance Reference Manual, INPO-06-003.
1‐35
‐ ‐
‐ ‐
‐
‐
DOE‐HDBK‐1208‐2012
The later in the work sequence the process barriers fall, noted by higher highlighted grey
numbers, the more significant or important the barrier is in preventing the undesired event
because it represents one of the last remaining barriers before a consequential event.
Section 41
A Systems View of Operating Performance
Products or results of processes Actions or inactions (i.e., using
(physical barriers) that create the or not using products of
right conditions for the worker to processes) by an individual
successfully & safely accomplish worker during the
tasks (e.g., engineered barriers, performance of a task
safety systems, procedures, tools, (procedure adherence,
readiness, etc.) protect barriers, etc.)
Programs and processes to
focus the org anization to
accomplish operational goals
while avoiding the
consequential accident.
Modifie d from INP O Human Pe rformance Re fe re nc e Manual, INP O 06‐003, 2006
“work as planned”
JOB‐SITE
CONDITIONS
OPERATIONAL
RESULTS
2
4
ORGANIZATIONAL
PROCESSES
& VALUES
1
“work as done”
WORKER
BEHAVIOR
Review Work
Post job Reviews
3
Pre job Brief Readiness
Authorization
Causal Factor Analysis
Independent Oversight
Qualifications
Assignments
QC Hold Points
Job Site Walk Down Mgt. Oversight
Independent Verification
Outcomes to the Plant as a
result of the worker ’s
behavior (e.g., events, TSR
violations, unplanned LCOs
events, etc.)
Leadership
High Standards Courage & Integrity Questioning
Attitude
Healthy Relationships
Open & Honest
Communications
Figure 1-13: Level III - High-Level Model for Examining Organizational Drift
1.10 Design of Accident Investigations
The organizational basis for the causes of accidents requires the accident investigators to develop
insights about organizational behavior, mental models and the factors that shape the environment
in which the incident occurred. This develops a better understanding of “what” in the
organizational system failed and “why” the organization allowed itself to degrade to the state that
resulted in an undesired consequence. The investigation progresses through the events in the
opposite order in which they occurred, as shown schematically in Figure 1-14.
1
.1
0
1‐36
DOE‐HDBK‐1208‐2012
Investigations to Determine Organizational Weaknesses
Unsafe acts
Local workplace
factors
Organizational
factors
Failed Defenses / Barriers
Active
failures
Latent
Conditions
Event
precursors
precursors
precursors
Even
t
In
vestigatio
n
Causal Factors Analysis starts
with the low consequence,
information‐rich event and
separates “What” happened
from “Why ” it happened.
This allows us to drill down to
find the:
1. Flawed defenses
2. Active failures (unsafe acts)
3. Human performance error
precursors
“What”
4. 4. latent conditions (local
workplace factors &
organizational factors).
“Why”
Adopted from Reason, Managing the Risks of Organiz ational Accidents
Figure 1-14: Factors Contributing to Organizational Drift
1.10.1 Primary Focus – Determine “What” Happened and “Why” It Happened
The basic steps and processes used for the Accident Investigation are:
Define the Scope of the Investigation and Select the Review Team
Collect the Evidence
Investigate “what happened”
Analyze “why it happened”
Define and Report the Judgments of Need and Corrective Actions
The purpose of an accident investigation is to determine:
1‐37
DOE‐HDBK‐1208‐2012
The “what” went wrong beginning by comparing “work-as-done” to planned work. The
purpose is to understand what was done, how it was planned, and identify unanticipated or
unforeseeable changes that may have intervened. Establishing the “what” was done tends to
result in a forward progression of the sequence of events that defines what barriers failed
and how they failed.
Section 42
The “why” things did not work according to plan comes from a cultural-based assessment of
the organization to understand why the employees thought it was OK to do what they did at
the time in question. Establishing the “why” tends to be a backwards regression identifying
the assumptions, motives, impetus, changes and inertia within the organization that may
reveal weaknesses and inadequacies of the barriers, barrier selection, and maintenance
processes. The objective is to understand the latent organizational weaknesses and cultural
factors that shaped unacceptable outcomes.
Investigative tools provided in this handbook are designed to determine the “what” and the
“why.” These investigative tools allow investigation teams to systematically explore what failed
in the systems used to ensure safety. Rooting out the deeper organizational issues reduces
degradation of any system modification put in place.
1.10.2 Determine Deeper Organizational Factors
Having determined “what” went wrong, the investigation team must attempt to use the theory
introduced in Chapter 1 to understand how extensive the issues discovered in the investigation
are throughout the organization, how long they have been undetected and uncorrected, and why
the culture of the organizations allowed this to occur. To answer these questions, the team needs
to determine the extent of conditions and causes, attempt to identify the Latent Organizational
Weaknesses (those management decisions made in the past that are now starting to set
employees up for errors) and attempt to identify underlying cultural issues that may have
contributed to these.
A learning organization must determine “what” did not work by performing a compliance-based
assessment and understand “why” the organization was allowed to get to this stage by
performing a cultural-based assessment. In the Federally-led accident investigation the
compliance based assessment is driven by DOE O 225.1B which requires the team investigate
policies, standards, and requirements that were applicable to the accident being investigated and
to investigate the safety management system that was to be in place to institutionalize the
resulting work practices to allow safe work (DOE Policy (P) 450.4A, Safety Management System
Policy). This is accomplished by reviewing work against the ISM Core Functions. The cultural-
based assessment is accomplished by examining three principal culture shaping factors
(leadership, employee engagement, organizational learning) which are developed from the ISM
Principles.
This output of the deeper organizational issues is much more subjective than previous sections
because it is based on the team assimilating information and making educated judgments as to
possible underlying organizational causes. The following sections are provided to frame the
deeper organizational part of the investigation and the results should be used in conjunction with
1‐38
DOE‐HDBK‐1208‐2012
the organizational mental model introduced earlier (Figure 1-13: Level III - High-Level Model
for Examining Organizational Drift).
1.10.3 Extent of Conditions and Cause
The team should determine how long conditions have existed without detection (hints that the
organization’s assessment and oversight processes are not very effective) and how extensive the
conditions are throughout the organization (hints which point to deeper management system
issues, indicating a higher level corrective action needed).
Section 43
As part of this effort, the team should also capture the missed opportunities to catch this event in
its early stages such that the event being investigated would not have occurred. A learning
organization should be taking every attempt to learn from previous mishaps or near misses
(including external lessons learned) and have sufficiently robust process to detect when things
are going wrong early in the process.
1.10.4 Latent Organizational Weaknesses
Latent organizational weaknesses are hidden deficiencies in management control processes (for
example, strategy, policies, work control, training, and resource allocation) or values (shared
beliefs, attitudes, norms, and assumptions) that create workplace conditions that can provoke
error (i.e., precursors) and degrade the integrity of defenses (flawed defenses). [Reason, pp. 10
18, 1997]14
Table 1-1 is a guide, to help identify latent organizational weaknesses - those factors in the
management control processes or associated values that influence errors or degrade defenses.
Consider work practices, resources, documentation, housekeeping, industrial safety, management
effectiveness, material availability, oversight, program controls, radiation employee practices,
security work practices, tools and equipment use, training and qualification, work planning and
execution, and work scheduling. For an expanded list of examples, see Attachment 1, ISM
Crosswalk and Safety Culture Lines of Inquiry.
1‐39
DOE‐HDBK‐1208‐2012
Table 1-1: Common Organizational Weaknesses
Category Weakness
Training Effectiveness of training on task qualification requirement for skill‐based tasks.
Focus is on lower level of cognitive knowledge.
Failure to involve management in training.
Training is inconsistent with company equipment, procedures, or process.
Communication Reinforcement of use of the phonetic alphabet in critical steps to preclude
misunderstanding of instructions.
Failure to reinforce use of 3‐way communications.
Failure to use specific unit ID numbers in procedures Unclear priorities or
expectations.
Unclear roles and responsibilities.
Planning and Provision for contingencies for failures.
Scheduling Failure to consider that multiple components may be out of service.
Failure to provide required materials or procedures.
Over scheduling of resources.
Failure to consider incorrect operation or damage to adjacent equipment.
Specific type of work not performed.
Specific type of issue not addressed Inadequate resources assigned.
Design or Process
Change
Involvement of users in design change implementation.
Inadequate training.
Inadequate contingencies in case a procedure goes wrong
Values, Priorities, Management policies on line input into adequacy of procedures or safety features.
Policies Too high a priority is placed on schedules.
Willingness to accept degraded conditions or performance.
Management failure to recognize the need for or importance of related program.
Section 44
Procedure Consideration of human factors in procedural development and implementation.
Development or Failure to perform procedural verification or validation.
Use
Failure to reference procedure during task performance.
Assumptions made in lieu of procedural guidance.
Omission of necessary functions in procedures.
1‐40
DOE‐HDBK‐1208‐2012
Category Weakness
Supervisory
Involvement
Performance of management observations and coaching.
Failure to correct poor performance or reinforce good performance.
Unassigned or fragmented responsibility and accountability.
Inadequate program oversight
Organizational
Interfaces
Interfaces for defining work priorities.
Lack of clear lines of communications between organizations.
Conflicting goals or requirements between programs Lack of self‐assessment
monitoring.
Lack of measurement tools for monitoring program performance.
Lack of interface between programs.
Work Practices Reinforcement of the use of established error prevention tools and techniques
(human performance tools).
1.10.5 Organizational Culture
Insights about safety culture may be inferred by considering aspects of leadership, employee
engagement and organizational learning. Observations about culture should be captured
reviewed and summarized to distill indicators of the most significant culture observations. These
are phrased as positive culture challenges in the report.
An organization’s culture, if not properly aligned with safety requirements, could result in
ignored safety requirements. A healthy culture exists when the “work-as-done” (culture artifacts
and behavior) overlap the “work-as-planned” (espoused beliefs and values) indicating an
alignment with the underlying assumptions (those factors felt important to management). A
misalignment between actual safety behavior and espoused safety beliefs indicates an unhealthy
culture or one in which the employees are not buying into the established safety system or one in
which the true underlying assumptions of management is focused on something besides safety
(Figure 1-15).
1‐41
39
DOE‐HDBK‐1208‐2012
Work‐as‐Imagined
Underlying
Assumptions
Espoused
Beliefs and
Values
Below the surface
Underlying assumptions must be
understood to properly interpret
artifacts and to create change
Work‐as‐Done Artifacts and
Behaviors
Misalignment hints at
deeper underlying
assumptions keeping the
organization from
attaining its desired
balance between
production and safety
Schein, Organizational Culture and Leadership, 2004
Figure 1-15: Assessing Organizational Culture
Safety culture factors offer important insights about event causation and prevention. Although
in-depth safety culture evaluations are beyond the doable scope of most accident investigations,
the DOE (with the help of EFCOG) has determined that examining three principal culture
shaping factors (leadership, employee engagement, organizational learning) will help to identify
cultural issues that contributed to the event. These factors were developed from the ISM
Principles by the EFCOG Safety Culture Working Group in 2007.
Leadership
Section 45
Leadership and culture are two sides of the same coin; neither can be realized without the other.
Leaders create and manage the safety culture in their organizations by maintaining safety as a
priority, communicating their safety expectations to the workers, setting the standard for safety
through actions not talk (walk the talk), leading needed change by defining the current state,
establishing a vision, developing a plan, and implementing the plan effectively. Leaders
cultivate trust to engender active participation in safety and to establish feedback on the
effectiveness of their organization’s safety efforts.
Leaders assure plans integrate safety into all aspects of an organization’s activities
considering the consequences of operational decisions for the entire life-cycle of operations
1‐42
DOE‐HDBK‐1208‐2012
and the safety impact on business processes, the organization, the public, and the
environment.
Leaders understand their business and ensure the systems employed provide the requisite
safety by identifying and minimizing hazards, proving the activity is safe, and not assuming
it is safe before operations commence.
Leaders consider safety implications in the change management processes.
Leaders model, coach, mentor, and reinforce their expectations and behaviors to improve
safe business performance.
Leaders value employee involvement, encourage individual questioning attitude, and instill
trust to encourage raising issues without fear of retribution.
Leaders assure employees are trained, experienced and have the resources, the time, and the
tools to complete their job safely.
Leaders hold personnel accountable for meeting standards and expectations to fulfill safety
responsibilities.
Leaders insist on conservative decision making with respect to the proven safety system and
recognize that production goals, if not properly considered and clearly communicated, can
send mixed signals on the importance of safety.
Leadership recognizes that humans make mistakes and take actions to mitigate this.
Leaders develop healthy, collaborative relationships within their own organization and
between their organization and regulators, suppliers, customers and contractors.
Employee/Worker Engagement
Safety is everyone’s responsibility. As such, employees understand and embrace the
organization’s safety behaviors, beliefs, and underlying assumptions. Employees understand and
embrace their responsibilities, maintain their proficiency so that they speak from experience,
challenge what is not right and help fix what is wrong and police the system to ensure them, their
co-workers, the environment, and the public remain safe.
Individuals team with leaders to commit to safety, to understand safety expectations, and to
meet expectations.
Individuals work with leaders to increase the level of trust and cooperation by holding each
other accountable for their actions with success evident by the openness to raise and resolve
issues in a timely fashion.
Everyone is personally responsible and accountable for safety, they learn their jobs, they
know the safety systems and they actively engage in protecting themselves, their co
workers, the public and the environment.
1‐43
DOE‐HDBK‐1208‐2012
Individuals develop healthy skepticism and constructively question deviations to the
established safety system and actively work to avoid complacency or arrogance based on
past successes.
Section 46
Individuals make conservative decisions with regards to the proven safety system and
consider the consequences of their decisions for the entire life-cycle of operations.
Individuals openly and promptly report errors and incidents and don’t rest until problems are
fully resolved and solutions proven sustainable.
Individuals instill a high level of trust by treating each other with dignity and respect and
avoiding harassment, intimidation, retaliation, and discrimination. Individuals welcome and
consider a diversity of thought and opposing views.
Individuals help develop healthy collaborative relationships within their organization and
between their organization and regulators, suppliers, customers and contractors.
Organizational Learning
The organization learns how to positively influence the desired behaviors, beliefs and
assumptions of their healthy safety culture. The organization acknowledges that errors are a way
to learn by rewarding those that report, sharing what is wrong, fixing what is broken and
addressing the organizational setup factors that led to employee error. This requires focusing on
reducing recurrences by correcting deeper, more systemic causal factors and systematically
monitoring performance and interpreting results to generate decision-making information on the
health of the system.
The organization establishes and cultivates a high level of trust; individuals are comfortable
raising, discussing and resolving questions or concerns.
The organization provides various methods to raise safety issues without fear of retribution,
harassment, intimidation, retaliation, or discrimination.
Leaders reward learning from minor problems to avoid more significant events.
Leaders promptly review, prioritize, and resolve problems, track long-term sustainability of
solutions, and communicate results back to employees.
The organization avoids complacency by cultivating a continuous learning/improvement
environment with the attitude that “it can happen here.”
Leaders systematically evaluate organizational performance using: workplace observations,
employee discussions, issue reporting, performance indicators, trend analysis, incident
investigations, benchmarking, assessments, and independent reviews.
The organization values learning from operational experience from both inside and outside
the organization.
1‐44
DOE‐HDBK‐1208‐2012
The organization willingly and openly engages in organizational learning activities.
1.11 Experiential Lessons for Successful Event Analysis
A fundamental shortcoming of some investigative techniques is that they do not address where
the physics could fail, based on perceptions of improbability due to lack of recent evidence (it
has happened before). People, equipment, and facilities only get hurt or damaged when energy
flows to where it does not belong. Investigations must determine where the physics could fail in
order to prevent potential bad consequences.
“System Optimism” is the belief that systems are well designed and well maintained, procedures
are complete and correct, designers can foresee and anticipate every situation, and that people
behave as they are expected to or as they were taught. This is the “work-as-imagined” by the
organizational management culture. In this view, people are a liability and deviation from the
“work-as-imagined” is seen as a threat to safety that needs to be eliminated. In other words, this
is the perception that errors are caused by the individuals who made them; correct or remove the
errant individual and the problem is fixed.
Section 47
“System Reality” is the belief that things go right because people learn to overcome design flaws
and functional glitches, adapt their performance to meet demands, interpret and apply procedures
to match conditions, and can detect and correct when things go wrong. In this view, people are
an asset and the deviation from the “work-as-imagined” is seen as how workers have to adapt to
successfully complete the work within the time and resources constraints that exist for that task.
In other words, if the worker is adapting incorrectly, the fault is in the conditions and methods
available to adapt.
Rather than simply judging a decision as wrong in retrospect, the decision needs to be evaluated
in the context of contributing factors that explain why the decision was made. If the
investigation stops with worker’s deviation as the cause, nothing is corrected. The next worker,
working in the same context, will eventually adapt in a similar fashion and deviate from “work
as imagined.” Performance variability is not limited to just the worker who triggers the accident.
People are involved in all aspects of the work, including variation in the actions of the co
workers, the expectations of the leaders, accuracy of the procedures, the effectiveness of the
defenses and barriers, or even the basic policies of the organization can influence an outcome.
This is reflected in the complex, non-linear accident model where unexpected combinations of
normal variability can result in the accident. Failure to follow up with lessons-to-be-learned and
validations of corrective actions and Judgment of Needs can certainly lead to a recurrence of an
event.
1‐45
DOE‐HDBK‐1208‐2012
1‐46
DOE‐HDBK‐1208‐2012
CHAPTER 2. THE ACCIDENT INVESTIGATION PROCESS
2. THE ACCIDENT INVESTIGATION PROCESS
2.1 Establishing the Federally Led Accident Investigation Board and
Its Authority
2.1.1 Accident Investigations’ Appointing Official
Section 2.1 primarily deals with the DOE Federal responsibilities under DOE O 225.1B. Upon
notification of an accident requiring a DOE Federal investigation, the Appointing Official selects
the AIB Chairperson. The Appointing Official, with the assistance of the Board Chairperson,
selects three to six other Board members, one of whom must be a trained DOE accident
investigator. All of the AIB members are DOE federal employees. To minimize conflicts of
interest influences, the Chairperson and the accident investigator must be from a different duty
station than the accident location. The Appointing Official for a Federal accident investigation is
the Head of Program Element, unless this responsibility is delegated to the Chief Health, Safety
and Security Officer (HS-1). The roles and responsibilities of the Appointing Official for
Accident Investigations, the Heads of Program Elements for Accident Investigations, and the
Heads of Field Elements for establishing and supporting AIBs are defined in the Table 2-1.
Table 2-1: DOE Federal Officials and Board Member Responsibilities
Participants Major Responsibilities
Appointing Official
for Accident
Investigations
Section 48
Formally appoints the Accident Investigation Board in writing within three days of
accident categorization
Establishes the scope of the Board’s authority, including the review of management
systems, policy, and line management oversight processes as possible causal factors
Briefs Board members within three days of their appointment
Ensures that notification is made to other agencies, if required by memoranda of
understanding, law, or regulation
Emphasizes the Board’s authority to investigate the causal roles of organizations,
management systems, and line management oversight up to and beyond the level of the
appointing official
Accepts the investigation report and the Board’s findings
Publishes and distributes the respective investigation report within seven calendar days
of report acceptance
Develops lessons learned for dissemination throughout the Department or the
organization for or the OSRs
Closes the investigation after the actions in DOE O 225.1B, Paragraph 4d, are completed
2
.1
2‐1
DOE‐HDBK‐1208‐2012
Participants Major Responsibilities
Serves as Appointing Official for Federal accident investigations for programs, offices and
Elements for
Heads of Program
facilities under their authority.
Accident Maintain a staff of trained and qualified personnel to serve in the capacity of Chairperson
Investigations and DOE Accident Investigators for AIBs and, upon request, provide them to support
other AIBs.
Ensure that DOE and contractor organizations are prepared to effectively accomplish
initial investigative actions and assist Accident Investigation Boards
Categorize the accident investigation in accordance with the criteria provided in
Attachment 2 of DOE O 225.1B
Report accident categorization and initial actions taken by DOE site teams to the Office
of Corporate Safety Programs (HS‐23)
Serve as the appointing official for Federal accident investigations
Ensure that readiness teams and emergency management personnel coordinate their
activities to facilitate an orderly transition of responsibilities for the accident scene
Develop lessons learned for Federal accident investigation
Require submittal of corrective action plans to address the Judgments of Need, approve
the implementation of those plans, and track the effective implementation of those
plans to closure.
Distribute accident investigation reports to all Heads of Field Elements under their
cognizance and direct that extent‐of‐condition reviews be conducted for issues identified
during accident investigations that are applicable to work locations and operations.
2‐2
Section 49
DOE‐HDBK‐1208‐2012
Participants Major Responsibilities
Heads of Field
Elements for
Accident
Investigations
Maintain a state of readiness to conduct investigations throughout the field element,
their operational facilities, and the DOE site teams
Ensure that sufficient numbers of site DOE and contractor staff understand and are
trained to conduct or support investigations
Procure appropriate equipment to support investigations
Maintain a current site list of DOE and contractor staff trained in conducting or
supporting investigations
Assist in coordinating investigation activities with accident mitigation measures taken by
emergency response personnel
Communicate and transfer information on accidents to the head of the Headquarters
program elements to whom they report
Communicate and transfer information to the Accident Investigation Board Chairperson
before and after his/her arrival on site
Coordinate corrective action planning and follow‐up with the head of the Headquarters
program element and coordinate comment resolution by reviewing parties
Facilitate distribution of lessons learned identified from accident investigations
Serve as liaison to the HSS AI Program Manager on accident investigation matters
Develop or provide assistance in developing lessons learned for accident investigations.
Require the submittal of contractor corrective action plans to address the Judgments of
Need, approve the implementation of those plans, and track the effective
implementation of those plans to closure
Conduct extent‐of‐condition reviews for specific issues resulting from accident
investigations that might be applicable to work locations or activities under the Heads of
Field Elements’ authority, and address applicable lessons learned from investigations
conducted at other DOE sites
2.1.2 Appointing the Accident Investigation Board
A list of prospective Chairpersons who meet minimum qualifications is available from the HSS
AI Program Manager and maintains a list of qualified Board members, consultants, advisors, and
support staff, including particular areas of expertise for potential Board members or
consultants/advisors. The Appointing Official, with the help of the HSS AI Program Manager,
and the selected AIB Chairperson, assess the potential scope of the investigation and identify
other board members needed to conduct the investigation. In selecting these individuals, the
chairperson and appointing official follow the criteria defined in DOE O 225.1B, which are
shown in Table 2-2.
2‐3
DOE‐HDBK‐1208‐2012
Table 2-2: DOE Federal Board Members Must Meet These Criteria
Role Qualifications
Chairperson Senior DOE manager
Preferably a member of the Senior Executive Service or at a senior
general service grade level deemed appropriate by the appointing
official
Demonstrated managerial competence
Knowledgeable of DOE accident investigation techniques
Experienced in conducting accident investigations through participation
in at least one Federal investigation, or equivalent experience
Section 50
Board Members DOE Federal employee
Subject matter expertise in areas related to the accident, including
knowledge of the Department’s safety management system policy and
integrated safety management system
Either the Chairperson or, at least one Board member, must be a DOE
accident investigator, who has participated in an accident investigation
course sponsored by the Office of Corporate Safety Programs
Board Advisor/Consultant Knowledgeable in evaluating management systems, the adequacy of
policy and its implementation, and the execution of line management
oversight
Industry working knowledge in the analytical techniques used to
determine accident causal factors
DOE O 225.1B establishes some additional restrictions concerning the selection of Board
members and Chairpersons. Members are not permitted to have:
A supervisor-subordinate relationship with another Board member
Any conflict of interest or direct or line management responsibility for day-to-day operation
or oversight of the facility, area, or activity involved in the accident.
Both the Chairperson and the DOE Accident Investigator must be selected from a different
duty station than the accident location.
Consultants, advisors, and support staff can be assigned to assist the Board where necessary,
particularly when DOE employees with necessary skills are not available. For example, advisory
staff may be necessary to provide knowledge of management systems or organizational concerns
or expertise on specific DOE policies. A dedicated and experienced administrative coordinator
(see Appendix C) is recommended. The Program Manager can help identify appropriate
personnel to support Accident Investigation Boards.
2‐4
DOE‐HDBK‐1208‐2012
The appointing official appoints the Accident Investigation Board within three calendar days
after the accident is categorized by issuing an appointment memorandum. The appointment
memorandum establishes the Board’s authority and releases all members of the AIB from their
normal responsibilities/duties for the period of time the Board is convened. The appointment
memorandum also includes the scope of the investigation, the names of the individuals being
appointed to the Board, a specified completion date for the final report (nominally 30 calendar
days), and any special provisions deemed appropriate.
The appointment memorandum should specify the scope of the investigation which includes:
Gathering facts;
Analyzing causes;
Developing conclusions and,
Developing Judgments of Need related to DOE and contractor organizations and
management systems that could or should have prevented the accident.
A Sample Appointment Memorandum may be found in Appendix D.
2.1.3 Briefing the Board
The appointing official is responsible for briefing all Board members as soon as possible (within
three days) after their appointment to ensure that they clearly understand their roles and
responsibilities. This briefing may be given via videoconference or teleconference. If it is
impractical to brief the entire Board, at least the Board Chairperson should receive the briefing
and then convey the contents of the briefing to the other Board members before starting the
investigation. The briefing emphasizes:
The scope of the investigation;
The Board’s authority to examine DOE and contractor organizations and management
systems, including line management oversight, as potential causes of an accident, up to and
beyond the level of the appointing official;
Section 51
The necessity for avoiding conflicts of interest;
Evaluation of the effectiveness of management systems, as defined by DOE P 450.4A;
Pertinent accident information and special concerns of the appointing official based on site
accident patterns or other considerations.
2‐5
DOE‐HDBK‐1208‐2012
2
.2
2.2 Organizing the Accident Investigation
The accident investigation is a complex project that involves a significant workload, time
constraints, sensitive issues, cooperation between team members, and dependence on others.
To finish the investigation within the time frame required, the AIB chairperson must exercise
good project management skills and promote teamwork. The Chairperson’s initial decisions and
actions will influence the tone, tempo, and degree of difficulty associated with the entire
investigation. This section provides the Board Chairperson with techniques and tools for
planning and organizing the investigation.
2.2.1 Planning
Project planning must occur early in the investigation. The Chairperson should begin developing
a plan for the investigation immediately after his/her appointment. The plan should include a
preliminary report outline, specific task assignments, and a schedule for completing the
investigation. It should also address the resources, logistical requirements, and protocols that
will be needed to conduct the investigation.
A tool for the Chairperson, the Accident Investigation Startup Activities List, is included in
Appendix D. The Chairperson and administrative coordinator can use this list to organize the
initial investigative activities.
2.2.2 Collecting Initial Site Information
Following appointment, the Chairperson is responsible for contacting the site/sponsoring
organization to obtain as many details on the accident as possible. The sponsoring organization,
which could include a DOE field program office, and/or contractor division point-of-contact, is
usually designated as the liaison with the Board. The Chairperson needs the details of the
accident to determine what resources, Board member expertise, and technical specialists will be
required. Furthermore, the Chairperson should request background information, including site
history, sitemaps, and organization charts. The Accident Investigation Information Request
Form (provided in Appendix D) can be used to document and track these and other information
requests throughout the investigation.
2.2.3 Determining Task Assignments
A useful strategy for determining and allocating tasks is to develop an outline of the accident
investigation report, including content and format, and use it to establish tasks for each Board
member. This outline helps to organize the investigation around important tasks and facilitates
getting the report writing started as early as possible in the investigation process. Board
members, advisors, and consultants are given specific assignments and responsibilities based on
their expertise in areas such as management systems, work planning and control, occupational
safety and health, training, and any other technical areas directly related to the accident. These
assignments include specific tasks related to gathering and analyzing facts, conducting
interviews, determining causal factors, developing Conclusions (CON) and JONs, and report
2‐6
Section 52
DOE‐HDBK‐1208‐2012
writing. Assigning designated Board members specific responsibilities ensures consistency
during the investigation.
2.2.4 Preparing a Schedule
The Chairperson also prepares a detailed schedule using the generic four-week accident
investigation cycle and any specific direction from the appointing official. The Chairperson
should establish significant milestones; working back from the appointing official’s designated
completion date. Table 2-3 shows a list of typical activities to schedule.
Table 2-3: These Activities should be Included in an Accident Investigation
Schedule
Interviews/Evidence Collection and Preliminary Analysis
Obtain needed site and/or facility/project background information, policies, procedures, and training
records
Assign investigation tasks and writing responsibilities
Initiate and complete first draft of accident chronology and facts
Select analytical methods (preliminary)
Complete interviews
Complete first analyses of facts using selected analytical tools; determine whether additional tools are
necessary
Obtain necessary photographs and complete illustrations for report
Internal Review Drafts
Complete first draft of report elements, up to and including facts and analysis section
Complete development and draft of direct, contributing, and root causes
Complete development and draft of Judgments of Need
Complete first draft of report for internal review
Complete draft analyses
Complete second draft of report for internal review
2‐7
DOE‐HDBK‐1208‐2012
External Review Drafts
Complete Classification/Privacy Act reviews
Conduct factual accuracy review and revise report based on input
Complete report for Quality Assurance review by HSS Office Corporate Safety Programs prior to
submission to the Appointing Official
Complete final draft of report
Prepare out‐brief materials
Brief relevant site/division and/or field office managers (depending on type of investigation) on findings
Leave site
Complete final production of report
The schedule developed by the Board Chairperson should include the activities to be conducted
and milestones for their completion. A sample schedule is included as Figure 2-1. The Accident
Investigation Day Planner: a Guide for Accident Investigation Board Chairpersons, available on
the AI Program website, can assist in the development of this schedule. Activities cover
nominally 30 days.
Figure 2-1: Typical Schedule of Accident Investigation
2.2.5 Acquiring Resources
From the first day, the Chairperson begins acquiring resources for the investigation. This
includes securing office space, a conference room or “command center”, office supplies, and
2‐8
DOE‐HDBK‐1208‐2012
computers through the Field Office Manager (FOM), a secured area for document storage, tools,
and personal protective equipment, if necessary. The site’s FOM should provide many of these
resources. The Accident Investigation Equipment Checklist (see Appendix D) is designed to help
identify resource needs and track resource status.
In addition, the Board Chairperson assures that contracting mechanisms exist and that funding is
available for the advisors and consultants required to support the investigation. These activities
are coordinated with the Appointing Official.
Section 53
2.2.6 Addressing Potential Conflicts of Interest
The Board Chairperson is responsible for resolving potential conflicts of interest regarding Board
members, advisors, and consultants. Each Board member, advisor, and consultant should certify
that he or she has no conflicts of interest by signing the Accident Investigation Individual
Conflict of Interest Certification Form (provided in Appendix D). If the Chairperson or any
individual has concern about the potential for or appearance of conflicts of interest, the
Chairperson should inform the Appointing Official and seek legal counsel input, if necessary.
The decision to allow the individual to participate in the investigation, and any restrictions on his
or her participation, shall be documented in a memorandum signed by the Board Chairperson to
the Appointing Official. If the Chairperson relies on the advice of legal counsel, the Chairperson
shall seek appropriate legal counsel concurrence through the Appointing Official. The
memorandum will become part of the Board’s permanent record.
2.2.7 Establishing Information Access and Release Protocols
The Chairperson is responsible for establishing protocols relating to information access and
release. These protocols are listed in Table 2-4. Information access and other control protocols
maintain the integrity of the investigation and preserve the privacy and confidentiality of
interviewees and other parties.
The Freedom of Information Act (FOIA) and Privacy Act may apply to information generated or
obtained during an investigation. These two laws dictate access to and release of government
records. The Chairperson should obtain guidance from a legal advisor or the FOIA/Privacy Act
contact person at the site, field office, or Headquarters regarding question of disclosure, or the
applicability of the FOIA or Privacy Act. The FOIA provides access to Federal agency records
except those protected from release by exemptions. Anyone can use the FOIA to request access
to government records.
The Board must ensure that the information it generates is accurate, relevant, complete, and up
to-date. For this reason, court reporters may be used in more serious investigations to record
interviews, and interviewees should be allowed to review and correct transcripts.
The Privacy Act protects government records on citizens and lawfully admitted permanent
residents from release without the prior written consent of the individual to whom the records
pertain.
2‐9
DOE‐HDBK‐1208‐2012
Specifically, when the Privacy Act is applicable, the Board is responsible for:
Informing interviewees why information about them is being collected and how it will be
used.
Ensuring that information subject to the Privacy Act is not disclosed without the consent of
the individual, except under the conditions prescribed by law. Information that can
normally not be disclosed includes name, present and past positions or “grade” (e.g., GS
13), annual salaries, duty station, and position description. Therefore, the Board should not
request this information unless it is relevant to the investigation.
Section 54
A Model Interview Opening Statement that addresses the provisions of both the FOIA and the
Privacy Act and their pertinence to interviews for DOE accident investigations is provided in
Appendix D. This statement should be read at the beginning of all applicable interviews. A
brief explanatory Reference Copy of 18USC Sec. 1001 for Information is provided to the
interviewer in Appendix D, in the event questions are raised by the opening statement.
2.2.8 Controlling the Release of Information to the Public
The Chairperson should instruct Board members not to communicate with the press or other
external organizations regarding the investigation. External communications are the
responsibility of the Board Chairperson until the final report is released. The Board
Chairperson should work closely with a person designated by the site to release other
information, such as statements to site employees and the public.
Table 2-4: The Chairperson Establishes Protocols for Controlling Information
Protocol Considerations
Information Security Keep all investigative evidence and documents locked in a secure area
accessible only to Board members, advisors, and support staff.
Press Releases
(if appropriate)
Board Chairpersons should coordinate with the official authorizing the
investigation or their normal chain of command for authority/guidance
on Press Releases.
Determine whether there is a designated contact to handle press
releases; if so, work with that person.
The Board is not obligated to release any information. However,
previous chairpersons have found that issuing an early press release can
be helpful.
The initial press release usually contains a general description of the
accident and the purpose of the investigation.
The Board chairperson should review and approve all press releases (in
addition to whatever review process at the parent organization).
2‐10
DOE‐HDBK‐1208‐2012
Protocol Considerations
Lines of Communication Establish liaison with field element management and/or with the
operating contractor at the site, facility, or area involved in the accident
to set up clear lines of communication and responsibility.
Format of Information
Releases
Determine the amount and format of information to be released to the
site contractor(s), union advisor, and local DOE office for internal
purposes.
Never release verbatim interview transcripts or tapes due to the
sensitivity of raw information.
Do not release preliminary results of analyses. These results can be
taken out of context and lead to premature conclusions by the site and
the media.
Consult with the appointing official before releasing any information.
Approvals for Information Assure that Board members, site contractors, and the local DOE office
Releases do not disseminate information concerning the Board’s activities,
findings, or products before obtaining the Chairperson’s approval. Brief
the Board on what they can reveal to others.
2.3 Managing the Investigation Process
Section 55
As an investigation proceeds, the Chairperson uses a variety of management techniques,
including guiding and directing, monitoring performance, providing feedback on performance,
and making decisions and changes required to meet the investigation’s objectives and schedule.
Because these activities are crucial, the Chairperson may designate an individual to oversee
management activities in case the Chairperson is not always immediately available.
2.3.1 Taking Control of the Accident Scene
Before arriving at the site, the Chairperson communicates with the point of contact or the
appropriate DOE site designee to assure that the scene and evidence are properly secured,
preserved, and documented and that preliminary witness information has been gathered. At the
accident scene, the Chairperson should:
Obtain briefings from all persons involved in managing the accident response.
Obtain all information and evidence gathered by the DOE site team.
Make a decision about how secure the accident scene must remain during the initial phases
of the investigation. If there are any concerns about loss or contamination of evidence, play
it safe and keep the scene restricted from use.
Assume responsibility only for activities directly related to the accident and investigation.
The Chairperson and Board members should not take responsibility for approving site
2
.3
2‐11
DOE‐HDBK‐1208‐2012
activities or procedures, or for recovery, rehabilitation, or mitigation activities. These
functions are the responsibility of line management.
2.3.2 Initial Meeting of the Accident Investigation Board
The Chairperson is responsible for ensuring that all Board members work as a team and share a
common approach to the investigation. As one of the Board’s first onsite activities, the
Chairperson typically holds a meeting to provide all Board members, advisors, consultants, and
support staff with an opportunity to introduce themselves and to give the Chairperson an
opportunity to brief the Board members on:
The scope of the investigation, including all levels of the organizations involved up to and
beyond the level of the appointing official;
An overview of the accident investigation process, with emphasis on:
Streamlined process and limited time frame to conduct the investigation (if applicable);
The schedule and plan for completing the investigation; and
The need to apply the components of DOE’s integrated safety management system during
the investigation as the means of evaluating management systems.
Potential analytical and testing techniques to be used;
The roles, responsibilities, and assignments for the Chairperson, the Board members, and
other participants;
Information control and release protocols; and
Administrative processes and logistics.
At the meeting, the Chairperson clearly communicates expectations and provides direction and
guidance for the investigation. In addition, at the meeting the Chairpersons should distribute
copies of local phone directories and a list of phone and fax numbers pertinent for the
investigation. The Board should also be briefed on procedures for:
Handling potential conflicts of interest resulting from using contractor-provided support and
obtaining support from other sources;
Storing investigative materials in a secured location and disposing of unneeded yet sensitive
materials;
Section 56
Using logbooks, inventory, checkout lists, or other methods to maintain control and
accountability of physical evidence, documents, photographs, and other material pertinent to
the investigation;
Recording and tracking incoming and outgoing correspondence; and
2‐12
DOE‐HDBK‐1208‐2012
Accessing the Board’s work area after hours.
2.3.3 Promoting Teamwork
The Board must work together as a team to finish the investigation within the time frame
established by the appointing official. To make this happen, the Board Chairperson should
ensure that strong-willed personalities do not dominate and influence the objectivity of the
investigation and that all viewpoints are heard and analyzed.
The Chairperson must capitalize on the synergy of the team’s collective skills and talents (i.e.,
the team is likely to make better decisions and provide a higher quality investigation than the
same group working individually), while allowing individual actions and decisions. It is
important that the Chairperson set the ground rules and provide guidance to the Board members
and other participants.
Friendship is not required, but poor relationships can impede the Board’s ability to conduct a
high-quality investigation. The Chairperson can encourage positive relationships by focusing
attention on each member’s strengths and downplaying weaknesses. The Chairperson can
facilitate this by arranging time to allow team members to get to know one another and learn
about each other’s credentials, strengths, and preferences. Effective interpersonal relationships
can save time and promote high-quality performance.
It is the Chairperson’s responsibility to make sure that all members get a chance to speak and
that no one member dominates conversations. The Chairperson should establish communication
guidelines and serve as an effective role model in terms of the following:
Be clear and concise; minimize the tendency to think out loud or tell “war stories.”
Be direct and make your perspective clear.
Use active listening techniques, such as focusing attention on the speaker, paraphrasing,
questioning, and refraining from interrupting.
Pay attention to non-verbal messages and attempt to verbalize what you observe.
Attempt to understand each speaker’s perspective.
Seek information and opinions from others, especially the less talkative members.
Consider all ideas and arguments.
Encourage diverse ideas and opinions.
Suggest ideas, approaches, and compromises.
Help keep discussions on track when they start to wander.
2‐13
DOE‐HDBK‐1208‐2012
The Chairperson should gain agreement in advance regarding how particular decisions will be
made. Decisions can b