Regular price
32.000 KD
inc. VAT
Couldn't load pickup availability
Table of contents
- Contentsv
- Forewordxi
- Prefacexiii
- Acknowledgementsxvii
- About the Authorsxix
- 1 Preliminaries1
- 1.1 Fault Classification2
- 1.2 Types of Redundancy3
- 1.3 Basic Measures of Fault Tolerance4
- 1.3.1 Traditional Measures5
- 1.3.2 Network Measures6
- 1.4 Outline of This Book7
- 1.5 Further Reading9
- References10
- 2 Hardware Fault Tolerance11
- 2.1 The Rate of Hardware Failures11
- 2.2 Failure Rate, Reliability, and Mean Time to Failure13
- 2.3 Canonical and Resilient Structures15
- 2.3.1 Series and Parallel Systems16
- 2.3.2 Non-Series/Parallel Systems17
- 2.3.3 M-of-N Systems20
- 2.3.4 Voters23
- 2.3.5 Variations on N-Modular Redundancy23
- 2.3.6 Duplex Systems27
- 2.4 Other Reliability Evaluation Techniques30
- 2.4.1 Poisson Processes30
- 2.4.2 Markov Models33
- 2.5 Fault-Tolerance Processor-Level Techniques36
- 2.5.1 Watchdog Processor37
- 2.5.2 Simultaneous Multithreading for Fault Tolerance39
- 2.6 Byzantine Failures41
- 2.6.1 Byzantine Agreement with Message Authentication46
- 2.7 Further Reading48
- 2.8 Exercises48
- References53
- 3 Information Redundancy55
- 3.1 Coding56
- 3.1.1 Parity Codes57
- 3.1.2 Checksum64
- 3.1.3 M-of-N Codes65
- 3.1.4 Berger Code66
- 3.1.5 Cyclic Codes67
- 3.1.6 Arithmetic Codes74
- 3.2 Resilient Disk Systems79
- 3.2.1 RAID Level 179
- 3.2.2 RAID Level 281
- 3.2.3 RAID Level 382
- 3.2.4 RAID Level 483
- 3.2.5 RAID Level 584
- 3.2.6 Modeling Correlated Failures84
- 3.3 Data Replication88
- 3.3.1 Voting: Non-Hierarchical Organization89
- 3.3.2 Voting: Hierarchical Organization95
- 3.3.3 Primary-Backup Approach96
- 3.4 Algorithm-Based Fault Tolerance99
- 3.5 Further Reading101
- 3.6 Exercises102
- References106
- 4 Fault-Tolerant Networks109
- 4.1 Measures of Resilience110
- 4.1.1 Graph-Theoretical Measures110
- 4.1.2 Computer Networks Measures111
- 4.2 Common Network Topologies and Their Resilience112
- 4.2.1 Multistage and Extra-Stage Networks112
- 4.2.2 Crossbar Networks119
- 4.2.3 Rectangular Mesh and Interstitial Mesh121
- 4.2.4 Hypercube Network124
- 4.2.5 Cube-Connected Cycles Networks128
- 4.2.6 Loop Networks130
- 4.2.7 Ad hoc Point-to-Point Networks132
- 4.3 Fault-Tolerant Routing135
- 4.3.1 Hypercube Fault-Tolerant Routing136
- 4.3.2 Origin-Based Routing in the Mesh138
- 4.4 Further Reading141
- 4.5 Exercises142
- References145
- 5 Software Fault Tolerance147
- 5.1 Acceptance Tests148
- 5.2 Single-Version Fault Tolerance149
- 5.2.1 Wrappers149
- 5.2.2 Software Rejuvenation152
- 5.2.3 Data Diversity155
- 5.2.4 Software Implemented Hardware Fault Tolerance (SIHFT)157
- 5.3 N-Version Programming160
- 5.3.1 Consistent Comparison Problem161
- 5.3.2 Version Independence162
- 5.4 Recovery Block Approach169
- 5.4.1 Basic Principles169
- 5.4.2 Success Probability Calculation169
- 5.4.3 Distributed Recovery Blocks171
- 5.5 Preconditions, Postconditions, and Assertions173
- 5.6 Exception-Handling173
- 5.6.1 Requirements from Exception-Handlers174
- 5.6.2 Basics of Exceptions and Exception-Handling175
- 5.6.3 Language Support177
- 5.7 Software Reliability Models178
- 5.7.1 Jelinski–Moranda Model178
- 5.7.2 Littlewood–Verrall Model179
- 5.7.3 Musa–Okumoto Model180
- 5.7.4 Model Selection and Parameter Estimation182
- 5.8 Fault-Tolerant Remote Procedure Calls182
- 5.8.1 Primary-Backup Approach182
- 5.8.2 The Circus Approach183
- 5.9 Further Reading184
- 5.10 Exercises186
- References188
- 6 Checkpointing193
- 6.1 What is Checkpointing?195
- 6.1.1 Why is Checkpointing Nontrivial?197
- 6.2 Checkpoint Level197
- 6.3 Optimal Checkpointing—An Analytical Model198
- 6.3.1 Time Between Checkpoints—A First-Order Approximation200
- 6.3.2 Optimal Checkpoint Placement201
- 6.3.3 Time Between Checkpoints—A More Accurate Model202
- 6.3.4 Reducing Overhead204
- 6.3.5 Reducing Latency205
- 6.4 Cache-Aided Rollback Error Recovery (CARER)206
- 6.5 Checkpointing in Distributed Systems207
- 6.5.1 The Domino Effect and Livelock209
- 6.5.2 A Coordinated Checkpointing Algorithm210
- 6.5.3 Time-Based Synchronization211
- 6.5.4 Diskless Checkpointing212
- 6.5.5 Message Logging213
- 6.6 Checkpointing in Shared-Memory Systems217
- 6.6.1 Bus-Based Coherence Protocol218
- 6.6.2 Directory-Based Protocol219
- 6.7 Checkpointing in Real-Time Systems220
- 6.8 Other Uses of Checkpointing223
- 6.9 Further Reading223
- 6.10 Exercises224
- References226
- 7 Case Studies229
- 7.1 NonStop Systems229
- 7.1.1 Architecture229
- 7.1.2 Maintenance and Repair Aids233
- 7.1.3 Software233
- 7.1.4 Modifications to the NonStop Architecture235
- 7.2 Stratus Systems236
- 7.3 Cassini Command and Data Subsystem238
- 7.4 IBM G5241
- 7.5 IBM Sysplex242
- 7.6 Itanium244
- 7.7 Further Reading246
- References247
- 8 Defect Tolerance in VLSI Circuits249
- 8.1 Manufacturing Defects and Circuit Faults249
- 8.2 Probability of Failure and Critical Area251
- 8.3 Basic Yield Models253
- 8.3.1 The Poisson and Compound Poisson Yield Models254
- 8.3.2 Variations on the Simple Yield Models256
- 8.4 Yield Enhancement Through Redundancy258
- 8.4.1 Yield Projection for Chips with Redundancy259
- 8.4.2 Memory Arrays with Redundancy263
- 8.4.3 Logic Integrated Circuits with Redundancy270
- 8.4.4 Modifying the Floorplan272
- 8.5 Further Reading276
- 8.6 Exercises277
- References281
- 9 Fault Detection in Cryptographic Systems285
- 9.1 Overview of Ciphers286
- 9.1.1 Symmetric Key Ciphers286
- 9.1.2 Public Key Ciphers295
- 9.2 Security Attacks Through Fault Injection296
- 9.2.1 Fault Attacks on Symmetric Key Ciphers297
- 9.2.2 Fault Attacks on Public (Asymmetric) Key Ciphers298
- 9.3 Countermeasures299
- 9.3.1 Spatial and Temporal Duplication300
- 9.3.2 Error-Detecting Codes300
- 9.3.3 Are These Countermeasures Sufficient?304
- 9.3.4 Final Comment307
- 9.4 Further Reading307
- 9.5 Exercises307
- References308
- 10 Simulation Techniques311
- 10.1 Writing a Simulation Program311
- 10.2 Parameter Estimation315
- 10.2.1 Point Versus Interval Estimation315
- 10.2.2 Method of Moments316
- 10.2.3 Method of Maximum Likelihood318
- 10.2.4 The Bayesian Approach to Parameter Estimation322
- 10.2.5 Confidence Intervals324
- 10.3 Variance Reduction Methods328
- 10.3.1 Antithetic Variables328
- 10.3.2 Using Control Variables330
- 10.3.3 Stratified Sampling331
- 10.3.4 Importance Sampling333
- 10.4 Random Number Generation341
- 10.4.1 Uniformly Distributed Random Number Generators342
- 10.4.2 Testing Uniform Random Number Generators345
- 10.4.3 Generating Other Distributions349
- 10.5 Fault Injection355
- 10.5.1 Types of Fault Injection Techniques356
- 10.5.2 Fault Injection Application and Tools358
- 10.6 Further Reading358
- 10.7 Exercises359
- References363
- Subject Index365
Book details
- Vendor Elsevier S & T
- SKU 9780120885251
- ISBN-13 9780080492681
- Author Koren, Israel; Krishna, C. Mani
- Category Computers
- Subject General
Do you have questions about this book?
There are many applications in which the reliability of the overall system must be far higher than the reliability of its individual components. In such cases, designers devise mechanisms and architectures that allow the system to either completely mask the effects of a component failure or recover from it so quickly that the application is not seriously affected. This is the work of fault-tolerant designers and their work is increasingly important and complex not only because of the increasing number of “mission critical applications, but also because the diminishing reliability of hardware means that even systems for non-critical applications will need to be designed with fault-tolerance in mind.
Reflecting the real-world challenges faced by designers of these systems, this book addresses fault tolerance design with a systems approach to both hardware and software. No other text on the market takes this approach, nor offers the comprehensive and up-to-date treatment Koren and Krishna provide. Students, designers and architects of high performance processors will value this comprehensive overview of the field.
* The first book on fault tolerance design with a systems approach
* Comprehensive coverage of both hardware and software fault tolerance, as well as information and time redundancy
* Incorporated case studies highlight six different computer systems with fault-tolerance techniques implemented in their design
* Available to lecturers is a complete ancillary package including online solutions manual for instructors and PowerPoint slides
Reflecting the real-world challenges faced by designers of these systems, this book addresses fault tolerance design with a systems approach to both hardware and software. No other text on the market takes this approach, nor offers the comprehensive and up-to-date treatment Koren and Krishna provide. Students, designers and architects of high performance processors will value this comprehensive overview of the field.
* The first book on fault tolerance design with a systems approach
* Comprehensive coverage of both hardware and software fault tolerance, as well as information and time redundancy
* Incorporated case studies highlight six different computer systems with fault-tolerance techniques implemented in their design
* Available to lecturers is a complete ancillary package including online solutions manual for instructors and PowerPoint slides
Instant delivery by email
Your access email arrives within minutes of checkout, with a sign-in link for each book — no shipping, no waiting.
Read on any device
Books open in VitalSource Bookshelf on your phone, tablet, or computer, online or offline. Your library is always available at aafaq.vitalsource.com — just log in with the email you used at checkout.
Lost the email?
Resend it to yourself in seconds from My eBook orders, or email cs@aafaqeducation.com and we'll help.