ArticleslgStudy

science

IT disaster recovery

IT disaster recovery is a science topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand IT disaster recovery rather than just read about it. In short: IT disaster recovery (also, simply disaster recovery (DR)) is the process of maintaining or reestablishing vital infrastructure and systems following a natural or human-induced disaster, such as a storm or battle. DR employs policies, tools, and procedures with a focus on IT systems supporting critical business functions.

IT disaster recovery — main illustration
IT disaster recovery — illustration

Key takeaways

  • IT disaster recovery belongs to science; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect IT disaster recovery to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of IT disaster recovery from memory before moving on to harder problems.

Reference excerpt

IT disaster recovery (also, simply disaster recovery (DR)) is the process of maintaining or reestablishing vital infrastructure and systems following a natural or human-induced disaster, such as a storm or battle. DR employs policies, tools, and procedures with a focus on IT systems supporting critical business functions. This involves keeping all essential aspects of a business functioning despite significant disruptive events; it can therefore be considered a subset of business continuity (BC). DR assumes that the primary site is not immediately recoverable and restores data and services to a secondary site.

IT service continuity

IT service continuity (ITSC) is a subset of BCP, which relies on the metrics (frequently used as key risk indicators) of recovery point/time objectives. It encompasses IT disaster recovery planning and the wider IT resilience planning. It also incorporates IT infrastructure and services related to communications, such as telephony and data communications.

Principles of backup sites

Planning includes arranging for backup sites, whether they are "hot" (operating prior to a disaster), "warm" (ready to begin operating), or "cold" (requires substantial work to begin operating), and standby sites with hardware as needed for continuity. In 2008, the British Standards Institution launched a specific standard supporting Business Continuity Standard BS 25999, titled BS25777, specifically to align computer continuity with business continuity. This was withdrawn following the publication in March 2011 of ISO/IEC 27301, "Security techniques — Guidelines for information and communication technology readiness for business continuity." ITIL has defined some of these terms.

Recovery Time Objective The Recovery Time Objective (RTO) is the targeted duration of time and a service level within which a business process must be restored after a disruption in order to avoid a break in business continuity. According to business continuity planning methodology, the RTO is established during the business impact analysis (BIA) by the owner(s) of the process, including identifying time frames for alternate or manual workarounds.

RTO is a complement of RPO. The limits of acceptable or "tolerable" ITSC performance are measured by RTO and RPO in terms of time lost from normal business process functioning and data lost or not backed up during that period.

Recovery Time Actual Recovery Time Actual (RTA) is the critical metric for business continuity and disaster recovery. The business continuity group conducts timed rehearsals (or actuals), during which RTA gets determined and refined as needed.

Recovery Point Objective A Recovery Point Objective (RPO) is the maximum acceptable interval during which transactional data is lost from an IT service. For example, if RPO is measured in minutes, then in practice, off-site mirrored backups must be continuously maintained as a daily off-site backup will not suffice.

Relationship to RTO A recovery that is not instantaneous restores transactional data over some interval without incurring significant risks or losses. RPO measures the maximum time in which recent data might have been permanently lost and not a direct measure of loss quantity. For instance, if the BC plan is to restore up to the last available backup, then the RPO is the interval between such backups. RPO is not determined by the existing backup regime. Instead BIA determines RPO for each service. When off-site data is required, the period during which data might be lost may start when backups are prepared, not when the backups are secured off-site.

Mean times The recovery metrics can be converted to/used alongside failure metrics. Common measurements include mean time between failures (MTBF), mean time to first failure (MTFF), mean time to repair (MTTR), and mean down time (MDT).

Data synchronization points A data synchronization point is the point at which a backup is completed. It halts update processing while a disk-to-disk copy is completed. The backup copy reflects the earlier version of the copy operation; not when the data is copied to tape or transmitted elsewhere.

System design RTO and the RPO must be balanced, taking business risk into account, along with other system design criteria. RPO is tied to the times backups are secured offsite. Sending synchronous copies to an offsite mirror allows for most unforeseen events. The use of physical transportation for tapes (or other transportable media) is common. Recovery can be activated at a predetermined site. Shared offsite space and hardware complete the package. For high volumes of high-value transaction data, hardware can be split across multiple sites.

History Planning for disaster recovery and information technology (IT) developed in the mid to late 1970s as computer center managers began to recognize the dependence of their organizations on their computer systems. At that time, most systems were batch-oriented mainframes. An offsite mainframe could be loaded from backup tapes pending recovery of the primary site; downtime was relatively less critical. The disaster recovery industry developed to provide backup computer centers. Sungard Availability Services was one of the earliest such centers, located in Sri Lanka (1978). During the 1980s and 90s, computing grew exponentially, including internal corporate timesharing, online data entry and real-time processing. Availability of IT systems became more important. Regulatory agencies became involved; availability objectives of 2, 3, 4 or 5 nines (99.999%) were often mandated, and high-availability solutions for hot-site facilities were sought. IT service continuity became essential as part of Business Continuity Management (BCM) and Information Security Management (ISM) as specified in ISO/IEC 27001 and ISO 22301 respectively. The rise of cloud computing since 2010 created new opportunities for system resiliency. Service providers absorbed the responsibility for maintaining high service levels, including availability and reliability. They offered highly resilient network designs. Recovery as a Service (RaaS) is widely available and promoted by the Cloud Security Alliance.

Classification Disasters can be the result of three broad categories of threats and hazards.

… excerpt ends here. Continue reading the full article.

Illustrations

IT disaster recovery: A modular data center connected to the power grid at a utility substation
A modular data center connected to the power grid at a utility substation

Worked examples

Example 1 — a first encounter with IT disaster recovery

Start with the simplest possible case. Write down what IT disaster recovery claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In science, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to IT disaster recovery before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about IT disaster recovery ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of IT disaster recovery

In research
IT disaster recovery appears in science research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses IT disaster recovery in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
IT disaster recovery is common in secondary-school and first-year university syllabi. It links to neighbouring topics Backup, Business continuity, Data management, so understanding it makes those chapters shorter.
In everyday life
Look for IT disaster recovery outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.
Ask Teacher Smith questions about this articleOpens your AI tutor with a question about “IT disaster recovery” →

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study IT disaster recovery in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what IT disaster recovery means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain IT disaster recovery out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is IT disaster recovery in simple terms?

IT disaster recovery (also, simply disaster recovery (DR)) is the process of maintaining or reestablishing vital infrastructure and systems following a natural or human-induced disaster, such as a storm or battle. DR employs policies, tools, and procedures with a focus on IT systems supporting crit…

Why does IT disaster recovery matter?

Because it connects several science ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study IT disaster recovery?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on IT disaster recovery.

Tags

  • Backup
  • Business continuity
  • Data management
  • Disaster recovery
  • IT risk management

Keep exploring