Backup and continuity

Can your backups actually be restored?

A backup can run successfully every night for months and still fail when the business actually needs it.

The problem is not always that data was never copied. The missing piece may be an unavailable encryption key, an inaccessible administrator account, a virtual machine that will not boot, an undocumented network dependency or a Microsoft 365 backup that does not contain the expected item.

The useful question is not only: “Did the backup complete?”

The useful question is: “What can we restore, in what order, with which access, and within what realistic timeframe?”

A successful backup and a successful recovery are not the same thing

A backup dashboard normally confirms that an agent copied data to a destination without reporting a critical error.

That matters, but it is incomplete.

Actual recovery may also require the team to:

  • locate the correct recovery point;
  • confirm that the copy is not corrupt;
  • retrieve the necessary credentials and encryption keys;
  • rebuild or start a target system;
  • reconnect storage, networking and identity services;
  • verify that applications can open their data;
  • restore users in a practical sequence;
  • confirm that restored systems do not reintroduce malware.

The required outcome is not a backup file. It is a usable business service.

1. The recovery point really exists

The test should select a specific date and workload, then confirm that the expected data is present.

For a file server, that may mean locating several folders and versions. For Microsoft 365, it may involve restoring an email, a OneDrive file or a SharePoint item. For a virtual machine, it may require a controlled boot in an isolated environment.

2. Required access is available

A backup may be intact but unusable when nobody has the necessary:

  • console credentials;
  • encryption keys;
  • cloud accounts;
  • storage access;
  • licence information;
  • server or hypervisor administrator passwords.

Those dependencies must be documented, protected and tested before the incident.

3. The copy is sufficiently independent

A backup accessible through the same accounts an attacker could compromise may be deleted or encrypted with the production environment.

The design should consider:

  • separation of administrative identities;
  • deletion permissions;
  • immutability where appropriate;
  • copies in more than one location;
  • MFA for backup consoles;
  • enough retention to return to a point before a late-discovered incident.

4. The restored system starts and works

Recovering files is only part of the job. A more complete test may need to validate:

  • operating-system startup;
  • Windows or Linux services;
  • databases;
  • line-of-business applications;
  • permissions;
  • scheduled tasks;
  • certificates;
  • communication with other servers;
  • functions that users actually need.

A machine that boots but cannot authenticate users is not a complete recovery.

5. The timeline matches the business requirement

A system recoverable in four days may be technically protected but operationally useless to an organization that can tolerate only four hours of interruption.

The test should measure:

  • time to locate the recovery point;
  • data-transfer time;
  • boot or rebuild time;
  • validation time;
  • time to return users to work.

That measured result must be compared with what the organization can realistically tolerate.

Which systems should be tested first?

Restoring the entire environment every month may not be practical. Start with the dependencies that determine whether the rest of the business can function.

For many small and midsize organizations, priorities include:

1. identity and administrative access; 2. internet, firewall and VPN connectivity; 3. file servers or core business applications; 4. essential virtual machines; 5. Microsoft 365 data; 6. documentation, passwords and licences; 7. workstations required for critical functions.

The order is organization-specific. A clinic, property operation, workshop and professional office do not have identical dependencies.

Microsoft 365 needs its own verification

Exchange Online, OneDrive, SharePoint and Teams provide retention and recovery mechanisms, but those controls do not automatically cover every loss, administrative error, deletion or long-term retention scenario.

A useful test should establish:

  • which workloads have an independent backup;
  • how long the data is retained;
  • who can initiate a restore;
  • whether an item can be restored to its original or an alternate location;
  • how former-employee data is handled;
  • how shared files and permissions are validated.

A backup and recovery service should be built around the organization’s actual data and recovery requirements, not around a product list.

Signs that a backup plan needs review

Review is justified when:

  • no recent restore is documented;
  • only one person understands the console;
  • production and backup use the same administrative identities;
  • passwords or keys are hard to find;
  • servers, VMs and Microsoft 365 are protected by disconnected systems;
  • nobody knows the real recovery time;
  • a provider or employee change left ownership unclear;
  • every copy is in one location;
  • data volume has grown substantially;
  • a new critical application was added without updating the recovery plan.

What to avoid during an incident

When a failure or attack occurs:

  • do not delete useful logs;
  • do not reset every account without preserving access to backup systems;
  • do not restore over the only available copy;
  • do not reconnect a potentially compromised system without validation;
  • do not assume the newest recovery point is the healthiest one;
  • do not run several undocumented repair tools against failing storage.

Rushed work can turn a recoverable incident into a larger loss. When a disk, RAID, NAS or image becomes unreadable, also review our data-recovery approach.

A practical validation cycle

An organization can use a simple cycle:

  • daily review of jobs and alerts;
  • monthly review of failures, capacity and retention;
  • regular file or Microsoft 365 item restoration;
  • quarterly or semiannual recovery of a representative workload;
  • a broader annual recovery exercise;
  • a new test after a migration, storage change or major redesign.

Frequency should reflect risk and how quickly the environment changes.

The expected result

After a useful test, the organization should be able to state clearly:

  • what data is protected;
  • where the copies are stored;
  • who can access them;
  • which recovery point was verified;
  • how long recovery took;
  • which dependencies failed;
  • which corrections are required;
  • when the next test will occur.

Review your recovery readiness

A backup and recovery readiness assessment can examine an agreed scope, approved samples and the restores actually tested. Conclusions remain limited to the systems, data and exercises examined.

UNITECH can own protection, monitoring, documentation and restore testing as part of a managed or co-managed IT service.

Discuss your recovery readiness

You do not need to wait for a failure to discover the limits of the backup design.

Bring a list of important systems, interruption constraints and what you already know about the existing copies. We can separate urgent risks from improvements that can be planned properly.

Contact Montreal IT