Home / Devices / SAN

SAN Data Recovery Leicester

Shared storage has one unpleasant property: when it stops, everything above it stops in the same second, and the call that follows has a particular tone. Pool metadata that no longer parses, LUNs that refuse to come online, datastores full of virtual machines that will not mount. All of it is rebuilt here for companies and IT providers across Leicester and the East Midlands, from a single shelf up to a full rack, as ordinary weekly work rather than a favour squeezed in around something else.

Every SAN that arrives here is examined at no charge. The figure quoted afterwards is put in writing and settled before anybody reaches for a screwdriver: from £500 + VAT on a SAN, the figure tracking the disk count.

Logical recoveries carry no fix, no fee. The named exceptions to it are electronic and mechanical failures, chip-level work, DVR jobs and forensic jobs; physical work takes 50% up front. The five bands are set out in full on the data recovery cost page.

// thirty faults, roughly in order of how often they arrive

Thirty ways it goes wrong, and what sits behind each

Tracing a symptom back to the failure underneath it is where the real work begins, and these thirty account for all but a handful of the boxes opened at the Cambridge bench. A fault missing from the list is not an unfamiliar one — describe it on the telephone and you will get a straight view of the odds before you spend anything on postage.

A LUN offline, or announcing itself as corrupt

Shared storage fails in one place and everything sitting on top of it fails in the same second, which is why the phone call has a particular tone to it. A LUN that has gone offline is usually a pool metadata problem rather than a set of dead disks, and the disks underneath are frequently in good order.

Several drives lost out of one shelf at once

Backplane faults, a failed expander and power problems all take out multiple disks together, which looks like a catastrophic array failure and often is not. The disks themselves are commonly undamaged. Every one of them gets imaged and the pool is reconstructed from the copies.

Damaged pool or vDisk metadata

Enterprise arrays put a virtualisation layer between the physical disks and the LUNs presented to hosts, and that layer keeps its own metadata. Damage there makes healthy disks useless without touching anything written to them. Reconstructing that layer is the substance of most SAN work.

A VMFS datastore that will not mount

VMware's file system has its own structures and its own failure modes, and a datastore that ESXi refuses to mount is normally intact underneath. The virtual machine disks are extracted directly from an image of the LUN, and there is a page on this site given over to virtual machine work.

A controller left dead by a firmware update

Dual-controller arrays update one controller at a time, and an update that goes wrong on the second one can leave the whole shelf offline. The disks are fine. Nothing about the recovery needs a working controller, because the layout is derived from the members themselves.

A LUN deleted or re-provisioned by mistake

Provisioning tools make removing a LUN a small number of clicks and the confirmation is easy to click past. What that writes is metadata rather than user data, so a shelf powered down promptly is usually recoverable in full. Every hour of continued service costs something.

Thin provisioning that ran out of real capacity

Thin LUNs present more space than physically exists, and when the pool behind them fills, writes start failing in ways the host does not expect. File systems on top become inconsistent quickly. It is recoverable, and it is one of the more common causes of a sudden shared-storage failure.

A snapshot chain that broke

Snapshots on enterprise arrays reference blocks in earlier states, and a broken chain can make the current view unreadable while every block still exists somewhere. Reconstructing the chain from the metadata is the work, and it usually returns everything.

Deduplication or compression metadata damaged

Where an array stores one copy of a block referenced from many places, damage to the reference table takes out far more than its size suggests. Reconstructing those tables is the difference between a shelf of readable disks and a shelf of readable rubbish.

A tiering arrangement where the fast tier failed

Automatic tiering moves hot blocks onto flash and leaves cold ones on spinning disks, so a LUN is spread across both. Losing the flash tier removes a share of nearly every file rather than a set of whole files. Both tiers have to arrive, labelled.

iSCSI targets that vanished from every host

A shelf that has disappeared from the network is not necessarily a shelf with a data fault. Switch failures, address conflicts and a change to the network that nobody documented all present identically. Ruling that out before anything is unplugged costs nothing and occasionally ends the whole matter.

Fibre channel zoning changed and the LUNs gone

Zoning and masking changes remove a host's view of storage without touching the storage itself. It looks exactly like a failure from the host end. It is worth confirming the configuration before assuming the worst, and it is the cheapest possible outcome.

An expansion shelf that was disconnected while running

Pulling a shelf out of a live pool leaves the array with half its members and the pool inconsistent. Reconnecting it afterwards does not always put things back, because the array may have written in the interval. Both shelves have to travel and both need labelling.

A pool built on top of another pool

Layered arrangements, where a host array sits on LUNs presented by a second array, need unpicking one layer at a time. It is slower rather than harder, and it needs describing on the booking form because guessing at the layering wastes a day.

Bad sectors on a large proportion of the members

Enterprise disks that have run for years without a scrub carry unreadable regions scattered across the set. No single member is complete and the good ground on each one has to be combined. That is ordinary work here rather than an unusual case.

A rebuild running against a degraded pool

Rebuilding an enterprise pool reads every member at full rate, which is the workload that finishes off a second marginal disk. A rebuild that has stalled leaves the pool in a state neither the array nor its owner can describe accurately. Stop it and power everything down.

Virtual machines that will not start after the storage came back

The LUN mounts, the datastore is visible, and the virtual disks are corrupt. That is a separate problem from the storage failure and it has its own tools. VMDK, VHDX and QCOW2 files are repaired individually once the volume beneath them is readable.

A database that will not open after an outage

Databases held open and written to constantly are the first casualty of a storage failure, and the file that will not open is frequently repairable from its own logs. Older consistent copies also survive in unallocated space more often than people expect.

An array that has been out of support for years

Older shelves from vendors who have moved on, or been bought, turn up regularly. Nothing about the recovery depends on the vendor still existing or on a support contract being live, because the reconstruction works from the disks. Age is rarely the obstacle it appears to be.

Encryption at the array level with the key gone

Self-encrypting drives and array-level encryption both put a key between the disks and the data. Where the key can be recovered from the controller or from a key manager, the data comes back. Where it genuinely cannot, that is said immediately rather than after an invoice.

Ransomware that reached the shared storage

Hosts with mounted LUNs pass an infection straight through to the storage, and the array itself is usually physically healthy afterwards. Snapshots sometimes survive. It is priced as ordinary media, from 500 pounds + VAT, and it is not forensic work and is never billed as such.

A shelf that was moved between buildings

Disks that have run continuously for years frequently do not survive being stopped and restarted, and a relocation is the point at which several marginal members give up together. That is the restart exposing wear rather than the journey causing it.

Disks pulled and reseated in the wrong order

Some arrays record their position in metadata and cope. Others do not. Photograph the front of every shelf before touching anything and label each disk with its slot, because reconstructing the order from parity is possible but it is work nobody needs to pay for.

A backup appliance that failed with the backups on it

Purpose-built backup targets are arrays with a deduplication layer on top, and when one fails the organisation loses its storage and its safety net in the same event. Those are reconstructed here, and the deduplication tables are usually the substance of the job.

An array where the initial synchronisation never finished

A pool that was put into service before its first parity build completed carries parity that does not match the data. Everything works until a member fails, at which point the reconstruction produces nonsense. This is identified during the assessment rather than found halfway through.

SAS expanders that failed and dropped a whole enclosure

One expander failure can remove twelve or twenty-four disks from view at once, which reads on the console as a total loss. The disks are typically fine. This is among the more encouraging diagnoses on this page and it is a common one.

A shelf where the documentation left with the last IT manager

No configuration records, no diagram, and a set of disks in a rack. That is a routine arrival. The layout is derived from the disks themselves and nothing needs to be produced from a filing cabinet. It goes quicker with documentation and it does not depend on it.

A LUN presented to two hosts that both wrote to it

Shared access without a clustered file system leads to two machines writing the same blocks, and the damage is quick and thorough. Some of it is repairable from the file system structures. How much depends on how long both hosts were live.

An IT provider who has already attempted a recovery

Second opinions on enterprise storage arrive here regularly and a good share of them come good. Be candid about what has been tried, because it decides the safest order to work in. The assessment costs nothing whether it is the first attempt or the fourth.

A company that cannot trade until this is back

Not a fault, a priority. Say it on the call and the job moves to the front of the list. The honest timings do not change: the free assessment takes two working days from arrival, and reconstruction runs from there. Anybody promising a rack back by tomorrow is guessing at your expense.

What makes a SAN different from an array

Underneath, a SAN is a set of disks in RAID groups and it is reconstructed the same way any other array would be: every member imaged read-only, the geometry derived from the data, nothing written back. What sits above that is the difference. Enterprise arrays put a virtualisation layer between the physical groups and the LUNs presented to hosts, and that layer keeps its own metadata describing which physical extents make up which volume. Thin provisioning, automatic tiering between flash and spinning disks, deduplication and snapshot chains all live there. Damage to that layer makes a shelf of perfectly healthy disks present nothing usable, which is why so many of these jobs turn out to be metadata problems rather than disk problems. Reconstructing it is the substance of most SAN work. Above that again sits a file system, most often VMFS on a datastore full of virtual machines, and a datastore that ESXi refuses to mount is normally intact underneath and gives up its virtual disks readily once the LUN is back.

What to send, what to leave, and what it costs

Disks only, each labelled with the shelf and the slot it came from, and photograph the front of every enclosure before a single drive is pulled. A photograph settles in five seconds what parity analysis otherwise has to work out. The controllers, the chassis, the caddies where they are awkward and the rails all stay with you, because the layout comes from the members and no replacement hardware is needed at this end. Where a cache device or a flash tier was part of the arrangement, that has to travel too and it has to be labelled as such, since a LUN spread across two tiers is not readable from one of them. A SAN is priced as an array, from £500 + VAT and rising with the member count, with the free assessment closing two working days after the disks are booked in. Logical work is covered by no fix, no fee; members that failed mechanically or electronically are not and take half up front. If a rebuild is running, stop it, and if a management tool has offered to create a new pool, do not accept, because both are recoverable situations that become considerably harder once acted on.

// what stands on the bench

The equipment involved, and why any of it matters

Shared storage is reconstructed from read-only images of its members, exactly as any other array is. What differs is the number of layers above the disks, and most of the equipment below exists to unpick those layers one at a time without disturbing what is underneath.

Bulk imaging for whole shelves

Twelve, twenty-four or forty-eight disks imaged in parallel with retry limits enforced in hardware. A shelf that has to be read one disk at a time is a schedule problem rather than a technical one, and doing it in parallel is what keeps enterprise work to a sensible timescale.

SAS, SATA, fibre channel and NVMe interfaces

Every interface an enterprise shelf has used in the last twenty years, all of it write-blocked. Older SCSI and fibre channel disks are still read here, which matters because the shelves that fail are rarely the newest ones in the building.

Pool and virtualisation layer reconstruction

The metadata that maps LUNs onto physical extents, rebuilt from the images. This is the layer that distinguishes a SAN from a plain array, and damage to it is the commonest reason a shelf of healthy disks presents nothing usable.

VMFS, NTFS, XFS and ext4 above the LUN

Once the LUN is reassembled the file system on it is a separate piece of work. VMFS in particular has its own structures, and a datastore that ESXi refuses to mount is normally intact underneath and gives up its virtual disks readily.

Virtual disk extraction and repair

VMDK, VHDX and QCOW2 files pulled out of a recovered datastore and repaired individually where they span damaged ground. A virtual machine that will not start after the storage is back is a separate problem with its own tools.

Deduplication and snapshot chain rebuilding

Reference tables and snapshot chains reconstructed so that blocks stored once and referenced many times can be found again. On a deduplicating appliance this is not a detail, it is the whole job.

// badges that arrive in the post

Enterprise storage handled here

Dell EMC, Unity and VNXDell PowerVault and CompellentHPE MSA, Nimble and 3PARNetApp FAS and E-SeriesIBM Storwize and DS seriesFujitsu EternusInfortrend and PromiseSupermicro JBOD shelvesiSCSI and fibre channel targetsSoftware-defined and hyperconverged nodes

What arrives from a rack

A SAN is priced as an array, from 500 pounds + VAT and rising with the member count, and the assessment costs nothing and closes two working days after the disks are booked in. Enterprise work is ordinary weekly work on this bench rather than a favour fitted in around something else, and a firm that has stopped trading goes to the front of the list on request. Send the disks only, labelled with their slots, and photograph the front of each shelf first. The controllers, the chassis and the rails stay with you, because the layout is derived from the members and no replacement card is needed. Logical faults are covered by no fix, no fee. Members that failed mechanically or electronically are not, and those take 50% up front. If a rebuild is running, stop it. If a management tool has offered to create a new pool, do not accept. Both of those are recoverable situations that become considerably harder once they have been acted on.

// getting it ready for the post

Before you tape the box shut — take the drive out if it comes out

Disks only, and each one labelled with the shelf and slot it came out of. Photograph the front of every enclosure before a single drive is pulled, because a photograph settles in five seconds what parity analysis otherwise has to work out. Leave the controllers, the chassis, the caddies where they are awkward to remove, and the rails: none of them are needed here and a rack unit in a parcel is an expensive way of protecting a set of drives. Wrap each disk so it cannot knock against its neighbours and use a box that holds its shape under weight. Send it tracked and insured to Cambridge Data Recovery, Compass House, Vision Park, Chivers Way, Cambridge CB24 9AD. Leicester to the lab is about seventy miles, the M1 south to Junction 19 then the A14 east, roughly an hour and a half, and the building sits two minutes off the A14 at Junction 32 with parking outside the door. Reception takes deliveries in person Monday to Friday, 9:00am to 5:30pm, and a courier you book yourself is equally welcome. Nothing is collected. Ring 0800 689 0668 before you pack it if the arrangement is unusual, and say on the call if the business is off the air.

// how the media reaches Cambridge

Sending a device — and the three exceptions

The post office does most of the work of getting a job here. A drive that is already unwell travels better boxed and insured than rattling around a car for a day of errands, and something dropped into a Leicestershire postbox this afternoon is generally logged in at Cambridge tomorrow.

The general rule is the drive travels and the machine stays behind — out of the laptop, out of the tower, out of the iMac, out of the recorder under the counter. This bench does not dismantle equipment, and a repair shop will do it while you wait. Three things are the other way round, and getting them wrong costs you the recovery: an external drive stays sealed in its own case, a NAS comes as a complete unit, and a WD My Passport or My Book travels whole with its cable, because on those the encryption key is held on the bridge board rather than on the disk — separate the two and the data becomes unreadable even to us. A Fusion Mac needs both of its drives, each labelled. The one thing nobody can work round is flash soldered onto the mainboard, as on Apple Silicon machines: if it will not come off, there is nothing to post.

  • A stiff box or a well-padded mailer, with enough packing that nothing moves when you shake it. Power supplies, docks and cables can stay at home unless the drive is one of the WD units above.
  • Running a RAID or a server? Send the member disks on their own, not the chassis or the controller, and write the bay order on each one — 1, 2, 3 and so on. Photograph the front of the unit before you pull anything, because that photograph occasionally saves a day of work.
  • Fill in the shipping and booking-in form (PDF) — a name, a number you actually answer, and a line on how the trouble started — and put it in the box.
  • Special Delivery is tracked and insured and is what most people use; your own courier is equally fine. Handing it over in person also works: reception at the Cambridge address takes devices across the counter, Mon–Fri 9:00am–5:30pm. What does not exist is a Leicester counter or anyone who comes to collect.
// write this on the label

Cambridge Data Recovery

Compass House
Vision Park, Chivers Way
Cambridge, CB24 9AD

↓ Print the shipping & booking-in form (PDF)

Address it to Cambridge Data Recovery. It is about seventy miles from Leicester if you fancy driving it — M1 south to Junction 19, then the A14 east — and the lab is two minutes off Junction 32 with parking at the door. Posting costs you a stamp and a day instead. Whichever you choose, you hear from us the moment it is booked in, and the free diagnostic closes two working days after that.

Not certain what belongs in the box? Ring 0800 689 0668 before you tape it up, or let the free online diagnostic ask the questions for you.

// SAN recovery questions

Common questions

From £500 + VAT, rising with the member count, exactly as any other array. The free assessment closes two working days after the disks arrive and produces one fixed figure in writing before work starts. Logical faults are covered by no fix, no fee. Members that failed mechanically or electronically fall outside it and take half as a deposit first, because the parts are bought for your particular disks.
Usually not. A datastore that ESXi refuses to mount is normally intact underneath, and the virtual disks are extracted directly from an image of the LUN rather than through the hypervisor. Where a VMDK spans damaged ground it is repaired as a separate piece of work afterwards. There is a page on this site given over to virtual machine recovery in more detail.
Usually a backplane fault, a failed expander or a power problem rather than the disks themselves, and one expander failure can remove twelve or twenty-four drives from view in a single moment. It reads on the console as a total loss and it very often is not. The disks are typically undamaged and the pool is reassembled from images of them.
No. Nothing about the reconstruction depends on a live support contract or on the vendor still existing, because the work is done from the disks. Older shelves running fibre channel or SAS turn up here regularly and the interfaces for them are on the shelf. Age is rarely the obstacle it appears to be from the outside.
// the rest of the week's work

What else lands on this bench

// where to read further

Pages that take it further

Whenever you are ready, the bench is.

The examination is free, one written figure follows it, and the band covering this page is from £500 + VAT on a SAN, the figure tracking the disk count.