Call us — 0191 406 1051
Mon–Fri · 9am–5:30pm · No fix, no fee
Start a free diagnostic →

Data Recovery Case File · NAS & Network Storage · The Arithmetic of Large Arrays

Why This Configuration Fails Exactly This Way

This enquiry came from an IT provider on behalf of a client, and it describes a failure that is close to predictable. An eight-bay network appliance, "eight 16TB disks in RAID 5. Two disks are reporting to have died and the data volume is now inaccessible." That is a hundred and twenty-eight terabytes of raw capacity protected by exactly one disk's worth of redundancy — because parity of that kind tolerates one failure regardless of how many members there are. Eight disks do not buy more tolerance than four. And on disks this large, the recovery step is itself the most dangerous thing the array ever does.

MediaEight-bay network storage appliance — eight 16TB disks in single-parity configuration; two members reported failed and the data volume inaccessible
Reported situationEight-disk single-parity array in service · two members reporting failure · data volume inaccessible · enquiry raised by an IT provider on behalf of a client
Fault classDual-member failure in a single-parity set — reconstruction dependent on failure sequence and on returning one member to a readable state
Equipment usedNo rebuild attempted · all members removed, labelled by bay and imaged individually write-blocked (Atola TaskForce 2) · failure sequence established from member metadata · set assembled offline from images with stale members excluded

The decode: why two failed, and the sequence that decides the outcome

Why member count does not increase protection: single-parity distributes one disk's worth of redundancy across the whole set, so any one member can be reconstructed from the others. That is true whether the set has four members or eight. What changes with more members is the likelihood of a second failure — more disks means more opportunities, and disks bought together, run together and aged together tend to reach the end of their lives together.

Why large disks make this worse, and this is the doctrine worth carrying: when one member fails, the array reconstructs it onto a replacement by reading every sector of every remaining member. On 16TB disks that is many hours and frequently days of continuous, maximum-intensity reading applied to seven disks of identical age and batch. It is by a wide margin the heaviest load the array will ever experience, and it arrives at the exact moment the set has no redundancy left. So a second failure during or shortly before a rebuild is not bad luck — it is the single most likely moment for it to happen, and it is why single parity on large modern disks so often produces precisely this enquiry.

What that means for what actually occurred: "two disks died" frequently turns out to be one disk failing some time ago, unnoticed or unacted upon, followed by a second later. Which is why the sequence matters more than the count.

Why the sequence is the whole case: a member that dropped out earlier holds data as it was at the moment it left — it is stale, and every write since has passed it by. Including a stale member in a reconstruction produces a volume that mounts and contains a mixture of current and outdated stripes, which is worse than no recovery at all because it looks successful. Array metadata carries event counters and timestamps on each member, so the order of departure can be established rather than guessed, and the more recently failed disk is the one to bring back.

Why no rebuild must be attempted: if the appliance can be persuaded to bring the volume up in a degraded state, it will offer to rebuild — repeating the exact load that caused the second failure, on six remaining disks of the same age, while writing. Nothing should be rebuilt on the original disks.

The recommendation for afterwards, said once: at this capacity, single parity is not an appropriate protection level. Dual parity survives two failures and covers the rebuild window, and neither replaces a backup — an array protects against disk failure, not against deletion, corruption or the appliance itself failing.

On the bench

No rebuild was attempted, since a rebuild repeats the exact load that caused the second failure while writing to the set. All eight members were removed, labelled by bay and imaged individually write-blocked on the Atola TaskForce 2. The failure sequence was established from member metadata — event counters and timestamps identifying which disk left the set first — since a stale member reintroduced into a reconstruction produces a volume that mounts and silently mixes current data with outdated stripes. The set was assembled offline from the images with stale members excluded.

The outcome

All members imaged individually, the failure sequence established from metadata rather than assumed, and the set assembled offline with stale members excluded. Free assessment, one fixed written figure including VAT; where a drive has to be opened, 50% of parts and labour is payable upfront with the balance only on success — otherwise no recovery, no fee. The decode, for anyone running large single-parity arrays: member count never bought more tolerance, since one disk's worth of redundancy covers one failure whether the set has four members or eight. Rebuilding after a failure reads every sector of every remaining disk for hours or days, at maximum load, with no redundancy left — which is why the second failure so often arrives then. Establish which disk failed first, and never rebuild.

Large array that has lost two disks

Don't let it rebuild, and find out which disk failed first. Single-parity protection covers exactly one failure no matter how many disks are in the set, so eight members never gave you more tolerance than four — just more opportunities to fail, on disks bought and aged together. And the rebuild itself is the dangerous part: reconstructing one failed member means reading every sector of every remaining disk, which on large modern drives takes hours or days of maximum-intensity reading at the precise moment the set has no redundancy left. That's why second failures cluster there. The sequence matters enormously, because a disk that dropped out earlier holds stale data, and including it produces a volume that mounts and silently mixes current files with outdated ones. Power down, label every disk by bay, and have them imaged individually.

Array down with two failed members?
Don't rebuild — call Newcastle Data Recovery on 0191 406 1051; every member imaged write-blocked, failure sequence established from metadata, set assembled offline with stale members excluded.
Request a quote online →

Our case files are drawn from genuine enquiries received by our laboratory over the past ten years, anonymised to protect client confidentiality. Each one describes the diagnostic and recovery procedure our engineers apply to that fault, using the equipment listed.