Repairing a Degraded RAID Array After a Drive Failure
A degraded RAID array means that one or more drives have failed or become unavailable, while the remaining disks are still presenting the volume and its data. The system may continue operating, but its redundancy has been reduced. A second failure during this period can make recovery much harder, particularly with RAID 5 or older disks that have similar wear.
The safest repair process is deliberate rather than hurried. Confirm which drive failed, protect the current data, check the RAID layout, and replace the disk with a suitable model before starting a rebuild. The menus differ between Synology DSM, QNAP QTS and QuTS hero, but the principles are similar for most NAS systems.
Confirm the Failure Before Removing Anything
Start in the NAS administration interface and record the storage pool, RAID level, volume name and failed disk bay. A degraded status does not always mean that the physical drive is dead. A loose tray, damaged SATA connection, overheating event or power interruption can cause a temporary disconnect. Review system logs and check whether the disk is marked failed, disconnected, critical or simply unavailable.
Avoid clicking “repair”, “initialise” or “create new volume” until the situation is clear. These commands can alter metadata or erase a disk that still contains useful information. If the array is accessible, copy important files to a separate backup first. Prioritise irreplaceable documents, photographs, business records and configuration exports rather than attempting to copy the entire NAS immediately.
Run a health check on the surviving drives using SMART data, the NAS vendor’s disk tool or a desktop utility. Look for reallocated sectors, pending sectors, excessive temperatures and large increases in error counts. A RAID 5 array with one failed disk and another disk showing warning signs should be treated as a high-risk recovery case, not as a routine drive swap.
Choose and Prepare the Replacement Drive
The replacement disk should normally be at least as large as the smallest member of the array. Capacity printed on a retail box can differ slightly between models, so selecting a drive with the same nominal size is not always enough. A 12 TB replacement may be rejected if its usable sector count is marginally smaller than the existing 12 TB members.
Use NAS-rated hard drives or enterprise models that suit the workload, vibration level and warranty requirements. Check the compatibility list for the Synology or QNAP model, but do not assume that every listed disk is currently easy to buy in Australia. Stock can vary between Sydney, Melbourne, Brisbane, Adelaide and Perth, and a replacement ordered from overseas may involve delays, freight damage or warranty complications.
Before installation, inspect the replacement disk with a full surface test when time allows. For a new disk, this can reveal early defects before it is added to the array. A secure erase or vendor preclear procedure may be appropriate, although the exact method depends on the NAS and RAID technology. Never preclear a surviving array member by mistake; verify the serial number and bay location several times.
For a NAS that uses a hot spare, confirm whether it is healthy and assigned to the correct storage pool. Automatic rebuilding can begin as soon as a failed disk is removed, so make sure the enclosure is stable and the correct drive has been identified before inserting or activating a replacement.
Select the Correct Recovery Path
The RAID level determines how much redundancy remains and what repair options are safe. RAID 1 can usually mirror data back to a new disk, while RAID 10 may tolerate a failure only if it occurs in a different mirror pair. RAID 6 can generally survive two failed members, but a degraded array still deserves urgent attention. RAID 5 has a single-drive fault tolerance and becomes particularly vulnerable during rebuild.
| RAID configuration | Typical single-drive tolerance | Usual repair approach | Main risk during rebuild |
|---|---|---|---|
| RAID 1 | One drive in a mirror | Replace the failed member and resynchronise | A second mirror failure |
| RAID 5 | One drive | Replace the disk and rebuild parity | Another unreadable sector or disk failure |
| RAID 6 | Two drives | Replace failed members one at a time | Long rebuild and additional hardware faults |
| RAID 10 | Usually one drive per mirror pair | Replace the failed member in its mirror | Failure of the other disk in the same pair |
| SHR or similar hybrid RAID | Depends on disk arrangement | Follow the NAS storage manager’s repair process | Misreading the underlying redundancy |
In Synology DSM, the Storage Manager generally provides a “Repair” or “Replace” action for a degraded storage pool. QNAP systems commonly expose a similar operation under Storage & Snapshots. Select the failed member or newly installed disk only after checking its serial number. If the interface offers both “rebuild” and “create new pool”, the rebuild or repair option is the relevant one.
Do not use a disk from another array without understanding its metadata, and do not force a drive online merely to make the warning disappear. If the NAS reports multiple failed members, repeated I/O errors or an inaccessible pool, stop and consider a sector-level clone or professional RAID recovery service. Repeated rebooting and experimental commands can reduce the chance of a clean recovery.
Control the Rebuild in Australian Conditions
A RAID rebuild can run for many hours or several days. During that time, reduce non-essential activity such as media transcoding, large downloads, virtual machines and cloud synchronisation. A quiet rebuild puts less stress on the remaining disks and usually completes sooner than a busy one. Keep the NAS connected to a reliable UPS, especially in areas affected by summer storms, bushfire-related outages or unstable local power.
Temperature management matters in Australian homes and offices. A NAS in a warm study in Perth or Adelaide may need more airflow during summer, while Brisbane and coastal environments can add humidity and dust concerns. Clean filters and vents, keep the enclosure away from direct sunlight, and check drive temperatures throughout the rebuild. Do not open the chassis unnecessarily while disks are operating.
Use these practical safeguards while the array is recovering:
- Pause intensive services, indexing and non-essential backup jobs.
- Keep at least one independent backup available before making changes.
- Monitor disk temperature, rebuild percentage and system alerts.
- Avoid moving the NAS or disconnecting network and power cables.
- Record the replacement disk’s model, serial number and installation time.
- Check whether the NAS has a current firmware and configuration backup.
A UPS with USB or network signalling can tell the NAS to shut down cleanly during an outage. This is especially useful for small businesses in regional New South Wales or Victoria, where a brief power interruption can otherwise turn a controlled rebuild into another degraded state. When buying locally, compare Australian warranty terms and GST-inclusive pricing rather than choosing solely on the lowest advertised drive price.
Verify the Array and Protect the Recovered Data
A successful rebuild does not prove that every file is healthy. Once the pool returns to a healthy state, review the NAS logs for unreadable sectors, checksum errors, bad blocks and resynchronisation warnings. Run the vendor’s scrub or data-scrubbing feature if supported, but schedule it after the rebuild and during a low-use period. A scrub reads data and parity to identify silent corruption that a normal status screen may not reveal.
Test representative files from every important share. Open office documents, play several videos, inspect recent photographs and restore a sample backup to another computer. If the NAS stores records for a property, workshop or hospitality operation, preserve configuration exports and documentation alongside the data. Keep unrelated archived references, such as an archived project site, separate from the recovery workflow and never treat an external website as a substitute for a verified backup.
RAID protects availability, not against accidental deletion, malware, theft, fire or file corruption. Apply a 3-2-1 backup approach: maintain three copies of important data, use at least two different storage types, and keep one copy off-site. An encrypted USB backup, a second NAS at another location or reputable cloud storage can provide the independent copy that a RAID array cannot.
After recovery, examine why the failure occurred. Replace a group of aging disks proactively when their error rates or operating hours indicate elevated risk, rather than waiting for sequential failures. Australian Consumer Law may support remedies for faulty products, but warranty handling still takes time, so retaining a tested spare drive can be worthwhile for a business-critical NAS.
The repair process is complete only when the pool is healthy, critical files have been checked and independent backups are current. Treat every degraded-array alert as a prompt to verify the entire storage strategy, not merely as a request to insert another disk. With careful identification, suitable hardware and controlled rebuilding, many single-drive failures can be resolved without losing the data the NAS was built to protect.