Skip to main content

Replacing a failed disk in a mirrored volume group with replacepv

The previous post covered replacing a disk in a mirrored rootvg, which is the hard case: boot image, boot list, dump device, and a long stretch where the machine runs on one disk. This is the case you will hit far more often — a failing disk in an ordinary mirrored volume group — and there is a single command for it that a surprising number of AIX administrators have never used.

replacepv

replacepv hdisk2 hdisk10

That is the whole thing. It moves every physical partition off the failing disk onto the replacement, removes the old disk from the volume group, and leaves the mirror intact the entire time. No unmirrorvg, no reducevg, no window where your data has one copy.

It is not a general-purpose tool and that is why people forget it exists. Six conditions have to hold, and if any of them fails you are back to the long procedure.



The first one is the killer. replacepv does not work on rootvg, full stop, which is exactly why the previous post is as long as it is.

Why it is worth checking those conditions

The difference between the two methods is not the number of commands. It is how long your data spends unprotected.


With unmirrorvg you deliberately throw away one copy, run without redundancy through the physical swap and the resync, and hope nothing happens to the surviving disk in the meantime. On a large volume group that resync can run for hours. With replacepv there is never a moment when a partition exists in only one place.

That is the entire argument. Spend two minutes checking whether you can use it.

The procedure

If the replacement disk is not yet visible to the system, present it and let AIX find it:

cfgmgr
lspv

lspv tells you what name the new disk got. Do not assume — on a busy system the number is whatever was free.

hdisk2   00c3e5f4a1b2c3d4   yourvg   active
hdisk3   00c3e5f4a1b2c3d5   yourvg   active
hdisk10  none               None

Then run the replacement:

replacepv hdisk2 hdisk10

When it finishes, hdisk2 no longer belongs to a volume group. Remove its definition and pull it:

rmdev -dl hdisk2

Note the serial number and location code before the disk leaves, the same as in the rootvg procedure:

lscfg -vl hdisk2

Do that before rmdev, obviously. Once the definition is gone so is the output.

When replacepv is interrupted

This is the part worth knowing before you need it. replacepv writes a recovery directory under /tmp while it works. If the command dies partway — the session drops, someone closes the terminal, the machine is rebooted — you do not start over. You resume:

replacepv -R /tmp/replacepv<something>

The exact directory name is printed when the command starts and again when it fails. Read the error, do not just re-run replacepv from scratch, and do not clear /tmp while a replacement is in flight. I would put that last one on a sticky note.

When you cannot use it

If any condition fails, the sequence is the familiar one:

unmirrorvg yourvg hdisk2
reducevg yourvg hdisk2
rmdev -dl hdisk2

Swap the disk, then:

cfgmgr
extendvg yourvg hdisk2
mirrorvg yourvg hdisk2
syncvg -v yourvg

There is one error here that stops people cold:

0516-050 Not enough descriptor space left in this volume group.

The volume group has no room in its descriptor area for another physical volume. Three ways out, in order of how much they cost you. Add a smaller disk, if the geometry allows it. Mirror onto a disk already in the volume group instead of adding a new one. Or convert the volume group to a format with more descriptor space using chvg — -B for Big, -G for Scalable.

Check the manual page for your AIX level before you plan that last one. Converting to Big typically needs free partitions on every disk in the group so the expanded VGDA has somewhere to live, and converting to Scalable needs the volume group varied off, which turns a disk swap into an outage. Neither is something to discover at two in the morning.

If it turns out to be rootvg after all

Then replacepv is out and there are three extra things the previous post covers in detail: moving the boot image and boot list before you start, dealing with the dedicated dump device, and one more that catches people out.

Removing the boot disk's device definition also removes the /dev/ipldevice hard link. Recreate it against the surviving disk:

ln /dev/rhdisk1 /dev/ipldevice
ls -i /dev/rhdisk1 /dev/ipldevice

Note the r prefix — it links against the raw device, not the block device. The ls -i confirms it: both entries must show the same i-node number.

There is also a case where mirrorvg will not do the job on rootvg. If the machine is an LPAR, a copy of hd5 was on the failed disk, and the replacement disk's adapter was configured into the LPAR dynamically since the last cold boot, then you have to rebuild copies with mklvcopy — starting with the boot logical volume so it lands on contiguous partitions — and the machine has to be shut down and activated rather than rebooted before it will boot from the new disk. That combination is rare and extremely annoying when you meet it.

Verify

Whichever path you took, four checks. They take a minute between them.



If lspv shows stale partitions after the sync has finished, run syncvg -v yourvg once more. If they come back, the replacement disk is bad and you are about to do this whole thing again — better to find out now than after you have signed the change off.

Comments

Popular posts from this blog

Diskless server iSCSI boot disk XCP-NG installation

 For Cisco UCS chassis, eg. 5108 for installation it's required to set FC or iSCSI disk for boot ( boot disk ). They are called diskless servers. Unfrtunatelly XCP-NG - 8.2.1 Jan 2025 - in my case cannot boot as it is from iSCSI. Here is a guide, how to do it.

AI's Democratic Era Ended in April 2026

For about three years, every big AI lab behaved like a teenager with a new car. Look what mine can do. Look at my benchmark. Look at my context window. Every few weeks somebody published a chart where their bar was taller than everybody else's bar, and we all went along with it, because the bars kept getting taller and the models kept getting better and it was fun to watch. Meanwhile China said almost nothing. Then once in a while DeepSeek dropped something that knocked everyone off their feet, everyone spent a week arguing about training costs, and the noise started again. That era is over. It ended in April 2026, and I don't think we noticed how completely. The race we thought we were in The mental model everybody used was Formula 1. Whoever builds the fastest car wins the season. Speed was the whole point, and speed was something you showed off, because showing off is how you raise the next round. Some people preferred the Manhattan Project comparison. I used to think that o...

Replacing a failing disk in AIX rootvg

A disk in a mirrored rootvg is the one disk you cannot just pull. It holds the boot image, it probably holds the primary dump device, and half of every logical volume on the system lives on it. Pull it without preparation and reducevg refuses, or worse, the machine comes up on nothing after the next reboot. This is the full sequence for a two-disk mirrored rootvg on AIX, replacing hdisk0 . Non-root volume groups are much simpler and I have noted the difference at the end. Decide whether you actually have a problem Disks rarely die cleanly. They complain first, and the complaints show up in the error log: errpt | more 1581762B 0727203502 T H hdisk0 DISK OPERATION ERROR 1581762B 0727203502 P H hdisk0 DISK OPERATION ERROR The third column is the one that matters. T is temporary, P is permanent. A single temporary error on a busy system is noise. A cluster of them over a few hours is a drive on its way out. A permanent error is not a warning, it is a report of ...