The previous post covered replacing a disk in a mirrored rootvg, which is the hard case: boot image, boot list, dump device, and a long stretch where the machine runs on one disk. This is the case you will hit far more often — a failing disk in an ordinary mirrored volume group — and there is a single command for it that a surprising number of AIX administrators have never used.
replacepv
replacepv hdisk2 hdisk10
That is the whole thing. It moves every physical partition off the failing disk onto the replacement, removes the old disk from the volume group, and leaves the mirror intact the entire time. No unmirrorvg, no reducevg, no window where your data has one copy.
It is not a general-purpose tool and that is why people forget it exists. Six conditions have to hold, and if any of them fails you are back to the long procedure.
The first one is the killer. replacepv does not work on rootvg, full stop, which is exactly why the previous post is as long as it is.
Why it is worth checking those conditions
The difference between the two methods is not the number of commands. It is how long your data spends unprotected.
With unmirrorvg you deliberately throw away one copy, run without redundancy through the physical swap and the resync, and hope nothing happens to the surviving disk in the meantime. On a large volume group that resync can run for hours. With replacepv there is never a moment when a partition exists in only one place.
That is the entire argument. Spend two minutes checking whether you can use it.
The procedure
If the replacement disk is not yet visible to the system, present it and let AIX find it:
cfgmgr
lspv
lspv tells you what name the new disk got. Do not assume — on a busy system the number is whatever was free.
hdisk2 00c3e5f4a1b2c3d4 yourvg active
hdisk3 00c3e5f4a1b2c3d5 yourvg active
hdisk10 none None
Then run the replacement:
replacepv hdisk2 hdisk10
When it finishes, hdisk2 no longer belongs to a volume group. Remove its definition and pull it:
rmdev -dl hdisk2
Note the serial number and location code before the disk leaves, the same as in the rootvg procedure:
lscfg -vl hdisk2
Do that before rmdev, obviously. Once the definition is gone so is the output.
When replacepv is interrupted
This is the part worth knowing before you need it. replacepv writes a recovery directory under /tmp while it works. If the command dies partway — the session drops, someone closes the terminal, the machine is rebooted — you do not start over. You resume:
replacepv -R /tmp/replacepv<something>
The exact directory name is printed when the command starts and again when it fails. Read the error, do not just re-run replacepv from scratch, and do not clear /tmp while a replacement is in flight. I would put that last one on a sticky note.
When you cannot use it
If any condition fails, the sequence is the familiar one:
unmirrorvg yourvg hdisk2
reducevg yourvg hdisk2
rmdev -dl hdisk2
Swap the disk, then:
cfgmgr
extendvg yourvg hdisk2
mirrorvg yourvg hdisk2
syncvg -v yourvg
There is one error here that stops people cold:
0516-050 Not enough descriptor space left in this volume group.
The volume group has no room in its descriptor area for another physical volume. Three ways out, in order of how much they cost you. Add a smaller disk, if the geometry allows it. Mirror onto a disk already in the volume group instead of adding a new one. Or convert the volume group to a format with more descriptor space using chvg — -B for Big, -G for Scalable.
Check the manual page for your AIX level before you plan that last one. Converting to Big typically needs free partitions on every disk in the group so the expanded VGDA has somewhere to live, and converting to Scalable needs the volume group varied off, which turns a disk swap into an outage. Neither is something to discover at two in the morning.
If it turns out to be rootvg after all
Then replacepv is out and there are three extra things the previous post covers in detail: moving the boot image and boot list before you start, dealing with the dedicated dump device, and one more that catches people out.
Removing the boot disk's device definition also removes the /dev/ipldevice hard link. Recreate it against the surviving disk:
ln /dev/rhdisk1 /dev/ipldevice
ls -i /dev/rhdisk1 /dev/ipldevice
Note the r prefix — it links against the raw device, not the block device. The ls -i confirms it: both entries must show the same i-node number.
There is also a case where mirrorvg will not do the job on rootvg. If the machine is an LPAR, a copy of hd5 was on the failed disk, and the replacement disk's adapter was configured into the LPAR dynamically since the last cold boot, then you have to rebuild copies with mklvcopy — starting with the boot logical volume so it lands on contiguous partitions — and the machine has to be shut down and activated rather than rebooted before it will boot from the new disk. That combination is rare and extremely annoying when you meet it.
Verify
Whichever path you took, four checks. They take a minute between them.
If lspv shows stale partitions after the sync has finished, run syncvg -v yourvg once more. If they come back, the replacement disk is bad and you are about to do this whole thing again — better to find out now than after you have signed the change off.
Comments
Post a Comment