Skip to main content

Replacing a failing disk in AIX rootvg

A disk in a mirrored rootvg is the one disk you cannot just pull. It holds the boot image, it probably holds the primary dump device, and half of every logical volume on the system lives on it. Pull it without preparation and reducevg refuses, or worse, the machine comes up on nothing after the next reboot.

This is the full sequence for a two-disk mirrored rootvg on AIX, replacing hdisk0. Non-root volume groups are much simpler and I have noted the difference at the end.

Decide whether you actually have a problem

Disks rarely die cleanly. They complain first, and the complaints show up in the error log:

errpt | more
1581762B   0727203502 T H hdisk0         DISK OPERATION ERROR
1581762B   0727203502 P H hdisk0         DISK OPERATION ERROR

The third column is the one that matters. T is temporary, P is permanent. A single temporary error on a busy system is noise. A cluster of them over a few hours is a drive on its way out. A permanent error is not a warning, it is a report of something that already failed.



Then check whether the mirror is still intact:

lsvg -l rootvg
rootvg:
LV NAME    TYPE    LPs   PPs   PVs  LV STATE        MOUNT POINT
hd5        boot    1     2     2    closed/syncd    N/A
hd6        paging  64    128   2    open/syncd      N/A
hd8        jfslog  1     2     2    open/stale      N/A
hd4        jfs     4     8     2    open/stale      /

stale means writes are no longer reaching both copies. At that point you are running on one disk whether you like it or not, and the mirror is giving you a false sense of protection. Fix it before you do anything else that week.

What the whole job looks like

Five stages. The middle three are the uncomfortable ones, because the machine is running unprotected on a single disk the entire time.



Plan for that. Do not start this at four in the afternoon on a Friday, and do not start it while the surviving disk is also logging errors.

Step 1: move everything that pins the disk

Three things make hdisk0 unremovable, and all three have to go first.



Write a boot image to the surviving disk and change the boot list so firmware prefers it:

bosboot -a -d hdisk1
bootlist -m normal hdisk1 hdisk0
bootlist -m normal -o

That last command reads the list back. Do it. A typo here is discovered at the worst possible moment.

Now the dump device. On a default install the primary dump lives in pdumplv on hdisk0, and you cannot remove a disk that holds it:

sysdumpdev -l
primary              /dev/pdumplv
secondary            /dev/sdumplv
copy directory       /var/adm/dump
forced copy flag     FALSE
always allow dump    TRUE
dump compression     ON

Point the primary at paging space, then delete the logical volume:

sysdumpdev -Pp /dev/hd6
rmlv pdumplv

Using hd6 as a temporary dump device is normal practice and costs you nothing except a dump that has to be copied out at the next boot. You will put pdumplv back at the end.

Step 2: break the mirror

There are two ways and they end in the same place.

The blunt one:

unmirrorvg rootvg hdisk0

It works most of the time. It does not work reliably when there are stale partitions, and it fails outright if pdumplv was mirrored — which it should not be by default, but check, because someone may have done it by hand. If you removed pdumplv in step 1 this is moot.

The controlled one, which I prefer when anything looks unhealthy, removes one copy at a time:

lsvg -l rootvg
rmlvcopy hd5 1 hdisk0
rmlvcopy hd6 1 hdisk0
rmlvcopy hd8 1 hdisk0
rmlvcopy hd4 1 hdisk0

The 1 is the number of copies you want to be left, not the number to remove. Run it for every logical volume the first command listed, then confirm the PVs column reads 1 across the board.

It is slower and it is worth it. When unmirrorvg fails halfway you get to work out which volumes it already touched; when rmlvcopy fails you know exactly where you are.

Step 3: write down the disk identity

Before the disk leaves the machine, capture what the engineer needs to find it in the rack:

lscfg -vl hdisk0
  DEVICE            LOCATION          DESCRIPTION
  hdisk0            10-88-00-8,0      16 Bit LVD SCSI Disk Drive
        Manufacturer............................IBM
        Machine Type and Model..................DDYS-T09170M
        FRU Number..............................00P1517
        Serial Number...........................4DFJY156

The location code and the serial number are what stop somebody pulling the wrong drive out of a running mirror. Send both to whoever is doing the physical swap, in writing.

Step 4: remove the disk from AIX

reducevg rootvg hdisk0
rmdev -dl hdisk0
lsvg -p rootvg
lspv

reducevg takes the disk out of the volume group, rmdev -dl removes its definition from the ODM. The last two commands are the confirmation — hdisk0 should be gone from both. If reducevg refuses, something is still on the disk and you missed a step above.

Now the disk can physically come out and the new one goes in.

Step 5: bring the new disk in

cfgmgr

Configuration Manager finds the new hardware and, on a two-disk system, hands it back the name hdisk0.

Then tell the machine that the old fault has been dealt with, otherwise it keeps flagging a disk that no longer exists. In diag, go to Task Selection, then Log Repair Action, then pick hdisk0. Exit with Esc 0. Check it registered:

errpt | more
2F3E09A4   0819110902 I H hdisk0         REPAIR ACTION

While you are in diag, run Certify the disk against the new drive. It takes a while on a large disk and it is the only thing standing between you and discovering that the replacement is also bad — after you have re-mirrored onto it.

Step 6: rebuild the mirror

extendvg rootvg hdisk0

Then mirror, again by whichever method you trust:

mirrorvg rootvg hdisk0
syncvg -v rootvg

Note that mirrorvg mirrors everything, including pdumplv if you have already recreated it. Do not recreate it yet — that is the next step, and it is deliberately last.

Per-volume, if you prefer control:

mklvcopy -k hd5 2 hdisk0
mklvcopy -k hd6 2 hdisk0
mklvcopy -k hd8 2 hdisk0
mklvcopy -k hd4 2 hdisk0
syncvg -v rootvg

The 2 is the target number of copies and -k synchronises as it goes. Confirm with lsvg -l rootvg that every PVs column reads 2 and every state reads syncd, not stale.

Step 7: put the dump device and boot image back

mklv -y pdumplv rootvg 40 hdisk0
sysdumpdev -Pp /dev/pdumplv

Forty logical partitions is the conventional size but do not copy it blindly — run sysdumpdev -e for the current estimated dump size and make the volume big enough for it. A dump device that is too small produces no dump at all, which you find out on the day you needed one.

Leave pdumplv unmirrored. If you used mirrorvg above and it picked the volume up, drop the copy:

rmlvcopy pdumplv 1 hdisk1

Finally, put the boot image and the boot list back the way they were:

bosboot -a -d hdisk0
bootlist -m normal hdisk0 hdisk1
bootlist -m normal -o

Non-root volume groups

Everything above exists because of what rootvg carries. For a data volume group there is no boot image, no boot list and no dump device, so the job collapses to unmirrorvg, reducevg, rmdev, physical swap, cfgmgr, extendvg, mirrorvg, syncvg. The diag repair action and certify steps are still worth doing.

The last thing to check is the one people skip: run lsvg -l rootvg one more time the next morning. A mirror that synchronised cleanly at midnight and shows stale partitions at eight is telling you the new disk is bad too, and that is much better to learn from a command than from a phone call.

Comments

Popular posts from this blog

Diskless server iSCSI boot disk XCP-NG installation

 For Cisco UCS chassis, eg. 5108 for installation it's required to set FC or iSCSI disk for boot ( boot disk ). They are called diskless servers. Unfrtunatelly XCP-NG - 8.2.1 Jan 2025 - in my case cannot boot as it is from iSCSI. Here is a guide, how to do it.

AI's Democratic Era Ended in April 2026

For about three years, every big AI lab behaved like a teenager with a new car. Look what mine can do. Look at my benchmark. Look at my context window. Every few weeks somebody published a chart where their bar was taller than everybody else's bar, and we all went along with it, because the bars kept getting taller and the models kept getting better and it was fun to watch. Meanwhile China said almost nothing. Then once in a while DeepSeek dropped something that knocked everyone off their feet, everyone spent a week arguing about training costs, and the noise started again. That era is over. It ended in April 2026, and I don't think we noticed how completely. The race we thought we were in The mental model everybody used was Formula 1. Whoever builds the fastest car wins the season. Speed was the whole point, and speed was something you showed off, because showing off is how you raise the next round. Some people preferred the Manhattan Project comparison. I used to think that o...