A disk in a mirrored rootvg is the one disk you cannot just pull. It holds the boot image, it probably holds the primary dump device, and half of every logical volume on the system lives on it. Pull it without preparation and reducevg refuses, or worse, the machine comes up on nothing after the next reboot.
This is the full sequence for a two-disk mirrored rootvg on AIX, replacing hdisk0. Non-root volume groups are much simpler and I have noted the difference at the end.
Decide whether you actually have a problem
Disks rarely die cleanly. They complain first, and the complaints show up in the error log:
errpt | more
1581762B 0727203502 T H hdisk0 DISK OPERATION ERROR
1581762B 0727203502 P H hdisk0 DISK OPERATION ERROR
The third column is the one that matters. T is temporary, P is permanent. A single temporary error on a busy system is noise. A cluster of them over a few hours is a drive on its way out. A permanent error is not a warning, it is a report of something that already failed.
Then check whether the mirror is still intact:
lsvg -l rootvg
rootvg:
LV NAME TYPE LPs PPs PVs LV STATE MOUNT POINT
hd5 boot 1 2 2 closed/syncd N/A
hd6 paging 64 128 2 open/syncd N/A
hd8 jfslog 1 2 2 open/stale N/A
hd4 jfs 4 8 2 open/stale /
stale means writes are no longer reaching both copies. At that point you are running on one disk whether you like it or not, and the mirror is giving you a false sense of protection. Fix it before you do anything else that week.
What the whole job looks like
Five stages. The middle three are the uncomfortable ones, because the machine is running unprotected on a single disk the entire time.
Plan for that. Do not start this at four in the afternoon on a Friday, and do not start it while the surviving disk is also logging errors.
Step 1: move everything that pins the disk
Three things make hdisk0 unremovable, and all three have to go first.
Write a boot image to the surviving disk and change the boot list so firmware prefers it:
bosboot -a -d hdisk1
bootlist -m normal hdisk1 hdisk0
bootlist -m normal -o
That last command reads the list back. Do it. A typo here is discovered at the worst possible moment.
Now the dump device. On a default install the primary dump lives in pdumplv on hdisk0, and you cannot remove a disk that holds it:
sysdumpdev -l
primary /dev/pdumplv
secondary /dev/sdumplv
copy directory /var/adm/dump
forced copy flag FALSE
always allow dump TRUE
dump compression ON
Point the primary at paging space, then delete the logical volume:
sysdumpdev -Pp /dev/hd6
rmlv pdumplv
Using hd6 as a temporary dump device is normal practice and costs you nothing except a dump that has to be copied out at the next boot. You will put pdumplv back at the end.
Step 2: break the mirror
There are two ways and they end in the same place.
The blunt one:
unmirrorvg rootvg hdisk0
It works most of the time. It does not work reliably when there are stale partitions, and it fails outright if pdumplv was mirrored — which it should not be by default, but check, because someone may have done it by hand. If you removed pdumplv in step 1 this is moot.
The controlled one, which I prefer when anything looks unhealthy, removes one copy at a time:
lsvg -l rootvg
rmlvcopy hd5 1 hdisk0
rmlvcopy hd6 1 hdisk0
rmlvcopy hd8 1 hdisk0
rmlvcopy hd4 1 hdisk0
The 1 is the number of copies you want to be left, not the number to remove. Run it for every logical volume the first command listed, then confirm the PVs column reads 1 across the board.
It is slower and it is worth it. When unmirrorvg fails halfway you get to work out which volumes it already touched; when rmlvcopy fails you know exactly where you are.
Step 3: write down the disk identity
Before the disk leaves the machine, capture what the engineer needs to find it in the rack:
lscfg -vl hdisk0
DEVICE LOCATION DESCRIPTION
hdisk0 10-88-00-8,0 16 Bit LVD SCSI Disk Drive
Manufacturer............................IBM
Machine Type and Model..................DDYS-T09170M
FRU Number..............................00P1517
Serial Number...........................4DFJY156
The location code and the serial number are what stop somebody pulling the wrong drive out of a running mirror. Send both to whoever is doing the physical swap, in writing.
Step 4: remove the disk from AIX
reducevg rootvg hdisk0
rmdev -dl hdisk0
lsvg -p rootvg
lspv
reducevg takes the disk out of the volume group, rmdev -dl removes its definition from the ODM. The last two commands are the confirmation — hdisk0 should be gone from both. If reducevg refuses, something is still on the disk and you missed a step above.
Now the disk can physically come out and the new one goes in.
Step 5: bring the new disk in
cfgmgr
Configuration Manager finds the new hardware and, on a two-disk system, hands it back the name hdisk0.
Then tell the machine that the old fault has been dealt with, otherwise it keeps flagging a disk that no longer exists. In diag, go to Task Selection, then Log Repair Action, then pick hdisk0. Exit with Esc 0. Check it registered:
errpt | more
2F3E09A4 0819110902 I H hdisk0 REPAIR ACTION
While you are in diag, run Certify the disk against the new drive. It takes a while on a large disk and it is the only thing standing between you and discovering that the replacement is also bad — after you have re-mirrored onto it.
Step 6: rebuild the mirror
extendvg rootvg hdisk0
Then mirror, again by whichever method you trust:
mirrorvg rootvg hdisk0
syncvg -v rootvg
Note that mirrorvg mirrors everything, including pdumplv if you have already recreated it. Do not recreate it yet — that is the next step, and it is deliberately last.
Per-volume, if you prefer control:
mklvcopy -k hd5 2 hdisk0
mklvcopy -k hd6 2 hdisk0
mklvcopy -k hd8 2 hdisk0
mklvcopy -k hd4 2 hdisk0
syncvg -v rootvg
The 2 is the target number of copies and -k synchronises as it goes. Confirm with lsvg -l rootvg that every PVs column reads 2 and every state reads syncd, not stale.
Step 7: put the dump device and boot image back
mklv -y pdumplv rootvg 40 hdisk0
sysdumpdev -Pp /dev/pdumplv
Forty logical partitions is the conventional size but do not copy it blindly — run sysdumpdev -e for the current estimated dump size and make the volume big enough for it. A dump device that is too small produces no dump at all, which you find out on the day you needed one.
Leave pdumplv unmirrored. If you used mirrorvg above and it picked the volume up, drop the copy:
rmlvcopy pdumplv 1 hdisk1
Finally, put the boot image and the boot list back the way they were:
bosboot -a -d hdisk0
bootlist -m normal hdisk0 hdisk1
bootlist -m normal -o
Non-root volume groups
Everything above exists because of what rootvg carries. For a data volume group there is no boot image, no boot list and no dump device, so the job collapses to unmirrorvg, reducevg, rmdev, physical swap, cfgmgr, extendvg, mirrorvg, syncvg. The diag repair action and certify steps are still worth doing.
The last thing to check is the one people skip: run lsvg -l rootvg one more time the next morning. A mirror that synchronised cleanly at midnight and shows stale partitions at eight is telling you the new disk is bad too, and that is much better to learn from a command than from a phone call.
Comments
Post a Comment