Restructured Documentation
Automatic Documentation Deployment / Sync Docs to https://kb.bunny-lab.io (push) Successful in 8s

This commit is contained in:
2026-09-05 14:08:43 -06:00
parent c4bd235eba
commit 289769a601
281 changed files with 5403 additions and 3563 deletions
@@ -0,0 +1,91 @@
---
tags:
- Virtualization and Storage
- Rebuild Failover Cluster Replication
---
## Purpose
If you run an environment with multiple Hyper-V: Failover Clusters, for the purpose of Hyper-V: Failover Cluster Replication via a `Hyper-V Replica Broker` role installed on a host within the Failover Cluster, sometimes a GuestVM will fail to replicate itself to the replica cluster, and in those cases, it may not be able to recover on its own. This guide attempts to outline the process to rebuild replication for GuestVMs on a one-by-one basis.
!!! note "Assumptions"
This guide assumes you have two Hyper-V Failover Clusters, for the sake of the guide, we will refer to the Production cluster as `CLUSTER-01` and the Replication cluster as `CLUSTER-02`. This guide also assumes that Replication was set up beforehand, and does not include instructions on how to deploy a Replica Broker (at this time).
## Production Cluster - CLUSTER-01
### Locate the GuestVM
You need to start by locating the GuestVM in the Production cluster, CLUSTER-01. You will know you found the VM if the "Replication Health" is either `Unhealthy`, `Warning`, or `Critical`.
### Remove Replication from GuestVM
- Within a node of the Hyper-V: Failover Cluster Manager
- Right-Click the GuestVM
- Navigate to "**Replication > Remove Replication**"
- Confirm the removal by clicking the "**Yes**" button. You will know if it removed replication when the "Replication State" of the GuestVM is `Not enabled`
## Replication Cluster - CLUSTER-02
### Note the storage GUID of the GuestVM in the replication cluster
- Within a node of the replication cluster's Hyper-V: Failover Cluster Manager
- Right-Click the same GuestVM and click "Manage..." `This will open Hyper-V Manager`
- Right-Click the GuestVM and click "Settings..."
- Navigate to "**ISCSI Controller**"
- Click on one of the Virtual Disks attached to the replica VM, and note the full folder path for later. e.g. `C:\ClusterStorage\Volume1\HYPER-V REPLICA\VIRTUAL HARD DISKS\020C9A30-EB02-41F3-8D8B-3561C4521182`
!!! warning "Noting the GUID of the GuestVM"
You need to note the folder location so you have the GUID. Without the GUID, cleaning up the old storage associated with the GuestVM replica files will be much more difficult / time-consuming. Note it down somewhere safe, and reference it later in this guide.
### Delete the GuestVM from the Replication Cluster
Now that you have noted the GUID of the storage folder of the GuestVM, we can safely move onto removing the GuestVM from the replication cluster.
- Within a node of the replication cluster's Hyper-V: Failover Cluster Manager
- Right-Click the GuestVM
- Navigate to "**Replication > Remove Replication**"
- Confirm the removal by clicking the "**Yes**" button. You will know if it removed replication when the "Replication State" of the GuestVM is `Not enabled`
- Right-Click the GuestVM (again) `You will see that "Enable Replication" is an option now, indicating it was successfully removed.`
!!! note "Replica Checkpoint Merges"
When you removed replication, there may have been replication checkpoints that automatically try to merge together with a `Merge in Progress` status. Just let it finish before moving forward.
- Within the same node of the replication cluster's Hyper-V: Failover Cluster Manager `Switch back from Hyper-V Manager`
- Right-Click the GuestVM and click "**Remove**"
- Confirm the action by clicking the "**Yes**" button
### Delete the GuestVM manually from Hyper-V Manager on all replication cluster hosts
At this point in time, we need to remove the GuestVM from all of the servers in the cluster. Just because we removed it from the Hyper-V: Failover Cluster did not remove it from the cluster's nodes. We can automate part of this work by opening Hyper-V Manager on the same Failover Node we have been working on thus far, and from there we can connect the rest of the replication nodes to the manager to have one place to connect to all of the nodes, avoiding hopping between servers.
- Open Hyper-V Manager
- Right-Click "Hyper-V Manager" on the left-hand navigation menu
- Click "Connect to Server..."
- Type the names of every node in the replication cluster to connect to each of them, repeating the two steps above for every node
- Remove GuestVM from the node it appears on
- On one of the replication cluster nodes, we will see the GuestVM listed, we are going to Right-Click the GuestVM and select "**Delete**"
### Delete the GuestVM's replicated VHDX storage from replication ClusterStorage
Now we need to clean up the storage left behind by the replication cluster.
- Within a node of the replication cluster
- Navigate to `C:\ClusterStorage\Volume1\HYPER-V REPLICA\VIRTUAL HARD DISKS`
- Delete the entire GUID folder noted in the previous steps. `e.g. 020C9A30-EB02-41F3-8D8B-3561C4521182`
## Production Cluster - CLUSTER-01
### Re-Enable Replication on GuestVM in Cluster-01 (Production Cluster)
At this point, we have disabled replication for the GuestVM and cleaned up traces of it in the replication cluster. Now we need to re-enable replication on the GuestVM back in the production cluster.
- Within a node of the production Hyper-V: Failover Cluster Manager
- Right-Click the GuestVM
- Navigate to "**Replication > Enable Replication...**"
- Click "Next"
- For the "**Replica Server**", enter the name of the role of the Hyper-V Replica Broker role in the (replication cluster's) Failover Cluster. `e.g. CLUSTER-02-REPL`, then click "Next"
- Click the "Select Certificate" button, since the Broker was configured with Certificate-based authentication instead of Kerberos (in this example environment). It will prompt you to accept the certificate by clicking "OK". (e.g. `HV Replica Root CA`), then click "Next"
- Make sure every drive you want replicated is checked, then click "Next"
- Replication Frequency: `5 Minutes`, then click "Next"
- Additional Recovery Points: `Maintain only the latest recovery point`, then click "Next"
- Initial Replication Method: `Send initial copy over the network`
- Schedule Initial Replication: `Start replication immediately`
- Click "Next"
- Click "Finish"
!!! success "Replication Enabled"
If everything was successful, you will see a dialog box named "Enable replication for `<GuestVM>`" with a message similar to the following: "Replica virtual machine `<GuestVM>` was successfully created on the specified Replica server `<Node-in-Replication-Cluster>`.
At this point, you can click "Close" to finish the process. Under the GuestVM details, you will see "Replication State": `Initial Replication in Progress`.
## Related Documentation
- [Related Virtualization and Storage Documentation](<../../../../reference/Virtualization and Storage/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,45 @@
---
tags:
- Virtualization and Storage
- Forcefully Stop GuestVM
---
## Purpose
If you have a GuestVM that will not stop gracefully either because the Hyper-V host is goofed-up or the VMMS service won't allow you to restart it. You can perform a hail-mary to forcefully stop the GuestVM's Hyper-V process.
!!! warning "May Cause GuestVM to be Inconsistent"
This is meant as a last-resort when there are no other options on-the-table. You may end up corrupting the GuestVM by doing this.
### Get the VMID of the GuestVM
```powershell
Get-VM SERVER-01 | Select VMName, VMId
# Example Output
# VMName VMId
# ------ ------------------------------------
# SERVER-01 3e4b6f91-6c6c-4075-9b7e-389d46315074
```
### Extrapolate Process ID
Now you need to hunt-down the process ID associated with the GuestVM.
```powershell
Get-CimInstance Win32_Process -Filter "Name='vmwp.exe'" |
Where-Object { $_.CommandLine -match "3e4b6f91-6c6c-4075-9b7e-389d46315074" } |
Select-Object ProcessId, CommandLine
# Example Output
# ProcessId CommandLine
# --------- ---------------------------------------------------------
# 12488 "C:\Windows\System32\vmwp.exe" -vmid 3e4b6f91-6c6c-4075-9b7e-389d46315074
```
### Terminate Process
Lastly, you terminate the process by its ID.
```powershell
Stop-Process -Id 12488 -Force
```
## Related Documentation
- [Related Virtualization and Storage Documentation](<../../../reference/Virtualization and Storage/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,44 @@
---
tags:
- Kerberos
- Virtualization and Storage
---
## Purpose
You may find that you want to be able to live-migrate guestVMs on a Hyper-V environment that is not clustered as a Hyper-V Failover Cluster, you will have permission issues. One way to work around this is to use CredSSP as the authentication mechanism, which is not ideal but useful in a pinch, or you can use Kerberos-based authentication.
This document will cover both scenarios.
=== "Kerberos Authentication (*Preferred*)"
- Log into a domain controller that both Hyper-V hosts are capable of communicating with
- Open "**Server Manager > Tools " Active Directory Users & Computers**"
- Locate the computer objects representing both of the Hyper-V servers and repeat the steps below for each Hyper-V computer object:
- Right-Click > "**Properties**"
- Click on the "**Delegation**" Tab
- Check the radiobox for the open "**Trust this computer for delegation to specified services only.**"
- Ensure that "**User Kerberos Only** is checked
- Click on the "**Add**" button
- Click the "**Users or Computers...**" button
- Within the object search field, type in the name of the Hyper-V server you want to delegate access to (this will be the opposite host. e.g. VIRT-NODE-02, then repeat these steps later to delegate access for VIRT-NODE-01, etc)
- You will see a list of services that you can allow delegation to, add the following services:
- `cisvc`
- `mcsvc`
- `cifs`
- `Virtual Machine Migration Service`
- `Microsoft Virtualization Console`
- Click the "**Apply**" button, then click the "**OK**" button to finalize these changes.
- Repeat the above steps for the opposite Hyper-V host. This way both hosts are delegated to eachother
- e.g. `VIRT-NODE-01 <---(delegation)---> VIRT-NODE-02`
=== "CredSSP Authentication"
- Log into both Hyper-V Hosts as the same administrative user. Preferrably a domain administrator
- From the Hyper-V host currently running the GuestVM that needs to be migrated, open Hyper-V Manager and right-click > "**Move**" the guestVM.
- Select the destination by providing the fully-qualified domain name of the destination server (or in some cases the shorthand hostname of the destination server)
- It should begin the migration process.
**Note**: Do not perform a "Pull" from source to the destination. You want to always "Push" the VM to its destination. It will generally fail if you try to "Pull" the VM to its destination due to the way that CredSSP works in this context.
## Related Documentation
- [Related Virtualization and Storage Documentation](<../../../reference/Virtualization and Storage/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,170 @@
---
tags:
- Linux
- Filesystems
---
## Purpose
The purpose of this workflow is to illustrate the process of expanding storage for a RHEL-based Linux server acting as a GuestVM. We want the VM to have more storage space, so this document will go over the steps to expand that usable space.
!!! info "Assumptions"
It is assumed you are using a RHEL variant of linux such as Rocky Linux. This should apply to any version of Linux, but was written in a Rocky Linux 9.4 lab environment.
This document also assumes you did not enable Logical Volume Management (LVM) when deploying your server. If you did, you will need to perform additional LVM-specific steps after increasing the space.
!!! abstract "Oracle Linux Disk / LVM Terminology Idiosyncrasy"
Oracle Linux refers to disks as `/dev/hda` and /dev/hda2` and not something like `/dev/sda` / `/dev/sda2`. You will see certain parts of this document mention `/dev/hda`, in those cases, you may need to switch it to a standard `/dev/sda<#>` in order to make it work in your particular environment.
## Increase GuestVM Virtual Disk Size
This part should be fairly straight-forward. Using whatever hypervisor is running the Linux GuestVM, expand the disk space of the disk to the desired size.
## Extend Partition Table
This step goes over how to increase the usable space of the virtual disk within the GuestVM itself after it was expanded within the hypervisor.
!!! warning "Be Careful"
When you follow these steps, you will be deleting the existing partition and immediately re-creating it. If you do not use the **EXACT SAME** starting sector for the new partition, you will destroy data. Be sure to read every annotation next to each command to fully understand what you are doing.
=== "Using GDISK"
```sh
sudo dnf install gdisk -y
gdisk /dev/<diskNumber> # (1)
p <ENTER> # (2)
d <ENTER> # (3)
4 <ENTER> # (4)
n <ENTER> # (5)
4 <ENTER> # (6)
<DEFAULT-FIRST-SECTOR-VALUE> (Just press ENTER) # (7)
<DEFAULT-LAST-SECTOR-VALUE> (Just press ENTER) # (8)
<FILESYSTEM-TYPE=8300 (Linux Filesystem)> (Just press ENTER) # (9)
w <ENTER> # (10)
```
??? info "Detailed Command Breakdown"
1. The first command needs you to enter the disk identifier. In most cases, this will likely be the first disk, such as `/dev/sda`. You do not need to indicate a partition number in this step, as you will be asked for one in a later step after identifying all of the partitions on this disk in the next command.
2. This will list all of the partitions on the disk.
3. This will ask you for a partition number to delete. Generally this is the last partition number listed. In the example below, you would type `4` then press ++enter++ to schedule the deletion of the partition.
4. See the previous annotation for details on what entering `4` does in this context.
5. This tells gdisk to create a new partition.
6. This tells gdisk to re-make partition 4 (the one we just deleted in the example).
7. We just want to leave this as the default. In my example, it would look like this:
`First sector (34-2147483614, default = 19826688) or {+-}size{KMGTP}: 19826688`
8. We just want to leave this as the default. In my example, it would look like this:
`Last sector (19826688-2147483614, default = 2147483614) or {+-}size{KMGTP}: 2147483614`
9. Just leave this as-is and press ++enter++ without entering any values. Assuming you are using XFS, as this guide was written for, the default "Linux Filesystem" is what you want for XFS.
10. This will write the changes to the partition table making them reality instead of just staging the changes.
!!! example "Example Output"
```text
Command (? for help): p
Disk /dev/sda: 2147483648 sectors, 1024.0 GiB
Model: Virtual Disk
Sector size (logical/physical): 512/4096 bytes
Disk identifier (GUID): 8A5C2469-B07B-42AC-8E57-E756E62D37D1
Partition table holds up to 128 entries
Main partition table begins at sector 2 and ends at sector 33
First usable sector is 34, last usable sector is 2147483614
Partitions will be aligned on 2048-sector boundaries
Total free space is 1073743838 sectors (512.0 GiB)
Number Start (sector) End (sector) Size Code Name
1 2048 1230847 600.0 MiB EF00 EFI System Partition
2 1230848 3327999 1024.0 MiB 8300
3 3328000 19826687 7.9 GiB 8200
4 19826688 1073741790 502.5 GiB 8300 Linux filesystem
```
=== "Using FDISK"
```sh
pvdisplay # (1)
fdisk /dev/hda # (2)
p <ENTER> # List Partitions
d <ENTER> # Delete a partition
2 <ENTER> # Delete Partition 2 (e.g. /dev/hda2)
n <ENTER> # Make a new Partition
p <ENTER> # Primary Partition Type
Starting Sector: <ENTER> # Use Default Value
Ending Sector: <ENTER> # Use Default Value
w <ENTER> # Commit all queued-up changes and write them to the disk
```
??? info "Detailed Command Breakdown"
1. Use pvdisplay to get the target disk identifier
2. Replace `/dev/hda` with the target disk identifier found in the previous step
**Point of No Return**:
When you press `w` in both cases of `gdisk` or `fdisk`, then ++enter++ the changes will be written to disk, meaning there is no turning back unless you have full GuestVM backups or a snapshot to rollback with, or something like Veeam Backup & Replication. Be certain the first and last sector values are correctly configured before proceeding. (Default values generally are good for this)
=== "Using GROWPART (Ubuntu)"
```sh
echo 1 | sudo tee /sys/class/block/sda/device/rescan
growpart /dev/sda 2
sudo resize2fs /dev/sda2 # Assuming ext4 filesystem, if unsure, run "df -Th"
```
## Detect the New Partition Sizes
At this point, the operating system wont detect the changes without a reboot, so we are going to force the operating system to detect them immediately with the following commands to avoid a reboot (if we can avoid it).
```sh
sudo partprobe /dev/<drive> # Drive Example: /dev/sda (Rocky) or /dev/hda (Oracle Linux)
sudo partx -u /dev/<diskNumber>
```
!!! bug "Partition Size Not Expanded? Reboot."
If you notice the partition still has not expanded to the desired size, you may have no choice but to reboot the server, then re-run the `gdisk` or `fdisk` commands a second time. In my lab environment, it didn't work until I rebooted. This might have been a hiccup on my end, but it's something to keep in mind if you run into the same issue of the size not changing.
```sh
sudo reboot
```
## Resize the Filesystem
=== "XFS Filesystem"
```sh
sudo xfs_growfs /
```
=== "Ext4 Filesystem"
```sh
resize2fs /dev/sda
```
=== "Ext4 Filesystem w/ LVM"
```sh
# Increase the Physical Volume Group Size
pvdisplay # Check the Current Size of the Physical Volume
pvresize /dev/hda2 # Enlarge the Physical Volume to Fit the New Partition Size
pvdisplay # Validatre the Size of the Physical Volume Increased to the New Size
# Increase the Logical Volume Group Size
lvextend -l +100%FREE /dev/VolGroup00/LogVol00 # Get this from running "lvdisplay" to find the correct Logical Volume Name
# Resize the Filesystem of the Disk to Fit the new Logical Volume
resize2fs /dev/VolGroup00/LogVol00
```
## Validate Storage Expansion
At this point, you can leverage `lsblk` or `df -h` to determine if the usable storage space was successfully increased or not. In this example, you can see that I increased my storage space from 512GB to 1TB.
!!! example "Example Command Output"
Command: `lsblk | grep "sda4"`
```text
└─sda4 8:4 0 1014.5G 0 part /
```
Command: `df -h | grep "sda4"`
```text
/dev/sda4 1015G 145G 871G 15% /
```
## Related Documentation
- [Related Virtualization and Storage Documentation](<../../../reference/Virtualization and Storage/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,51 @@
---
tags:
- ZFS
- iSCSI
- Linux
- Filesystems
---
## Purpose
The purpose of this workflow is to illustrate the process of expanding storage for a Linux server that uses an iSCSI-based ZFS storage. We want the VM to have more storage space, so this document will go over the steps to expand that usable space.
!!! info "Assumptions"
It is assumed you are using an Ubuntu based operating system, as these commands may not be the same on other distributions of Linux.
This document also assumes you did not enable Logical Volume Management (LVM) when deploying your server. If you did, you will need to perform additional LVM-specific steps after increasing the space.
## Increase iSCSI Disk Size
This part should be fairly straight-forward. Using whatever hypervisor / storage appliance hosting the iSCSI target, expand the disk space of the LUN to the desired size.
## Extend ZFS Pool
This step goes over how to increase the usable space of the ZFS pool within the server itself after it was expanded.
```sh
iscsiadm -m session --rescan # (1)
lsblk # (2)
parted /dev/sdX # (3)
unit TB # (4)
resizepart X XXTB # (5)
zpool list # (6)
zpool online -e <POOL-NAME> /dev/sdX # (7)
zpool scrub <POOL-NAME> # (8)
```
1. Re-scan iSCSI targets for changes.
2. Leverage `lsblk` to ensure that the storage size increase from the hypervisor / storage appliance reflects correctly.
3. Open partitioning utility on the ZFS volume / LUN / iSCSI disk. Replace `dev/sdX` with the actual device name.
4. Self-explanatory storage measurement.
5. Resizes whatever partition is given to fit the new storage capacity. Replace `X` with the partition number. Replace `XXTB` with a valid value, such as `10TB`.
6. This will allow you to list all ZFS pools that are available for the next command.
7. Brings the ZFS Pool back online. Replace `<POOL-NAME>` with the actual name of the ZFS pool.
8. This tells the system to scan the ZFS pool for any errors or corruption and correct them. Think of it as a form of housekeeping.
## Check on Scrubbing Progress
At this point, the ZFS pool has been expanded and a scrub task has been started. The scrubbing task can take several hours / days to run, so to keep track of it, you can run the following command to check the status of the ZFS pool / scrubbing task.
```sh
zpool status
```
## Related Documentation
- [Related Virtualization and Storage Documentation](<../../../reference/Virtualization and Storage/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,678 @@
---
tags:
- Virtualization and Storage
- Detect and Remove Orphaned VM Disks
---
## Purpose
This procedure describes how to identify and safely remove orphaned Proxmox VE virtual machine disks from shared iSCSI-backed LVM storage.
It is intended for environments where:
- Proxmox VE is clustered.
- Multiple Proxmox nodes access the same shared iSCSI LUN.
- The shared storage is exposed to Proxmox as LVM storage.
- VM disks are stored as LVM logical volumes.
- Some volumes may remain after VM disk deletion, failed migrations, failed resizes, storage UI inconsistencies, or manual recovery work.
The goal is to reclaim storage space without accidentally deleting disks that are still attached to running or stopped VMs.
---
## Scope
This document focuses on the following storage type:
```text
Proxmox storage type: lvm
Backing storage: shared iSCSI
Volume group: vg_proxmox_iscsi
Storage ID example: iscsi-cluster-lvm
```
Adjust the storage ID and volume group names as needed for your environment.
---
## Safety Requirements
!!! danger "Never delete based on the Storage UI alone"
The Proxmox storage UI may show a volume as belonging to a VM because its name follows the pattern `vm-<vmid>-disk-<n>`. That does not prove the disk is currently attached to the VM.
```text
Always verify against VM configuration files and active QEMU processes before deleting.
```
!!! warning "Run the audit before running any cleanup commands"
The audit scripts in this document are read-only. The cleanup commands are destructive. Do not run cleanup commands until the audit output has been reviewed.
!!! warning "Snapshot volumes require extra caution"
Volumes named like the following may be part of a snapshot chain:
````text
```text
snap_vm-<vmid>-disk-<n>_<snapshot-name>
```
Do not remove snapshot volumes manually unless you have verified that the VM and snapshot are no longer known to Proxmox, no backing chain references them, and no QEMU process has them open.
````
!!! note "Shared storage does not mean shared config"
In some cluster layouts, each node may only show VM config files for VMs assigned to that node. Therefore, an audit run from only one node can falsely report disks from other nodes as orphaned.
```text
Run the confirmation script on every node in the cluster.
```
---
## Terms
| Term | Meaning |
| -------------- | ------------------------------------------------------------------------------------------------------------------------------------------ |
| Attached disk | A disk volume referenced in a VM config, such as `scsi0`, `sata0`, `virtio0`, `efidisk0`, or `tpmstate0`. |
| Orphan disk | A storage volume that exists on shared storage but is not referenced by any VM config on any node and is not opened by any active process. |
| Volume ID | Proxmox storage identifier, such as `iscsi-cluster-lvm:vm-107-disk-1.qcow2`. |
| LV | LVM logical volume, such as `/dev/vg_proxmox_iscsi/vm-107-disk-1.qcow2`. |
| Snapshot chain | A chain of qcow2 backing files or Proxmox snapshot volumes. |
---
## Phase 1: Identify Storage Names
Run this on any Proxmox node:
```bash
pvesm status
cat /etc/pve/storage.cfg
vgs
```
Identify the shared iSCSI/LVM storage.
Example:
```text
Storage ID: iscsi-cluster-lvm
VG name: vg_proxmox_iscsi
```
For the rest of this document, replace these values if your environment differs:
```bash
STORAGE="iscsi-cluster-lvm"
VG="vg_proxmox_iscsi"
```
---
## Phase 2: Run the Storage Orphan Audit
Run the following script on one node that can see the shared LVM storage.
This script does not delete anything.
```bash
cat > /root/pve-iscsi-orphan-audit.sh <<'EOF'
#!/usr/bin/env bash
set -u
STORAGE="iscsi-cluster-lvm"
VG="vg_proxmox_iscsi"
OUT="/root/pve-iscsi-orphan-audit-$(hostname)-$(date +%Y%m%d-%H%M%S).txt"
{
echo "===== PVE ISCSI ORPHAN AUDIT ====="
echo "Host: $(hostname)"
echo "Date: $(date)"
echo "Storage: ${STORAGE}"
echo "VG: ${VG}"
echo
echo "===== STORAGE STATUS ====="
pvesm status 2>&1 | egrep "^(Name|${STORAGE})" || true
vgs "${VG}" 2>&1 || true
echo
echo "===== ALL VOLUMES IN ${STORAGE} ====="
pvesm list "${STORAGE}" 2>&1 || true
echo
echo "===== ALL LVs IN ${VG} ====="
lvs -a -o lv_name,lv_path,lv_size,lv_attr,devices "${VG}" 2>&1 || true
echo
echo "===== LOCAL VM CONFIG FILES ====="
for conf in /etc/pve/qemu-server/*.conf; do
[ -e "$conf" ] || continue
echo
echo "----- $conf -----"
cat "$conf"
done
echo
echo "===== REFERENCE ANALYSIS - LOCAL CONFIG FILES ONLY ====="
printf '%-55s | %-8s | %-10s | %-8s | %-8s | %s\n' \
"volume" "vmid" "referenced" "open" "size" "path"
{
pvesm list "${STORAGE}" 2>/dev/null | awk 'NR>1 {print $1}' | sed "s#^${STORAGE}:##"
lvs --noheadings -o lv_name "${VG}" 2>/dev/null | awk '{print $1}'
} | sort -u | while read -r vol; do
[ -n "$vol" ] || continue
case "$vol" in
vm-*-disk-*|snap_vm-*-disk-*) ;;
*) continue ;;
esac
vmid="unknown"
if [[ "$vol" =~ ^vm-([0-9]+)-disk- ]]; then
vmid="${BASH_REMATCH[1]}"
elif [[ "$vol" =~ ^snap_vm-([0-9]+)-disk- ]]; then
vmid="${BASH_REMATCH[1]}"
fi
ref="no"
if grep -R -Fq "$vol" /etc/pve/qemu-server 2>/dev/null; then
ref="yes"
fi
open="no"
if lsof 2>/dev/null | grep -Fq "$vol"; then
open="yes"
fi
size="$(lvs --noheadings -o lv_size "${VG}/${vol}" 2>/dev/null | awk '{$1=$1;print}')"
path="$(lvs --noheadings -o lv_path "${VG}/${vol}" 2>/dev/null | awk '{$1=$1;print}')"
printf '%-55s | %-8s | %-10s | %-8s | %-8s | %s\n' \
"$vol" "$vmid" "$ref" "$open" "$size" "$path"
done
echo
echo "===== DONE ====="
} | tee "$OUT"
echo
echo "Saved audit to: $OUT"
EOF
chmod +x /root/pve-iscsi-orphan-audit.sh
/root/pve-iscsi-orphan-audit.sh
```
The script writes a file similar to:
```text
/root/pve-iscsi-orphan-audit-<node>-<timestamp>.txt
```
---
## How to Read the First Audit
The most important section is:
```text
REFERENCE ANALYSIS - LOCAL CONFIG FILES ONLY
```
Example:
```text
volume | vmid | referenced | open | size
vm-107-disk-0.qcow2 | 107 | no | no | 4.00m
vm-107-disk-1.qcow2 | 107 | no | no | 256.04g
vm-107-disk-2.qcow2 | 107 | yes | no | 4.00m
vm-107-disk-3.qcow2 | 107 | yes | no | 256.04g
```
Interpretation:
| Field | Meaning |
| ---------------- | ------------------------------------------------------------------------------------------------------- |
| `referenced=yes` | The volume appears in a local VM config file. Do not delete. |
| `referenced=no` | The volume does not appear in local VM configs. It may be orphaned, but confirm across all nodes first. |
| `open=yes` | A process has the volume open. Do not delete. |
| `open=no` | No process on this node has the volume open. Still confirm across all nodes. |
!!! warning "Local reference analysis is not enough"
If a VM runs on another cluster node, its config may not appear on the node where you ran the audit. This can make valid disks look orphaned.
```text
Continue to Phase 3 before deleting anything.
```
---
## Phase 3: Run Cluster-Wide Confirmation
Run the following script on **every Proxmox node** in the cluster.
This script is read-only.
```bash
cat > /root/pve-cluster-vm-confirm.sh <<'EOF'
#!/usr/bin/env bash
set -u
STORAGE="iscsi-cluster-lvm"
VG="vg_proxmox_iscsi"
OUT="/root/pve-cluster-vm-confirm-$(hostname)-$(date +%Y%m%d-%H%M%S).txt"
{
echo "===== NODE ====="
hostname
date
echo
echo "===== CLUSTER RESOURCES - VMS ====="
pvesh get /cluster/resources --type vm 2>&1 || true
echo
echo "===== LOCAL QM LIST ====="
qm list 2>&1 || true
echo
echo "===== QEMU CONFIG FILES PRESENT ====="
ls -la /etc/pve/qemu-server/ 2>&1 || true
echo
echo "===== QEMU CONFIG FILE CONTENTS ====="
for conf in /etc/pve/qemu-server/*.conf; do
[ -e "$conf" ] || continue
echo
echo "----- $conf -----"
cat "$conf"
done
echo
echo "===== ALL STORAGE VOLUMES ====="
pvesm list "${STORAGE}" 2>&1 || true
echo
echo "===== ALL LVs ====="
lvs -a -o lv_name,lv_path,lv_size,lv_attr,devices "${VG}" 2>&1 || true
} | tee "$OUT"
echo
echo "Saved to: $OUT"
EOF
chmod +x /root/pve-cluster-vm-confirm.sh
/root/pve-cluster-vm-confirm.sh
```
Collect the output file from each node.
Example for a three-node cluster:
```text
/root/pve-cluster-vm-confirm-cluster-node-01-YYYYMMDD-HHMMSS.txt
/root/pve-cluster-vm-confirm-cluster-node-02-YYYYMMDD-HHMMSS.txt
/root/pve-cluster-vm-confirm-cluster-node-03-YYYYMMDD-HHMMSS.txt
```
---
## How to Read the Cluster Confirmation
For each suspicious volume, search all three outputs.
Example candidate:
```text
vm-107-disk-1.qcow2
```
Check whether it appears in any VM config:
```bash
grep -R "vm-107-disk-1.qcow2" /etc/pve/qemu-server/ || true
```
If reviewing output files manually, look for config lines such as:
```text
scsi0: iscsi-cluster-lvm:vm-107-disk-1.qcow2
sata0: iscsi-cluster-lvm:vm-107-disk-1.qcow2
virtio0: iscsi-cluster-lvm:vm-107-disk-1.qcow2
efidisk0: iscsi-cluster-lvm:vm-107-disk-1.qcow2
tpmstate0: iscsi-cluster-lvm:vm-107-disk-1.qcow2
```
If a volume appears in any of those lines, it is attached to a VM and must not be deleted.
---
## Phase 4: Classify Candidate Volumes
Use the following decision table.
| Condition | Classification | Action |
| --------------------------------------------------------------------- | ------------------- | ---------------------------------------------- |
| Volume appears in any VM config on any node | In use | Do not delete |
| Volume is opened by QEMU or another process | In use or unsafe | Do not delete |
| Volume is a `snap_vm-*` snapshot volume | Snapshot-chain item | Inspect snapshot/backing chain before deletion |
| Volume does not appear in any VM config and is not open | Orphan candidate | Eligible for final verification |
| VMID no longer exists in cluster resources and disk is not referenced | Strong orphan | Eligible for cleanup |
---
## Example: Valid VM Disks
If VM `107` has this config:
```text
efidisk0: iscsi-cluster-lvm:vm-107-disk-2.qcow2
sata0: iscsi-cluster-lvm:vm-107-disk-3.qcow2
```
Then these disks are valid and must not be deleted:
```text
vm-107-disk-2.qcow2
vm-107-disk-3.qcow2
```
If storage also contains:
```text
vm-107-disk-0.qcow2
vm-107-disk-1.qcow2
```
and neither appears in any config file on any node, those are orphan candidates.
---
## Phase 5: Final Verification Before Deletion
For each candidate volume, run the following checks on a node that can see the shared storage.
Replace the volume name as appropriate.
```bash
VOL="vm-107-disk-1.qcow2"
STORAGE="iscsi-cluster-lvm"
VG="vg_proxmox_iscsi"
echo "===== Check all cluster config references ====="
grep -R "$VOL" /etc/pve/qemu-server/ || true
echo
echo "===== Check Proxmox storage listing ====="
pvesm list "$STORAGE" | grep "$VOL" || true
echo
echo "===== Check LVM volume ====="
lvs -a -o lv_name,lv_path,lv_size,lv_attr,devices "$VG" | grep "$VOL" || true
echo
echo "===== Check whether open by any process ====="
lsof | grep "$VOL" || true
echo
echo "===== Check qemu-img metadata if device path exists ====="
LVPATH="$(lvs --noheadings -o lv_path "${VG}/${VOL}" 2>/dev/null | awk '{$1=$1;print}')"
if [ -n "$LVPATH" ] && [ -e "$LVPATH" ]; then
qemu-img info --backing-chain "$LVPATH"
else
echo "LV path missing or inactive: $LVPATH"
fi
```
Safe deletion pattern:
```text
grep -R ... no output
pvesm list ... shows the volume
lvs ... shows the volume
lsof ... no output
qemu-img info ... no unexpected backing file dependency
```
!!! danger "Stop if grep finds a reference"
If the candidate volume appears in any `/etc/pve/qemu-server/*.conf` file, do not delete it.
!!! danger "Stop if lsof finds a process"
If `lsof` shows the volume is open, do not delete it.
---
## Phase 6: Cleanup Commands
## Preferred Method: Proxmox Storage Layer
Use `pvesm free` first.
Example:
```bash
pvesm free iscsi-cluster-lvm:vm-107-disk-0.qcow2
pvesm free iscsi-cluster-lvm:vm-107-disk-1.qcow2
```
Then verify:
```bash
pvesm list iscsi-cluster-lvm | grep "vm-107-disk" || true
lvs -a -o lv_name,lv_size,lv_attr,devices vg_proxmox_iscsi | grep "vm-107-disk" || true
vgs vg_proxmox_iscsi
pvesm status | egrep '^(Name|iscsi-cluster-lvm)'
```
Expected result:
```text
vm-107-disk-0.qcow2 gone
vm-107-disk-1.qcow2 gone
vm-107-disk-2.qcow2 still present
vm-107-disk-3.qcow2 still present
```
---
## Fallback Method: Direct LVM Removal
Only use this if `pvesm free` refuses and the final verification confirms the volume is not referenced and not open.
```bash
lvremove /dev/vg_proxmox_iscsi/vm-107-disk-0.qcow2
lvremove /dev/vg_proxmox_iscsi/vm-107-disk-1.qcow2
```
Then refresh device nodes and verify:
```bash
vgscan --mknodes
udevadm settle
pvesm list iscsi-cluster-lvm | grep "vm-107-disk" || true
lvs -a -o lv_name,lv_size,lv_attr,devices vg_proxmox_iscsi | grep "vm-107-disk" || true
vgs vg_proxmox_iscsi
```
!!! warning "Prefer pvesm free over lvremove"
`pvesm free` lets Proxmox remove the volume through its storage abstraction. Use direct `lvremove` only when Proxmox refuses and the orphan status is already proven.
---
## Phase 7: Post-Cleanup Validation
After deleting orphan volumes, validate storage and VM health.
```bash
pvesm status
vgs vg_proxmox_iscsi
lvs -a -o lv_name,lv_size,lv_attr,devices vg_proxmox_iscsi
```
Check the affected VM’s config:
```bash
qm config 107
```
Confirm the VM still starts or remains healthy:
```bash
qm status 107
```
If the VM is running, confirm its active QEMU process only references expected disks:
```bash
ps auxww | grep "kvm -id 107" | grep -o "/dev/vg_proxmox_iscsi/[^ ,\"]*" | sort -u
```
Expected example:
```text
/dev/vg_proxmox_iscsi/vm-107-disk-2.qcow2
/dev/vg_proxmox_iscsi/vm-107-disk-3.qcow2
```
---
## Snapshot Volume Handling
Snapshot volumes require additional review.
Examples:
```text
snap_vm-105-disk-0_Fresh_Install.qcow2
snap_vm-106-disk-0_Fresh_Install_FullyUpdated.qcow2
```
Before deleting a snapshot volume, check:
```bash
qm config <vmid>
qm listsnapshot <vmid>
grep -R "snap_vm-<vmid>" /etc/pve/qemu-server/ || true
qemu-img info --backing-chain /dev/vg_proxmox_iscsi/vm-<vmid>-disk-<n>.qcow2
```
If the VM still has a `parent:` line or `qm listsnapshot` shows the snapshot, remove it through Proxmox first:
```bash
qm delsnapshot <vmid> <snapshot-name>
```
Only consider manual removal if:
- the VM no longer references the snapshot,
- no backing chain references the snapshot volume,
- no QEMU process has it open,
- and Proxmox cannot delete it normally.
!!! danger "Do not manually delete active snapshot-chain volumes"
Deleting an active snapshot backing volume can corrupt the VM disk chain.
---
## Example Cleanup Walkthrough
## Scenario
VM `107` has this config:
```text
efidisk0: iscsi-cluster-lvm:vm-107-disk-2.qcow2
sata0: iscsi-cluster-lvm:vm-107-disk-3.qcow2
```
Storage contains:
```text
vm-107-disk-0.qcow2
vm-107-disk-1.qcow2
vm-107-disk-2.qcow2
vm-107-disk-3.qcow2
```
`disk-0` and `disk-1` do not appear in any config and are not open by any process.
## Verify
```bash
grep -R "vm-107-disk-0.qcow2" /etc/pve/qemu-server/ || true
grep -R "vm-107-disk-1.qcow2" /etc/pve/qemu-server/ || true
lsof | grep "vm-107-disk-0.qcow2" || true
lsof | grep "vm-107-disk-1.qcow2" || true
```
Expected output:
```text
no output
```
## Delete
```bash
pvesm free iscsi-cluster-lvm:vm-107-disk-0.qcow2
pvesm free iscsi-cluster-lvm:vm-107-disk-1.qcow2
```
## Validate
```bash
pvesm list iscsi-cluster-lvm | grep "vm-107-disk"
lvs -a -o lv_name,lv_size,lv_attr,devices vg_proxmox_iscsi | grep "vm-107-disk"
vgs vg_proxmox_iscsi
```
Expected remaining volumes:
```text
vm-107-disk-2.qcow2
vm-107-disk-3.qcow2
```
---
## Technician Checklist
Use this checklist before removing any orphan disk.
- [ ] I ran the storage orphan audit.
- [ ] I ran the cluster confirmation script on every Proxmox node.
- [ ] I confirmed the candidate volume is not referenced in any VM config.
- [ ] I confirmed the candidate volume is not open by any process.
- [ ] I confirmed the candidate volume is not part of an active snapshot chain.
- [ ] I confirmed the VMID relationship is understood.
- [ ] I used `pvesm free` first.
- [ ] I used `lvremove` only if Proxmox refused and the volume was proven orphaned.
- [ ] I validated storage state after cleanup.
- [ ] I validated the affected VM still references only expected disks.
---
## Quick Reference Commands
## List shared storage volumes
```bash
pvesm list iscsi-cluster-lvm
```
## List LVs
```bash
lvs -a -o lv_name,lv_path,lv_size,lv_attr,devices vg_proxmox_iscsi
```
## Search VM configs
```bash
grep -R "vm-<vmid>-disk-<n>" /etc/pve/qemu-server/ || true
```
## Check open files
```bash
lsof | grep "vm-<vmid>-disk-<n>" || true
```
## Check image metadata
```bash
qemu-img info --backing-chain /dev/vg_proxmox_iscsi/vm-<vmid>-disk-<n>.qcow2
```
## Delete via Proxmox
```bash
pvesm free iscsi-cluster-lvm:vm-<vmid>-disk-<n>.qcow2
```
## Delete via LVM fallback
```bash
lvremove /dev/vg_proxmox_iscsi/vm-<vmid>-disk-<n>.qcow2
```
## Verify storage usage
```bash
pvesm status
vgs vg_proxmox_iscsi
```
## Related Documentation
- [Related Proxmox Documentation](<../../../reference/Virtualization and Storage/Proxmox/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,24 @@
---
tags:
- Proxmox
- Virtualization and Storage
---
## Purpose
Sometimes in some very specific situations, you will find that an LVM / VG just won't come online in ProxmoxVE. If this happens, you can run the following commands (and replace the placeholder location) to manually bring the storage online.
```sh
lvchange -an local-vm-storage/local-vm-storage
lvchange -an local-vm-storage/local-vm-storage_tmeta
lvchange -an local-vm-storage/local-vm-storage_tdata
vgchange -ay local-vm-storage
```
!!! info "Be Patient"
It can take some time for everything to come online.
!!! success
If you see something like this: `6 logical volume(s) in volume group "local-vm-storage" now active`, then you successfully brought the volume online.
## Related Documentation
- [Related Proxmox Documentation](<../../../reference/Virtualization and Storage/Proxmox/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,18 @@
---
tags:
- Proxmox
- Virtualization and Storage
---
## Purpose
The purpose of this document is to outline common tasks that you may need to run in your cluster to perform various tasks.
## Delete Node from Cluster
Sometimes you may need to delete a node from the cluster if you have re-built it or had issues and needed to destroy it. In these instances, you would run the following command (assuming you have a 3-node quorum in your cluster).
```sh
pvecm delnode promox-node-01
```
## Related Documentation
- [Related Proxmox Documentation](<../../../reference/Virtualization and Storage/Proxmox/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,71 @@
---
tags:
- Proxmox
- Veeam
- Backup
- Disaster Recovery
---
## Purpose
When you migrate virtual machines from Hyper-V (and possibly other platforms) to ProxmoxVE, you may run into several issues, from the disk formats being in `.raw` format instead of `.qcow2`, among other things. One thing in particular, which is the reason for this document, is that if you migrate Rocky Linux from Hyper-V into ProxmoxVE using Veeam Backup & Replication, it will break the storage system so badly that the operating system will not boot.
### Fixing Boot Issues
Some high-level things to do to fix this are listed below:
- Switch the VM processor type to `host`.
- The socket and core counts are reversed, so a single socket CPU with 16 cores will appear like 16 sockets with one core each, flip these around to correct this issue.
- The storage controller needs to be set to `VirtIO iSCSI`
- The display driver needs set to `Default`
#### Dracut Emergency Shell
If you start the VM and you reach a "dracut" prompt, then the bootloader got nuked and needs to be regenerated. Follow the steps below to work through this process:
- Boot from a Rocky Linux 9.5+ installation ISO in the broken Rocky Linux VM
- Select "**Troubleshooting -->**" in the boot menu
- Select "**Rescue a Rocky Linux System**"
- Press through the prompt with value `1` and `Continue` to select the automatic mounting of the detected operating system of the virtual machine
- Press **<ENTER>** to enter the shell, then run the following commands to fix the booting issues
```sh
chroot /mnt/sysroot
dracut --force --regenerate-all
grub2-mkconfig -o /boot/grub2/grub.cfg
exit
exit
```
!!! info "Boot Fix May Trigger Reboot Twice"
During the process, you may notice that the VM reboots itself a second-time. This is normal and can be left alone. The VM will eventually reach the login screen. Once you get this far, you can login and fix the networking issues in the VM to get it stabilized.
### Fixing Network Issues
The VM will lose the adapter name of `eth0` and put something else like `ens18` that needs to be reconfigured manually to get networking functional again:
- Type `ethtool ens18`, and if the link speed is `Unknown!`, then poweroff the VM and switch the network adapter from `VirtIO (paravirtualized)` to `Intel E1000`, then boot the VM back up.
- Run the following commands to assign the new `ens18` interface as a networking interface for the VM to use:
```sh
# Create the Interface (Replace the IP & DNS Variables)
nmcli connection add type ethernet ifname ens18 con-name ens18 ipv4.method manual ipv4.addresses 192.168.3.21/24 ipv4.gateway 192.168.3.1 ipv4.dns "1.1.1.1 1.0.0.1"
# Bring the Connection Online
nmcli connection up ens18
```
!!! success "VM Successfully Fixed"
At this point, the virtual machine should be booting, and have network access, bringing it back into production use.
### Convert VM Disk from `.RAW` to `.QCOW2`
Given that the migration process via Veeam Backup & Replication ignores the destination disk format (at the time of writing this), it is necessary to convert the format of the disk from `.raw` to `.qcow2` so that you can perform things like VM snapshots, which are essential during updates, development, and testing.
Open a shell onto the ProxmoxVE server that is currently holding the VM that you need to convert the disks for, then locate the disks (this is not explained here, yet), and run the following commands to convert them.
```sh
# Convert a Single Disk
qemu-img convert -f raw -O qcow2 source.raw destination.qcow2
# Convert All Disks in a Given Directory
find . -type f -name "*.raw" -exec sh -c 'qemu-img convert -f raw -O qcow2 "$1" "${1%.raw}.qcow2"' _ {} \;
```
## Related Documentation
- [Related Proxmox Documentation](<../../../reference/Virtualization and Storage/Proxmox/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,68 @@
---
tags:
- Virtualization and Storage
- Repair iSCSI Connections After Reboot
---
## Purpose
You may enounter an issue where when you reboot your ProxmoxVE cluster node, it will fail to start virtual machines because of an error related to being unable to see the underlying LVM disks for the GuestVM. This is generally bandaid-fixed by running `iscsiadm -m node --login` on the node, which makes it reconnect to the cluster's iSCSI storage. This is not a viable long-term solution.
### Configure Automatic Startup of iSCSI Services
Run these commands on every ProxmoxVE cluster node:
```sh
systemctl enable --now iscsid
systemctl enable --now open-iscsi
```
### Discover Targets (To Create Necessary iSCSI Records)
```sh
iscsiadm -m discovery -t sendtargets -p 192.168.3.3:3260
```
### Configure Automatic iSCSI Target Connection Behavior
Run these commands on every ProxmoxVE cluster node:
```sh
iscsiadm -m node \
-T iqn.2026-01.io.bunny-lab:storage:iscsi-cluster-storage \
-p 192.168.3.3 \
--op update \
-n node.startup \
-v automatic
iscsiadm -m node \
-T iqn.2026-01.io.bunny-lab:storage:iscsi-cluster-storage \
-p 192.168.3.3 \
--op update \
-n node.conn[0].startup \
-v automatic
```
### Verify Configuration
You will want to ensure that `node.startup = automatic` and `node.conn[0].startup = automatic` when you run the following command.
```sh
iscsiadm -m node -o show | grep -E 'node.name|node.conn\[0\].address|node.startup|node.conn\[0\].startup'
```
### Login/Mount iSCSI Targets
```sh
iscsiadm -m node \
-T iqn.2026-01.io.bunny-lab:storage:iscsi-cluster-storage \
-p 192.168.3.3:3260 \
--login
```
!!! success "Example Output"
If everything worked correctly, you should see output like the example below:
```sh
node.name = iqn.2026-01.io.bunny-lab:storage:iscsi-cluster-storage
node.startup = automatic
node.conn[0].address = 192.168.3.3
node.conn[0].startup = automatic
```
## Related Documentation
- [Related Proxmox Documentation](<../../../reference/Virtualization and Storage/Proxmox/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,48 @@
---
tags:
- Proxmox
- Virtualization and Storage
---
## Purpose
There are a few steps you have to take when upgrading ProxmoxVE from 8.4.1+ to 9.0+. The process is fairly straightforward, so just follow the instructions seen below.
!!! info "GuestVM Assumptions"
It is assumed that if you are running a ProxmoxVE cluster, you will migrate all GuestVMs to another cluster node. If this is a standalone ProxmoxVE server, you will shut down all GuestVMs safely before proceeding.
!!! warning "Perform `pve8to9` Readiness Check"
It's critical that you run the `pve8to9` command to ensure that your ProxmoxVE server meets all of the requirements and doesn't have any failures or potentially server-breaking warnings. If the `pve8to9` command is unknown, then run `apt update && apt dist-upgrade` in the shell then try again. Warnings should be addressed ad-hoc, but *CPU Microcode warnings can be safely ignored*.
**Example pve8to9 Summary Output**:
```sh
= SUMMARY =
TOTAL: 48
PASSED: 39
SKIPPED: 8
WARNINGS: 1
FAILURES: 0
```
### Update Repositories from `bookworm` to `trixie`
```sh
sed -i 's/bookworm/trixie/g' /etc/apt/sources.list
sed -i 's/bookworm/trixie/g' /etc/apt/sources.list.d/pve-install-repo.list
apt update
```
### Upgrade to ProxmoxVE 9.0
!!! warning "Run Upgrade Commands in iLO/iDRAC/IPMI"
At this point, its very likely that if you are using SSH, it may unexpectedly have the session terminated, so you absolutely want to use a local or remote console to the server to run the commands below, both to ensure you maintain access to the console, as well as to see if any issues arise during POST after the reboot.
```sh
apt dist-upgrade -y
reboot
```
!!! note "Disable `pve-enterprise` Repository"
At this point, the ProxmoxVE server should be running on v9.0+, you will want to disable the `pve-enterprise` repository as it will goof up future updates if you don't disable it.
## Related Documentation
- [Related Proxmox Documentation](<../../../reference/Virtualization and Storage/Proxmox/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,50 @@
---
tags:
- TrueNAS
- Storage
- Hardware
---
## Purpose
This document acts as a workflow to understand how to replace a drive on TrueNAS Core when it is hosted on an HPE Proliant server with HBA / IT Mode enabled. This enables you to hot-swap drives without rebooting TrueNAS Core.
### Offline the Disk
- You will log into the TrueNAS Core [WebUI](http://192.168.3.3).
- Navigate to "**Storage > Disks**"
- Look for the drive that is having issues / faults / unavailable and reference it's `da` number to reference later. (e.g. `da3`)
- Confirm the serial number of the drive and correlate that to the physical location in the [Disk Arrays](<../../../reference/Lab Map/Hardware/Storage Node 01 TrueNAS Core Disk Layout.md>) document,
- Navigate to "**Storage > Pools**"
- Look for the gear icon to the right of the storage pool and click on it
- Click on "**Status**"
- Locate the failing / failed drive and click on the "**...**" elipsis menu button
- Proceed to "**Offline**" the disk. This ensures that TrueNAS Core stops trying to use the disk.
### Physical Disk Replacement
At this point, we need to physically go to the server and pull out the failing drive and replace it.
- Take note of the new serial number on the replacement drive and update the [Disk Arrays](<../../../reference/Lab Map/Hardware/Storage Node 01 TrueNAS Core Disk Layout.md>) document accordingly.
- Insert the replacement drive back into the TrueNAS Core server
### Trigger Disk Re-Scan
Now we need to tell TrueNAS / FreeBSD to re-scan all disks to locate the new one.
- Within the TrueNAS Core WebUI, navigate to "**Shell**" and run the following command: `camcontrol rescan all`
- Navigate (back) to "**Storage > Pools > Status**"
- Locate the failed drive via it's `da` number again, and click the "**...**" elipsis menu button
- Proceed to "**Replace**" the disk, and when given a dropdown menu, only the new replacement disk should appear with the same `da` number
!!! success "Resilvering Started"
At this point, TrueNAS core will start taking parity data from the rest of the drives in the storage pool to reconstruct the replaced drive. This may take an hour or two depending on the speed of the drives and used capacity within the pool itself.
It is recommended to run a SCRUB right after resilvering to ensure that all data is accurate and healthy.
!!! info "Checking on Resilvering Process via CLI"
If you feel so inclined, you can check on the resilvering process by running the following command:
```sh
zpool status | grep "to go"
```
## Related Documentation
- [Storage Node 01 Disk Layout](<../../../reference/Lab Map/Hardware/Storage Node 01 TrueNAS Core Disk Layout.md>) — Match the serial number and physical slot before replacement.
- [Related Virtualization and Storage Documentation](<../../../reference/Virtualization and Storage/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,19 @@
---
tags:
- Virtualization and Storage
- Workflows
- Documentation
---
# Virtualization and Storage
## Purpose
Find workflows for virtualization and storage. Follow the subject guide to choose the relevant environment and connect this material to the other document types.
## Includes
- Hyper-V
- Linux
- Proxmox
- TrueNAS
## Follow the Subject
[Virtualization and Storage](<../../reference/Virtualization and Storage/index.md>) explains the relationships and offers starting points for the documented tasks.