Restructured Documentation
Automatic Documentation Deployment / Sync Docs to https://kb.bunny-lab.io (push) Successful in 8s

This commit is contained in:
2026-09-05 14:08:43 -06:00
parent c4bd235eba
commit 289769a601
281 changed files with 5403 additions and 3563 deletions
@@ -0,0 +1,20 @@
---
tags:
- Rocket.Chat
- Communication
---
## Purpose
When someone types a message that includes a ticket number (e.g. `T00000000.0000`) we want to replace that text with an API-friendly URL that leverages Markdown language as well.
From RocketChat, navigate to the "Marketplace" and look for "**Word Replacer**". You can find the application's [GitHub Page](https://github.com/Dimsday/WordReplacer) for additional information / source code review. Proceed to install the application. Once it has been installed, use the following RegEx filter / string in the application's settings:
```json
[{"search": "T(\\d{8}\\.\\d{4})", "replace": "[$&](https://ww15.autotask.net/Autotask/AutotaskExtend/ExecuteCommand.aspx?Code=OpenTicketDetail&TicketNumber=$&)"}]
```
!!! success
Now everything should be functional and replacing ticket numbers with valid links that open the ticket in Autotask.
## Related Documentation
- [Related Applications Documentation](<../../../../reference/Applications/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,87 @@
---
tags:
- Microsoft Exchange
- Email
---
## Purpose
If you operate an Exchange Database Availability Group (DAG) with 2 or more servers, you may need to do maintenance to one of the members, and during that maintenance, it's possible that one of the databases of the server that was rebooted etc will be out-of-date. In case this happens, it may suspend the database replication to one of the DAG's member servers.
!!! warning "Exchange Version Context"
This page preserves the older database-copy and content-index repair examples. The Exchange SE rolling-update guide documents a different search architecture and service-validation process. Confirm the Exchange version and failure state before selecting a repair command.
## Checking DAG Database Replication Status
You will want to first log into one of the DAG servers and open the *"Exchange Management Shell"*. From there, run the following command to get the status of database replication. An example of the kind of output you would see is below the command.
```powershell
Get-MailboxDatabaseCopyStatus * | Format-Table Name, Status, CopyQueueLength, ReplayQueueLength, ContentIndexState
```
| **Name** | **Status** | **CopyQueueLength** | **ReplayQueueLength** | **ContentIndexState** |
| :--- | ---: | ---: | ---: | ---: |
| DB01\MX-DAG-01 | Mounted | 0 | 0 | Healthy |
| DB01\MX-DAG-02 | Healthy | 0 | 0 | Healthy |
!!! info "Example Output Breakdown"
In the above example output, you can see that there are two member servers in the DAG, `MX-DAG-01` and `MX-DAG-02`. Then you will see that there is a status of `Mounted`, this means that `MX-DAG-01` is the active production server; this means that it is handling all mailflow and web requests / webmail.
**CopyQueueLength**: This is a number of database "*transaction logs*" that have taken place since a replica database stopped getting updates. This is the queue of all database transactions that are being copied from the production (mounted) database to replica databases. This data is not immediately written to the replica database(s).
**CopyReplayLength**: This represents the queue of all data that was successfully copied from the production database to the replica database on the given DAG member that still needs to process on the replica database. The "**CopyQueueLength**" will need to reach zero before the "**CopyReplayLength**" will start making meaningful progress to reaching zero.
When both the "**CopyQueueLength**" and "**CopyReplayLength**" queues have reached zero, the replica database(s) will have reached 100% parity with the production (active/mounted) database.
## Changing Active/Mounted DAG Member
You may find that you need to perform work on one of the DAG members, and that requires you to failover the responsibility of hosting the Exchange environment to one of the other members of the DAG. You can generally do this with one command, seen below:
```powershell
Move-ActiveMailboxDatabase -Identity "DB01" -ActivateOnServer "MX-DAG-02" -MountDialOverride BestAvailability
```
!!! info "Argument Breakdown"
`-MountDialOverride`
Specifies how tolerant Exchange should be to database copy health when mounting a database on the target server. This setting controls the level of availability Exchange requires before mounting the mailbox database after the move.
`-MountDialOverride`
Instructs Exchange to mount the database as long as at least one healthy copy is available. This option maximizes uptime by allowing a database to mount even if some copies are unhealthy, prioritizing availability over strict health checks.
## Troubleshooting
You may run into issues where either the `Status` or `ContentIndexState` are either Unhealthy, Suspended, or Failed. If this happens, you need to resume replication of the database from the production active/mounted server to the server that is having issues. In the worst-case, you would re-seed the replica database from-scratch.
### If `Status` is Unhealthy or Suspended
If one of the DAG members has a status of "**Unhealthy**", you can run the following command to attempt to resume replication.
```powershell
Resume-MailboxDatabaseCopy -Identity "DB01\MX-DAG-02"
```
If this fails to cause replication to resume, you can try telling the database to just focus on replication, which tells it to copy the queues and replay them on the replica database, while avoiding interacting with the "**ContentIndexState**" which can be individually fixed in the commands below:
```powershell
Resume-MailboxDatabaseCopy -Identity "DB01\MX-DAG-02" -ReplicationOnly
```
### If `Status` is `ServiceDown`
If you see this, it generally means that the Exchange Services for some reason or another are not running. You can remediate this with a powershell script. You will then have to double-check your work to ensure that all "Microsoft Exchange" services that have a startup mode of "Automatic" are running, if not, manually start them, then check on the status of the DAG again to see if the status changes from `ServiceDown` to `Healthy`. Depending on the speed of the Exchange server, it may take a few minutes, 5-10 minutes, for the services to fully initialize and be ready to handle requests. Go get a coffee and come back and check on the status of the DAG at that time.
[:material-powershell: Start Exchange Services Script](<../../../../scripts/Applications/Email/Microsoft Exchange/Start Exchange Services.md>){ .md-button }
### If `ContentIndexState` is Unhealthy or Suspended
If you see that the "ContentIndexState" is unhappy, you can run the following command to force it to re-seed / rebuild itself. (This is non-destructive this this is happening on a replica database).
```powershell
Update-MailboxDatabaseCopy "DB01\MX05" -CatalogOnly -BeginSeed
```
### If Replica Database is FUBAR
If the replica database just is not playing nice, you can take the *nuclear option* of completely rebuilding the replica database.
!!! warning
This will destroy the replica database, so be careful to ensure you have a backup (if possible) before you do this. The following command will completely replace the replica database and replicate the data from the production active/mounted database to the newly-created replica database.
```powershell
Update-MailboxDatabaseCopy -Identity "DB01\MX-DAG-02" -SourceServer "MX-DAG-01"
```
## Related Documentation
- [Related Email Documentation](<../../../../reference/Applications/Email/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,933 @@
---
tags:
- Exchange Server
- Database Availability Group
- Maintenance
- Windows Server
- PowerShell
---
## Purpose
This workflow applies Exchange Server Subscription Edition (SE), Windows Server, and approved prerequisite updates to a three-member database availability group (DAG) by updating one DAG member at a time. The procedure drains client and transport activity, moves active mailbox databases, places the target member into maintenance mode, installs updates, validates the updated member, and restores the intended database placement before the next cycle begins.
All organization names, hostnames, FQDNs, DAG names, database names, and example values in this page describe the fictional Bunny Lab environment. The examples use the `bunny-lab.io` DNS namespace and do not identify another organization.
!!! warning "Update One DAG Member at a Time"
Only one DAG member may be in maintenance mode at a time. Do not begin the next cycle until the previous member has returned to service, all database copies are healthy, all copy and replay queues have drained, transport queues are clear, and replication health passes across the entire DAG.
## Assumptions and Risk Boundaries
- The Exchange organization is running Exchange Server SE on three Mailbox servers in one DAG.
- Every mailbox database has at least two healthy passive copies before maintenance begins.
- A current backup and a tested Exchange recovery path exist. DAG replication provides availability, but it is not a substitute for backup.
- The operator has the Exchange and local administrative permissions required by the selected update. A CU that extends the schema or prepares Active Directory may require additional directory permissions before the first server is upgraded.
- The administrative shell host remains online and is not the current maintenance target.
- The Exchange DAG maintenance scripts are available through `$ExScripts`, and the administrative shell host has the Failover Clustering management tools installed.
- The current Exchange release notes, prerequisites, known issues, and update-specific manual actions have been reviewed.
- The required Exchange update media and Windows updates are approved and staged before the maintenance window begins.
- Any load balancer, monitoring platform, backup platform, mail gateway, or third-party application that targets an individual Exchange server has an established drain and restore procedure.
- All DAG members are returned to the same Exchange CU, SU, and HU level during the rolling update window.
!!! danger "Do Not Use a Snapshot as the Only Recovery Plan"
Do not begin the rolling update without an Exchange-aware backup and documented recovery method. If an update fails, keep the affected server isolated in maintenance mode and repair that server before continuing to another DAG member.
## Example Bunny Lab Environment
### Exchange Topology
| **Object** | **Example Value** |
| :--- | :--- |
| Organization | `Bunny Lab` |
| Active Directory DNS domain | `bunny-lab.io` |
| DAG | `BL-DAG-01` |
| Exchange version | Exchange Server Subscription Edition |
| Update staging directory | `C:\ExchangeUpdates` |
| Support scripts directory | `C:\Scripts` |
### DAG Members
| **Short Name** | **FQDN** | **Normal Role** |
| :--- | :--- | :--- |
| `EXCH-SE-01` | `EXCH-SE-01.bunny-lab.io` | DAG member and primary administrative shell host |
| `EXCH-SE-02` | `EXCH-SE-02.bunny-lab.io` | DAG member and alternate administrative shell host |
| `EXCH-SE-03` | `EXCH-SE-03.bunny-lab.io` | DAG member |
Exchange cmdlets in this page normally use the Exchange server object name, such as `EXCH-SE-01`. Commands that require an FQDN, including `Redirect-Message -Target`, use the corresponding `bunny-lab.io` FQDN.
### Intended Database Placement
| **Active Database** | **Intended Active Server** | **Passive Copy Servers** |
| :--- | :--- | :--- |
| `BL-MBX-01` | `EXCH-SE-01` | `EXCH-SE-02`, `EXCH-SE-03` |
| `BL-ARC-01` | `EXCH-SE-01` | `EXCH-SE-02`, `EXCH-SE-03` |
| `BL-MBX-02` | `EXCH-SE-02` | `EXCH-SE-01`, `EXCH-SE-03` |
| `BL-ARC-02` | `EXCH-SE-02` | `EXCH-SE-01`, `EXCH-SE-03` |
| `BL-MBX-03` | `EXCH-SE-03` | `EXCH-SE-01`, `EXCH-SE-02` |
### Rolling Upgrade Plan
| **Cycle** | **Administrative Shell Host** | **Maintenance Target** | **Transport Redirect Target** | **Temporary Database Placement** |
| ---: | :--- | :--- | :--- | :--- |
| `1` | `EXCH-SE-01` | `EXCH-SE-03` | `EXCH-SE-01.bunny-lab.io` | `BL-MBX-03` to `EXCH-SE-01` |
| `2` | `EXCH-SE-01` | `EXCH-SE-02` | `EXCH-SE-03.bunny-lab.io` | `BL-MBX-02` to `EXCH-SE-01`; `BL-ARC-02` to `EXCH-SE-03` |
| `3` | `EXCH-SE-02` | `EXCH-SE-01` | `EXCH-SE-03.bunny-lab.io` | `BL-MBX-01` to `EXCH-SE-02`; `BL-ARC-01` to `EXCH-SE-03` |
This order is specific to the fictional topology above. In another environment, choose an order that preserves quorum, keeps an administrative shell host available, and distributes active databases across healthy remaining members.
## Prepare the Exchange Updates
### Determine the Required Exchange Build
From Exchange Management Shell on a healthy DAG member, record the base Exchange build stored in Active Directory:
```powershell
$DagMembers = @("EXCH-SE-01", "EXCH-SE-02", "EXCH-SE-03")
$DagMembers | ForEach-Object {
Get-ExchangeServer -Identity $_
} | Format-Table Name, Edition, AdminDisplayVersion -Auto
```
`AdminDisplayVersion` identifies the base release or CU, but it does not reliably identify the installed SU or HU. Query the local `ExSetup.exe` file on every DAG member to record the full installed file version:
```powershell
Invoke-Command -ComputerName $DagMembers -ScriptBlock {
$VersionInfo = (Get-Item (Join-Path $env:ExchangeInstallPath "bin\ExSetup.exe")).VersionInfo
[PSCustomObject]@{
Server = $env:COMPUTERNAME
ProductVersion = $VersionInfo.ProductVersion
FileVersion = $VersionInfo.FileVersion
}
} | Format-Table -Auto
```
Compare the results with the current Microsoft build table and the release article for the intended update. Install the latest supported CU required for the target release, then install the latest applicable SU or HU according to that release article. Do not install an SU or HU that was built for a different CU.
!!! warning "Release-Specific Instructions Take Precedence"
Review the selected update's prerequisites, Active Directory preparation requirements, known issues, and manual post-installation actions before changing the first DAG member. This generic workflow does not replace release-specific Microsoft instructions.
### Run the Exchange Health Checker
Download and stage the current Microsoft Exchange Health Checker script before the maintenance window. From an elevated PowerShell session on the administrative shell host, run it against every DAG member:
```powershell
$DagMembers | ForEach-Object {
& "C:\Scripts\HealthChecker.ps1" -Server $_
}
```
Resolve update-blocking findings before continuing. Preserve the generated reports with the maintenance record so the pre-update and post-update states can be compared.
### Stage Update Media
Create the local staging directory on every DAG member:
```powershell
Invoke-Command -ComputerName $DagMembers -ScriptBlock {
New-Item -Path "C:\ExchangeUpdates" -ItemType Directory -Force | Out-Null
}
```
Copy the approved Exchange CU media, SU or HU package, prerequisite installers, and any required scripts to `C:\ExchangeUpdates` on every DAG member. Use the exact package linked by the applicable Microsoft release article and verify that the file is complete before the maintenance window begins.
## Select the Current Upgrade Cycle
Set these variables in Exchange Management Shell on the administrative shell host before running the common workflow. Run only the block for the current cycle.
### Cycle 1: Update `EXCH-SE-03`
Run from Exchange Management Shell on `EXCH-SE-01`:
```powershell
$DagName = "BL-DAG-01"
$DagMembers = @("EXCH-SE-01", "EXCH-SE-02", "EXCH-SE-03")
$AdminHost = "EXCH-SE-01"
$TargetServer = "EXCH-SE-03"
$RedirectTargetFqdn = "EXCH-SE-01.bunny-lab.io"
```
### Cycle 2: Update `EXCH-SE-02`
Run from Exchange Management Shell on `EXCH-SE-01` after Cycle 1 is fully validated:
```powershell
$DagName = "BL-DAG-01"
$DagMembers = @("EXCH-SE-01", "EXCH-SE-02", "EXCH-SE-03")
$AdminHost = "EXCH-SE-01"
$TargetServer = "EXCH-SE-02"
$RedirectTargetFqdn = "EXCH-SE-03.bunny-lab.io"
```
### Cycle 3: Update `EXCH-SE-01`
Run from Exchange Management Shell on `EXCH-SE-02` after Cycle 2 is fully validated:
```powershell
$DagName = "BL-DAG-01"
$DagMembers = @("EXCH-SE-01", "EXCH-SE-02", "EXCH-SE-03")
$AdminHost = "EXCH-SE-02"
$TargetServer = "EXCH-SE-01"
$RedirectTargetFqdn = "EXCH-SE-03.bunny-lab.io"
```
Confirm that the current Exchange Management Shell session is running on `$AdminHost` and that `$AdminHost` is not equal to `$TargetServer`:
```powershell
[PSCustomObject]@{
CurrentComputer = $env:COMPUTERNAME
AdminHost = $AdminHost
TargetServer = $TargetServer
}
```
Do not continue if `CurrentComputer` does not match `AdminHost`, or if the administrative shell host is not fully healthy.
### Capture a Pre-Update Service Baseline
Capture the Exchange-related service state on each server before the first maintenance action. Run this block once for the current target server from an elevated PowerShell session:
```powershell
Invoke-Command -ComputerName $TargetServer -ScriptBlock {
Get-CimInstance Win32_Service |
Where-Object { $_.Name -like "MSExchange*" -or $_.Name -eq "FMS" } |
Select-Object Name, DisplayName, State, StartMode, ExitCode |
Export-Csv -Path "C:\ExchangeUpdates\PreUpdate-Services.csv" -NoTypeInformation
}
```
Do not replace this baseline with a service list copied from another organization. Exchange service startup modes can differ because of installed roles, enabled protocols, product version, and local design decisions.
## Validate DAG Health Before Each Cycle
### Check Database Copy Health
From the administrative shell host, check every database copy:
```powershell
Get-MailboxDatabaseCopyStatus * |
Sort-Object Name |
Format-Table Name, Status, CopyQueueLength, ReplayQueueLength, ContentIndexState -Auto
```
The preflight passes only when:
- Every active copy reports `Mounted`
- Every passive copy reports `Healthy`
- No copy reports `Failed`, `Suspended`, `Disconnected`, `ServiceDown`, or another unresolved failure state
- Every copy queue is `0`
- Every replay queue is `0`, or is low and demonstrably draining before any state-changing action
- `ContentIndexState` matches the expected Exchange SE state; `NotApplicable` is normal for the BigFunnel search architecture and is not, by itself, a failure
### Check Replication Health
Run replication health against every DAG member:
```powershell
$DagMembers | ForEach-Object {
Test-ReplicationHealth -Identity $_
}
```
Every applicable check must return `Passed`, and the `Error` field must be blank. Investigate any failure before moving a database or entering maintenance mode.
### Check Cluster State and Primary Active Manager
Confirm every cluster node is up and identify the current Primary Active Manager (PAM):
```powershell
Get-ClusterNode | Format-Table Name, State, NodeWeight, DynamicWeight -Auto
Get-DatabaseAvailabilityGroup -Identity $DagName -Status |
Format-List Name, PrimaryActiveManager, OperationalServers
```
All DAG members must be operational before the cycle begins. The maintenance script will move critical DAG functionality away from the target and pause its cluster node.
### Check Exchange Service Health
Check Exchange service health on the target and the administrative shell host:
```powershell
Test-ServiceHealth -Server $TargetServer
Test-ServiceHealth -Server $AdminHost
```
Resolve any required service failure before continuing.
### Check Active Database Placement
List the currently mounted copies:
```powershell
Get-MailboxDatabaseCopyStatus * |
Where-Object { $_.Status -eq "Mounted" } |
Sort-Object ActiveDatabaseCopy, Name |
Format-Table Name, Status, ActiveDatabaseCopy, CopyQueueLength, ReplayQueueLength -Auto
```
Compare the result with the intended placement table. A database may be temporarily active on another healthy member because of an earlier event, but its current state and all candidate copies must be understood before the maintenance move begins.
### Check Transport Queues
Inspect transport queues on the target before draining it:
```powershell
Get-Queue -Server $TargetServer |
Sort-Object MessageCount -Descending |
Format-Table Identity, Status, MessageCount, NextHopDomain -Auto
```
Investigate growing, retrying, or unreachable queues before maintenance. Redirecting a queue does not correct an underlying transport or name-resolution failure.
!!! warning "Preflight Stop Conditions"
Stop the cycle if any database copy is failed or suspended, replication health does not pass, a required Exchange service is unavailable, a cluster node is down, quorum is at risk, transport queues are persistently growing, the update prerequisites are unresolved, or the recovery path is unavailable.
## Move Active Databases Away From the Target
### Confirm Candidate Copies
Before each move, confirm that the chosen destination server holds a healthy passive copy with drained queues:
```powershell
Get-MailboxDatabaseCopyStatus * |
Sort-Object Name |
Format-Table Name, Status, CopyQueueLength, ReplayQueueLength, ActiveDatabaseCopy -Auto
```
Confirm the selected destination copy reports `Healthy` with `CopyQueueLength` and `ReplayQueueLength` equal to `0` before moving the active database.
!!! warning "Run Only the Current Cycle's Move Block"
The following move blocks are cycle-specific. Do not run move commands for a different maintenance target.
### Cycle 1 Database Move
Move `BL-MBX-03` from `EXCH-SE-03` to `EXCH-SE-01`:
```powershell
Move-ActiveMailboxDatabase -Identity "BL-MBX-03" -ActivateOnServer "EXCH-SE-01" -Confirm:$false
```
### Cycle 2 Database Moves
Distribute the two active databases from `EXCH-SE-02` across the remaining healthy members:
```powershell
Move-ActiveMailboxDatabase -Identity "BL-MBX-02" -ActivateOnServer "EXCH-SE-01" -Confirm:$false
Move-ActiveMailboxDatabase -Identity "BL-ARC-02" -ActivateOnServer "EXCH-SE-03" -Confirm:$false
```
### Cycle 3 Database Moves
Distribute the two active databases from `EXCH-SE-01` across the remaining healthy members:
```powershell
Move-ActiveMailboxDatabase -Identity "BL-MBX-01" -ActivateOnServer "EXCH-SE-02" -Confirm:$false
Move-ActiveMailboxDatabase -Identity "BL-ARC-01" -ActivateOnServer "EXCH-SE-03" -Confirm:$false
```
### Validate the Database Moves
Confirm that no active database remains on the target:
```powershell
Get-MailboxDatabaseCopyStatus * |
Where-Object { $_.Status -eq "Mounted" -and $_.Name -like "*\$TargetServer" } |
Format-Table Name, Status, ActiveDatabaseCopy, CopyQueueLength, ReplayQueueLength -Auto
```
Expected output:
```text
<no output>
```
Review each move result and confirm that `Status` is `Succeeded`, `NumberOfLogsLost` is `0`, and `MountStatusAtMoveEnd` is `Mounted`. If any move fails or reports log loss, stop the cycle and investigate before entering maintenance mode.
## Place the Target DAG Member Into Maintenance Mode
### Drain External Client Traffic
If the environment uses a load balancer, reverse proxy, monitoring probe, backup job, or third-party connector that targets individual Exchange servers, drain or disable the target member according to that platform's documented procedure. Confirm that healthy remaining members are serving the traffic before continuing.
### Drain Hub Transport
From Exchange Management Shell on the administrative shell host, set the target Hub Transport component to draining:
```powershell
Set-ServerComponentState -Identity $TargetServer -Component HubTransport -State Draining -Requester Maintenance
```
Restart the Microsoft Exchange Transport service on the target to initiate queue draining:
```powershell
Invoke-Command -ComputerName $TargetServer -ScriptBlock {
Restart-Service MSExchangeTransport
}
```
### Run the DAG Maintenance Script
From Exchange Management Shell on the administrative shell host, run the Exchange-provided DAG maintenance script:
```powershell
Set-Location $ExScripts
.\StartDagServerMaintenance.ps1 -ServerName $TargetServer -MoveComment "Rolling Exchange update" -PauseClusterNode
```
The script blocks database activation, pauses the target cluster node, moves any remaining active databases, and moves critical DAG functionality away from the target. The command can take time without producing continuous console output.
### Redirect Pending Transport Messages
Redirect messages still pending on the target to the healthy server selected for the current cycle:
```powershell
Redirect-Message -Server $TargetServer -Target $RedirectTargetFqdn -Confirm:$false
```
### Set the Server-Wide Maintenance State
Place the target server into Exchange server-wide maintenance mode:
```powershell
Set-ServerComponentState -Identity $TargetServer -Component ServerWideOffline -State Inactive -Requester Maintenance
```
### Validate Maintenance Mode
Check effective Exchange component states:
```powershell
Get-ServerComponentState -Identity $TargetServer |
Format-Table Component, State -Auto
```
`ServerWideOffline` must report `Inactive`. In the standard maintenance state, only `Monitoring` and `RecoveryActionsEnabled` should remain `Active`; investigate any other component that remains active before rebooting the server.
Confirm the database activation policy is blocked:
```powershell
Get-MailboxServer -Identity $TargetServer |
Format-List Name, DatabaseCopyAutoActivationPolicy
```
Confirm that the target cluster node is paused:
```powershell
Get-ClusterNode -Name $TargetServer |
Format-List Name, State
```
Confirm that no active databases remain on the target:
```powershell
Get-MailboxDatabaseCopyStatus * |
Where-Object { $_.Status -eq "Mounted" -and $_.Name -like "*\$TargetServer" } |
Format-Table Name, Status, ActiveDatabaseCopy -Auto
```
Expected output:
```text
<no output>
```
Confirm that the target transport queues have drained:
```powershell
Get-Queue -Server $TargetServer |
Sort-Object MessageCount -Descending |
Format-Table Identity, Status, MessageCount, NextHopDomain -Auto
```
Confirm that the PAM is hosted by another DAG member:
```powershell
Get-DatabaseAvailabilityGroup -Identity $DagName -Status |
Format-List PrimaryActiveManager, OperationalServers
```
!!! warning "Do Not Reboot Until Every Maintenance Gate Passes"
Do not install updates or reboot the target until it has no mounted databases, its database activation policy is `Blocked`, its cluster node is `Paused`, `ServerWideOffline` is `Inactive`, its transport queues are drained, and the PAM is hosted by another member.
## Install Exchange and Windows Updates
### Reboot Before Installing the Exchange Update
A clean reboot before Exchange Setup or an Exchange SU or HU reduces failures caused by pending file handles and services that do not stop cleanly. From the administrative shell host, reboot the target:
```powershell
Restart-Computer -ComputerName $TargetServer -Force
```
Wait until the server is reachable, then revalidate that maintenance state persisted:
```powershell
Get-ServerComponentState -Identity $TargetServer |
Where-Object { $_.Component -eq "ServerWideOffline" } |
Format-Table Server, Component, State -Auto
Get-ClusterNode -Name $TargetServer |
Format-List Name, State
Get-MailboxDatabaseCopyStatus * |
Where-Object { $_.Status -eq "Mounted" -and $_.Name -like "*\$TargetServer" }
```
Do not launch the update if `ServerWideOffline` is not `Inactive`, the cluster node is not `Paused`, or a database mounted on the target after reboot.
### Install an Exchange CU or Build Upgrade
Mount the correct Exchange installation media on the target server. From an elevated Command Prompt on the target, run Setup by using the full media path so Windows does not invoke the installed `Setup.exe` from the Exchange binary directory.
=== "Diagnostic Data Off"
```cmd
D:\Setup.exe /Mode:Upgrade /IAcceptExchangeServerLicenseTerms_DiagnosticDataOFF
```
=== "Diagnostic Data On"
```cmd
D:\Setup.exe /Mode:Upgrade /IAcceptExchangeServerLicenseTerms_DiagnosticDataON
```
Replace `D:` with the actual mounted-media drive. Wait for Exchange Setup to complete successfully and review `C:\ExchangeSetupLogs\ExchangeSetup.log` before continuing.
!!! danger "A CU Upgrade Is Not Reversible by Uninstalling It"
Do not treat a CU as an ordinary removable patch. If the CU fails, keep the server in maintenance mode and repair or recover that server by following Microsoft Exchange recovery guidance.
### Install an Exchange SU or HU
Run the exact update package specified by the applicable release article from an elevated Command Prompt on the target server:
```cmd
C:\ExchangeUpdates\<EXCHANGE_UPDATE_FILENAME>.exe
```
Replace `<EXCHANGE_UPDATE_FILENAME>` with the staged package name. Do not launch the installer from a non-elevated shell or by double-clicking it in File Explorer.
Do not stop IIS, WMI, Windows Event Log, or other Windows services preemptively unless the release article or a matching Microsoft troubleshooting procedure explicitly instructs you to do so. Wait for the installer to report successful completion before continuing.
### Install Approved Windows Updates
Install the approved Windows Server updates while the target remains in maintenance mode. Apply prerequisites in the order required by the Exchange release notes, and continue rebooting as required until the server has no remaining approved updates or pending restart.
### Perform the Final Maintenance Reboot
After Exchange and Windows updates have completed, reboot the target one final time:
```powershell
Restart-Computer -ComputerName $TargetServer -Force
```
Wait for Windows, Active Directory connectivity, the Cluster service, and Exchange services to initialize before beginning return-to-service validation.
## Return the Updated DAG Member to Service
### Verify the Installed Exchange Build
From Exchange Management Shell on the administrative shell host, confirm the base build:
```powershell
Get-ExchangeServer -Identity $TargetServer |
Format-Table Name, Edition, AdminDisplayVersion -Auto
```
Query the local `ExSetup.exe` file on the target to identify the installed SU or HU file version:
```powershell
Invoke-Command -ComputerName $TargetServer -ScriptBlock {
$VersionInfo = (Get-Item (Join-Path $env:ExchangeInstallPath "bin\ExSetup.exe")).VersionInfo
[PSCustomObject]@{
Server = $env:COMPUTERNAME
ProductVersion = $VersionInfo.ProductVersion
FileVersion = $VersionInfo.FileVersion
}
}
```
Compare both values with the intended Microsoft build. An unchanged `AdminDisplayVersion` does not, by itself, prove that an SU or HU failed to install.
### Validate Exchange Services
From the administrative shell host, run:
```powershell
Test-ServiceHealth -Server $TargetServer
```
Inspect Exchange-related services on the target:
```powershell
Invoke-Command -ComputerName $TargetServer -ScriptBlock {
Get-CimInstance Win32_Service |
Where-Object { $_.Name -like "MSExchange*" -or $_.Name -eq "FMS" } |
Select-Object Name, DisplayName, State, StartMode, ExitCode
} | Format-Table -Auto
```
Compare the result with `C:\ExchangeUpdates\PreUpdate-Services.csv` from that same server. Do not enable a service merely because it is stopped; protocol services and role-specific services may intentionally be manual or stopped.
!!! warning "Keep the Server in Maintenance Mode During Repair"
If required Exchange services are disabled, fail to start, or report dependency errors, keep the target in maintenance mode. Repair the service state and complete the failed update before running the return-to-service commands.
### Remove the Server-Wide Maintenance State
From Exchange Management Shell on the administrative shell host, restore the server-wide component state:
```powershell
Set-ServerComponentState -Identity $TargetServer -Component ServerWideOffline -State Active -Requester Maintenance
```
### Run the DAG Return-to-Service Script
Run the Exchange-provided DAG maintenance exit script:
```powershell
Set-Location $ExScripts
.\StopDagServerMaintenance.ps1 -ServerName $TargetServer
```
The script resumes the cluster node, restores database activation policy, and resumes database copies hosted by the member.
### Resume Hub Transport
Return Hub Transport to active state:
```powershell
Set-ServerComponentState -Identity $TargetServer -Component HubTransport -State Active -Requester Maintenance
```
Restart transport on the target:
```powershell
Invoke-Command -ComputerName $TargetServer -ScriptBlock {
Restart-Service MSExchangeTransport
}
```
### Validate the Restored Member
Confirm Exchange component states:
```powershell
Get-ServerComponentState -Identity $TargetServer |
Format-Table Component, State -Auto
```
Confirm database activation is unrestricted:
```powershell
Get-MailboxServer -Identity $TargetServer |
Format-List Name, DatabaseCopyAutoActivationPolicy
```
Confirm the cluster node is up:
```powershell
Get-ClusterNode -Name $TargetServer |
Format-List Name, State
```
Run service and replication health:
```powershell
Test-ServiceHealth -Server $TargetServer
Test-ReplicationHealth -Identity $TargetServer
```
Inspect transport queues:
```powershell
Get-Queue -Server $TargetServer |
Sort-Object MessageCount -Descending |
Format-Table Identity, Status, MessageCount, NextHopDomain -Auto
```
Do not restore the target to a load balancer or other external traffic source until required Exchange components are active, required services are running, the cluster node is up, replication health passes, and transport queues are processing normally.
### Restore External Client Traffic
Return the target to its load balancer pool, monitoring platform, backup schedule, mail gateway, and any other system that was intentionally drained. Confirm the external health checks recognize the server as healthy before restoring active mailbox databases to it.
## Restore the Intended Database Placement
### Wait for Database Copies to Catch Up
Before moving an active database back to the updated member, confirm every copy is healthy and all queues have drained:
```powershell
Get-MailboxDatabaseCopyStatus * |
Sort-Object Name |
Format-Table Name, Status, CopyQueueLength, ReplayQueueLength, ContentIndexState -Auto
```
Do not activate a copy on the updated member while it is `Failed`, `Suspended`, `Disconnected`, `ServiceDown`, or building a persistent queue.
!!! warning "Run Only the Current Cycle's Restore Block"
The following restore blocks are cycle-specific. Do not move databases for a different cycle.
### Cycle 1 Database Restore
Move `BL-MBX-03` back to `EXCH-SE-03`:
```powershell
Move-ActiveMailboxDatabase -Identity "BL-MBX-03" -ActivateOnServer "EXCH-SE-03" -Confirm:$false
```
### Cycle 2 Database Restore
Move both databases back to `EXCH-SE-02`:
```powershell
Move-ActiveMailboxDatabase -Identity "BL-MBX-02" -ActivateOnServer "EXCH-SE-02" -Confirm:$false
Move-ActiveMailboxDatabase -Identity "BL-ARC-02" -ActivateOnServer "EXCH-SE-02" -Confirm:$false
```
### Cycle 3 Database Restore
Move both databases back to `EXCH-SE-01`:
```powershell
Move-ActiveMailboxDatabase -Identity "BL-MBX-01" -ActivateOnServer "EXCH-SE-01" -Confirm:$false
Move-ActiveMailboxDatabase -Identity "BL-ARC-01" -ActivateOnServer "EXCH-SE-01" -Confirm:$false
```
### Validate the Restored Placement
Confirm active database placement:
```powershell
Get-MailboxDatabaseCopyStatus * |
Where-Object { $_.Status -eq "Mounted" } |
Sort-Object ActiveDatabaseCopy, Name |
Format-Table Name, Status, ActiveDatabaseCopy, CopyQueueLength, ReplayQueueLength -Auto
```
Expected Bunny Lab placement:
```text
Name Status ActiveDatabaseCopy CopyQueueLength ReplayQueueLength
---- ------ ------------------ --------------- -----------------
BL-ARC-01\EXCH-SE-01 Mounted EXCH-SE-01 0 0
BL-MBX-01\EXCH-SE-01 Mounted EXCH-SE-01 0 0
BL-ARC-02\EXCH-SE-02 Mounted EXCH-SE-02 0 0
BL-MBX-02\EXCH-SE-02 Mounted EXCH-SE-02 0 0
BL-MBX-03\EXCH-SE-03 Mounted EXCH-SE-03 0 0
```
Temporary queues may appear immediately after activation. Wait until all copy and replay queues return to `0` before declaring the cycle complete.
## Validate the Completed Cycle
Run replication health against every DAG member:
```powershell
$DagMembers | ForEach-Object {
Test-ReplicationHealth -Identity $_
}
```
Check all database copies:
```powershell
Get-MailboxDatabaseCopyStatus * |
Sort-Object Name |
Format-Table Name, Status, CopyQueueLength, ReplayQueueLength, ContentIndexState -Auto
```
Run the Exchange Health Checker against the updated member:
```powershell
& "C:\Scripts\HealthChecker.ps1" -Server $TargetServer
```
The cycle is complete only when:
- The target reports the intended Exchange build
- Required Exchange services pass `Test-ServiceHealth`
- The target cluster node is `Up`
- Database activation policy is `Unrestricted`
- Required Exchange components are active
- Every applicable replication-health check returns `Passed`
- Every active database copy reports `Mounted`
- Every passive database copy reports `Healthy`
- Every copy and replay queue is `0`
- Transport queues are processing normally
- The intended active database placement is restored
- External health checks and monitoring report the target as healthy
- The post-update Health Checker report contains no unresolved update-blocking finding
Do not begin the next cycle until all conditions are satisfied. Repeat the common workflow with the next cycle's variables and database move block.
## Final DAG Validation
### Confirm a Consistent Exchange Build
After all three cycles are complete, confirm the base build recorded for every DAG member:
```powershell
$DagMembers | ForEach-Object {
Get-ExchangeServer -Identity $_
} | Format-Table Name, Edition, AdminDisplayVersion -Auto
```
Confirm the local `ExSetup.exe` file version on every member:
```powershell
Invoke-Command -ComputerName $DagMembers -ScriptBlock {
$VersionInfo = (Get-Item (Join-Path $env:ExchangeInstallPath "bin\ExSetup.exe")).VersionInfo
[PSCustomObject]@{
Server = $env:COMPUTERNAME
ProductVersion = $VersionInfo.ProductVersion
FileVersion = $VersionInfo.FileVersion
}
} | Sort-Object Server | Format-Table -Auto
```
All three servers must report the same intended Exchange build unless Microsoft explicitly documents a temporary mixed-build state during the active maintenance window.
### Validate Database Health and Placement
Run:
```powershell
Get-MailboxDatabaseCopyStatus * |
Sort-Object Name |
Format-Table Name, Status, CopyQueueLength, ReplayQueueLength, ContentIndexState -Auto
Get-MailboxDatabaseCopyStatus * |
Where-Object { $_.Status -eq "Mounted" } |
Sort-Object ActiveDatabaseCopy, Name |
Format-Table Name, Status, ActiveDatabaseCopy, CopyQueueLength, ReplayQueueLength -Auto
```
Confirm every active copy is `Mounted`, every passive copy is `Healthy`, all queues are `0`, and placement matches the Bunny Lab topology table.
### Validate Replication, Services, Components, and Cluster State
Run:
```powershell
$DagMembers | ForEach-Object {
Test-ServiceHealth -Server $_
Test-ReplicationHealth -Identity $_
}
Get-ClusterNode |
Format-Table Name, State, NodeWeight, DynamicWeight -Auto
$DagMembers | ForEach-Object {
Get-ServerComponentState -Identity $_ |
Where-Object { $_.State -ne "Active" } |
Select-Object Server, Component, State
}
```
Investigate every unexpected non-active component state. Protocol components that are deliberately disabled must match the documented Bunny Lab design rather than being assumed healthy merely because they were disabled before the update.
### Validate Mail Flow
From Exchange Management Shell on each source server, test mail flow to another DAG member. The following examples cover all three members:
```powershell
Test-Mailflow -Identity "EXCH-SE-01" -TargetMailboxServer "EXCH-SE-02"
Test-Mailflow -Identity "EXCH-SE-02" -TargetMailboxServer "EXCH-SE-03"
Test-Mailflow -Identity "EXCH-SE-03" -TargetMailboxServer "EXCH-SE-01"
```
Each test must return `Success`. Also validate inbound and outbound mail flow through the organization's real mail gateways and confirm the client-access namespaces used by the environment are healthy.
### Run the Final Health Checker Audit
Run the current Exchange Health Checker against every DAG member:
```powershell
$DagMembers | ForEach-Object {
& "C:\Scripts\HealthChecker.ps1" -Server $_
}
```
Archive the final reports with the change record.
## Troubleshooting
### The Installer Reports Files in Use or Cannot Stop Services
Do not select an installer option that ignores locked files, and do not terminate the Windows Event Log service. Exit the installer, confirm the server remains in maintenance mode, reboot it, and rerun the update from an elevated Command Prompt.
If the problem persists:
- Confirm the update package matches the installed CU
- Confirm the server was rebooted immediately before the installation attempt
- Review the selected update's known issues and antivirus exclusion guidance
- Preserve `C:\ExchangeSetupLogs`
- Use the Microsoft SetupAssist or Setup Log Reviewer tooling appropriate to the failure
- Follow the matching procedure in [Fix Failed Exchange Server Updates](https://learn.microsoft.com/en-us/troubleshoot/exchange/client-connectivity/exchange-security-update-issues)
Do not reuse process IDs from an earlier attempt or another server. Process IDs are transient, and terminating an unidentified Windows or Exchange process can leave the installation in a worse state.
### Exchange Services Are Disabled After the Update
Compare the current service configuration with `C:\ExchangeUpdates\PreUpdate-Services.csv` from the same server:
```powershell
Invoke-Command -ComputerName $TargetServer -ScriptBlock {
$Baseline = Import-Csv "C:\ExchangeUpdates\PreUpdate-Services.csv"
$Current = Get-CimInstance Win32_Service |
Where-Object { $_.Name -like "MSExchange*" -or $_.Name -eq "FMS" }
foreach ($Before in $Baseline) {
$After = $Current | Where-Object { $_.Name -eq $Before.Name }
if ($After -and ($After.State -ne $Before.State -or $After.StartMode -ne $Before.StartMode)) {
[PSCustomObject]@{
Name = $Before.Name
BeforeState = $Before.State
AfterState = $After.State
BeforeStartMode = $Before.StartMode
AfterStartMode = $After.StartMode
}
}
}
} | Format-Table -Auto
```
Restore only a service whose required startup mode is confirmed by the server's baseline, current Exchange role, and Microsoft guidance:
```powershell
Set-Service -Name <SERVICE_NAME> -StartupType Automatic
Start-Service -Name <SERVICE_NAME>
```
If the service fails with dependency error `1068`, inspect its required services:
```powershell
Get-Service -Name <SERVICE_NAME> -RequiredServices |
Format-Table Name, DisplayName, Status, StartType -Auto
```
Correct the failed dependency before retrying the dependent service. Do not automatically enable IMAP, POP, EdgeSync, or another optional service that was intentionally disabled before the update.
### An Exchange Component Remains Inactive
Inspect the effective and requester-specific component states:
```powershell
Get-ServerComponentState -Identity $TargetServer |
Format-Table Component, State -Auto
Get-ServerComponentState -Identity $TargetServer -Component <COMPONENT_NAME> |
Format-List Component, State, LocalStates, RemoteStates
```
Correct the requester that actually holds the component inactive. Do not repeatedly issue `Requester Maintenance` commands when the inactive state belongs to `Functional`, `HealthAPI`, or another requester.
If a failed Exchange update left `ServerWideOffline`, `Monitoring`, or `RecoveryActionsEnabled` inactive under the `Functional` requester, and the failed installation has already been repaired, restore only the affected states:
```powershell
Set-ServerComponentState -Identity $TargetServer -Component ServerWideOffline -State Active -Requester Functional
Set-ServerComponentState -Identity $TargetServer -Component Monitoring -State Active -Requester Functional
Set-ServerComponentState -Identity $TargetServer -Component RecoveryActionsEnabled -State Active -Requester Functional
```
Rerun `Test-ServiceHealth` and `Test-ReplicationHealth` after the correction.
### High Availability Remains Offline After Maintenance Ends
If a database move fails with an error stating that the `HighAvailability` component is offline, inspect its requester-specific state:
```powershell
Get-ServerComponentState -Identity $TargetServer -Component HighAvailability |
Format-List Component, State, LocalStates, RemoteStates
```
Confirm `StopDagServerMaintenance.ps1` completed successfully, the cluster node is `Up`, database activation is `Unrestricted`, and `MSExchangeRepl` is running. If every requester reports `Active` but Active Manager state remains stale, restart the replication service on the target:
```powershell
Invoke-Command -ComputerName $TargetServer -ScriptBlock {
Restart-Service MSExchangeRepl
Get-Service MSExchangeRepl
}
```
Rerun replication health and retry the database move only after every applicable check passes.
### Exchange Setup Reports Success but the Build Does Not Change
Confirm Setup was launched from the mounted media by using an absolute path such as `D:\Setup.exe` or, from PowerShell in the media root, `.\Setup.exe`. Running only `Setup.exe` can invoke the installed copy under the Exchange binary path instead of the intended installation media.
Review `C:\ExchangeSetupLogs\ExchangeSetup.log`, correct the launch path, and rerun Setup while the server remains in maintenance mode.
### Outlook on the Web or the Exchange Admin Center Fails After the Update
First confirm the Exchange update completed successfully and required services are running. Review the matching symptom in Microsoft's failed-update guidance rather than running a generic post-update command sequence.
For an IIS state that specifically requires a service restart, run from an elevated PowerShell session on the affected server:
```powershell
Restart-Service -Name WAS, W3SVC
```
Do not make `UpdateCas.ps1`, `UpdateConfigFiles.ps1`, `IISADMIN`, ADSI changes, or forced process termination part of the normal update workflow. Use a repair command only when an authoritative troubleshooting procedure identifies the same failure condition and explains the required validation.
## Recovery and Stop Conditions
If a target server cannot be returned to service:
- Keep `ServerWideOffline` inactive and keep the cluster node paused
- Keep database activation blocked on the failed member
- Leave active databases on healthy DAG members
- Do not begin maintenance on another member
- Preserve Exchange Setup logs, Windows event logs, Health Checker reports, and the pre-update service baseline
- Repair the update or perform the documented Exchange server recovery procedure
- Revalidate quorum, database redundancy, transport, client access, and backup status before resuming the rolling upgrade
The rolling update is complete only after all three members report the intended Exchange build, every required health check passes, the documented database placement is restored, and Bunny Lab mail flow and client access are validated end to end.
## Reference Documentation
- [Manage Database Availability Groups in Exchange Server](https://learn.microsoft.com/en-us/exchange/high-availability/manage-ha/manage-dags)
- [Upgrade Exchange to the Latest Cumulative Update](https://learn.microsoft.com/en-us/exchange/plan-and-deploy/install-cumulative-updates)
- [Use Unattended Mode in Exchange Setup](https://learn.microsoft.com/en-us/exchange/plan-and-deploy/deploy-new-installations/unattended-installs)
- [Exchange Server Build Numbers and Release Dates](https://learn.microsoft.com/en-us/exchange/new-features/build-numbers-and-release-dates)
- [Exchange Server Update FAQ](https://learn.microsoft.com/en-us/exchange/plan-and-deploy/post-installation-tasks/security-best-practices/exchange-server-update-faq)
- [Fix Failed Exchange Server Updates](https://learn.microsoft.com/en-us/troubleshoot/exchange/client-connectivity/exchange-security-update-issues)
- [Exchange Server Health Checker](https://microsoft.github.io/CSS-Exchange/Diagnostics/HealthChecker/)
## Related Documentation
- [Related Email Documentation](<../../../../reference/Applications/Email/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,126 @@
---
tags:
- Microsoft Exchange
- Email
---
## Purpose
This document is meant to be an abstract guide on what to do before installing Cumulative Updates on Microsoft Exchange Server. There are a few considerations that need to be made ahead of time. This list was put together through shere brute-force while troubleshooting an update issue for a server on 12/16/2024.
!!! abstract "Overview"
We are looking to add an administrative user to several domain security groups, adjust local security policy to put them into the "Manage Auditing and Security Logs" security policy, and run the setup.exe included on the Cumulative Update ISO images within a `SeSecurityPrivilege` operational context.
## Domain Group Membership
You have to be logged in with a domain user that possesses the following domain group memberships, if these group memberships are missing, the upgrade process will fail.
- `Enterprise Admins`
- `Schema Admins`
- `Organization Management`
## User Rights Management
You have to be part of the "**Local Policies > User Rights Assignment > "Manage Auditing and Security Logs**" security policy. You can set this via group policy management or locally on the Exchange server via `secpol.msc`. This is required for the "Monitoring Tools" portion of the upgrade.
It's recommended to reboot the server after making this change to be triple-sure that everything was applied correctly.
!!! note "Security Policy Only Required on Exchange Server"
While the `Enterprise Admins`, `Schema Admins`, and `Organization Management` security group memberships are required on a domain-wide level, the security policy membership for "Manage Auditing and Security Logs" mentioned above is only required on the Exchange Server itself. You can create a group policy that only targets the Exchange Server to add this, or you can make your user a domain-wide member of "Manage Auditing and Security Logs" (Optional). If no existing policies are in-place affecting the Exchange server, you can just use `secpol.msc` to manually add your user to this security policy for the duration of the upgrade/update (or leave it there for future updates).
## Running Updater within `SeSecurityPrivilege` Operational Context
At this point, you would technically be ready to invoke `setup.exe` on the Cumulative Update ISO image to launch the upgrade process, but we are going to go the extra mile to manually "Enable" the `SeSecurityPrivilege` within a Powershell session, then use that same session to invoke the `setup.exe` so the updater runs within that context. This is not really necessary, but something I added as a "hail mary" to make the upgrade successful.
### Open Powershell ISE
The first thing we are going to do, is open the Powershell ISE so we can copy/paste the following powershell script, this script will explicitely enable `SeSecurityPrivilege` for anyone who holds that privilege within the powershell session.
!!! warning "Run Powershell ISE as Administrator"
In order for everything to work correctly, the ISE has to be launched by right-clicking "Run as Administrator", otherwise it is guarenteed that the updater application will fail at some point.
```powershell title="SeSecurityPrivilege Enablement Script"
# Create a Privilege Adjustment
$definition = @"
using System;
using System.Runtime.InteropServices;
public class Privilege
{
const int SE_PRIVILEGE_ENABLED = 0x00000002;
const int TOKEN_ADJUST_PRIVILEGES = 0x0020;
const int TOKEN_QUERY = 0x0008;
const string SE_SECURITY_NAME = "SeSecurityPrivilege";
[DllImport("advapi32.dll", SetLastError = true)]
public static extern bool OpenProcessToken(IntPtr ProcessHandle, int DesiredAccess, out IntPtr TokenHandle);
[DllImport("advapi32.dll", SetLastError = true, CharSet = CharSet.Unicode)]
public static extern bool LookupPrivilegeValue(string lpSystemName, string lpName, out long lpLuid);
[DllImport("advapi32.dll", SetLastError = true)]
public static extern bool AdjustTokenPrivileges(IntPtr TokenHandle, bool DisableAllPrivileges, ref TOKEN_PRIVILEGES NewState, int BufferLength, IntPtr PreviousState, IntPtr ReturnLength);
[StructLayout(LayoutKind.Sequential, Pack = 1)]
public struct TOKEN_PRIVILEGES
{
public int PrivilegeCount;
public long Luid;
public int Attributes;
}
public static bool EnablePrivilege()
{
IntPtr tokenHandle;
TOKEN_PRIVILEGES tokenPrivileges;
if (!OpenProcessToken(System.Diagnostics.Process.GetCurrentProcess().Handle, TOKEN_ADJUST_PRIVILEGES | TOKEN_QUERY, out tokenHandle))
return false;
if (!LookupPrivilegeValue(null, SE_SECURITY_NAME, out tokenPrivileges.Luid))
return false;
tokenPrivileges.PrivilegeCount = 1;
tokenPrivileges.Attributes = SE_PRIVILEGE_ENABLED;
return AdjustTokenPrivileges(tokenHandle, false, ref tokenPrivileges, 0, IntPtr.Zero, IntPtr.Zero);
}
}
"@
Add-Type -TypeDefinition $definition
[Privilege]::EnablePrivilege()
```
### Validate Privilege
At this point, we now have a powershell session operating with the `SeSecurityPrivilege` privilege enabled. We want to confirm this by running the following commands:
```powershell
whoami # (1)
whoami /priv # (2)
```
1. Output will appear similar to "bunny-lab\nicole.rappe", prefixing the username of the person running the command with the domain they belong to.
2. Reference the privilege table seen below to validate the output of this command matches what you see below.
| **Privilege Name** | **Description** | **State** |
| :--- | :--- | :--- |
| `SeSecurityPrivilege` | Manage auditing and security log | Enabled |
### Execute `setup.exe`
Finally, at the last stage, we mount the ISO file for the Cumulative Update ISO (e.g. 6.6GB ISO image), and using this powershell session we made above, we navigate to the drive it is running on, and invoke setup.exe, causing it to run under the `SeSecurityPrivilege` operational state.
```powershell
D: <ENTER> # (1)
.\Setup.EXE /m:upgrade /IAcceptExchangeServersLicenseTerms_DiagnosticDataON # (2)
```
1. Replace this drive letter with whatever letter was assigned when you mounted the ISO image for the Exchange Updater.
2. This launches the Exchange updater application. Be patient and give it time to launch. At this point, you should be good to proceed with the update. You can optionally change the argument to `/IAcceptExchangeServersLicenseTerms_DiagnosticDataOFF` if you do not need diagnostic data.
!!! success "Ready to Proceed with Updating Exchange"
At this point, after doing the three sections above, you should be safe to do the upgrade/update of Microsoft Exchange Server. The installer will run its own readiness checks for other aspects such as IIS Rewrite Modules and will give you a link to download / upgrade it separately, then giving you the option to "**Retry**" after installing the module for the installer to re-check and proceed.
## Post-Update Health Checks
After the update(s) are installed, you will likely want to check to ensure things are healthy and operational, validating mail flow in both directions, running `Get-Queue` to check for backlogged emails, etc.
!!! note "Under Construction"
This section is under construction and will be based on some feedback from others to help build the section out.
## Related Documentation
- [Related Email Documentation](<../../../../reference/Applications/Email/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,329 @@
---
tags:
- Proxmox Mail Gateway
- Mailcow
- Email
---
## Purpose
Use this workflow when PMG can receive mail but delivery or quarantine release to Mailcow is rejected by sender validation or relay trust settings. It assumes the PMG and Mailcow integration already exists.
## Scope
This workflow describes how to configure Mailcow to trust Proxmox Mail Gateway (PMG) as an upstream SMTP relay after PMG has been placed in front of Mailcow for inbound filtering. It prevents Mailcow from applying sender-domain validation and duplicate spam filtering to messages that PMG has already accepted, including messages released from the PMG quarantine.
In the current Bunny Lab design, this trust applies to the inbound delivery path from `192.168.3.15` to `192.168.3.61`. Mailcow continues to send outbound mail directly to the Internet and does not relay outbound mail through PMG.
!!! info "Assumptions"
- Mailcow is already deployed at `192.168.3.61`.
- PMG is already deployed at `192.168.3.15`.
- Public inbound SMTP port `25` is forwarded to PMG.
- PMG relays accepted inbound mail to Mailcow on `192.168.3.61:25`.
- Mailcow remains responsible for mailbox hosting, authenticated submission, DKIM signing, and outbound delivery.
- You have `root` access to the Mailcow host.
!!! warning "Trust Only the PMG Host"
Add only `192.168.3.15/32` to Mailcow's trusted networks. Do not trust the entire `192.168.3.0/24` subnet unless every system on that subnet is authorized to relay mail through Mailcow without authentication.
## Architecture
```text
Internet
|
v
pfSense WAN :25
|
v
PMG 192.168.3.15:25
|
v
Mailcow 192.168.3.61:25
|
v
User Mailbox
```
PMG is the primary inbound spam and virus filtering authority. Mailcow should accept the SMTP session from PMG as a trusted relay and avoid repeating edge filtering or rejecting PMG-generated envelope senders.
## Symptoms
Use this workflow when PMG receives and processes mail correctly, but Mailcow rejects or defers the final delivery.
Common symptoms include:
- A released PMG quarantine message remains in the PMG deferred queue
- Clicking **Flush** in PMG does not deliver the message
- PMG repeatedly attempts delivery to `192.168.3.61:25`
- Mailcow rejects a PMG-generated envelope sender such as `postmaster@lab-mail-gw-01.bunny-lab.io`
- PMG reports an SMTP response similar to:
```text
450 4.1.8 Sender address rejected: Domain not found
```
The visible message sender may be valid even when Mailcow rejects the PMG-generated envelope sender used during quarantine release.
## Confirm the PMG Source Address
Mailcow must trust the source address it actually sees on the SMTP connection. Confirm that address before changing the configuration.
On the Mailcow host, run:
```sh
cd /opt/mailcow-dockerized
docker compose logs --since 5m postfix-mailcow
```
Flush or resend a test message from PMG, then locate the connection line.
Expected connection source:
```text
connect from lab-mail-gw-01[192.168.3.15]
```
If Mailcow sees a different address because of NAT, a load balancer, or another SMTP proxy, use that observed address instead of `192.168.3.15`.
## Configure Mailcow Forwarding Hosts
Configure PMG as a trusted forwarding host in the Mailcow WebUI.
- Navigate to "**System > Configuration > Options > Forwarding Hosts**"
- Add `192.168.3.15`
- Set the forwarding-host spam filter option to **Inactive**
- Save the configuration
!!! note "Forwarding Host Behavior"
The forwarding-host entry tells Mailcow that PMG is the immediate upstream SMTP relay. Keeping the forwarding-host spam filter inactive avoids applying a second aggressive spam-filtering layer to mail that PMG has already inspected.
## Configure Postfix Trusted Networks
The forwarding-host entry does not replace Postfix `mynetworks`. Add PMG to `mynetworks` so Postfix treats SMTP sessions from PMG as trusted and evaluates `permit_mynetworks` before sender-domain restrictions.
### Inspect the Existing Trusted Networks
Before modifying `mynetworks`, determine the currently active value. Mailcow automatically populates `mynetworks` with its loopback and Docker networks. Defining your own value replaces that automatically generated configuration, so you must preserve the existing entries.
On the Mailcow host, run:
```sh
cd /opt/mailcow-dockerized
docker compose exec postfix-mailcow postconf mynetworks
```
Example output:
```text
mynetworks = 127.0.0.0/8 172.22.1.0/24 [::1]/128
```
At this point, append the PMG address as a single-host CIDR. Do not remove any existing networks.
Expected result:
```text
mynetworks = 127.0.0.0/8 172.22.1.0/24 [::1]/128 192.168.3.15/32
```
### Update the Persistent Postfix Override
Edit the Postfix override file:
```sh
nano /opt/mailcow-dockerized/data/conf/postfix/extra.cf
```
If `mynetworks` config line already exists, append the PMG address while preserving the existing values. If the key does not exist, create it using the value discovered in the previous step.
```ini title="/opt/mailcow-dockerized/data/conf/postfix/extra.cf"
mynetworks = 127.0.0.0/8 172.22.1.0/24 [::1]/128 192.168.3.15/32
```
!!! warning "Do Not Replace Existing Networks"
The `mynetworks` directive replaces Mailcow's automatically generated value. Removing the existing loopback or Docker networks may prevent Postfix from functioning correctly. Always preserve the existing entries and append the PMG address.
### Restart Mailcow Postfix
Apply the persistent override:
```sh
cd /opt/mailcow-dockerized
docker compose restart postfix-mailcow
```
## Validate the Trusted Relay Configuration
Confirm the active Postfix configuration after the restart:
```sh
cd /opt/mailcow-dockerized
docker compose exec postfix-mailcow postconf mynetworks
```
Expected result:
```text
mynetworks = 127.0.0.0/8 172.22.1.0/24 [::1]/128 192.168.3.15/32
```
Confirm the active SMTP restriction order includes `permit_mynetworks`:
```sh
cd /opt/mailcow-dockerized
docker compose exec postfix-mailcow postconf \
smtpd_sender_restrictions \
smtpd_recipient_restrictions \
smtpd_relay_restrictions
```
The exact restriction lists may change between Mailcow releases. Confirm that `permit_mynetworks` remains present and that the effective configuration recognizes `192.168.3.15/32` as trusted.
## Retry Deferred PMG Mail
After Mailcow trusts PMG, retry the deferred message from the PMG WebUI.
- Navigate to "**Administration > Queue Administration > Deferred Mail**"
- Select the deferred message
- Click **Flush**
Alternatively, run the following on PMG:
```sh
postqueue -f
```
To retry one specific queue item:
```sh
postsuper -r <QUEUE_ID>
postqueue -f
```
## Validate Mail Delivery
Monitor Mailcow while PMG retries the message:
```sh
cd /opt/mailcow-dockerized
docker compose logs -f postfix-mailcow
```
Mailcow should accept the SMTP transaction and return a successful queue response similar to:
```text
250 2.0.0 Ok: queued as <MAILCOW_QUEUE_ID>
```
On PMG, confirm the deferred queue no longer contains the message:
```sh
postqueue -p
```
Confirm the message is present in the destination Mailcow mailbox.
!!! success "Trusted PMG Delivery Confirmed"
The workflow is complete when Mailcow accepts mail from `192.168.3.15`, the PMG deferred queue clears, and the destination mailbox receives the message.
## Security Boundaries
Adding `192.168.3.15/32` to `mynetworks` permits PMG to relay mail through Mailcow without SMTP authentication. This is appropriate only while PMG remains a controlled gateway host.
Maintain the following boundaries:
- Restrict the trusted entry to `192.168.3.15/32`
- Prevent other hosts from impersonating the PMG source address
- Keep Mailcow port `25` restricted to expected SMTP sources where practical
- Confirm PMG is not configured as an unrestricted open relay
- Continue routing authenticated client submission directly to Mailcow on ports `465` and `587`
- Continue routing public inbound SMTP port `25` to PMG rather than directly to Mailcow
## Troubleshooting
### Mailcow Still Returns `450 4.1.8`
Confirm Mailcow is using the updated configuration:
```sh
cd /opt/mailcow-dockerized
docker compose exec postfix-mailcow postconf mynetworks
```
If `192.168.3.15/32` is missing, verify `/opt/mailcow-dockerized/data/conf/postfix/extra.cf` and restart `postfix-mailcow`.
### Mailcow Sees a Different Source Address
Inspect the Mailcow Postfix logs:
```sh
cd /opt/mailcow-dockerized
docker compose logs --since 5m postfix-mailcow
```
Use the IP address shown in the `connect from ...[IP_ADDRESS]` line. Do not assume Mailcow sees the PMG management address when NAT or an SMTP proxy exists between the systems.
### PMG Flushes the Message but the Mailbox Does Not Receive It
Determine whether Mailcow accepted the message.
If Mailcow returned `250 2.0.0`, the PMG-to-Mailcow relay succeeded. Continue troubleshooting inside Mailcow by reviewing Postfix, Rspamd, and Dovecot logs.
```sh
cd /opt/mailcow-dockerized
docker compose logs --since 10m postfix-mailcow rspamd-mailcow dovecot-mailcow
```
If Mailcow returned a `4xx` or `5xx` response, use the complete SMTP response as the controlling error and resolve that policy or recipient failure before retrying again.
### Mailcow Applies Spam Filtering Again
Confirm the PMG entry under "**System > Configuration > Options > Forwarding Hosts**" has its spam filter set to **Inactive**.
PMG should remain the primary inbound filtering authority in this deployment.
### Mail Clients or Outbound Delivery Stop Working
This workflow does not change client access or outbound delivery.
Confirm the existing service ownership remains:
```text
Inbound SMTP: Internet -> PMG -> Mailcow
Outbound SMTP: Mailcow -> Internet
SMTP Submission: Clients -> Mailcow
IMAP and POP3: Clients -> Mailcow
Webmail and Admin: Internet -> Traefik -> Mailcow
```
Do not redirect ports `465`, `587`, `993`, `995`, `110`, `143`, or `4190` to PMG.
## Rollback
Remove the PMG forwarding-host entry from the Mailcow WebUI only when PMG is no longer the upstream SMTP gateway.
Then edit:
```sh
nano /opt/mailcow-dockerized/data/conf/postfix/extra.cf
```
Restore the previous `mynetworks` value:
```ini title="/opt/mailcow-dockerized/data/conf/postfix/extra.cf"
mynetworks = 127.0.0.0/8 172.22.1.0/24 [::1]/128
```
Restart Postfix:
```sh
cd /opt/mailcow-dockerized
docker compose restart postfix-mailcow
```
!!! warning "Coordinate the SMTP Path Before Rollback"
Do not remove PMG trust while public inbound SMTP still routes through PMG. Mailcow may begin rejecting legitimate messages relayed from PMG.
## Confirmed Final State
The validated Bunny Lab configuration is:
```text
PMG Address: 192.168.3.15
Mailcow Address: 192.168.3.61
Mailcow Forwarder: 192.168.3.15
Forwarder Spam Check: Inactive
Postfix mynetworks: 127.0.0.0/8 172.22.1.0/24 [::1]/128 192.168.3.15/32
```
The resulting mail flow is:
```text
Inbound SMTP:
Internet -> pfSense WAN :25 -> PMG 192.168.3.15:25 -> Mailcow 192.168.3.61:25
Outbound SMTP:
Mailcow -> Internet
```
## Related Documentation
- [PMG and Mailcow Integration](<../../../../deployments/Applications/Email/Proxmox Mail Gateway/Integrate PMG with Mailcow.md>) — Confirm the intended gateway and relay topology before changing trust settings.
- [Related Email Documentation](<../../../../reference/Applications/Email/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,45 @@
---
tags:
- IredMail
- SMTP
- Email
---
## Purpose
You may need to troubleshoot the outgoing SMTP email queue / active sessions in iRedMail for one reason or another. This can provide useful insight into the reason why emails are not being delivered, etc.
### Overall Queue Backlog
You can run the following command to get the complete backlog of all email senders in the queue. This can be useful for tracking the queue's "drainage" over-time.
```sh
# List the total number of queued messages
postqueue -p | egrep -c '^[A-F0-9]'
# Itemize and count the queued messages based on sender.
postqueue -p | awk '/^[A-F0-9]/ {id=$1} /from=<[^>]+>/ && $0 !~ /from=<>/ {print id; exit}'
```
!!! example "Example Output"
- 10392 problematic@bunny-lab.io
- 301 prettybad@bunny-lab.io
- 39 infrastructure@bunny-lab.io
- 20 nicole.rappe@bunny-lab.io
### Investigating Individual Emails
You can run the following command to list all queued messages: `postqueue -p`. You can then run `postcat -vq <message-ID>` to read detailed information on any specific queued SMTP message:
```sh
postqueue -p
postcat -vq 4dgHry5LZnzH6x08 # (1)
```
1. Example message ID gathered from the previous `postqueue -p` command.
### Attempt to Gracefully Reload Postfix
You may want to try to unstick things by gracefully "reloading" the postfix service via `postfix reload`. This will ensure that we don't drop / disconnect / lose all of the active outgoing SMTP sessions in the queue. It may not help resolve issues, but it's worth noting down:
### Reattempt Delivery
You can attempt redelivery via running `postqueue -f` to try to free up the queue. Postfix will immediately re-attempt delivery of all queued messages instead of waiting for their scheduled retry time. It does not override remote rejections or fix underlying delivery errors; it only accelerates the next delivery attempt.
## Related Documentation
- [Related Email Documentation](<../../../../reference/Applications/Email/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,445 @@
---
tags:
- Rclone
- Synchronization
- PowerShell
---
## Purpose
Configure, run, and recover the documented rclone bisync pair between a local Windows path and a Google Drive remote. The examples use PowerShell and were written for rclone `v1.75.0`.
!!! danger "Bisync Can Delete or Overwrite Data"
Confirm both paths and keep an independent backup before applying changes. Review dry-run output before removing `--dry-run`.
## Understand Bisync State
Bisync is a stateful two-way synchronization command. It retains listings of Path1 and Path2 from the prior successful run and compares those listings against the current state during the next run.
Bisync does not maintain a file-version archive or historical copies of every changed file. The listings record synchronization state and metadata needed to determine whether a file is new, changed, or deleted relative to the previous run.
!!! warning "Protect the Bisync Work Directory"
Use a stable, explicit `--workdir` for every run of a given bisync pair. Do not change the work directory, reverse the order of Path1 and Path2, delete the listing files, or run the same pair with different filters without understanding that the existing state may no longer be valid.
On Windows, bisync otherwise stores its state beneath the profile of the account running rclone. This can cause scheduled and interactive executions to use different state directories when they run as different accounts.
## Define the Example Environment
Replace the example values before running the commands in this document.
```powershell
$Rclone = "C:\Path\To\rclone.exe"
$LocalPath = "C:\Path\To\Local"
$RemotePath = "GoogleDrive:Path/To/Remote"
$WorkDir = "C:\Path\To\BisyncState"
$LogDir = "C:\Path\To\Logs"
$FilterFile = "C:\Path\To\bisync-filters.txt"
```
Create the state and log directories:
```powershell
New-Item -Path $WorkDir,$LogDir -ItemType Directory -Force | Out-Null
```
Create a consistent filters file:
```text title="C:\Path\To\bisync-filters.txt"
- *.lnk
```
The `*.lnk` rule excludes Windows shortcut files at any depth beneath the synchronization root.
!!! warning "Filter Changes Require a Reviewed Rebaseline"
Bisync compares current listings against prior listings. Adding, removing, or changing a filter can make previously tracked files disappear from the new listings and appear to have been deleted.
Treat filter changes as a state-changing event. Stop scheduled runs, review the new scope, and perform a dry-run resync before committing the change.
## Validate the Paths
Confirm the installed rclone version:
```powershell
& $Rclone version
```
Confirm that the local root already exists:
```powershell
Test-Path -LiteralPath $LocalPath
```
Expected result:
```text
True
```
Confirm that the configured remote path is reachable:
```powershell
& $Rclone lsd $RemotePath
```
!!! warning "Do Not Automatically Create an Unexpected Root"
Stop if either path is missing or points to an unexpected location. Automatically creating an empty root can conceal a path, authentication, drive-mount, or configuration failure and can cause bisync to interpret an entire data set as deleted.
## Configure the Access Check
The `--check-access` flag verifies that matching files named `RCLONE_TEST` are visible in the same relative locations on both sides. This provides additional protection against an unavailable mount, incorrect remote root, or incomplete listing being interpreted as mass deletion.
Create the access-check file at the root of the local path and copy it to the root of the remote path:
```powershell
$AccessFile = Join-Path $LocalPath "RCLONE_TEST"
New-Item -Path $AccessFile -ItemType File -Force | Out-Null
& $Rclone copyto $AccessFile "$RemotePath/RCLONE_TEST"
```
Confirm that the file appears on both sides:
```powershell
& $Rclone lsf $LocalPath --include "RCLONE_TEST"
& $Rclone lsf $RemotePath --include "RCLONE_TEST"
```
Expected result from both commands:
```text
RCLONE_TEST
```
!!! warning "Do Not Delete the Access-Check File"
The `RCLONE_TEST` file must remain visible on both sides while `--check-access` is enabled. Bisync will abort when the file cannot be found.
## Audit the Existing Data
Before establishing or rebuilding the bisync state, compare the current file sets without changing either side:
```powershell
& $Rclone check $LocalPath $RemotePath --filter-from $FilterFile --drive-skip-gdocs --combined (Join-Path $LogDir "pre-bisync-check.txt") --log-level INFO --log-file (Join-Path $LogDir "pre-bisync-check.log")
```
Review all source-only, destination-only, differing, and unreadable files before proceeding. Determine whether the expected recovery is to merge both sides, prefer one authoritative side, or preserve selected files manually.
## Initialize a Bisync Pair
A new bisync pair requires an initial resync to establish its Path1 and Path2 listings. A preliminary `sync --update` operation is not required and can unnecessarily delete destination-only files.
!!! danger "Resync Is Not a Database-Only Repair"
A resync reconciles the actual contents of both paths while rebuilding the listings.
Files that exist on only one side are copied to the other side. When the same relative path contains different files on both sides, `--resync-mode` selects one version as the winner and overwrites the other version.
The conflict flags used during normal bisync runs do not apply during a resync. A resync does not rename the losing version into a conflict copy.
### Select the Resync Policy
Choose the resync policy according to which data is authoritative:
| **Mode** | **Use When** |
| :--- | :--- |
| `path1` | Path1 is authoritative and must win every same-path difference |
| `path2` | Path2 is authoritative and must win every same-path difference |
| `newer` | Modification times are trustworthy and the newest same-path version should win |
| `larger` | File size is a more reliable winner than modification time |
A bare `--resync` is equivalent to `--resync-mode path1`. Do not use a bare `--resync` unless Path1 is intentionally authoritative.
The examples below use `--resync-mode newer`. This is appropriate only when both sides provide trustworthy modification times and the intended policy is to preserve the newer same-path file.
### Preview the Initial Resync
Run the initial resync as a dry run:
```powershell
& $Rclone bisync $LocalPath $RemotePath --workdir $WorkDir --resync-mode newer --create-empty-src-dirs --compare size,modtime,checksum --check-access --max-delete 0 --filters-file $FilterFile --drive-skip-gdocs --fix-case --dry-run --verbose --log-file (Join-Path $LogDir "bisync-resync-dry-run.log")
```
Review the complete log for:
- Files copied from Path1 to Path2
- Files copied from Path2 to Path1
- Same-path files that would be replaced
- Unexpected paths
- Access or checksum errors
- Proposed deletion behavior
!!! note "Dry-Run Deletion Messages"
A bisync dry run may display confusing deletion messages because the simulated copy that would normally precede the deletion did not actually occur. Review the complete sequence rather than evaluating an isolated deletion line.
Do not dismiss an unexpected deletion unless the log clearly shows that it is a dry-run artifact associated with a preceding simulated copy.
The `--max-delete 0` setting blocks ordinary file deletions during this recovery run. It does not prevent a same-path losing version from being overwritten during resync, so a separate backup or snapshot remains required.
### Perform the Initial Resync
After reviewing and approving the dry-run output, repeat the command without `--dry-run`:
```powershell
& $Rclone bisync $LocalPath $RemotePath --workdir $WorkDir --resync-mode newer --create-empty-src-dirs --compare size,modtime,checksum --check-access --max-delete 0 --filters-file $FilterFile --drive-skip-gdocs --fix-case --verbose --log-file (Join-Path $LogDir "bisync-resync-live.log")
```
A successful run should end with:
```text
Bisync successful
```
### Validate the Reconciled Data
Compare both sides after the live resync:
```powershell
& $Rclone check $LocalPath $RemotePath --filter-from $FilterFile --drive-skip-gdocs --combined (Join-Path $LogDir "post-resync-check.txt") --log-level INFO --log-file (Join-Path $LogDir "post-resync-check.log")
```
The expected result is:
```text
0 differences found
```
## Run Normal Bisync Operations
After the baseline has been established, all normal runs must omit `--resync` and `--resync-mode`.
Preview the first normal run:
```powershell
& $Rclone bisync $LocalPath $RemotePath --workdir $WorkDir --create-empty-src-dirs --conflict-resolve newer --conflict-loser num --compare size,modtime,checksum --resilient --recover --max-lock 2m --check-access --max-delete 10 --filters-file $FilterFile --drive-skip-gdocs --fix-case --dry-run --log-level INFO --log-file (Join-Path $LogDir "bisync-normal-dry-run.log")
```
A healthy unchanged pair should report:
```text
No changes found
Updating listings
Bisync successful
```
After reviewing the dry run, perform one controlled live run:
```powershell
& $Rclone bisync $LocalPath $RemotePath --workdir $WorkDir --create-empty-src-dirs --conflict-resolve newer --conflict-loser num --compare size,modtime,checksum --resilient --recover --max-lock 2m --check-access --max-delete 10 --filters-file $FilterFile --drive-skip-gdocs --fix-case --log-level INFO --log-file (Join-Path $LogDir "bisync.log")
```
This normal command can be placed into the scheduled automation after it has completed successfully under the same Windows account and execution context that will run the scheduled task.
## Understand the Normal Bisync Flags
| **Flag** | **Behavior** |
| :--- | :--- |
| `--workdir` | Stores the prior Path1 and Path2 listings in an explicit, stable directory |
| `--create-empty-src-dirs` | Synchronizes the creation and deletion of empty directories |
| `--compare size,modtime,checksum` | Uses size, modification time, and checksums when determining file state |
| `--conflict-resolve newer` | Prefers the newer file when the same file changed independently on both sides since the prior run |
| `--conflict-loser num` | Preserves the losing conflict version under a numbered `.conflictN` name |
| `--resilient` | Allows certain less-serious errors to be retried during future runs rather than immediately requiring resync |
| `--recover` | Uses a backup listing to recover from some interrupted or ungracefully terminated runs |
| `--max-lock 2m` | Allows a stale bisync lock to expire after two minutes |
| `--check-access` | Aborts when matching `RCLONE_TEST` files cannot be found on both sides |
| `--max-delete 10` | Aborts when more than 10 percent of tracked files appear deleted on either side |
| `--filters-file` | Applies one consistent set of exclusions to both bisync listings |
| `--drive-skip-gdocs` | Makes native Google Docs, Sheets, Slides, and other Google-native documents invisible to rclone |
| `--fix-case` | Allows supported case-only filename corrections between filesystems |
| `--log-level INFO` | Records normal operations, changes, warnings, and errors rather than errors alone |
### Understand Conflict Resolution
A bisync conflict occurs when the same relative file is new or changed on both sides compared with the prior successful run and the current file contents are not identical.
The `--conflict-resolve newer` option does not mean that rclone blindly compares every file and always keeps whichever modification time is newest. It applies specifically to a detected two-sided conflict.
With `--conflict-loser num`, the winner retains the original filename and the losing version is preserved with a numbered conflict suffix, such as:
```text
Document.docx
Document.docx.conflict1
```
This is safer than `--conflict-loser delete`, which permanently removes the losing conflict version.
### Understand the Deletion Limit
The `--max-delete 10` value is a conservative example rather than a universal requirement. Select a threshold appropriate for the size of the data set and the organization's normal deletion patterns.
A large folder rename may appear as many deletions and many new files and can exceed the configured threshold.
!!! danger "Do Not Use `--force` as a Recurring Option"
The `--force` flag bypasses the `--max-delete` safety check. Do not include it in a normal scheduled command.
When a legitimate change exceeds the threshold, stop the scheduled operation, review the proposed changes with `--dry-run`, and use a one-time explicitly approved threshold instead of permanently disabling the protection.
### Understand Google-Native Documents
The `--drive-skip-gdocs` flag removes native Google documents from rclone listings. These objects are not downloaded, exported, compared, or synchronized while the flag is enabled.
This flag does not merely ignore local files with extensions such as `.gdoc` or `.gsheet`. It makes the corresponding native Google Drive objects effectively invisible to rclone.
Remove this flag only when Google-native documents must be exported and synchronized and the desired export formats have been deliberately configured. Changing this behavior on an established bisync pair requires a reviewed rebaseline because it changes which objects appear in the listings.
## Repair a Broken Bisync Pair
A request for `--resync` does not automatically mean that the data is damaged. Bisync may be unable to locate or trust its prior state because the work directory changed, the command used different paths or filters, the execution account changed, or a prior run ended with a critical error.
### Stop Automated Runs
Stop all scheduled or looping bisync processes before investigating or repairing the state. Confirm that another rclone process is not modifying either path or the bisync listings.
### Review the Logs
Search the most recent log for the first critical error rather than relying only on the final resync message.
```powershell
Get-Content (Join-Path $LogDir "bisync.log") -Tail 300
```
Look for:
- Authentication or remote-access errors
- Missing-path errors
- Missing or invalid listing files
- Access-check failures
- Excessive deletion warnings
- File-copy or file-move failures
- Lock-file errors
- Filter changes
- `Bisync aborted`
- `Bisync critical error`
### Verify the Existing Configuration
Confirm that the repair command uses:
- The original Path1 and Path2 in the original order
- The original `--workdir`
- The original filter rules
- The intended rclone configuration file
- The correct remote account and storage location
- The same Google Docs behavior
- The same comparison settings
List the current bisync state directory:
```powershell
Get-ChildItem -LiteralPath $WorkDir
```
The directory should contain Path1 and Path2 `.lst` files for the configured pair.
### Test Remote and Local Access
Confirm that both roots are present and readable:
```powershell
Test-Path -LiteralPath $LocalPath
& $Rclone lsd $RemotePath
```
Confirm that the access-check file remains visible:
```powershell
& $Rclone lsf $LocalPath --include "RCLONE_TEST"
& $Rclone lsf $RemotePath --include "RCLONE_TEST"
```
### Compare the Current Data
Run a read-only comparison before deciding to resync:
```powershell
& $Rclone check $LocalPath $RemotePath --filter-from $FilterFile --drive-skip-gdocs --combined (Join-Path $LogDir "repair-comparison.txt") --log-level INFO --log-file (Join-Path $LogDir "repair-check.log")
```
Review whether one-sided files are expected, whether one side is authoritative, and whether same-path differences should be resolved manually.
### Attempt Normal Recovery When Appropriate
When the prior run was only interrupted and the normal command already uses `--recover`, a normal controlled run may recover without requiring a resync.
Run it first as a dry run:
```powershell
& $Rclone bisync $LocalPath $RemotePath --workdir $WorkDir --create-empty-src-dirs --conflict-resolve newer --conflict-loser num --compare size,modtime,checksum --resilient --recover --max-lock 2m --check-access --max-delete 10 --filters-file $FilterFile --drive-skip-gdocs --fix-case --dry-run --log-level INFO --log-file (Join-Path $LogDir "bisync-recovery-dry-run.log")
```
Proceed with a live normal run only when the output is understood and expected.
### Rebuild the State Only When Required
Use a resync only when:
- The pair has never been initialized
- The listings are missing or cannot be trusted
- A critical bisync error explicitly requires a resync
- The path scope or filter rules intentionally changed
- You are deliberately establishing a new authoritative baseline
Select the correct `--resync-mode`, perform a dry run, review every proposed change, and then follow the initialization and validation sequence documented above.
!!! danger "Do Not Describe Resync as Non-Destructive"
Resync can overwrite a same-path file and can restore files that were intentionally deleted by copying one-sided files back to the other side.
The `--conflict-resolve` and `--conflict-loser` flags are ignored during resync. Maintain an independent backup and do not proceed until the chosen resync policy is understood.
## Validate Bidirectional Synchronization
After the initial setup or a repair, validate both directions with disposable files.
- Create a test text file beneath the local root.
- Run a normal bisync and confirm that the file appears remotely.
- Create a second test text file beneath the remote root.
- Run another normal bisync and confirm that the file appears locally.
- Delete both test files.
- Run another normal bisync and confirm that the deletions propagate as intended.
- Confirm that the final run ends with `Bisync successful`.
- Run `rclone check` and confirm that no file differences remain.
Remove all test artifacts when validation is complete.
## Troubleshooting
### Google Drive Reports Shared Drive Not Found
An error such as `404: Shared Drive not found` normally means that the authenticated Google account cannot access the configured Shared Drive ID.
Verify:
- The correct Google account completed OAuth
- The account still has access to the Shared Drive
- The configured Shared Drive ID is correct
- The Shared Drive was not deleted and recreated under a new ID
Do not replace the drive ID with another visible Shared Drive merely because authentication succeeded.
### Bisync Cannot Find Its Listings
Confirm that the command uses the original `--workdir` and that the scheduled task runs under the expected account.
Do not perform an immediate resync when the actual problem is that rclone is looking in the wrong state directory.
### The Deletion Limit Was Exceeded
Stop the operation and determine why so many files appear deleted. Common causes include:
- An unavailable local mount
- An inaccessible remote path
- An incorrect synchronization root
- A large folder rename
- Changed filter rules
- An actual mass deletion
Do not use `--force` until the reported deletions have been independently reviewed and approved.
### Conflict Files Are Appearing
Files ending in `.conflict1`, `.conflict2`, or another numbered suffix indicate that both sides changed independently and bisync preserved the losing version.
Review the contents, retain the correct version, and remove the obsolete conflict copy after confirming that it is no longer needed.
### Google Docs Are Missing
Native Google Docs are intentionally absent when `--drive-skip-gdocs` is enabled. Remove or change this behavior only as a deliberate configuration change with a reviewed resync.
### Dry Run Appears to Delete a Newly Copied File
Review the complete dry-run sequence. Bisync can display an apparent deletion because the preceding simulated copy did not actually create the file on the other side.
Do not ignore unrelated or unexplained deletion messages.
### Google Drive Reports Duplicate Objects
Google Drive can contain multiple objects with the same name in one folder. List duplicate names with:
```powershell
& $Rclone dedupe list $RemotePath
```
Use the interactive resolver only after reviewing the file sizes, modification times, and hashes:
```powershell
& $Rclone dedupe interactive $RemotePath
```
When the correct version cannot be determined confidently, rename and preserve both objects rather than deleting one automatically.
## Reference Documentation
- [Rclone Command Overview](https://rclone.org/commands/)
- [Rclone Bisync](https://rclone.org/bisync/)
- [Rclone Copy](https://rclone.org/commands/rclone_copy/)
- [Rclone Sync](https://rclone.org/commands/rclone_sync/)
- [Rclone Check](https://rclone.org/commands/rclone_check/)
- [Rclone Filtering](https://rclone.org/filtering/)
- [Rclone Google Drive Backend](https://rclone.org/drive/)
- [Rclone Changelog](https://rclone.org/changelog/)
## Related Documentation
- [Related Files and Collaboration Documentation](<../../../reference/Applications/Files and Collaboration/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,41 @@
---
tags:
- DFS
- Active Directory
- Troubleshooting
---
## Purpose
Refresh a DFS Management console that shows inconsistent namespaces or replication groups between member servers. Check Active Directory replication before restarting DFS Replication or clearing the console cache.
## Repair the Console
Sometimes the GUI for managing DFS becomes "inconsistent" whereas the namespaces and replication groups are different between member servers, and may be missing namspaces or missing replication groups. DFS Management is an MMC snap-in. MMC persists per-user console state under `%APPDATA%\Microsoft\MMC\`. If that state gets out of sync (common after service hiccups or server crashes), the snap-in can render partial/incorrect namespace/replication trees even when DFS itself is fine. Deleting the cached dfsmgmt* console forces a fresh enumeration. We will also include a few extra commands for extra thouroughness.
Before anything, we want to make sure that active directory itself is not having replication issues, as this would be a deeper, more complicated issue. Run the following command on one of your domain controllers:
```powershell
repadmin /syncall /AdeP
repadmin /replsummary
```
If AD-level replication is successful and timely, you can proceed to run the commands below (one-line-at-a-time):
```sh
# Pull-Down DFS Configuration from Active Directory & Restart DFS
dfsrdiag pollad
net stop dfsr
net start dfsr
# Clear DFS Management Snap-In Cache
taskkill /im mmc.exe /f
del "%appdata%\Microsoft\MMC\dfsmgmt*"
dfsmgmt.msc
```
!!! success "DFS Management GUI Restored"
At this point, the DFS Management snap-in (should) be successfully showing all of the DFS namespaces and replication groups when you re-open "DFS Management".
## Related Documentation
- [DFS Deployment](<../../../../deployments/Applications/Files and Collaboration/Windows Server/DFS Namespaces with Replication.md>) — Confirm the expected namespace and replication configuration.
- [DFS Configuration Report](<../../../../scripts/Applications/Files and Collaboration/DFS/Report DFS Namespaces and Replication.md>) — Inspect the objects after reopening the management console.
- [Related Files and Collaboration Documentation](<../../../../reference/Applications/Files and Collaboration/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,58 @@
---
tags:
- Applications
- Access Another User OneDrive Data
---
## Purpose
You may find that you need to access data from a user's personal OneDrive account, and since that data cannot be accessed directly via Office365's Admin Portal, you have to do some legwork in Powershell.
### Connect to Sharepoint
```powershell
Install-Module Microsoft.Online.SharePoint.PowerShell
Import-Module Microsoft.Online.SharePoint.PowerShell
Connect-SPOService -Url https://<companyname>-admin.sharepoint.com # Login with your Office365 Admin Credentials when Prompted
```
### Display List of Personal Sharepoint Sites (OneDrive)
```powershell
Get-SPOSite -IncludePersonalSite $true -Limit All
```
### Check OneDrive Usage of the Given User
```powershell
Get-SPOSite -Identity "https://<companyname>-my.sharepoint.com/personal/username_companyname_com" | Select Url, Owner, StorageUsageCurrent, StorageQuota, LastContentModifiedDate
```
### Assign Yourself Permissions to Their OneDrive
```powershell
Set-SPOUser `
-Site "https://<companyname>-my.sharepoint.com/personal/username_companyname_com" `
-LoginName "i:0#.f|membership|admin@companyname.com" `
-IsSiteCollectionAdmin $true
```
### Navigate to Webpage
At this point, you now have permissions to access the OneDrive data, so open a web browser and navigate to the SPOSite URL seen previously, seen below. From here, you can download, upload, and manage the data however you need.
- `https://<companyname>-my.sharepoint.com/personal/username_companyname_com`
### Remove Permissions
At this point, when the work is done, revoke your permissions to lock-down the OneDrive data once again by running the following command:
```powershell
Set-SPOUser `
-Site "https://<companyname>-my.sharepoint.com/personal/username_companyname_com" `
-LoginName "i:0#.f|membership|admin@companyname.com" `
-IsSiteCollectionAdmin $false
```
### Logout and Cleanup Auth Tokens and Remove Module
```powershell
Disconnect-SPOService
Remove-Module Microsoft.Online.SharePoint.PowerShell
# Close Powershell Window
```
## Related Documentation
- [Related Files and Collaboration Documentation](<../../../../reference/Applications/Files and Collaboration/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,34 @@
---
tags:
- Bash
- Netcat
- File Transfer
- Scripting
- Linux
---
## Purpose
You may find that you need to transfer a file, such as a public SSH key, or some other kind of file between two devices. In this scenario, we assume both devices have the `netcat` command available to them. By putting a network listener on the device recieving the file, then sending the file to that device's IP and port, you can successfully transfer data between computers without needing to set up SSH, FTP, or anything else to establish initial trust between the devices. [Original Reference Material](https://www.youtube.com/shorts/1j17UBGqSog).
!!! warning
The data being transferred will not be encrypted. If you are transferring relatively-safe files such as public SSH keys, etc, this should be fine.
### Destination Computer
Run the following command on the computer that will be recieving the file.
```sh
netcat -l <random-port> > /tmp/OUTPUT-AS-FILE.txt
```
### Source Computer
Run the following command on the computer that will be sending the file to the destination computer.
```sh
cat INPUT-DATA.txt | netcat <IP-of-Destination-Computer> <Port-of-Destination-Computer> -q 0
```
!!! info
The `-q 0` command argument causes the netcat connection to close itself automatically when the transfer is complete.
## Related Documentation
- [Related Files and Collaboration Documentation](<../../../reference/Applications/Files and Collaboration/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,37 @@
---
tags:
- Tuya
- Networking
---
## Purpose
Connect the recorded Tuya devices with the local integration tooling and correlate their DHCP reservations. Obtain the correct local key for each device before using the integration.
### pfSense DHCP Reservations for Tuya-Based Smart Devices
| **Description** | **IP Address** | **MAC Address** | **Hostname** | **Device ID** | **Local Key** |
| :--- | :--- | :--- | :--- | :--- | :--- |
| Bottom of Stairs | 10.0.0.200 | bcddc29072bf | ESP\_9072BF | 50316010bcddc29072bf | REDACTED |
| Right Monitor | 10.0.0.201 | bcddc2901aef | ESP\_901AEF | 50316010bcddc2901aef | REDACTED |
| Downstairs Light | 10.0.0.202 | bcddc28fe4c4 | ESP\_8FE4C4 | 74160333bcddc28fe4c4 | REDACTED |
| Right TV Light | 10.0.0.203 | b4e62d4bc3fe | ESP\_4BC3FE | 36087764b4e62d4bc3fe | REDACTED |
| Nightstand | 10.0.0.204 | b4e62d4bc3cb | ESP\_4BC3CB | 36087764b4e62d4bc3cb | REDACTED |
| Top of Stairs | 10.0.0.205 | bcddc2904ed9 | ESP\_904ED9 | 50316010bcddc2904ed9 | REDACTED |
| Bathroom | 10.0.0.206 | 2cf432220421 | ESP\_220421 | 105480752cf432220421 | REDACTED |
| Front Porch | 10.0.0.207 | bcddc2947aae | ESP\_947AAE | 50316010bcddc2947aae | REDACTED |
| Left Monitor | 10.0.0.208 | 2cf43221af1a | ESP\_21AF1A | 105480752cf43221af1a | REDACTED |
| Puppy Nook | 10.0.0.209 | cc50e3feaa2b | ESP\_FEAA2B | 76380710cc50e3feaa2b | REDACTED |
| TV | 10.0.0.210 | cc50e378bab9 | ESP\_78BAB9 | 35138222cc50e378bab9 | REDACTED |
| Left TV Light | 10.0.0.211 | cc50e378916d | ESP\_78916D | 10548075cc50e378916d | REDACTED |
| Bedroom Light | 10.0.0.212 | bcddc2907645 | | 50316010bcddc2907645 | REDACTED |
| Garden Water Pump | 10.0.0.213 | 98f4abef7c2c | | 2182401498f4abef7c2c | REDACTED |
| wifi Water Timer | 10.0.0.214 | | | eb54e6ae7c7536a4bbawyw | REDACTED |
| Tech Room Light Strips | 10.0.0.215 | | | eb7dd0deaab376a2ffsqwl | REDACTED |
| Irrigation Hub | 10.0.0.216 | 10d5615ab16b | | eb3b374a82a993d252j7v9 | REDACTED |
| Front Lawn Sprinkler | | | | ebace67e93f8fde4ccdv7p | 2a72efa9b2f36437 |
### Misc Color Profile Notes
- **1 NAME 3 4 0 255 2 5 1500 8000**
- 20 NAME 22 23 29 1000 21 24 2700 6500
## Related Documentation
- [Related Applications Documentation](<../../../reference/Applications/index.md>) — Find the connected deployments, procedures, and references for this subject.
+19
View File
@@ -0,0 +1,19 @@
---
tags:
- Applications
- Workflows
- Documentation
---
# Applications
## Purpose
Find workflows for applications. Follow the subject guide to choose the relevant environment and connect this material to the other document types.
## Includes
- Communication
- Email
- Files and Collaboration
- Home Automation
## Follow the Subject
[Applications](<../../reference/Applications/index.md>) explains the relationships and offers starting points for the documented tasks.
@@ -6,13 +6,16 @@ tags:
- Automation
---
## Purpose
Configure the AWX execution environment and workflow used to authenticate to domain-joined Windows targets with Kerberos. The AWX Operator deployment and Windows remote-management prerequisites must already be in place.
## Kerberos Implementation
You may find that you need to be able to run playbooks on domain-joined Windows devices using Kerberos. You need to go through some extra steps to set this up after you have successfully fully deployed AWX Operator into Kubernetes.
### Configure Windows Devices
You will need to prepare the Windows devices to allow them to be remotely controlled by Ansible playbooks. Run the following powershell script on all of the devices that will be managed by the Ansible AWX environment.
- [WinRM Prerequisite Setup Script](../enable-winrm-on-windows-devices.md)
- [WinRM Prerequisite Setup Script](<../../Identity and Certificates/Windows/Enable WinRM over HTTPS.md>)
### Create an AWX Instance Group
At this point, we need to make an "Instance Group" for the AWX Execution Environments that will use both a Keytab file and custom DNS servers defined by configmap files created below. Reference information was found [here](https://github.com/kurokobo/awx-on-k3s/blob/main/tips/use-kerberos.md#create-container-group). This group allows for persistence across playbooks/templates, so that if you establish a Kerberos authentication in one playbook, it will persist through the entire job's workflow.
@@ -57,8 +60,9 @@ Create the following files in the `/awx` folder on the AWX Operator server you d
bunny-lab.io = BUNNY-LAB.IO
```
Then we apply these configmaps to the AWX namespace with the following commands:
``` sh
Then we apply these configmaps to the AWX namespace with the following commands:
```sh
cd /awx
kubectl -n awx create configmap awx-kerberos-config --from-file=/awx/krb5.conf
kubectl apply -f custom_dns_records.yml
@@ -67,6 +71,7 @@ kubectl apply -f custom_dns_records.yml
- Open AWX UI and click on "**Instance Groups**" under the "**Administration**" section, then press "**Add > Add container group**".
- Enter a descriptive name as you like (e.g. `Kerberos`) and click the toggle "**Customize Pod Specification**".
- Put the following YAML string in "**Custom pod spec**" then press the "**Save**" button
```yaml title="Custom Pod Spec"
apiVersion: v1
kind: Pod
@@ -109,17 +114,19 @@ spec:
name: custom-dns
```
### Job Template & Inventory Examples
### Job Template and Inventory Examples
At this point, you need to adjust your exist Job Template(s) that need to communicate via Kerberos to domain-joined Windows devices to use the "Instance Group" of "**Kerberos**" while keeping the same Execution Environment you have been using up until this point. This will change the Execution Environment to include the Kerberos Keytab file in the EE at playbook runtime. When the playbook has completed running, (or if you are chain-loading multiple playbooks in a workflow job template), it will cease to exist. The kerberos keytab data will be regenerated at the next runtime.
Also add the following variables to the job template you have associated with the playbook below:
``` yaml
```yaml
---
kerberos_user: nicole.rappe@BUNNY-LAB.IO
kerberos_password: <DomainPassword>
```
You will want to ensure your inventory file is configured to use Kerberos Authentication as well, so the following example is a starting point:
```ini
virt-node-01 ansible_host=virt-node-01.bunny-lab.io
bunny-node-02 ansible_host=bunny-node-02.bunny-lab.io
@@ -137,17 +144,18 @@ ansible_winrm_server_cert_validation=ignore
#kerberos_user=nicole.rappe@BUNNY-LAB.IO #Optional, if you define this in the Job Template, it is not necessary here.
#kerberos_password=<DomainPassword> #Optional, if you define this in the Job Template, it is not necessary here.
```
!!! failure "Usage of Fully-Quality Domain Names"
It is **critical** that you define Kerberos-authenticated devices with fully qualified domain names. This is just something I found out from 4+ hours of troubleshooting. If the device is Linux or you are using NTLM authentication instead of Kerberos authentication, you can skip this warning. If you do not define the inventory using FQDNs, it will fail to run the commands against the targeted device(s).
In this example, the host is defined via FQDN: `virt-node-01 ansible_host=virt-node-01.bunny-lab.io`
### Kerberos Connection Playbook
At this point, you need a playbook that you can run in a Workflow Job Template (to keep things modular and simplified) to establish a connection to an Active Directory Domain Controller via Kerberos before running additional playbooks/templates against the actual devices.
At this point, you need a playbook that you can run in a Workflow Job Template (to keep things modular and simplified) to establish a connection to an Active Directory Domain Controller via Kerberos before running additional playbooks/templates against the actual devices.
You can visualize the connection workflow below:
``` mermaid
```mermaid
graph LR
A[Update AWX Project] --> B[Update Project Inventory]
B --> C[Establish Kerberos Connection]
@@ -175,7 +183,7 @@ The following playbook is an example pulled from https://git.bunny-lab.io
{{ kerberos_password }}
wkt /tmp/krb5.keytab
quit
EOF
EOF
environment:
KRB5_CONFIG: /etc/krb5.conf
register: generate_keytab_result
@@ -192,7 +200,7 @@ The following playbook is an example pulled from https://git.bunny-lab.io
- name: Acquire Kerberos ticket using keytab
ansible.builtin.shell: |
kinit -kt /tmp/krb5.keytab {{ kerberos_user }}
kinit -kt /tmp/krb5.keytab {{ kerberos_user }}
environment:
KRB5_CONFIG: /etc/krb5.conf
register: kinit_result
@@ -208,3 +216,5 @@ The following playbook is an example pulled from https://git.bunny-lab.io
when: kinit_result.rc == 0
```
## Related Documentation
- [Related AWX Documentation](<../../../reference/Automation/AWX/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -6,7 +6,8 @@ tags:
- Automation
---
**Purpose**: Once AWX is deployed, you will want to connect Gitea at https://git.bunny-lab.io. The reason for this is so we can pull in our playbooks, inventories, and templates automatically into AWX, making it more stateless overall and more resilient to potential failures of either AWX or the underlying Kubernetes Cluster hosting it.
## Purpose
Once AWX is deployed, you will want to connect Gitea at https://git.bunny-lab.io. The reason for this is so we can pull in our playbooks, inventories, and templates automatically into AWX, making it more stateless overall and more resilient to potential failures of either AWX or the underlying Kubernetes Cluster hosting it.
## Obtain Gitea Token
You already have this documented in Vaultwarden's password notes for awx.bunny-lab.io, but in case it gets lost, go to the [Gitea Token Page](https://git.bunny-lab.io/user/settings/applications) to set up an application token with read-only access for AWX, with a descriptive name.
@@ -72,4 +73,7 @@ Now you will want to connect this inventory to the inventory file(s) hosted in t
You want to make sure that the checkboxes for "**Overwrite**" and "**Overwrite Variables**" are checked. This ensures that if devices and/or group variables are removed from the inventory file in Gitea, they will also be removed from the inventory inside AWX.
## Webhooks
Optionally, set up webhooks in Gitea to trigger inventory updates in AWX upon changes in the repository. This section is not documented yet, but will eventually be documented.
Optionally, set up webhooks in Gitea to trigger inventory updates in AWX upon changes in the repository. This section is not documented yet, but will eventually be documented.
## Related Documentation
- [Related AWX Documentation](<../../../reference/Automation/AWX/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,76 @@
---
tags:
- Ansible
- AWX
- Automation
---
## Purpose
This document records the procedure for repair upgrades beyond awx operator 2.10.0. Follow the environment assumptions and commands below.
## Upgrading from 2.10.0 to 2.19.1+
There is a known issue with upgrading / install AWX Operator beyond version 2.10.0, because of how the PostgreSQL database upgrades from 13.0 to 15.0, and has changed permissions. The following workflow will help get past that and adjust the permissions in such a way that allows the upgrade to proceed successfully. If this is a clean installation, you can also perform this step if the fresh install of 2.19.1 is not working yet. (It wont work out of the box because of this bug). `The developers of AWX seem to just not care about this issue, and have not implemented an official fix themselves at this time).
### Create a Temporary Pod to Adjust Permissions
We need to create a pod that will mount the PostgreSQL PVC, make changes to permissions, then destroy the v15.0 pod to have the AWX Operator automatically regenerate it.
```yaml title="/awx/temp-pod.yml"
apiVersion: v1
kind: Pod
metadata:
name: temp-pod
namespace: awx
spec:
containers:
- name: temp-container
image: busybox
command: ['sh', '-c', 'sleep 3600']
volumeMounts:
- mountPath: /var/lib/pgsql/data
name: postgres-data
volumes:
- name: postgres-data
persistentVolumeClaim:
claimName: postgres-15-awx-postgres-15-0
restartPolicy: Never
```
```sh
# Deploy Temporary Pod
kubectl apply -f /awx/temp-pod.yaml
# Open a Shell in the Temporary Pod
kubectl exec -it temp-pod -n awx -- sh
# Adjust Permissions of the PostgreSQL 15.0 Database Folder
chown -R 26:root /var/lib/pgsql/data
exit
# Delete the Temporary Pod
kubectl delete pod temp-pod -n awx
# Delete the Crashlooped PostgreSQL 15.0 Pod to Regenerate It
kubectl delete pod awx-postgres-15-0 -n awx
# Track the Migration
kubectl get pods -n awx
kubectl logs -n awx awx-postgres-15-0
```
!!! warning "Be Patient"
This upgrade may take a few minutes depending on the speed of the node it is running on. Be patient and wait until the output looks something similar to this:
```text
root@awx:/awx# kubectl get pods -n awx
NAME READY STATUS RESTARTS AGE
awx-migration-24.6.1-bh5vb 0/1 Completed 0 9m55s
awx-operator-controller-manager-745b55d94b-2dhvx 2/2 Running 0 25m
awx-postgres-15-0 1/1 Running 0 12m
awx-task-7946b46dd6-7z9jm 4/4 Running 0 10m
awx-web-9497647b4-s4gmj 3/3 Running 0 10m
```
If you see a migration pod, like seen in the above example, you can feel free to delete it with the following command: `kubectl delete pod awx-migration-24.6.1-bh5vb -n awx`.
## Related Documentation
- [Related AWX Documentation](<../../../reference/Automation/AWX/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,90 @@
---
tags:
- Gitea
- Docker
- GitOps
---
## Purpose
Configure the Docker-based Gitea runner and repository workflow described in the May 2025 GitOps experiment. The example synchronizes repository files into a bind-mounted destination and sends an ntfy notification.
!!! info "Dated Docker Runner Example"
This is the configuration from the May 2025 runner experiment. Its Docker execution model and destination mounts differ from the Zensical host-runner workflow.
## Deploy the Docker Runner
When it comes to deploying a runner, (*assuming you want to use a docker-based runner*) it has a few simple things that need to be configured, the `docker-compose.yml` and the `.env` files. These tell the runner to reach out to Gitea server to register the runner with the given repository that you generated a registration token on.
```yaml title="docker-compose.yml"
version: "3.8"
services:
app:
image: docker.io/gitea/act_runner:latest
environment:
CONFIG_FILE: /config.yaml
GITEA_INSTANCE_URL: "${INSTANCE_URL}"
GITEA_RUNNER_REGISTRATION_TOKEN: "${REGISTRATION_TOKEN}"
GITEA_RUNNER_NAME: "${RUNNER_NAME}"
GITEA_RUNNER_LABELS: "${RUNNER_NAME}" # This can be anything, and is referenced by the workflow task(s) later.
volumes:
- /srv/containers/gitea-runner-mkdocs/config.yaml:/config.yaml # You have to manually make this file before you start the container
- /srv/containers/material-mkdocs/docs/docs:/Gitops_Destination # This is where the repository data will be copied to
```
```sh title=".env"
INSTANCE_URL=https://git.bunny-lab.io
RUNNER_NAME=gitea-runner-mkdocs
REGISTRATION_TOKEN=<Generated Here: https://git.bunny-lab.io/bunny-lab/docs/settings/actions/runners>
```
### Creating the `config.yaml`
The oddball thing about the way that I configured the Gitea Act Runner was telling it to run the container in "host mode" which tells it to run the tasks / workflows directly on the container itself instead of spinning up an instanced container (referred to as "*Docker-in-Docker*"). This keeps things simpler, but requires us to add a line to the `config.yaml` located at `/srv/containers/gitea-runner-mkdocs/config.yaml`. You can use your preferred text editor to add the following to the file's contents. This tells the runner to use itself for the tasks instead of an instanced docker container.
```yaml title="/srv/containers/gitea-runner-mkdocs/config.yaml"
container_engine: ""
```
!!! info "Quick Config Command"
```sh
mkdir -p /srv/containers/gitea-runner-mkdocs
echo 'container_engine: ""' > "/srv/containers/gitea-runner-mkdocs/config.yaml"
```
### Runner Workflow Task Files
When it comes to telling the runner what to do and how to do it, you create what are called runner "**Workflows**". These files reside within `<RepoRoot>/.gitea/workflows` and are `.yaml` format. If you have any familiarity with Ansible, the similarities are staggaring. You can have multiple workflows for one repository, with different flows that fire-off on different runners. An example of the flow used to replace Git-Repo-Updater's functionality can be seen below.
In the workflow below, it spins up a runner within the Alpine Linux environment that the `docker.io/gitea/act_runner:latest` uses, then installs NodeJS, Git, and Rsync for the core functionality that mirrors Git-Repo-Updater:
```yaml title=".gitea/workflows/gitops-automatic-deployment.yml"
name: GitOps Automatic Deployment
on:
push:
branches: [ main ]
jobs:
GitOps Automatic Deployment:
runs-on: gitea-runner-mkdocs
steps:
- name: Install Node.js, git, rsync, and curl
run: |
apk add --no-cache nodejs npm git rsync curl
- name: Checkout Repository
uses: actions/checkout@v3
- name: Copy Repository Data to Production Server
run: |
rsync -a --delete --exclude='.git/' --exclude='.gitea/' . /Gitops_Destination/
- name: Notify via NTFY
run: |
curl -d "https://docs.bunny-lab.io - Workflow Completed" https://ntfy.bunny-lab.io/gitea-runners
```
!!! note "`runs-on` Variable"
In this example workflow file, we are targeting the previously-mentioned `gitea-runner-mkdocs` runner, which we gave that "label" in the docker-compose.yaml file's `GITEA_RUNNER_LABELS` variable. You can name these labels whatever you want, as a way of organizing which runners run which workflows associated with a repository when changes are made to the repository.
## Related Documentation
- [Related Gitea Workflows](<../../../reference/Automation/Gitea Configuration Delivery.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,147 @@
---
tags:
- Gitea
- Zensical
- GitOps
---
## Purpose
Install and register the host-based Gitea Actions runner that synchronizes this documentation repository into `/srv/zensical/docs`. The Zensical account, watchdog service, and destination directory must already exist from the Zensical deployment.
## Install the Host Runner
Now is time for the arguably most-important stage of deployment, which is setting up a [Gitea Act Runner](https://docs.gitea.com/usage/actions/act-runner). This is how document changes in a Gitea repository will propagate automatically into Zensical's `/srv/zensical/docs` folder.
```sh
# Install Dependencies
sudo apt install -y nodejs npm git rsync curl
# Create dedicated Gitea runner service account
sudo useradd --system --create-home --home /var/lib/gitea_runner --shell /usr/sbin/nologin gitearunner || true
# Allow the runner to write documentation changes
sudo usermod -aG zensical gitearunner
# Allow the runner to start and stop Zensical Watchdog Service
sudo tee /etc/sudoers.d/gitearunner-systemctl > /dev/null <<'EOF'
gitearunner ALL=NOPASSWD: /usr/bin/systemctl start zensical-watchdog.service, /usr/bin/systemctl stop zensical-watchdog.service
EOF
sudo chmod 440 /etc/sudoers.d/gitearunner-systemctl
sudo chown root:root /etc/sudoers.d/gitearunner-systemctl
sudo visudo -c
# Download Newest Gitea Runner Binary (https://gitea.com/gitea/act_runner/releases)
cd /tmp
wget https://gitea.com/gitea/act_runner/releases/download/v0.2.13/act_runner-0.2.13-linux-amd64
sudo install -m 0755 act_runner-0.2.13-linux-amd64 /usr/local/bin/gitea_runner
gitea_runner --version
# Generate Gitea Runner Configuration
sudo mkdir -p /etc/gitea_runner
sudo chown gitearunner:gitearunner /etc/gitea_runner
sudo -u gitearunner gitea_runner generate-config > /etc/gitea_runner/config.yaml
```
### Configure Registration Token
- Navigate to: "**<Gitea Repo> > Settings > Actions > Runners**"
- If you don't see this, it needs to be enabled. Navigate to: "**<Gitea Repo> > Settings > "Enable Repository Actions: Enabled" > Update Settings**"
- Click the "**Create New Runner**" button on the top-right of the page and copy the registration token somewhere temporarily.
- Navigate back to the GuestVM running Zensical and run the following commands.
```sh
# Start Token Registration Process
sudo -u gitearunner env HOME=/var/lib/gitea_runner /usr/local/bin/gitea_runner register --config /etc/gitea_runner/config.yaml
# Gitea Instance URL: https://git.bunny-lab.io
# Gitea Runner Token: <Gitea-Runner-Token>
# Runner Name: zensical-docs-runner
# Move Runner Config to Correct Location & Configure Permissions
sudo mv /tmp/.runner /var/lib/gitea_runner/.runner
sudo chown gitearunner:gitearunner /var/lib/gitea_runner/.runner
sudo chmod 600 /var/lib/gitea_runner/.runner
```
### Create Service
Now we need to configure the Gitea runner to start automatically via a service just like the Zensical Watchdog service.
```sh
# Create Gitea Runner Service
sudo tee /etc/systemd/system/gitea-runner.service > /dev/null <<'EOF'
[Unit]
Description=Gitea Actions Runner (gitea_runner)
After=network-online.target
Wants=network-online.target
[Service]
Environment=HOME=/var/lib/gitea_runner
User=gitearunner
Group=gitearunner
WorkingDirectory=/var/lib/gitea_runner
ExecStart=/usr/local/bin/gitea_runner daemon --config /etc/gitea_runner/config.yaml
Restart=always
RestartSec=2
[Install]
WantedBy=multi-user.target
EOF
# Remove Container-Based Configurations to Force Runner to Run in Host Mode
sudo sed -i \
'/^[[:space:]]*labels:/,/^[[:space:]]*cache:/{
/^[[:space:]]*labels:/c\ labels:\n - "zensical-host:host"
/^[[:space:]]*cache:/!d
}' \
/etc/gitea_runner/config.yaml
# Enable and Start the Service
sudo systemctl daemon-reload
sudo systemctl enable --now gitea-runner.service
```
### Repository Workflow
Place the following file into your documentation repository at the given location and this will enable the runner to execute when changes happen to the repository data.
```yaml title="gitea/workflows/automatic-deployment.yml"
name: Automatic Documentation Deployment
on:
push:
branches: [ main ]
jobs:
zensical_deploy:
name: Sync Docs to https://kb.bunny-lab.io
runs-on: zensical-host
steps:
- name: Checkout Repository
uses: actions/checkout@v3
- name: Stop Zensical Service
run: sudo /usr/bin/systemctl stop zensical-watchdog.service
- name: Sync repository into /srv/zensical/docs
run: |
rsync -rlD --delete \
--exclude='.git/' \
--exclude='.gitea/' \
--exclude='assets/' \
--exclude='schema/' \
--exclude='stylesheets/' \
--exclude='schema.json' \
--chmod=D2775,F664 \
. /srv/zensical/docs/
- name: Start Zensical Service
run: sudo /usr/bin/systemctl start zensical-watchdog.service
- name: Notify via NTFY
if: always()
run: |
curl -d "https://kb.bunny-lab.io - Zensical job status: ${{ job.status }}" https://ntfy.bunny-lab.io/gitea-runners
```
## Related Documentation
- [Zensical Deployment](<../../../deployments/automation/Documentation/Zensical.md>) — Prepare the service account, watchdog, and destination directory first.
- [Related Gitea Workflows](<../../../reference/Automation/Gitea Configuration Delivery.md>) — Find the connected deployments, procedures, and references for this subject.
+17
View File
@@ -0,0 +1,17 @@
---
tags:
- Automation
- Workflows
- Documentation
---
# Automation
## Purpose
Find workflows for automation. Follow the subject guide to choose the relevant environment and connect this material to the other document types.
## Includes
- AWX
- Gitea
## Follow the Subject
[Automation](<../../reference/Automation/index.md>) explains the relationships and offers starting points for the documented tasks.
@@ -5,7 +5,8 @@ tags:
- Disaster Recovery
---
**Purpose**: You may find that you need to adopt a device that was onboarded by a different Veeam Backup & Replication server. Maybe the old server died, or maybe you are restructuring your backup infrastructure, and want a new server taking over the backup responsibilities for the device.
## Purpose
You may find that you need to adopt a device that was onboarded by a different Veeam Backup & Replication server. Maybe the old server died, or maybe you are restructuring your backup infrastructure, and want a new server taking over the backup responsibilities for the device.
If this happens, Veeam will complain that the device is managed by a different server. To circumvent this, perform the following changes in the Windows Registry based on the version of Veeam Backup & Replication you are currently using, then try to Update the Agent / Backup the agent again, and it should be successful after the registry changes are made.
@@ -14,14 +15,17 @@ https://forums.veeam.com/servers-workstations-f49/how-do-we-move-agent-to-associ
=== "VBR v11"
```jsx title="HKEY_LOCAL_MACHINE\SOFTWARE\Veeam\Veeam Backup and Replication"
```text title="HKEY_LOCAL_MACHINE\SOFTWARE\Veeam\Veeam Backup and Replication"
AgentDiscoveryIgnoreOwnership
REG_DWORD (32-bit) Value: 1
```
=== "VBR v12"
```jsx title="HKEY_LOCAL_MACHINE\SOFTWARE\Veeam\Veeam Backup and Replication"
```text title="HKEY_LOCAL_MACHINE\SOFTWARE\Veeam\Veeam Backup and Replication"
ProtectionGroupIgnoreOwnership
REG_DWORD (32-bit) Value: 1
```
```
## Related Documentation
- [Related Backup and Recovery Documentation](<../../../reference/Backup and Recovery/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -5,6 +5,9 @@ tags:
- Disaster Recovery
---
## Purpose
Repair the documented Veeam Cloud Connect certificate failure when a backup reports that no cloud gateways are available because gateway certificates cannot be validated.
### Symptoms
When you try to run a backup to a remote backup server, the backup job fails and gives the following error:
@@ -33,4 +36,7 @@ At this point, the endpoint should immediately trust the new certificate from th
```batch
::Connect to the backup server and download current configuration settings.
"C:\Program Files\Veeam\Endpoint Backup\Veeam.Agent.Configurator.exe" -syncnow
```
```
## Related Documentation
- [Related Backup and Recovery Documentation](<../../../reference/Backup and Recovery/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -5,8 +5,8 @@ tags:
- Disaster Recovery
---
**Purpose**:
There may come a time that you need to free up space in a Veeam Backup & Replication backup repository because you are running out of space. In these cases, you need to manually trim the older backups in a specific way to ensure this is non-destructive.
## Purpose
There may come a time that you need to free up space in a Veeam Backup & Replication backup repository because you are running out of space. In these cases, you need to manually trim the older backups in a specific way to ensure this is non-destructive.
## Manual Removal of Backup Data
You need to perform these steps to carefully delete the oldest full backup chain w/ incrementals.
@@ -30,4 +30,7 @@ At this point, you have deleted the backup files and re-scanned the backup repos
- Navigate to "**Home > Backups > Disk**"
- Locate the backup job associated with the device's backup files you deleted
- Right-click the associated backup job > "**Properties...**"
- In the Backup Properties window, right-click the missing restore point(s) and click "**Forget**" > "**All Unavailable Backups**"
- In the Backup Properties window, right-click the missing restore point(s) and click "**Forget**" > "**All Unavailable Backups**"
## Related Documentation
- [Related Backup and Recovery Documentation](<../../../reference/Backup and Recovery/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -15,4 +15,7 @@ If you find that you need to migrate cloud backups that are being sent to a serv
- Log back into Veeam Backup & Replication Console and re-scan the new repositories
### Move Backup Job Location in VSPC Portal
At this point, we need to migrate point the backup job(s) that were affected to the new location. This is a job-level change, not company-level change.
At this point, we need to migrate point the backup job(s) that were affected to the new location. This is a job-level change, not company-level change.
## Related Documentation
- [Related Backup and Recovery Documentation](<../../../reference/Backup and Recovery/index.md>) — Find the connected deployments, procedures, and references for this subject.
+16
View File
@@ -0,0 +1,16 @@
---
tags:
- Backup and Recovery
- Workflows
- Documentation
---
# Backup and Recovery
## Purpose
Find workflows for backup and recovery. Follow the subject guide to choose the relevant environment and connect this material to the other document types.
## Includes
- Veeam
## Follow the Subject
[Backup and Recovery](<../../reference/Backup and Recovery/index.md>) explains the relationships and offers starting points for the documented tasks.
@@ -0,0 +1,82 @@
---
tags:
- Containers
- Docker
- Containerization
---
## Purpose
This document will outline the general workflow of using Visual Studio Code to author and update custom containers and push them to a container registry hosted in Gitea. This will be referencing the `git-repo-updater` project throughout.
!!! note "Assumptions"
This document assumes you are authoring the containers in Microsoft Windows, and does not include the fine-tuning necessary to work in Linux or MacOS environments. You are on your own if you want to author containers in Linux.
## Install Visual Studio Code
The management of the Gitea repositories, Dockerfile building, and pushing container images to the Gitea container registry will all involve using just Visual Studio Code. You can download Visual Studio Code from this [direct download link](https://code.visualstudio.com/docs/?dv=win64user).
## Configure Required Docker Extensions
You will need to locate and install the `Dev Containers`, `Docker`, and `WSL` extensions in Visual Studio Code to move forward. This may request that you install Docker Desktop onto your computer as part of the installation process. Proceed to do so, then when the Docker "Engine" is running, you can proceed to the next step.
!!! warning
You need to have Docker Desktop "Engine" running whenever working with containers, as it is necessary to build the images. VSCode will complain if it is not running.
## Add Gitea Container Registry
At this point, we need to add a registry to Visual Studio Code so it can proceed with pulling down the repository data.
- Click the Docker icon on the left-hand toolbar
- Under "**Registries**", click "**Connect Registry...**"
- In the dropdown menu that appears, click "**Generic Registry V2**"
- Enter `https://git.bunny-lab.io/container-registry`
- Registry Username: `nicole.rappe`
- Registry Password or Personal Access Token: `Personal Access API Token You Generated in Gitea`
- You will now see a sub-listing named "**Generic Registry V2**"
- If you click the dropdown, you will see "**https://git.bunny-lab.io/container-registry**"
- Under this section, you will see any containers in the registry that you have access to, in this case, you will see `container-registry/git-repo-updater`
## Add Source Control Repository
Now it is time to pull down the repository where the container's core elements are stored on Gitea.
- Click the "**Source Control**" button on the left-hand menu then click the "**Clone Repository**" button
- Enter `https://git.bunny-lab.io/container-registry/git-repo-updater.git`
- Click the dropdown menu option "**Clone from URL**" then choose a location to locally store the repository on your computer
- When prompted with "**Would you like to open the cloned repository**", click the "**Open**" button
## Making Changes
You will be presented with four files in this specific repository. `.env`, `docker-compose.yml`, `Dockerfile`, and `repo_watcher.sh`
- `.env` is the environment variables passed to the container to tell it which ntfy server to talk to, which credentials to use with Gitea, and which repositories to download and push into production servers
- `docker-compose.yml` is an example docker-compose file that can be used in Portainer to deploy the server along with the contents of the `.env` file
- `Dockerfile` is the base of the container, telling docker what operating system to use and how to start the script in the container
- `repo_watcher.sh` is the script called by the `Dockerfile` which loops checking for updates in Gitea repositories that were configured in the `.env` file
### Push to Repository
When you make any changes, you will need to first commit them to the repository
- Save all of the edited files
- Click the "**Source Control**" button in the toolbar
- Write a message about what you changed in the commit description field
- Click the "**Commit**" button
- Click the "**Sync Changes**" button that appears
- You may be presented with various dialogs, just click the equivalant of "**Yes/OK**" to each of them
### Build the Dockerfile
At this point, we need to build the dockerfile, which takes all of the changes and packages it into a container image
- Navigate back to the file explorer inside of Visual Studio Code
- Right-click the `Dockerfile`, then click "**Build Image...**"
- In the "Tag Image As..." window, type in `git.bunny-lab.io/container-registry/git-repo-updater:latest`
- When you navigate back to the Docker menu, you will see a new image appear under the "**Images**" section
- You should see something similar to "Latest - X Seconds Ago` indicating this is the image you just built
- Delete the older image(s) by right-clicking on them and selecting "**Remove...**"
- Push the image to the container registry in Gitea by right-clicking the latest image, and selecting "**Push...**"
- In the dropdown menu that appears, enter `git.bunny-lab.io/container-registry/git-repo-updater:latest`
- You can confirm if it was successful by navigating to the [Gitea Container Webpage](https://git.bunny-lab.io/container-registry/-/packages/container/git-repo-updater/latest) and seeing if it says "**Published Now**" or "**Published 1 Minute Ago**"
!!! warning "CRLF End of Line Sequences"
When you are editing files in the container's repository, you need to ensure that Visual Studio Code is editing that file in "**LF**" mode and not "**CRLF**". You can find this toggle at the bottom-right of the VSCode window. Simply clicking on the letters "**CRLF**" will let you toggle the file to "**LF**". If you do not make this change, the container will misunderstand the dockerfile and/or scripts inside of the container and have runtime errors.
## Deploy the Container
You can now use the `.env` file along with the `docker-compose.yml` file inside of Portainer to deploy a stack using the container you just built / updated.
## Related Documentation
- [Related Containers Documentation](<../../../reference/Containers/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,31 @@
---
tags:
- Docker
- Macvlan
- Networking
---
## Purpose
You may find that you only have one network adapter on a server / VM and need to have multiple virtual networks associated with it. For example, Home Assistant exists on the `192.168.3.0/24` network but it needs to also access devices on the `192.168.4.0/24` surveillance network. To facilitate this, we will make a MACVLAN Sub-Interface. This will make a virtual interface that is parented to the actual physical interface.
!!! info "Assumptions"
It is assumed that you are running Rocky Linux (or CentOS / RedHat).
## Create the Permanent Sub-Interface
You will begin with making a new interface, it will have the name `macvlan0`.
```sh
nmcli connection add type macvlan ifname surveillance dev ens18 mode bridge ipv4.method manual ipv4.addresses 192.168.4.100/24
nmcli connection up macvlan-surveillance
nmcli connection show
```
## Bind a Docker Network to the Sub-Interface
Now you need to run the following command to allow docker to use this interface for the `surveillance_network`
```sh
docker network create -d macvlan --subnet=192.168.4.0/24 --gateway=192.168.4.1 -o parent=surveillance surveillance_network
```
## Related Documentation
- [Related Containers Documentation](<../../../reference/Containers/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,60 @@
---
tags:
- Containers
- Docker
- Bash
- Scripting
- Linux
---
## Purpose
If you find that you need to migrate a container, along with any supporting files, permissions, etc from an old server to a new server, rsync helps make this as painless as possible.
Be sure to perform the following steps to make sure that you can copy the container's files.
!!! warning
You need to stop the running containers on the old server before copying their data over, otherwise the state of the data may be unstable. Once you have migrated the data, you can spin up the containers on the new server and confirm they work before deleting the data on the old server.
On the destination (new) server, the directory needs to exist and be writable via the person copying the data over SSH:
## Copying Data Between the Old and New Servers
=== "Safe Method"
```sh
sudo mkdir -p /srv/containers/example
sudo chmod 740 /srv/containers/example
sudo chown nicole:nicole /srv/containers/example
```
=== "Quick & Dirty Method"
```sh
sudo mkdir -p /srv/containers
sudo chmod 777 /srv/containers
```
On the source (old) server, perform an rsync over to the new server, authenticating yourself as you will be prompted to do so:
=== "Safe Method"
```sh
rsync -avz -e ssh --progress /srv/containers/example/* nicole@192.168.3.30:/srv/containers/example
```
=== "Quick & Dirty Method"
```sh
rsync -avz -e ssh --progress /srv/containers/example nicole@192.168.3.30:/srv/containers
```
=== "Quick & Dirty w/ Provided SSH Key Method"
This method assumes that you have the private key for your SSH-based authentication locally on the server somewhere safe with permissions `chmod 600` applied to it. In this example, I placed the private key at `/tmp/id_rsa_OpenSSH`.
```sh
rsync -avz -e "ssh -i /tmp/id_rsa_OpenSSH" --progress /srv/containers/pihole nicole@192.168.3.62:/srv/containers
```
## Spinning Up Docker / Portainer Stack
Once everything has been moved over, copy the `docker-compose` and `.env` (environment variables) from the old server to the new one, pointing to the same location since we maintained the same folder structure, and the container should spin up like nothing ever happened.
## Related Documentation
- [Related Containers Documentation](<../../../reference/Containers/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -5,11 +5,11 @@ tags:
- Containerization
---
# Migrating `docker-compose.yml` to Rancher RKE2 Cluster
You may be comfortable operating with Portainer or `docker-compose`, but there comes a point where you might want to migrate those existing workloads to a Kubernetes cluster as easily-as-possible. Lucklily, there is a way to do this using a tool called "**Kompose**'. Follow the instructions seen below to convert and deploy your existing `docker-compose.yml` into a Kubernetes cluster such as Rancher RKE2.
## Purpose
Convert the example Docker Compose workload into Kubernetes manifests and expose it through the documented Rancher RKE2 and Traefik environment.
!!! info "RKE2 Cluster Deployment"
This document assumes that you have an existing Rancher RKE2 cluster deployed. If not, you can deploy one following the [Deploy RKE2 Cluster](../../../../deployments/platforms/containerization/kubernetes/deployment/rancher-rke2.md) documentation.
This document assumes that you have an existing Rancher RKE2 cluster deployed. If not, you can deploy one following the [Deploy RKE2 Cluster](<../../../deployments/Containers/Kubernetes/Rancher RKE2.md>) documentation.
We also assume that the cluster name within Rancher RKE2 is named `local`, which is the default cluster name when setting up a Kubernetes Cluster in the way seen in the above documentation.
@@ -24,7 +24,7 @@ This will attempt to convert the `docker-compose.yml` file into a Kubernetes man
=== "(Original) docker-compose.yml"
``` yaml
```yaml
version: "2.1"
services:
ntfy:
@@ -57,7 +57,7 @@ This will attempt to convert the `docker-compose.yml` file into a Kubernetes man
=== "(Converted) ntfy-k8s.yaml"
``` yaml
```yaml
---
apiVersion: v1
kind: Service
@@ -182,7 +182,7 @@ At this point, we need to create a namespace. This basically isolates the netwo
- The name for the namespace should be named based on its operational-context, such as `prod-ntfy` or `dev-ntfy`.
### Import Converted YAML Manifest into Namespace
At this point, we can now proceed to import the YAML file we generated in the beginning of this document.
At this point, we can now proceed to import the YAML file we generated in the beginning of this document.
- Navigate to: **Clusters > `local` > Cluster > Projects/Namespaces**
- At the top-right of the screen will be an upload / up-arrow button with tooltip text stating "Import YAML" > Click on this button
@@ -221,13 +221,13 @@ If you were able to successfully verify access to the service when talking to it
!!! info "Section Considerations"
This section of the document does not (*currently*) cover the process of setting up health checks to ensure that the load-balanced server destinations in the reverse proxy are online before redirecting traffic to them. This is on my to-do list of things to implement to further harden the deployment process.
This section also does not cover the process of setting up a reverse proxy. If you want to follow along with this document, you can deploy a Traefik reverse proxy via the [Traefik](../../../../deployments/services/edge/traefik.md) deployment documentation.
This section also does not cover the process of setting up a reverse proxy. If you want to follow along with this document, you can deploy a Traefik reverse proxy via the [Traefik](<../../../deployments/Networking and Access/Reverse Proxies/Traefik.md>) deployment documentation.
With the above considerations in-mind, we just need to make some small changes to the existing Traefik configuration file to ensure that it load-balanced across every node of the cluster to ensure high-availability functions as-expected.
=== "(Original) ntfy.bunny-lab.io.yml"
``` yaml
```yaml
http:
routers:
ntfy:
@@ -237,7 +237,7 @@ With the above considerations in-mind, we just need to make some small changes t
certResolver: letsencrypt
service: ntfy
rule: Host(`ntfy.bunny-lab.io`)
services:
ntfy:
loadBalancer:
@@ -248,7 +248,7 @@ With the above considerations in-mind, we just need to make some small changes t
=== "(Updated) ntfy.bunny-lab.io.yml"
``` yaml
```yaml
http:
routers:
ntfy:
@@ -273,3 +273,5 @@ With the above considerations in-mind, we just need to make some small changes t
!!! success "Verify Access via Reverse Proxy"
If everything worked, you should be able to access the service at https://ntfy.bunny-lab.io, and if one of the cluster nodes goes offline, Rancher will automatically migrate the load to another cluster node which will take over the web request.
## Related Documentation
- [Related Containers Documentation](<../../../reference/Containers/index.md>) — Find the connected deployments, procedures, and references for this subject.
+17
View File
@@ -0,0 +1,17 @@
---
tags:
- Containers
- Workflows
- Documentation
---
# Containers
## Purpose
Find workflows for containers. Follow the subject guide to choose the relevant environment and connect this material to the other document types.
## Includes
- Docker
- Kubernetes
## Follow the Subject
[Containers](<../../reference/Containers/index.md>) explains the relationships and offers starting points for the documented tasks.
@@ -0,0 +1,37 @@
---
tags:
- Active Directory
- Group Policy
- Authentication
---
## Purpose
To deploy a shortcut to the desktop pointing to a network share's root path. (e.g. `\\storage.bunny-lab.io`). There is a quirk with how Windows handles network shares and shortcuts and doesn't like when you point the shortcut to a root UNC path.
### Group Policy Location
```mermaid
graph LR
A[Create Group Policy] --> B[User Configuration]
B --> C[Preferences]
C --> D[Windows Settings]
D --> E[Shortcuts]
```
### Group Policy Settings
- **Action**: `Update`
- **Name**: `<FriendlyName>`
- **Target Type**: `File System Object`
- **Location**: `Desktop`
- **Target Path**: `C:\windows\explorer.exe`
- **Arguments**: `\\storage.bunny-lab.io`
- **Start In**: `<Blank>`
- **Shortcut Key**: `<None>`
- **Run**: `Normal Window`
- **Icon File Path**: `%SystemRoot%\System32\SHELL32.dll`
- **Icon Index**: `9`
### Additional Notes
Navigate to the "**Common**" tab in the properties of the shortcut, and check the "**Run in logged-on user's security context (user policy option)**".
## Related Documentation
- [Related Identity and Certificates Documentation](<../../../reference/Identity and Certificates/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,17 @@
---
tags:
- Active Directory
- Authentication
---
## Purpose
If you have a device that lost trust in the domain for some reason, and won't let you login using domain credentials, run the following command as a local administrator on the device to repair trust.
```powershell
Test-ComputerSecureChannel -Repair -Credential (Get-Credential)
```
If it outputs `True`, go ahead and log out then try to login again with the domain credentials.
## Related Documentation
- [Related Identity and Certificates Documentation](<../../../reference/Identity and Certificates/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -5,25 +5,30 @@ tags:
- SSL
---
**Purpose**: Sometimes you may find that you need to convert a `.crt` or `.pem` certificate file into a `.pfx` file that Microsoft IIS Server Manager can import for something like Exchange Server or another custom IIS-based server.
## Purpose
Sometimes you may find that you need to convert a `.crt` or `.pem` certificate file into a `.pfx` file that Microsoft IIS Server Manager can import for something like Exchange Server or another custom IIS-based server.
# Download the Certificate Files
## Download the Certificate Files
This step will vary based on how you are obtaining the certificates. The primary thing to focus on is making sure you have the certificate file and the private key.
```jsx title="Certificate Folder Structure"
```text title="Certificate Folder Structure"
certificate.crt
certificate.pem
gd-g2_iis_intermediates.p7b
private.key
```
# Convert using OpenSSL
## Convert using OpenSSL
You will need a linux machine such as Ubuntu 22.04LTS, or to download the Windows equivelant of OpenSSL in order to run the necessary commands to convert and package the files into a `.pfx` file that IIS Server Manager can use.
!!! note
You need to make sure that all of the certificate files as well as private key are in the same folder (to keep things simple) during the conversion process. **It will prompt you to enter a password for the PFX file, choose anything you want.**
```jsx title="OpenSSL Conversion Command"
```sh title="OpenSSL Conversion Command"
openssl pkcs12 -export -out IIS-Certificate.pfx -inkey private.key -in gd-g2_iis_intermediates.p7b -in certificate.crt
```
!!! tip
You can rename the files anything you want for organizational purposes. Afterall, they are just plaintext files. For example, you could rename `gd-g2_iis_intermediates.p7b` to `intermediate.bundle` and it would still work without issue in the command. During the import phase in IIS Server Manager, you can check a box to enable Exporting the certificate, effectively reverse-engineering it back into a certificate and private key.
## Related Documentation
- [Related Identity and Certificates Documentation](<../../../reference/Identity and Certificates/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,51 @@
---
tags:
- Active Directory
- LDAPS
- Certificates
---
## Purpose
Export the CA chain or domain-controller certificate required by an application that connects to Active Directory over LDAPS. Use the certificate objects from the documented CA and domain-controller environment.
## Export the LDAPS Certificate for Third-Party Applications
Some applications do not automatically trust your internal PKI and require you to manually install the certificate used by your domain controllers for LDAPS. In most cases, you should export the issuing CA certificates rather than the individual domain controller certificate. Only export the domain controller certificate if the third-party application explicitly requires it.
### Export the Root and Subordinate CA Certificates
The Root CA and Subordinate CA certificates establish trust for every domain controller certificate issued by your PKI.
From any domain-joined system:
- Launch `certlm.msc`
- Navigate to "**Trusted Root Certification Authorities > Certificates**"
- Locate your Root CA certificate
- Right-click the certificate and select "**All Tasks > Export...**"
- Select "**No, do not export the private key**"
- Export the certificate as either:
- `DER encoded binary X.509 (.CER)`, or
- `Base-64 encoded X.509 (.CER)`
- Navigate to "**Intermediate Certification Authorities > Certificates**"
- Locate your Subordinate CA certificate
- Repeat the export process
Import both certificates into the trusted certificate store required by the third-party application.
### Export a Domain Controller Certificate
If the application requires the LDAPS server certificate itself:
- On the target domain controller, launch `certlm.msc`
- Navigate to "**Personal > Certificates**"
- Locate the certificate issued to the domain controller's FQDN that includes **Server Authentication** as an intended purpose
- If the certificate's intended purpose looks like `Client Authentication, Server Authentication, Smart Card Logon, KDC Authentication` this cert may be more versatile for you.
- Right-click the certificate and select "**All Tasks > Export...**"
- Select "**No, do not export the private key**"
- Export the certificate as either:
- `DER encoded binary X.509 (.CER)`, or
- `Base-64 encoded X.509 (.CER)`
!!! warning "Do Not Export the Private Key"
Third-party LDAPS clients require only the public certificate. Do **not** export the certificate as a `.pfx` file or include the private key unless the vendor explicitly documents that requirement.
## Related Documentation
- [Certificate Services Deployment](<../../../deployments/Identity and Certificates/Active Directory/Certificate Services.md>) — Identify the CA chain used by the LDAPS clients.
- [Related Identity and Certificates Documentation](<../../../reference/Identity and Certificates/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,77 @@
---
tags:
- Active Directory
- Certificate Services
- PKI
---
## Purpose
Publish the root and subordinate CA revocation lists for the documented two-tier Active Directory certificate environment. The Root CA and HTTP distribution point must already be configured as described in the certificate deployment.
## CRL Publishing and Maintenance
CRLs must be generated and published on a recurring basis. If a CRL expires, certificate validation may fail even if the CA services themselves are running.
### Root CA CRL Publishing
Because the Root CA is offline, periodically bring it online only long enough to generate a new CRL and copy it to the HTTP distribution point.
On `LAB-CA-01`:
```powershell
certutil -crl
```
Copy the generated CRL from:
```text
C:\Windows\System32\CertSrv\CertEnroll\
```
to the IIS publication directory on `LAB-CA-02`:
```text
C:\inetpub\wwwroot\pki\
```
Validate:
```powershell
Invoke-WebRequest http://pki.bunny-lab.io/pki/BunnyLab-RootCA.crl
```
### Subordinate CA CRL Publishing
On `LAB-CA-02`:
```powershell
certutil -crl
```
Copy or confirm the Subordinate CA CRL exists in:
```text
C:\inetpub\wwwroot\pki\
```
Validate the URL from a domain-joined system.
### Operational Monitoring
Monitor CRL expiration and publication. Certificate validation failures can occur if CRLs expire, even if certificates themselves have not expired.
Recommended operational tasks:
- Track Root CA CRL expiration.
- Track Subordinate CA CRL expiration.
- Verify HTTP CRL URLs after each publication.
- Keep the Root CA offline except during controlled maintenance windows.
- Document the expected CRL filenames generated in `C:\Windows\System32\CertSrv\CertEnroll\`.
!!! abstract "Raw Unprocessed/Unimplemented Steps"
Publish CRLs regularly, configure overlap periods, and monitor expiration. Enable Delta CRLs on the Subordinate CA, but not on the Root.
Security Recommendations
- Harden CA servers; limit access to PKI admins.
- Use BitLocker or HSM for key protection.
- Monitor issuance and renewals with audit logs and scripts.
## Related Documentation
- [Certificate Services Deployment](<../../../deployments/Identity and Certificates/Active Directory/Certificate Services.md>) — Confirm the CA names, publication paths, and HTTP distribution point.
- [Related Identity and Certificates Documentation](<../../../reference/Identity and Certificates/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,27 @@
---
tags:
- Gitea
- Keycloak
- OAuth2
- Authentication
---
## Purpose
Configure the documented Gitea OAuth2 client settings after Keycloak and Gitea are deployed.
### OAuth2 Configuration
These are variables referenced by the associated service to connect its authentication system to [Keycloak](<../../../deployments/Identity and Certificates/Keycloak/Deploy Keycloak.md>).
| **Parameter** | **Value** |
| :--- | :--- |
| Authentication Name | `auth-bunny-lab-io` |
| OAuth2 Provider | `OpenID Connect` |
| Client ID (Key) | `git-bunny-lab-io` |
| Client Secret | `https://auth.bunny-lab.io > Clients > git-bunny-lab-io > Credentials > Client Secret` |
| OpenID Connect Auto Discovery URL | `https://auth.bunny-lab.io/realms/master/.well-known/openid-configuration` |
| Skip Local 2FA | Yes |
## Related Documentation
- [Gitea Deployment](<../../../deployments/automation/Gitea/Gitea.md>) — Prepare the application before configuring OAuth2.
- [Keycloak Integrations](<../../../reference/Identity and Certificates/Keycloak Integrations.md>) — Review the identity-provider deployment and other integrations.
- [Related Identity and Certificates Documentation](<../../../reference/Identity and Certificates/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,30 @@
---
tags:
- Portainer
- Keycloak
- OAuth2
- Authentication
---
## Purpose
Configure the documented Portainer OAuth2 settings after Keycloak and Portainer are deployed.
### OAuth2 Configuration
These are variables referenced by the associated service to connect its authentication system to [Keycloak](<../../../deployments/Identity and Certificates/Keycloak/Deploy Keycloak.md>).
| **Parameter** | **Value** |
| :--- | :--- |
| Client ID | `container-node-01` |
| Client Secret | `https://auth.bunny-lab.io > Clients > container-node-01 > Credentials > Client Secret` |
| Authorization URL | `https://auth.bunny-lab.io/realms/master/protocol/openid-connect/auth` |
| Access Token URL | `https://auth.bunny-lab.io/realms/master/protocol/openid-connect/token` |
| Resource URL | `https://auth.bunny-lab.io/realms/master/protocol/openid-connect/userinfo` |
| Redirect URL | `https://192.168.3.19:9443` |
| Logout URL | `https://auth.bunny-lab.io/realms/master/protocol/openid-connect/logout` |
| User Identifier | `email` |
| Scopes | `email openid profile` |
## Related Documentation
- [Portainer Deployment](<../../../deployments/Containers/Docker/Deploy Portainer.md>) — Prepare the application before configuring OAuth2.
- [Keycloak Integrations](<../../../reference/Identity and Certificates/Keycloak Integrations.md>) — Review the identity-provider deployment and other integrations.
- [Related Identity and Certificates Documentation](<../../../reference/Identity and Certificates/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,19 @@
---
tags:
- MFA
- Identity and Certificates
---
## Purpose
Sometimes you may need to change the MFA on an account, by adding a new email or phone number for SMS-based MFA. This can be done fairly quickly and only involves a few steps:
- Navigate to the [Azure Web Portal](https://portal.azure.com) and log in using your Office365 admin credentials.
- Navigate to the [Azure Active Directory (Microsoft Entra ID) Users List](https://portal.azure.com/#view/Microsoft_AAD_UsersAndTenants/UserManagementMenuBlade/~/AllUsers)
- Click on the User Account that needs their MFA information changed / wiped
- On the left-hand navigation menu, click on "**Authentication Methods**" at the bottom
- Make adjustments to existing methods or click on "**+ Add Authentication Method**"
- Valid options generally are Phone Numbers, Email Addresses, and a "**Temporary Access Pass**"
- Save the changes by clicking the "**Add**" button, then have the user attempt to log in again using their MFA method configured
## Related Documentation
- [Related Identity and Certificates Documentation](<../../../reference/Identity and Certificates/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -6,10 +6,10 @@ tags:
- Automation
---
**Purpose**:
## Purpose
You will need to enable secure WinRM management of the Windows devices you are running playbooks against, as compared to the Linux devices. The following powershell script needs to be ran on every Windows device you intend to run Ansible playbooks on. This script can also be useful for simply enabling / resetting WinRM configurations for Hyper-V hosts in general, just omit the Powershell script remote signing section if you dont plan on using it for Ansible.
``` powershell
```powershell
# Script to configure WinRM over HTTPS on the Hyper-V host
# Ensure WinRM is enabled
@@ -77,3 +77,6 @@ Set-ExecutionPolicy RemoteSigned -Force
Write-Host "Configuration complete. The Hyper-V host is ready for remote management over HTTPS with Kerberos authentication."
```
## Related Documentation
- [Related Identity and Certificates Documentation](<../../../reference/Identity and Certificates/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,20 @@
---
tags:
- Identity and Certificates
- Workflows
- Documentation
---
# Identity and Certificates
## Purpose
Find workflows for identity and certificates. Follow the subject guide to choose the relevant environment and connect this material to the other document types.
## Includes
- Active Directory
- Certificates
- Keycloak
- Microsoft 365
- Windows
## Follow the Subject
[Identity and Certificates](<../../reference/Identity and Certificates/index.md>) explains the relationships and offers starting points for the documented tasks.
@@ -0,0 +1,37 @@
---
tags:
- Linux
- Networking
---
## Purpose
This is a scaffold document outlining the high level of changing an IP address of a server in either Debian or RHEL based operating systems.
=== "Ubuntu / Debian"
```sh
# Edit Netplan File
nano /etc/netplan/<name-of-netplan-file>
# <edit existing networking> --> <save file>
# Apply Netplan Changes
netplan apply
```
=== "Rocky / Fedora / RHEL"
```sh
# Modify the Existing Connection via nmcli
nmcli connection modify ens18-connection \
ipv4.addresses 192.168.3.13/24 \
ipv4.gateway 192.168.3.1 \
ipv4.dns "192.168.3.25,192.168.3.26" \
ipv4.method manual
# Bring the Connection Online
sudo nmcli connection down ens18-connection
sudo nmcli connection up ens18-connection
```
## Related Documentation
- [Related Networking and Access Documentation](<../../../reference/Networking and Access/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,30 @@
---
tags:
- SSH
- Bash
- Authentication
- Scripting
- Linux
---
## Purpose
*Purpose*: Sometimes you need two linux computers to be able to talk to eachother without requiring a password. Passwordless SSH can be achieved by running the following commands:
!!! note "Non-Root Key Storage Considerations"
When you generate SSH keys, they will be stored in a specific user's profile, the one currently executing the commands. If you want to have passwordless SSH, you would run the commands from a non-root user (e.g. `nicole`).
```sh
ssh-keygen # (1)
ssh-copy-id -i /home/nicole/.ssh/id_rsa.pub nicole@192.168.3.18 # (2)
ssh -i /home/nicole/.ssh/id_rsa nicole@192.168.3.18 # (3)
```
1. Just leave all of the default options and do not put a password on the SSH key. )
2. Change the directories to account for your given username, and change the destination to the user@IP corresponding to the remote server. You will be prompted to enter the password once to store the SSH public key on the remote computer.
3. This command is to validate that everything worked. If the remote user is the same as the local user (e.g. `nicole`) then you dont need to add the `-i /home/nicole/.ssh/id_rsa` section to the SSH command.
!!! warning "Run before configuring Global SSH Infrastructure Key"
There is a global automation that leverages a [Global Infrastructure Public SSH Key](https://git.bunny-lab.io/Infrastructure/LinuxServer_SSH_PublicKey). If this runs before you run the commands above, you will be unable to configure SSH key relationships and it will need to be done manually.
## Related Documentation
- [Related Networking and Access Documentation](<../../../reference/Networking and Access/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,171 @@
---
tags:
- Sophos
- IPsec
- VPN
- Firewall
- Routing
---
## Purpose
Generally speaking, when you have site-to-site VPN tunnels, you have to ensure that the *health* of the tunnel is operating as-expected. Sometimes VPN tunnels will report that they are online and connected, but in reality, no traffic is flowing to the remote side of the tunnel. In these instances, we can create a script that pings a device on the remote end, and if it does not respond in a timely manner, the script restart the VPN tunnel automatically.
!!! note "Assumptions"
This document assumes that you will be running a powershell script on a Windows environment. The `curl` commands can be used interchangably in Linux, but the example script provided here will be using `curl.exe` within a powershell script, and instead of running on a schedule using crontab, it will be using Windows Task Scheduler.
I will attempt to provide Linux-equivalant commands where-possible.
## Sophos Environment
### Configure Sophos XGS Firewall ACLs
You need to configure a user account that will be specifically used for leveraging the API controls that allow resetting the VPN tunnel(s). At this stage, you need to log into your Sophos XGS Firewall. For this example, we will assume you can reach your firewall at https://172.16.16.16:4444 and log in as the administrator.
### Create API Access Profile
You need to create a profile that the API User will leverage to issue commands to the firewall's VPN settings. Without this profile, the user may have either not enough, or too much access.
- Navigate to **System > Profiles > Device Access > "Add"**
- Profile Name: `VPNTunnelAPI`
- Check the radio box column named "**None**" to Deny all permissions to all areas of the firewall
- Expand the "**VPN**" section of the permission tree, and check the box for "**Read-Write**" next to "**Connect Tunnel**"
- Click the "**Save**" button to save the access profile
### Create API Access User
Now we need to make a user account that we will use inside the script to authenticate against the firewall using the previously-mentioned access profile
- Navigate to **Configure > Authentication > Users > "Add"**
- Username: `TunnelCheckerAPIUser`
- Name: `TunnelCheckerAPIUser`
- User Type: `Administrator`
- Profile: `VPNTunnelAPI`
- Password: `01_placeholder_PASSWORD_here_02`
- Group: `Open Group`
- Click the "**Save**" button to save the API user account
### Create Device Access ACL
Now we need to configure an ACL within the Firewall to allow API access from the specific server we will be using in the next section.
- Navigate to **Administration > Device Access > Local service ACL exception rule > "Add"**
- Rule Name: `API Access (IPSec Tunnel Heartbeat Script)`
- Source Zone: `The Zone of the Server/Device that will be used to run the script, such as a server network.
- Source Network/Host: `<IP_HOST_OF_DEVICE_RUNNING_SCRIPT>`
- Destination Host: `XGS Firewall (Local IP)` (*This is an IP host pointing to the internal IP of the Firewall*)
- Services: `HTTPS`
- Action: `Accept`
### Configure API Access via IP
Lastly, you need to configure the API access to allow communication from the IP of the device. I know this seems redundant to the previous "Device Access ACL" but its required for this to work, otherwise you will get an `Sophos API Operations are not allowed from the requester IP address` error when running the script.
- Navigate to **System > Backup & Firmware > API > API Configuration**
- Add the IP of the Server/Device
- Click the "**Apply** button
## Server Environment
### Choose a Server
It is important to choose a server/device that is able to communicate with the devices on the remote end of the tunnel. If it cannot ping the remote device(s), it will assume that the tunnel is offline and do an infinite loop of restarting the VPN tunnel.
### Prepare the Script Folder
You need a place to put the script (and if on Windows, `curl.exe`). Follow the instructions specific to your platform below:
=== "Windows"
Download `curl.exe` from this location: [Download](https://curl.se/windows/dl-8.10.0_1/curl-8.10.0_1-win64-mingw.zip) and place it somewhere on the operating system, such as `C:\Scripts\VPN_Tunnel_Checker`. Then copy this script into that same folder and call it `Tunnel_Checker.ps1` with the content below:
!!! note "Curl Files Extraction"
You will want to extract all of the files included in the zip file's `bin` folder. Specifically, copy the following files into the `C:\Scripts\VPN_Tunnel_Checker` folder:
- `curl.exe`
- `curl-ca-bundle`
- `libcurl-x64.def`
- `libcurl-x64.dll`
```powershell
function Reset-VPN-Tunnel {
Write-Host "VPN Tunnel Broken - Bringing VPN Tunnel Down..."
.\curl -k https://172.16.16.16:4444/webconsole/APIController?reqxml=<Request><Login><Username>TunnelCheckerAPIUser</Username><Password>01_placeholder_PASSWORD_here_02</Password></Login><Set><VPNIPSecConnection><DeActive><Name>VPN_TUNNEL_NAME</Name></DeActive></VPNIPSecConnection></Set></Request>
Start-Sleep -Seconds 5
Write-Host "Bringing VPN Tunnel Up..."
.\curl -k https://172.16.16.16:4444/webconsole/APIController?reqxml=<Request><Login><Username>TunnelCheckerAPIUser</Username><Password>01_placeholder_PASSWORD_here_02</Password></Login><Set><VPNIPSecConnection><Active><Name>VPN_TUNNEL_NAME</Name></Active></VPNIPSecConnection></Set></Request>
}
function Check-VPN-Tunnel {
# Server Connectivity Check
Write-Host "Checking Tunnel Connection to PLACEHOLDER..."
if (-not (Test-Connection '10.0.0.29' -Quiet)) {
Reset-VPN-Tunnel
}
# Server Connectivity Check
Write-Host "Checking Tunnel Connection to PLACEHOLDER..."
if (-not (Test-Connection '10.0.0.30' -Quiet)) {
Reset-VPN-Tunnel
}
}
function Trace-VPN-Tunnel {
Write-Host "Tracing Path to PLACEHOLDER:"
pathping -n -w 500 -p 100 10.0.0.29
Write-Host "Tracing Path to PLACEHOLDER:"
pathping -n -w 500 -p 100 10.0.0.30
}
CD "C:\Scripts\VPN_Tunnel_Checker"
Check-VPN-Tunnel
#Write-Host "Checking Tunnel Quality After Running Script..."
#Trace-VPN-Tunnel
```
!!! note "Optional Reporting"
You may find that you want some extra logging enabled so you can track the script doing its job to ensure its working. You can add the following to the script above to add that functionality.
Add the following to the bottom of each server in the `Check-VPN-Tunnel` function, directly below the `Reset-VPN-Tunnel` function.
```powershell
Add-Content -Path "C:\Scripts\VPN_Tunnel_Checker\Tunnel.log" -Value "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') PLACEHOLDER Connection Down"
```
Lastly, change the very end of the script under where the `Check-IHS-Tunnel` function is being called to look like this if you want to log heartbeats and not just when a VPN tunnel is down. The purpose of this is to show the script is actually running. I recommend only temporarily implementing it during initial deployment.
```powershell
CD "C:\Scripts\VPN_Tunnel_Checker"
Check-VPN-Tunnel
Add-Content -Path "C:\Scripts\VPN_Tunnel_Checker\Tunnel.log" -Value "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') Heartbeat"
```
=== "Linux"
```sh
PLACEHOLDER
```
### Create Scheduled Task
At this point, you need this script to run automatically on its own every 5 minutes or so, so you need to create a task in the Windows Task Scheduler in order to achieve this.
=== "Windows"
- Open "**Task Scheduler**" on the device
- Expand "**Task Scheduler Library**" in the tree on the left-hand side
- Right-click anywhere in the task list and select "**Create New Task...**"
- **General**:
- Name: `Check VPN Tunnel Every 5 Minutes`
- When running this task, use the following user account: `SYSTEM`
- **Triggers**:
- Click "**New...**"
- Begin the Task: `On a Schedule`
- Settings: `Daily`
- Advanced Settings > Repeat Task Every: `5 Minutes` > for a duration of `1 Day`
- **Actions**:
- Click "**New...**"
- Action: `Start a Program`
- Program/Script: `C:\Windows\System32\WindowsPowershell\v1.0\powershell.exe`
- Add Arguments: `-ExecutionPolicy Bypass -File "C:\Scripts\VPN_Tunnel_Checker\Tunnel_Checker.ps1"`
- Press the "**OK**" button to save the scheduled task, then wait for 5 minutes to ensure it triggers as-expected.
=== "Linux"
- PLACEHOLDER
- PLACEHOLDER
## Related Documentation
- [Related Networking and Access Documentation](<../../../reference/Networking and Access/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,41 @@
---
tags:
- Sophos
- Firewall
- Routing
- LAN
- Networking
---
## Purpose
You may have a Sophos XGS appliance and need more than one interface to act as additional LAN ports. You can achieve this with bridges.
!!! info "Assumptions"
It is assumed that your Sophos XGS appliance has at least 3 interfaces, one for `WAN`, one for `LAN`, and a third one that will act as a member of the bridge. You can have as many member interfaces of the bridge as needed, but you need at least one.
## Login to the Firewall
You will need to access the firewall either directly on the local network at `https://<IP-of-Firewall>:4444` or remotely in Sophos Central.
## Configure a LAN bridge
Navigate to "**Configure > Network > Interfaces > "Add Interface" > "Add Bridge"**"
| **Field** | **Value** |
| :--- | :--- |
| Name | `LAN Bridge` |
| Hardware | `br0` |
| Enable routing on this bridge pair | `<Unchecked>` |
| Member Interfaces | `<Interfaces-of-Additional-Ports> / Zone: "LAN"` |
!!! warning
The LAN interface itself needs to be a member of the bridge. If it is not, the Sophos Appliance will not allow you to use the same IP address as the existing LAN interface.
### IPv4 Configuration
| **Field** | **Value** |
| :--- | :--- |
| IP Assignment | `Static` |
| IPv4/netmask | `<IP-of-LAN-Interface> / <CIDR-of-LAN-Interface>` |
| Gateway IP | `<Blank>` |
| Member Interfaces | `<Interfaces-of-Additional-Ports> / Zone: "LAN"` |
## Related Documentation
- [Related Networking and Access Documentation](<../../../reference/Networking and Access/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,46 @@
---
tags:
- Sophos
- RDP
- SSL VPN
- VPN
- Firewall
---
## Purpose
This document exists to outline the generalized process to configuring remote access in a Sophos XGS Firewall to allow a VPN user to RDP into a workstation. *Setting up Remote SSL VPN Access is not covered in this document.*
### Create MAC Host for Destination Device
The first step in the process is to create a MAC address host for the device being RDP'd into, that way if it's IP rotates, the firewall rule will continue to work correctly.
- Navigate to **Sophos XGS Firewall > [System] Hosts and Services**
- Click on the **Mac Host** tab > "**Add**"
- Name: `<Device-Hostname>`
- Description: `<Workstation Remote Access for (username)>`
- Type: `Mac Address`
- MAC Address: `<mac address of device>`
Click **Save**
### Configure Firewall Rule
- Navigate to **[Protect] Rules and Policies > Add Firewall Rule (New Firewall Rule)**
- Rule Name: `Remote Workstation Access for (username)`
- Source Zone: `VPN`
- Source Networks and Devices: `Any`
- Destination Zone: `LAN`
- Destination Networks: `<MAC Host We Previously Made>`
- Services > Add New Item > `RDP`
- If `RDP` does not exist, click "Add", `Services`
- Name: `RDP`
- Description: `Remote Desktop Protocol`
- Type: `TCP/UDP`
- Protocol: `TCP`
- Source Port: `1:65535`
- Destination Port: `3389`
Click **Save**
- Check **Match Known Users**
- Under "Users or Groups" click "Add New Item"
- Search for the username of the person using the VPN that needs to access the workstation (e.g. `nicole.rappe@bunny-lab.io`)
- Click the **Save** button and have the user try to connect to the VPN, then RDP into their workstation.
## Related Documentation
- [Related Networking and Access Documentation](<../../../reference/Networking and Access/index.md>) — Find the connected deployments, procedures, and references for this subject.
+17
View File
@@ -0,0 +1,17 @@
---
tags:
- Networking and Access
- Workflows
- Documentation
---
# Networking and Access
## Purpose
Find workflows for networking and access. Follow the subject guide to choose the relevant environment and connect this material to the other document types.
## Includes
- Linux
- Sophos
## Follow the Subject
[Networking and Access](<../../reference/Networking and Access/index.md>) explains the relationships and offers starting points for the documented tasks.
@@ -1,12 +0,0 @@
---
tags:
- Proxmox
---
**Purpose**: The purpose of this document is to outline common tasks that you may need to run in your cluster to perform various tasks.
## Delete Node from Cluster
Sometimes you may need to delete a node from the cluster if you have re-built it or had issues and needed to destroy it. In these instances, you would run the following command (assuming you have a 3-node quorum in your cluster).
```
pvecm delnode promox-node-01
```
@@ -1,9 +1,11 @@
---
tags:
- Documentation
- Virtualization and Storage
- Rebuild Failover Cluster Replication
---
**Purpose**: If you run an environment with multiple Hyper-V: Failover Clusters, for the purpose of Hyper-V: Failover Cluster Replication via a `Hyper-V Replica Broker` role installed on a host within the Failover Cluster, sometimes a GuestVM will fail to replicate itself to the replica cluster, and in those cases, it may not be able to recover on its own. This guide attempts to outline the process to rebuild replication for GuestVMs on a one-by-one basis.
## Purpose
If you run an environment with multiple Hyper-V: Failover Clusters, for the purpose of Hyper-V: Failover Cluster Replication via a `Hyper-V Replica Broker` role installed on a host within the Failover Cluster, sometimes a GuestVM will fail to replicate itself to the replica cluster, and in those cases, it may not be able to recover on its own. This guide attempts to outline the process to rebuild replication for GuestVMs on a one-by-one basis.
!!! note "Assumptions"
This guide assumes you have two Hyper-V Failover Clusters, for the sake of the guide, we will refer to the Production cluster as `CLUSTER-01` and the Replication cluster as `CLUSTER-02`. This guide also assumes that Replication was set up beforehand, and does not include instructions on how to deploy a Replica Broker (at this time).
@@ -11,11 +13,12 @@ tags:
## Production Cluster - CLUSTER-01
### Locate the GuestVM
You need to start by locating the GuestVM in the Production cluster, CLUSTER-01. You will know you found the VM if the "Replication Health" is either `Unhealthy`, `Warning`, or `Critical`.
### Remove Replication from GuestVM
- Within a node of the Hyper-V: Failover Cluster Manager
- Right-Click the GuestVM
- Navigate to "**Replication > Remove Replication**"
- Confirm the removal by clicking the "**Yes**" button. You will know if it removed replication when the "Replication State" of the GuestVM is `Not enabled`
- Confirm the removal by clicking the "**Yes**" button. You will know if it removed replication when the "Replication State" of the GuestVM is `Not enabled`
## Replication Cluster - CLUSTER-02
### Note the storage GUID of the GuestVM in the replication cluster
@@ -34,7 +37,7 @@ Now that you have noted the GUID of the storage folder of the GuestVM, we can sa
- Within a node of the replication cluster's Hyper-V: Failover Cluster Manager
- Right-Click the GuestVM
- Navigate to "**Replication > Remove Replication**"
- Confirm the removal by clicking the "**Yes**" button. You will know if it removed replication when the "Replication State" of the GuestVM is `Not enabled`
- Confirm the removal by clicking the "**Yes**" button. You will know if it removed replication when the "Replication State" of the GuestVM is `Not enabled`
- Right-Click the GuestVM (again) `You will see that "Enable Replication" is an option now, indicating it was successfully removed.`
!!! note "Replica Checkpoint Merges"
@@ -45,7 +48,7 @@ Now that you have noted the GUID of the storage folder of the GuestVM, we can sa
- Confirm the action by clicking the "**Yes**" button
### Delete the GuestVM manually from Hyper-V Manager on all replication cluster hosts
At this point in time, we need to remove the GuestVM from all of the servers in the cluster. Just because we removed it from the Hyper-V: Failover Cluster did not remove it from the cluster's nodes. We can automate part of this work by opening Hyper-V Manager on the same Failover Node we have been working on thus far, and from there we can connect the rest of the replication nodes to the manager to have one place to connect to all of the nodes, avoiding hopping between servers.
At this point in time, we need to remove the GuestVM from all of the servers in the cluster. Just because we removed it from the Hyper-V: Failover Cluster did not remove it from the cluster's nodes. We can automate part of this work by opening Hyper-V Manager on the same Failover Node we have been working on thus far, and from there we can connect the rest of the replication nodes to the manager to have one place to connect to all of the nodes, avoiding hopping between servers.
- Open Hyper-V Manager
- Right-Click "Hyper-V Manager" on the left-hand navigation menu
@@ -82,4 +85,7 @@ At this point, we have disabled replication for the GuestVM and cleaned up trace
!!! success "Replication Enabled"
If everything was successful, you will see a dialog box named "Enable replication for `<GuestVM>`" with a message similar to the following: "Replica virtual machine `<GuestVM>` was successfully created on the specified Replica server `<Node-in-Replication-Cluster>`.
At this point, you can click "Close" to finish the process. Under the GuestVM details, you will see "Replication State": `Initial Replication in Progress`.
At this point, you can click "Close" to finish the process. Under the GuestVM details, you will see "Replication State": `Initial Replication in Progress`.
## Related Documentation
- [Related Virtualization and Storage Documentation](<../../../../reference/Virtualization and Storage/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -1,3 +1,9 @@
---
tags:
- Virtualization and Storage
- Forcefully Stop GuestVM
---
## Purpose
If you have a GuestVM that will not stop gracefully either because the Hyper-V host is goofed-up or the VMMS service won't allow you to restart it. You can perform a hail-mary to forcefully stop the GuestVM's Hyper-V process.
@@ -16,6 +22,7 @@ Get-VM SERVER-01 | Select VMName, VMId
### Extrapolate Process ID
Now you need to hunt-down the process ID associated with the GuestVM.
```powershell
Get-CimInstance Win32_Process -Filter "Name='vmwp.exe'" |
Where-Object { $_.CommandLine -match "3e4b6f91-6c6c-4075-9b7e-389d46315074" } |
@@ -29,6 +36,10 @@ Select-Object ProcessId, CommandLine
### Terminate Process
Lastly, you terminate the process by its ID.
```powershell
Stop-Process -Id 12488 -Force
```
```
## Related Documentation
- [Related Virtualization and Storage Documentation](<../../../reference/Virtualization and Storage/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -1,9 +1,10 @@
---
tags:
- Kerberos
- Virtualization and Storage
---
**Purpose**:
## Purpose
You may find that you want to be able to live-migrate guestVMs on a Hyper-V environment that is not clustered as a Hyper-V Failover Cluster, you will have permission issues. One way to work around this is to use CredSSP as the authentication mechanism, which is not ideal but useful in a pinch, or you can use Kerberos-based authentication.
This document will cover both scenarios.
@@ -37,4 +38,7 @@ This document will cover both scenarios.
- Select the destination by providing the fully-qualified domain name of the destination server (or in some cases the shorthand hostname of the destination server)
- It should begin the migration process.
**Note**: Do not perform a "Pull" from source to the destination. You want to always "Push" the VM to its destination. It will generally fail if you try to "Pull" the VM to its destination due to the way that CredSSP works in this context.
**Note**: Do not perform a "Pull" from source to the destination. You want to always "Push" the VM to its destination. It will generally fail if you try to "Pull" the VM to its destination due to the way that CredSSP works in this context.
## Related Documentation
- [Related Virtualization and Storage Documentation](<../../../reference/Virtualization and Storage/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -4,7 +4,7 @@ tags:
- Filesystems
---
**Purpose**:
## Purpose
The purpose of this workflow is to illustrate the process of expanding storage for a RHEL-based Linux server acting as a GuestVM. We want the VM to have more storage space, so this document will go over the steps to expand that usable space.
!!! info "Assumptions"
@@ -26,7 +26,7 @@ This step goes over how to increase the usable space of the virtual disk within
=== "Using GDISK"
``` sh
```sh
sudo dnf install gdisk -y
gdisk /dev/<diskNumber> # (1)
p <ENTER> # (2)
@@ -39,6 +39,7 @@ This step goes over how to increase the usable space of the virtual disk within
<FILESYSTEM-TYPE=8300 (Linux Filesystem)> (Just press ENTER) # (9)
w <ENTER> # (10)
```
??? info "Detailed Command Breakdown"
1. The first command needs you to enter the disk identifier. In most cases, this will likely be the first disk, such as `/dev/sda`. You do not need to indicate a partition number in this step, as you will be asked for one in a later step after identifying all of the partitions on this disk in the next command.
2. This will list all of the partitions on the disk.
@@ -54,7 +55,8 @@ This step goes over how to increase the usable space of the virtual disk within
10. This will write the changes to the partition table making them reality instead of just staging the changes.
!!! example "Example Output"
```
```text
Command (? for help): p
Disk /dev/sda: 2147483648 sectors, 1024.0 GiB
Model: Virtual Disk
@@ -72,9 +74,10 @@ This step goes over how to increase the usable space of the virtual disk within
3 3328000 19826687 7.9 GiB 8200
4 19826688 1073741790 502.5 GiB 8300 Linux filesystem
```
=== "Using FDISK"
``` sh
```sh
pvdisplay # (1)
fdisk /dev/hda # (2)
p <ENTER> # List Partitions
@@ -86,8 +89,9 @@ This step goes over how to increase the usable space of the virtual disk within
Ending Sector: <ENTER> # Use Default Value
w <ENTER> # Commit all queued-up changes and write them to the disk
```
??? info "Detailed Command Breakdown"
1. Use pvdisplay to get the target disk identifier
2. Replace `/dev/hda` with the target disk identifier found in the previous step
@@ -97,15 +101,16 @@ This step goes over how to increase the usable space of the virtual disk within
=== "Using GROWPART (Ubuntu)"
``` sh
```sh
echo 1 | sudo tee /sys/class/block/sda/device/rescan
growpart /dev/sda 2
sudo resize2fs /dev/sda2 # Assuming ext4 filesystem, if unsure, run "df -Th"
```
## Detect the New Partition Sizes
At this point, the operating system wont detect the changes without a reboot, so we are going to force the operating system to detect them immediately with the following commands to avoid a reboot (if we can avoid it).
``` sh
```sh
sudo partprobe /dev/<drive> # Drive Example: /dev/sda (Rocky) or /dev/hda (Oracle Linux)
sudo partx -u /dev/<diskNumber>
```
@@ -113,23 +118,26 @@ sudo partx -u /dev/<diskNumber>
!!! bug "Partition Size Not Expanded? Reboot."
If you notice the partition still has not expanded to the desired size, you may have no choice but to reboot the server, then re-run the `gdisk` or `fdisk` commands a second time. In my lab environment, it didn't work until I rebooted. This might have been a hiccup on my end, but it's something to keep in mind if you run into the same issue of the size not changing.
``` sh
```sh
sudo reboot
```
## Resize the Filesystem
=== "XFS Filesystem"
``` sh
```sh
sudo xfs_growfs /
```
=== "Ext4 Filesystem"
``` sh
```sh
resize2fs /dev/sda
```
=== "Ext4 Filesystem w/ LVM"
``` sh
```sh
# Increase the Physical Volume Group Size
pvdisplay # Check the Current Size of the Physical Volume
pvresize /dev/hda2 # Enlarge the Physical Volume to Fit the New Partition Size
@@ -142,17 +150,21 @@ sudo partx -u /dev/<diskNumber>
resize2fs /dev/VolGroup00/LogVol00
```
## Validate Storage Expansion
At this point, you can leverage `lsblk` or `df -h` to determine if the usable storage space was successfully increased or not. In this example, you can see that I increased my storage space from 512GB to 1TB.
!!! example "Example Command Output"
Command: `lsblk | grep "sda4"`
```
```text
└─sda4 8:4 0 1014.5G 0 part /
```
Command: `df -h | grep "sda4"`
```
```text
/dev/sda4 1015G 145G 871G 15% /
```
## Related Documentation
- [Related Virtualization and Storage Documentation](<../../../reference/Virtualization and Storage/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -6,7 +6,7 @@ tags:
- Filesystems
---
**Purpose**:
## Purpose
The purpose of this workflow is to illustrate the process of expanding storage for a Linux server that uses an iSCSI-based ZFS storage. We want the VM to have more storage space, so this document will go over the steps to expand that usable space.
!!! info "Assumptions"
@@ -20,7 +20,7 @@ This part should be fairly straight-forward. Using whatever hypervisor / storag
## Extend ZFS Pool
This step goes over how to increase the usable space of the ZFS pool within the server itself after it was expanded.
``` sh
```sh
iscsiadm -m session --rescan # (1)
lsblk # (2)
parted /dev/sdX # (3)
@@ -45,4 +45,7 @@ At this point, the ZFS pool has been expanded and a scrub task has been started.
```sh
zpool status
```
```
## Related Documentation
- [Related Virtualization and Storage Documentation](<../../../reference/Virtualization and Storage/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,678 @@
---
tags:
- Virtualization and Storage
- Detect and Remove Orphaned VM Disks
---
## Purpose
This procedure describes how to identify and safely remove orphaned Proxmox VE virtual machine disks from shared iSCSI-backed LVM storage.
It is intended for environments where:
- Proxmox VE is clustered.
- Multiple Proxmox nodes access the same shared iSCSI LUN.
- The shared storage is exposed to Proxmox as LVM storage.
- VM disks are stored as LVM logical volumes.
- Some volumes may remain after VM disk deletion, failed migrations, failed resizes, storage UI inconsistencies, or manual recovery work.
The goal is to reclaim storage space without accidentally deleting disks that are still attached to running or stopped VMs.
---
## Scope
This document focuses on the following storage type:
```text
Proxmox storage type: lvm
Backing storage: shared iSCSI
Volume group: vg_proxmox_iscsi
Storage ID example: iscsi-cluster-lvm
```
Adjust the storage ID and volume group names as needed for your environment.
---
## Safety Requirements
!!! danger "Never delete based on the Storage UI alone"
The Proxmox storage UI may show a volume as belonging to a VM because its name follows the pattern `vm-<vmid>-disk-<n>`. That does not prove the disk is currently attached to the VM.
```text
Always verify against VM configuration files and active QEMU processes before deleting.
```
!!! warning "Run the audit before running any cleanup commands"
The audit scripts in this document are read-only. The cleanup commands are destructive. Do not run cleanup commands until the audit output has been reviewed.
!!! warning "Snapshot volumes require extra caution"
Volumes named like the following may be part of a snapshot chain:
````text
```text
snap_vm-<vmid>-disk-<n>_<snapshot-name>
```
Do not remove snapshot volumes manually unless you have verified that the VM and snapshot are no longer known to Proxmox, no backing chain references them, and no QEMU process has them open.
````
!!! note "Shared storage does not mean shared config"
In some cluster layouts, each node may only show VM config files for VMs assigned to that node. Therefore, an audit run from only one node can falsely report disks from other nodes as orphaned.
```text
Run the confirmation script on every node in the cluster.
```
---
## Terms
| Term | Meaning |
| -------------- | ------------------------------------------------------------------------------------------------------------------------------------------ |
| Attached disk | A disk volume referenced in a VM config, such as `scsi0`, `sata0`, `virtio0`, `efidisk0`, or `tpmstate0`. |
| Orphan disk | A storage volume that exists on shared storage but is not referenced by any VM config on any node and is not opened by any active process. |
| Volume ID | Proxmox storage identifier, such as `iscsi-cluster-lvm:vm-107-disk-1.qcow2`. |
| LV | LVM logical volume, such as `/dev/vg_proxmox_iscsi/vm-107-disk-1.qcow2`. |
| Snapshot chain | A chain of qcow2 backing files or Proxmox snapshot volumes. |
---
## Phase 1: Identify Storage Names
Run this on any Proxmox node:
```bash
pvesm status
cat /etc/pve/storage.cfg
vgs
```
Identify the shared iSCSI/LVM storage.
Example:
```text
Storage ID: iscsi-cluster-lvm
VG name: vg_proxmox_iscsi
```
For the rest of this document, replace these values if your environment differs:
```bash
STORAGE="iscsi-cluster-lvm"
VG="vg_proxmox_iscsi"
```
---
## Phase 2: Run the Storage Orphan Audit
Run the following script on one node that can see the shared LVM storage.
This script does not delete anything.
```bash
cat > /root/pve-iscsi-orphan-audit.sh <<'EOF'
#!/usr/bin/env bash
set -u
STORAGE="iscsi-cluster-lvm"
VG="vg_proxmox_iscsi"
OUT="/root/pve-iscsi-orphan-audit-$(hostname)-$(date +%Y%m%d-%H%M%S).txt"
{
echo "===== PVE ISCSI ORPHAN AUDIT ====="
echo "Host: $(hostname)"
echo "Date: $(date)"
echo "Storage: ${STORAGE}"
echo "VG: ${VG}"
echo
echo "===== STORAGE STATUS ====="
pvesm status 2>&1 | egrep "^(Name|${STORAGE})" || true
vgs "${VG}" 2>&1 || true
echo
echo "===== ALL VOLUMES IN ${STORAGE} ====="
pvesm list "${STORAGE}" 2>&1 || true
echo
echo "===== ALL LVs IN ${VG} ====="
lvs -a -o lv_name,lv_path,lv_size,lv_attr,devices "${VG}" 2>&1 || true
echo
echo "===== LOCAL VM CONFIG FILES ====="
for conf in /etc/pve/qemu-server/*.conf; do
[ -e "$conf" ] || continue
echo
echo "----- $conf -----"
cat "$conf"
done
echo
echo "===== REFERENCE ANALYSIS - LOCAL CONFIG FILES ONLY ====="
printf '%-55s | %-8s | %-10s | %-8s | %-8s | %s\n' \
"volume" "vmid" "referenced" "open" "size" "path"
{
pvesm list "${STORAGE}" 2>/dev/null | awk 'NR>1 {print $1}' | sed "s#^${STORAGE}:##"
lvs --noheadings -o lv_name "${VG}" 2>/dev/null | awk '{print $1}'
} | sort -u | while read -r vol; do
[ -n "$vol" ] || continue
case "$vol" in
vm-*-disk-*|snap_vm-*-disk-*) ;;
*) continue ;;
esac
vmid="unknown"
if [[ "$vol" =~ ^vm-([0-9]+)-disk- ]]; then
vmid="${BASH_REMATCH[1]}"
elif [[ "$vol" =~ ^snap_vm-([0-9]+)-disk- ]]; then
vmid="${BASH_REMATCH[1]}"
fi
ref="no"
if grep -R -Fq "$vol" /etc/pve/qemu-server 2>/dev/null; then
ref="yes"
fi
open="no"
if lsof 2>/dev/null | grep -Fq "$vol"; then
open="yes"
fi
size="$(lvs --noheadings -o lv_size "${VG}/${vol}" 2>/dev/null | awk '{$1=$1;print}')"
path="$(lvs --noheadings -o lv_path "${VG}/${vol}" 2>/dev/null | awk '{$1=$1;print}')"
printf '%-55s | %-8s | %-10s | %-8s | %-8s | %s\n' \
"$vol" "$vmid" "$ref" "$open" "$size" "$path"
done
echo
echo "===== DONE ====="
} | tee "$OUT"
echo
echo "Saved audit to: $OUT"
EOF
chmod +x /root/pve-iscsi-orphan-audit.sh
/root/pve-iscsi-orphan-audit.sh
```
The script writes a file similar to:
```text
/root/pve-iscsi-orphan-audit-<node>-<timestamp>.txt
```
---
## How to Read the First Audit
The most important section is:
```text
REFERENCE ANALYSIS - LOCAL CONFIG FILES ONLY
```
Example:
```text
volume | vmid | referenced | open | size
vm-107-disk-0.qcow2 | 107 | no | no | 4.00m
vm-107-disk-1.qcow2 | 107 | no | no | 256.04g
vm-107-disk-2.qcow2 | 107 | yes | no | 4.00m
vm-107-disk-3.qcow2 | 107 | yes | no | 256.04g
```
Interpretation:
| Field | Meaning |
| ---------------- | ------------------------------------------------------------------------------------------------------- |
| `referenced=yes` | The volume appears in a local VM config file. Do not delete. |
| `referenced=no` | The volume does not appear in local VM configs. It may be orphaned, but confirm across all nodes first. |
| `open=yes` | A process has the volume open. Do not delete. |
| `open=no` | No process on this node has the volume open. Still confirm across all nodes. |
!!! warning "Local reference analysis is not enough"
If a VM runs on another cluster node, its config may not appear on the node where you ran the audit. This can make valid disks look orphaned.
```text
Continue to Phase 3 before deleting anything.
```
---
## Phase 3: Run Cluster-Wide Confirmation
Run the following script on **every Proxmox node** in the cluster.
This script is read-only.
```bash
cat > /root/pve-cluster-vm-confirm.sh <<'EOF'
#!/usr/bin/env bash
set -u
STORAGE="iscsi-cluster-lvm"
VG="vg_proxmox_iscsi"
OUT="/root/pve-cluster-vm-confirm-$(hostname)-$(date +%Y%m%d-%H%M%S).txt"
{
echo "===== NODE ====="
hostname
date
echo
echo "===== CLUSTER RESOURCES - VMS ====="
pvesh get /cluster/resources --type vm 2>&1 || true
echo
echo "===== LOCAL QM LIST ====="
qm list 2>&1 || true
echo
echo "===== QEMU CONFIG FILES PRESENT ====="
ls -la /etc/pve/qemu-server/ 2>&1 || true
echo
echo "===== QEMU CONFIG FILE CONTENTS ====="
for conf in /etc/pve/qemu-server/*.conf; do
[ -e "$conf" ] || continue
echo
echo "----- $conf -----"
cat "$conf"
done
echo
echo "===== ALL STORAGE VOLUMES ====="
pvesm list "${STORAGE}" 2>&1 || true
echo
echo "===== ALL LVs ====="
lvs -a -o lv_name,lv_path,lv_size,lv_attr,devices "${VG}" 2>&1 || true
} | tee "$OUT"
echo
echo "Saved to: $OUT"
EOF
chmod +x /root/pve-cluster-vm-confirm.sh
/root/pve-cluster-vm-confirm.sh
```
Collect the output file from each node.
Example for a three-node cluster:
```text
/root/pve-cluster-vm-confirm-cluster-node-01-YYYYMMDD-HHMMSS.txt
/root/pve-cluster-vm-confirm-cluster-node-02-YYYYMMDD-HHMMSS.txt
/root/pve-cluster-vm-confirm-cluster-node-03-YYYYMMDD-HHMMSS.txt
```
---
## How to Read the Cluster Confirmation
For each suspicious volume, search all three outputs.
Example candidate:
```text
vm-107-disk-1.qcow2
```
Check whether it appears in any VM config:
```bash
grep -R "vm-107-disk-1.qcow2" /etc/pve/qemu-server/ || true
```
If reviewing output files manually, look for config lines such as:
```text
scsi0: iscsi-cluster-lvm:vm-107-disk-1.qcow2
sata0: iscsi-cluster-lvm:vm-107-disk-1.qcow2
virtio0: iscsi-cluster-lvm:vm-107-disk-1.qcow2
efidisk0: iscsi-cluster-lvm:vm-107-disk-1.qcow2
tpmstate0: iscsi-cluster-lvm:vm-107-disk-1.qcow2
```
If a volume appears in any of those lines, it is attached to a VM and must not be deleted.
---
## Phase 4: Classify Candidate Volumes
Use the following decision table.
| Condition | Classification | Action |
| --------------------------------------------------------------------- | ------------------- | ---------------------------------------------- |
| Volume appears in any VM config on any node | In use | Do not delete |
| Volume is opened by QEMU or another process | In use or unsafe | Do not delete |
| Volume is a `snap_vm-*` snapshot volume | Snapshot-chain item | Inspect snapshot/backing chain before deletion |
| Volume does not appear in any VM config and is not open | Orphan candidate | Eligible for final verification |
| VMID no longer exists in cluster resources and disk is not referenced | Strong orphan | Eligible for cleanup |
---
## Example: Valid VM Disks
If VM `107` has this config:
```text
efidisk0: iscsi-cluster-lvm:vm-107-disk-2.qcow2
sata0: iscsi-cluster-lvm:vm-107-disk-3.qcow2
```
Then these disks are valid and must not be deleted:
```text
vm-107-disk-2.qcow2
vm-107-disk-3.qcow2
```
If storage also contains:
```text
vm-107-disk-0.qcow2
vm-107-disk-1.qcow2
```
and neither appears in any config file on any node, those are orphan candidates.
---
## Phase 5: Final Verification Before Deletion
For each candidate volume, run the following checks on a node that can see the shared storage.
Replace the volume name as appropriate.
```bash
VOL="vm-107-disk-1.qcow2"
STORAGE="iscsi-cluster-lvm"
VG="vg_proxmox_iscsi"
echo "===== Check all cluster config references ====="
grep -R "$VOL" /etc/pve/qemu-server/ || true
echo
echo "===== Check Proxmox storage listing ====="
pvesm list "$STORAGE" | grep "$VOL" || true
echo
echo "===== Check LVM volume ====="
lvs -a -o lv_name,lv_path,lv_size,lv_attr,devices "$VG" | grep "$VOL" || true
echo
echo "===== Check whether open by any process ====="
lsof | grep "$VOL" || true
echo
echo "===== Check qemu-img metadata if device path exists ====="
LVPATH="$(lvs --noheadings -o lv_path "${VG}/${VOL}" 2>/dev/null | awk '{$1=$1;print}')"
if [ -n "$LVPATH" ] && [ -e "$LVPATH" ]; then
qemu-img info --backing-chain "$LVPATH"
else
echo "LV path missing or inactive: $LVPATH"
fi
```
Safe deletion pattern:
```text
grep -R ... no output
pvesm list ... shows the volume
lvs ... shows the volume
lsof ... no output
qemu-img info ... no unexpected backing file dependency
```
!!! danger "Stop if grep finds a reference"
If the candidate volume appears in any `/etc/pve/qemu-server/*.conf` file, do not delete it.
!!! danger "Stop if lsof finds a process"
If `lsof` shows the volume is open, do not delete it.
---
## Phase 6: Cleanup Commands
## Preferred Method: Proxmox Storage Layer
Use `pvesm free` first.
Example:
```bash
pvesm free iscsi-cluster-lvm:vm-107-disk-0.qcow2
pvesm free iscsi-cluster-lvm:vm-107-disk-1.qcow2
```
Then verify:
```bash
pvesm list iscsi-cluster-lvm | grep "vm-107-disk" || true
lvs -a -o lv_name,lv_size,lv_attr,devices vg_proxmox_iscsi | grep "vm-107-disk" || true
vgs vg_proxmox_iscsi
pvesm status | egrep '^(Name|iscsi-cluster-lvm)'
```
Expected result:
```text
vm-107-disk-0.qcow2 gone
vm-107-disk-1.qcow2 gone
vm-107-disk-2.qcow2 still present
vm-107-disk-3.qcow2 still present
```
---
## Fallback Method: Direct LVM Removal
Only use this if `pvesm free` refuses and the final verification confirms the volume is not referenced and not open.
```bash
lvremove /dev/vg_proxmox_iscsi/vm-107-disk-0.qcow2
lvremove /dev/vg_proxmox_iscsi/vm-107-disk-1.qcow2
```
Then refresh device nodes and verify:
```bash
vgscan --mknodes
udevadm settle
pvesm list iscsi-cluster-lvm | grep "vm-107-disk" || true
lvs -a -o lv_name,lv_size,lv_attr,devices vg_proxmox_iscsi | grep "vm-107-disk" || true
vgs vg_proxmox_iscsi
```
!!! warning "Prefer pvesm free over lvremove"
`pvesm free` lets Proxmox remove the volume through its storage abstraction. Use direct `lvremove` only when Proxmox refuses and the orphan status is already proven.
---
## Phase 7: Post-Cleanup Validation
After deleting orphan volumes, validate storage and VM health.
```bash
pvesm status
vgs vg_proxmox_iscsi
lvs -a -o lv_name,lv_size,lv_attr,devices vg_proxmox_iscsi
```
Check the affected VM’s config:
```bash
qm config 107
```
Confirm the VM still starts or remains healthy:
```bash
qm status 107
```
If the VM is running, confirm its active QEMU process only references expected disks:
```bash
ps auxww | grep "kvm -id 107" | grep -o "/dev/vg_proxmox_iscsi/[^ ,\"]*" | sort -u
```
Expected example:
```text
/dev/vg_proxmox_iscsi/vm-107-disk-2.qcow2
/dev/vg_proxmox_iscsi/vm-107-disk-3.qcow2
```
---
## Snapshot Volume Handling
Snapshot volumes require additional review.
Examples:
```text
snap_vm-105-disk-0_Fresh_Install.qcow2
snap_vm-106-disk-0_Fresh_Install_FullyUpdated.qcow2
```
Before deleting a snapshot volume, check:
```bash
qm config <vmid>
qm listsnapshot <vmid>
grep -R "snap_vm-<vmid>" /etc/pve/qemu-server/ || true
qemu-img info --backing-chain /dev/vg_proxmox_iscsi/vm-<vmid>-disk-<n>.qcow2
```
If the VM still has a `parent:` line or `qm listsnapshot` shows the snapshot, remove it through Proxmox first:
```bash
qm delsnapshot <vmid> <snapshot-name>
```
Only consider manual removal if:
- the VM no longer references the snapshot,
- no backing chain references the snapshot volume,
- no QEMU process has it open,
- and Proxmox cannot delete it normally.
!!! danger "Do not manually delete active snapshot-chain volumes"
Deleting an active snapshot backing volume can corrupt the VM disk chain.
---
## Example Cleanup Walkthrough
## Scenario
VM `107` has this config:
```text
efidisk0: iscsi-cluster-lvm:vm-107-disk-2.qcow2
sata0: iscsi-cluster-lvm:vm-107-disk-3.qcow2
```
Storage contains:
```text
vm-107-disk-0.qcow2
vm-107-disk-1.qcow2
vm-107-disk-2.qcow2
vm-107-disk-3.qcow2
```
`disk-0` and `disk-1` do not appear in any config and are not open by any process.
## Verify
```bash
grep -R "vm-107-disk-0.qcow2" /etc/pve/qemu-server/ || true
grep -R "vm-107-disk-1.qcow2" /etc/pve/qemu-server/ || true
lsof | grep "vm-107-disk-0.qcow2" || true
lsof | grep "vm-107-disk-1.qcow2" || true
```
Expected output:
```text
no output
```
## Delete
```bash
pvesm free iscsi-cluster-lvm:vm-107-disk-0.qcow2
pvesm free iscsi-cluster-lvm:vm-107-disk-1.qcow2
```
## Validate
```bash
pvesm list iscsi-cluster-lvm | grep "vm-107-disk"
lvs -a -o lv_name,lv_size,lv_attr,devices vg_proxmox_iscsi | grep "vm-107-disk"
vgs vg_proxmox_iscsi
```
Expected remaining volumes:
```text
vm-107-disk-2.qcow2
vm-107-disk-3.qcow2
```
---
## Technician Checklist
Use this checklist before removing any orphan disk.
- [ ] I ran the storage orphan audit.
- [ ] I ran the cluster confirmation script on every Proxmox node.
- [ ] I confirmed the candidate volume is not referenced in any VM config.
- [ ] I confirmed the candidate volume is not open by any process.
- [ ] I confirmed the candidate volume is not part of an active snapshot chain.
- [ ] I confirmed the VMID relationship is understood.
- [ ] I used `pvesm free` first.
- [ ] I used `lvremove` only if Proxmox refused and the volume was proven orphaned.
- [ ] I validated storage state after cleanup.
- [ ] I validated the affected VM still references only expected disks.
---
## Quick Reference Commands
## List shared storage volumes
```bash
pvesm list iscsi-cluster-lvm
```
## List LVs
```bash
lvs -a -o lv_name,lv_path,lv_size,lv_attr,devices vg_proxmox_iscsi
```
## Search VM configs
```bash
grep -R "vm-<vmid>-disk-<n>" /etc/pve/qemu-server/ || true
```
## Check open files
```bash
lsof | grep "vm-<vmid>-disk-<n>" || true
```
## Check image metadata
```bash
qemu-img info --backing-chain /dev/vg_proxmox_iscsi/vm-<vmid>-disk-<n>.qcow2
```
## Delete via Proxmox
```bash
pvesm free iscsi-cluster-lvm:vm-<vmid>-disk-<n>.qcow2
```
## Delete via LVM fallback
```bash
lvremove /dev/vg_proxmox_iscsi/vm-<vmid>-disk-<n>.qcow2
```
## Verify storage usage
```bash
pvesm status
vgs vg_proxmox_iscsi
```
## Related Documentation
- [Related Proxmox Documentation](<../../../reference/Virtualization and Storage/Proxmox/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -1,6 +1,7 @@
---
tags:
- Proxmox
- Virtualization and Storage
---
## Purpose
@@ -17,4 +18,7 @@ vgchange -ay local-vm-storage
It can take some time for everything to come online.
!!! success
If you see something like this: `6 logical volume(s) in volume group "local-vm-storage" now active`, then you successfully brought the volume online.
If you see something like this: `6 logical volume(s) in volume group "local-vm-storage" now active`, then you successfully brought the volume online.
## Related Documentation
- [Related Proxmox Documentation](<../../../reference/Virtualization and Storage/Proxmox/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,18 @@
---
tags:
- Proxmox
- Virtualization and Storage
---
## Purpose
The purpose of this document is to outline common tasks that you may need to run in your cluster to perform various tasks.
## Delete Node from Cluster
Sometimes you may need to delete a node from the cluster if you have re-built it or had issues and needed to destroy it. In these instances, you would run the following command (assuming you have a 3-node quorum in your cluster).
```sh
pvecm delnode promox-node-01
```
## Related Documentation
- [Related Proxmox Documentation](<../../../reference/Virtualization and Storage/Proxmox/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -6,10 +6,9 @@ tags:
- Disaster Recovery
---
**Purpose**:
## Purpose
When you migrate virtual machines from Hyper-V (and possibly other platforms) to ProxmoxVE, you may run into several issues, from the disk formats being in `.raw` format instead of `.qcow2`, among other things. One thing in particular, which is the reason for this document, is that if you migrate Rocky Linux from Hyper-V into ProxmoxVE using Veeam Backup & Replication, it will break the storage system so badly that the operating system will not boot.
### Fixing Boot Issues
Some high-level things to do to fix this are listed below:
@@ -26,6 +25,7 @@ If you start the VM and you reach a "dracut" prompt, then the bootloader got nuk
- Select "**Rescue a Rocky Linux System**"
- Press through the prompt with value `1` and `Continue` to select the automatic mounting of the detected operating system of the virtual machine
- Press **<ENTER>** to enter the shell, then run the following commands to fix the booting issues
```sh
chroot /mnt/sysroot
dracut --force --regenerate-all
@@ -33,6 +33,7 @@ grub2-mkconfig -o /boot/grub2/grub.cfg
exit
exit
```
!!! info "Boot Fix May Trigger Reboot Twice"
During the process, you may notice that the VM reboots itself a second-time. This is normal and can be left alone. The VM will eventually reach the login screen. Once you get this far, you can login and fix the networking issues in the VM to get it stabilized.
@@ -41,6 +42,7 @@ The VM will lose the adapter name of `eth0` and put something else like `ens18`
- Type `ethtool ens18`, and if the link speed is `Unknown!`, then poweroff the VM and switch the network adapter from `VirtIO (paravirtualized)` to `Intel E1000`, then boot the VM back up.
- Run the following commands to assign the new `ens18` interface as a networking interface for the VM to use:
```sh
# Create the Interface (Replace the IP & DNS Variables)
nmcli connection add type ethernet ifname ens18 con-name ens18 ipv4.method manual ipv4.addresses 192.168.3.21/24 ipv4.gateway 192.168.3.1 ipv4.dns "1.1.1.1 1.0.0.1"
@@ -55,11 +57,15 @@ nmcli connection up ens18
### Convert VM Disk from `.RAW` to `.QCOW2`
Given that the migration process via Veeam Backup & Replication ignores the destination disk format (at the time of writing this), it is necessary to convert the format of the disk from `.raw` to `.qcow2` so that you can perform things like VM snapshots, which are essential during updates, development, and testing.
Open a shell onto the ProxmoxVE server that is currently holding the VM that you need to convert the disks for, then locate the disks (this is not explained here, yet), and run the following commands to convert them.
Open a shell onto the ProxmoxVE server that is currently holding the VM that you need to convert the disks for, then locate the disks (this is not explained here, yet), and run the following commands to convert them.
```sh
# Convert a Single Disk
qemu-img convert -f raw -O qcow2 source.raw destination.qcow2
# Convert All Disks in a Given Directory
find . -type f -name "*.raw" -exec sh -c 'qemu-img convert -f raw -O qcow2 "$1" "${1%.raw}.qcow2"' _ {} \;
```
```
## Related Documentation
- [Related Proxmox Documentation](<../../../reference/Virtualization and Storage/Proxmox/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,68 @@
---
tags:
- Virtualization and Storage
- Repair iSCSI Connections After Reboot
---
## Purpose
You may enounter an issue where when you reboot your ProxmoxVE cluster node, it will fail to start virtual machines because of an error related to being unable to see the underlying LVM disks for the GuestVM. This is generally bandaid-fixed by running `iscsiadm -m node --login` on the node, which makes it reconnect to the cluster's iSCSI storage. This is not a viable long-term solution.
### Configure Automatic Startup of iSCSI Services
Run these commands on every ProxmoxVE cluster node:
```sh
systemctl enable --now iscsid
systemctl enable --now open-iscsi
```
### Discover Targets (To Create Necessary iSCSI Records)
```sh
iscsiadm -m discovery -t sendtargets -p 192.168.3.3:3260
```
### Configure Automatic iSCSI Target Connection Behavior
Run these commands on every ProxmoxVE cluster node:
```sh
iscsiadm -m node \
-T iqn.2026-01.io.bunny-lab:storage:iscsi-cluster-storage \
-p 192.168.3.3 \
--op update \
-n node.startup \
-v automatic
iscsiadm -m node \
-T iqn.2026-01.io.bunny-lab:storage:iscsi-cluster-storage \
-p 192.168.3.3 \
--op update \
-n node.conn[0].startup \
-v automatic
```
### Verify Configuration
You will want to ensure that `node.startup = automatic` and `node.conn[0].startup = automatic` when you run the following command.
```sh
iscsiadm -m node -o show | grep -E 'node.name|node.conn\[0\].address|node.startup|node.conn\[0\].startup'
```
### Login/Mount iSCSI Targets
```sh
iscsiadm -m node \
-T iqn.2026-01.io.bunny-lab:storage:iscsi-cluster-storage \
-p 192.168.3.3:3260 \
--login
```
!!! success "Example Output"
If everything worked correctly, you should see output like the example below:
```sh
node.name = iqn.2026-01.io.bunny-lab:storage:iscsi-cluster-storage
node.startup = automatic
node.conn[0].address = 192.168.3.3
node.conn[0].startup = automatic
```
## Related Documentation
- [Related Proxmox Documentation](<../../../reference/Virtualization and Storage/Proxmox/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -1,6 +1,7 @@
---
tags:
- Proxmox
- Virtualization and Storage
---
## Purpose
@@ -13,6 +14,7 @@ There are a few steps you have to take when upgrading ProxmoxVE from 8.4.1+ to 9
It's critical that you run the `pve8to9` command to ensure that your ProxmoxVE server meets all of the requirements and doesn't have any failures or potentially server-breaking warnings. If the `pve8to9` command is unknown, then run `apt update && apt dist-upgrade` in the shell then try again. Warnings should be addressed ad-hoc, but *CPU Microcode warnings can be safely ignored*.
**Example pve8to9 Summary Output**:
```sh
= SUMMARY =
@@ -40,4 +42,7 @@ reboot
```
!!! note "Disable `pve-enterprise` Repository"
At this point, the ProxmoxVE server should be running on v9.0+, you will want to disable the `pve-enterprise` repository as it will goof up future updates if you don't disable it.
At this point, the ProxmoxVE server should be running on v9.0+, you will want to disable the `pve-enterprise` repository as it will goof up future updates if you don't disable it.
## Related Documentation
- [Related Proxmox Documentation](<../../../reference/Virtualization and Storage/Proxmox/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,50 @@
---
tags:
- TrueNAS
- Storage
- Hardware
---
## Purpose
This document acts as a workflow to understand how to replace a drive on TrueNAS Core when it is hosted on an HPE Proliant server with HBA / IT Mode enabled. This enables you to hot-swap drives without rebooting TrueNAS Core.
### Offline the Disk
- You will log into the TrueNAS Core [WebUI](http://192.168.3.3).
- Navigate to "**Storage > Disks**"
- Look for the drive that is having issues / faults / unavailable and reference it's `da` number to reference later. (e.g. `da3`)
- Confirm the serial number of the drive and correlate that to the physical location in the [Disk Arrays](<../../../reference/Lab Map/Hardware/Storage Node 01 TrueNAS Core Disk Layout.md>) document,
- Navigate to "**Storage > Pools**"
- Look for the gear icon to the right of the storage pool and click on it
- Click on "**Status**"
- Locate the failing / failed drive and click on the "**...**" elipsis menu button
- Proceed to "**Offline**" the disk. This ensures that TrueNAS Core stops trying to use the disk.
### Physical Disk Replacement
At this point, we need to physically go to the server and pull out the failing drive and replace it.
- Take note of the new serial number on the replacement drive and update the [Disk Arrays](<../../../reference/Lab Map/Hardware/Storage Node 01 TrueNAS Core Disk Layout.md>) document accordingly.
- Insert the replacement drive back into the TrueNAS Core server
### Trigger Disk Re-Scan
Now we need to tell TrueNAS / FreeBSD to re-scan all disks to locate the new one.
- Within the TrueNAS Core WebUI, navigate to "**Shell**" and run the following command: `camcontrol rescan all`
- Navigate (back) to "**Storage > Pools > Status**"
- Locate the failed drive via it's `da` number again, and click the "**...**" elipsis menu button
- Proceed to "**Replace**" the disk, and when given a dropdown menu, only the new replacement disk should appear with the same `da` number
!!! success "Resilvering Started"
At this point, TrueNAS core will start taking parity data from the rest of the drives in the storage pool to reconstruct the replaced drive. This may take an hour or two depending on the speed of the drives and used capacity within the pool itself.
It is recommended to run a SCRUB right after resilvering to ensure that all data is accurate and healthy.
!!! info "Checking on Resilvering Process via CLI"
If you feel so inclined, you can check on the resilvering process by running the following command:
```sh
zpool status | grep "to go"
```
## Related Documentation
- [Storage Node 01 Disk Layout](<../../../reference/Lab Map/Hardware/Storage Node 01 TrueNAS Core Disk Layout.md>) — Match the serial number and physical slot before replacement.
- [Related Virtualization and Storage Documentation](<../../../reference/Virtualization and Storage/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,19 @@
---
tags:
- Virtualization and Storage
- Workflows
- Documentation
---
# Virtualization and Storage
## Purpose
Find workflows for virtualization and storage. Follow the subject guide to choose the relevant environment and connect this material to the other document types.
## Includes
- Hyper-V
- Linux
- Proxmox
- TrueNAS
## Follow the Subject
[Virtualization and Storage](<../../reference/Virtualization and Storage/index.md>) explains the relationships and offers starting points for the documented tasks.
@@ -1,3 +1,9 @@
---
tags:
- Windows and Linux
- Restrict Monitors on Plasma Login Screen
---
## Purpose
I wrote this document because CachyOS's desktop environment has this dumb issue where the login screen tries presenting itself on every monitor, and due to my unique configuration that includes an invisible display created from my soundbar, I have to run the commands below to configure the Plasma login screen to mimic my environment.
@@ -6,6 +12,7 @@ The first thing you want to do is configure the displays so they are arranged /
### Copy Display Configuration into Plasma Login Screen
Run the following commands as your normal user, do not run them as `root`.
```sh
sudo install -d -o plasmalogin -g plasmalogin /var/lib/plasmalogin/.config
sudo cp ~/.config/kwinoutputconfig.json /var/lib/plasmalogin/.config/kwinoutputconfig.json
@@ -13,4 +20,7 @@ sudo chown plasmalogin:plasmalogin /var/lib/plasmalogin/.config/kwinoutputconfig
```
### Logout and Verify
At this point, the lock screen / login screen should be correctly configured to mirror the normal desktop environment and display arrangement.
At this point, the lock screen / login screen should be correctly configured to mirror the normal desktop environment and display arrangement.
## Related Documentation
- [Related Windows and Linux Documentation](<../../../../reference/Windows and Linux/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -11,11 +11,14 @@ You may need to install flatpak packages like Signal in your workstation environ
```sh
# Usually already installed
sudo dnf install flatpak
sudo dnf install flatpak
# Add Flathub Repo
flatpak --user remote-add --if-not-exists flathub https://flathub.org/repo/flathub.flatpakrepo
flatpak --user remote-add --if-not-exists flathub https://flathub.org/repo/flathub.flatpakrepo
# Install Signal
flatpak install flathub org.signal.Signal
```
```
## Related Documentation
- [Related Windows and Linux Documentation](<../../../../reference/Windows and Linux/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -5,7 +5,7 @@ tags:
- Workstation
---
**Purpose**:
## Purpose
If you want to upgrade Fedora Workstation to a new version (e.g. 41 --> 42) you can run the following commands to do so. The overall process is fairly straightforward and requires a reboot.
```sh
@@ -15,4 +15,7 @@ sudo dnf system-upgrade reboot
```
**Additional Documentation**:
https://docs.fedoraproject.org/en-US/quick-docs/upgrading-fedora-new-release/
https://docs.fedoraproject.org/en-US/quick-docs/upgrading-fedora-new-release/
## Related Documentation
- [Related Windows and Linux Documentation](<../../../../reference/Windows and Linux/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,18 @@
---
tags:
- Linux
- Bash
- Scripting
---
## Purpose
This document records the procedure for repair displaylink usb authorization. Follow the environment assumptions and commands below.
```sh
xrandr --auto
xrandr --setprovideroutputsource 4 0
xrandr --output HDMI-1 --primary --mode 1920x1080 --rate 75.00 --output DVI-I-1-1 --mode 1920x1080 --rate 60.00 --right-of HDMI-1 --output -eDP-1 --off
```
## Related Documentation
- [Related Windows and Linux Documentation](<../../../reference/Windows and Linux/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -1,15 +1,18 @@
---
tags:
- Windows
- Windows and Linux
---
# Changing Windows Editions
### Changing Editions:
## Purpose
Record the commands and generic installation-key reference used to change a Windows edition. Select the command and edition that match the target operating system.
### Changing Editions
Windows Server: `DISM /ONLINE /set-edition:serverstandard /productkey:AAAAA-BBBBB-CCCCC-DDDDD-EEEEE /AcceptEula`
Windows (Home/Pro): `DISM /ONLINE /set-edition:professional /productkey:AAAAA-BBBBB-CCCCC-DDDDD-EEEEE /AcceptEula`
### Force Activation / Edition Switcher:
### Force Activation / Edition Switcher
`irm https://get.activated.win | iex`
## Generic Install Keys
@@ -39,19 +42,18 @@ Windows (Home/Pro): `DISM /ONLINE /set-edition:professional /productkey:AAAAA-B
| Windows 10 Enterprise N LTSB 2016 | RW7WN-FMT44-KRGBK-G44WK-QV7YK | QFFDN-GRT3P-VKWWX-X7T3R-8B639 |
| Windows 10 Enterprise LTSC 2019 | | M7XTQ-FN8P6-TTKYV-9D4CC-J462D |
| Windows 10 Enterprise N LTSC 2019 | | 92NFX-8DJQP-P6BBQ-THF9C-7CG2H |
| Windows 10 Home | 37GNV-YCQVD-38XP9-T848R-FC2HD | | |
| Windows 10 Home N | 33CY4-NPKCC-V98JP-42G8W-VH636 | | |
| Windows 10 Pro | NF6HC-QH89W-F8WYV-WWXV4-WFG6P | | |
| Windows 10 Pro N | NH7W7-BMC3R-4W9XT-94B6D-TCQG3 | | |
| Windows 10 SL | NTRHT-XTHTG-GBWCG-4MTMP-HH64C | | |
| Windows 10 CHN SL | 7B6NC-V3438-TRQG7-8TCCX-H6DDY | | |
| Windows 10 Home | 46J3N-RY6B3-BJFDY-VBFT9-V22HG | | |
| Windows 10 Home N | PGGM7-N77TC-KVR98-D82KJ-DGPHV | | |
| Windows 10 Pro | RHGJR-N7FVY-Q3B8F-KBQ6V-46YP4 | | |
| Windows 10 Pro N | 2KMWQ-NRH27-DV92J-J9GGT-TJF9R | | |
| Windows 10 SL | GH37Y-TNG7X-PP2TK-CMRMT-D3WV4 | | |
| Windows 10 CHN SL | 68WP7-N2JMW-B676K-WR24Q-9D7YC | | |
| Windows 10 Home | 37GNV-YCQVD-38XP9-T848R-FC2HD | | |
| Windows 10 Home N | 33CY4-NPKCC-V98JP-42G8W-VH636 | | |
| Windows 10 Pro | NF6HC-QH89W-F8WYV-WWXV4-WFG6P | | |
| Windows 10 Pro N | NH7W7-BMC3R-4W9XT-94B6D-TCQG3 | | |
| Windows 10 SL | NTRHT-XTHTG-GBWCG-4MTMP-HH64C | | |
| Windows 10 CHN SL | 7B6NC-V3438-TRQG7-8TCCX-H6DDY | | |
| Windows 10 Home | 46J3N-RY6B3-BJFDY-VBFT9-V22HG | | |
| Windows 10 Home N | PGGM7-N77TC-KVR98-D82KJ-DGPHV | | |
| Windows 10 Pro | RHGJR-N7FVY-Q3B8F-KBQ6V-46YP4 | | |
| Windows 10 Pro N | 2KMWQ-NRH27-DV92J-J9GGT-TJF9R | | |
| Windows 10 SL | GH37Y-TNG7X-PP2TK-CMRMT-D3WV4 | | |
| Windows 10 CHN SL | 68WP7-N2JMW-B676K-WR24Q-9D7YC | | |
### Windows Server
| Windows Edition | RTM Generic Key (Retail) | [**KMS Client Setup Key**](https://docs.microsoft.com/en-us/previous-versions/windows/it-pro/windows-server-2012-R2-and-2012/jj612867(v%3dws.11)) |
@@ -66,6 +68,9 @@ Windows (Home/Pro): `DISM /ONLINE /set-edition:professional /productkey:AAAAA-B
| Windows Server 2022 Datacenter Azure | | NTBV8-9K7Q8-V27C6-M2BTV-KHMXV |
| Windows Server 2022 Datacenter | | WX4NM-KYWYW-QJJR4-XV3QB-6VM33 |
## Additional Reference Documentation:
## Additional Reference Documentation
https://www.tenforums.com/tutorials/95922-generic-product-keys-install-windows-10-editions.html
[https://learn.microsoft.com/en-us/windows-server/get-started/kms-client-activation-keys](https://learn.microsoft.com/en-us/windows-server/get-started/kms-client-activation-keys)
[https://learn.microsoft.com/en-us/windows-server/get-started/kms-client-activation-keys](https://learn.microsoft.com/en-us/windows-server/get-started/kms-client-activation-keys)
## Related Documentation
- [Related Windows and Linux Documentation](<../../../reference/Windows and Linux/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -1,10 +1,11 @@
---
tags:
- Windows
- Windows and Linux
---
**Purpose**:
Sometimes you are running a virtual machine and are running out of space, and want to expand the operating system disk. However, there is a recovery partition to-the-right of the operating system partition. When this happens, you have to delete that partition in order to expand the storage space for the operating system.
## Purpose
Sometimes you are running a virtual machine and are running out of space, and want to expand the operating system disk. However, there is a recovery partition to-the-right of the operating system partition. When this happens, you have to delete that partition in order to expand the storage space for the operating system.
These commands can be run in a headless environment using just powershell.
@@ -12,6 +13,7 @@ These commands can be run in a headless environment using just powershell.
In my example codeblock, I assume the OS drive is `0` and the recovery partition is `4`. Please validate your own drive and partition numbers with the supplied `list disk` and `list partition` commands. Failure to identify the correct drive and/or partition could result in the unintended destruction of data.
**From within the VM** > Open a powershell window and run the following commands:
```powershell
diskpart # (1)
list disk # (2)
@@ -36,15 +38,20 @@ exit # (9)
## Free Space Validation
From this point, you might want to verify the free space has been accounted for, so you can run the following command to check for free space:
```powershell
Get-Volume | Select-Object DriveLetter, FileSystem, @{Name="FreeSpace(GB)"; Expression={"{0:N2}" -f ($_.SizeRemaining / 1GB)}}, @{Name="TotalSize(GB)"; Expression={"{0:N2}" -f ($_.Size / 1GB)}}
```
!!! example "Output Example"
```
```text
DriveLetter FileSystem FreeSpace(GB) TotalSize(GB)
----------- ---------- ------------- -------------
C NTFS 398.40 476.20
FAT32 0.06 0.09
NTFS 0.11 0.63
```
```
## Related Documentation
- [Related Windows and Linux Documentation](<../../../reference/Windows and Linux/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -1,3 +1,9 @@
---
tags:
- Windows and Linux
- Uninstall Updates via DISM
---
## Purpose
You may find that you have a machine that does not want to boot because of a recent Windows Update. In these cases, you only viable option sometimes may be to simply uninstall the update. You can do this via a bootable Windows installer (ISO or USB), or you can try to use the built-in recovery menu in windows.
@@ -40,4 +46,7 @@ DISM /Image:D:\ /Remove-Package /PackageName:Package_for_ServicingStack_4763~31b
At this point, you should be safe to close the command prompt and attempt to reboot the computer. If it still doesn't boot, you can try alternative offline repair methods such as `DISM /Image:D:\ /Cleanup-Image /RestoreHealth` and `sfc /scannow /offbootdir=D:\ /offwindir=D:\Windows`. If you have an install ISO mounted / accessible to the problematic device, you can tell DISM to exclusively use that install media to repair the OS:
- `DISM /Image:D:\ /Cleanup-Image /RestoreHealth /Source:wim:X:\sources\install.wim:1 /LimitAccess`
- `DISM /Image:D:\ /Cleanup-Image /RestoreHealth /Source:esd:X:\sources\install.esd:1 /LimitAccess`
- `DISM /Image:D:\ /Cleanup-Image /RestoreHealth /Source:esd:X:\sources\install.esd:1 /LimitAccess`
## Related Documentation
- [Related Windows and Linux Documentation](<../../../reference/Windows and Linux/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -0,0 +1,42 @@
---
tags:
- Windows
- VSS
- Backup
---
## Purpose
There are times when you may need to delete shadow copies (Volume Shadow Copies) from a drive, commonly to free up disk space. While this is usually straightforward, you may encounter scenarios where shadow copies cannot be deleted through normal means. The following methods provide ways to forcibly remove all shadow copies from a specific volume.
!!! warning
The examples below will **permanently delete all shadow copies** on the specified drive. The examples use drive `D:` > Adjust the drive letter as needed.
## Method 1: Delete Shadow Copies Using `vssadmin`
The `vssadmin` utility is the standard tool for managing shadow copies. It is typically safe and handles deletions gracefully.
However, some antivirus or endpoint protection software may block its execution due to its similarity to behavior used by ransomware. If `vssadmin` fails, use the `diskshadow` method described below.
```cmd
vssadmin delete shadows /for=D: /all /quiet
```
- `/for=D:` specifies the target volume.
- `/all` removes all shadow copies on that volume.
- `/quiet` suppresses confirmation prompts.
## Method 2: Delete Shadow Copies Using `diskshadow`
`diskshadow` is a more direct and lower-level tool than `vssadmin`. It should be used as a fallback option if `vssadmin` fails or is blocked.
```cmd
diskshadow
set context persistent nowriters
delete shadows volume D:
exit
```
Explanation:
- `set context persistent nowriters` ensures the command does not involve writer components (e.g., for backups), reducing the chance of interference.
- `delete shadows volume D:` removes all persistent shadow copies for volume `D:`.
## Related Documentation
- [Related Windows and Linux Documentation](<../../../../reference/Windows and Linux/index.md>) — Find the connected deployments, procedures, and references for this subject.
@@ -5,7 +5,8 @@ tags:
- User Accounts
---
**Purpose:** You may find that Windows 11 does not allow you to install it with a local account. This is a documented case of Microsoft attempting to push Microsoft accounts, and can be bypassed by following the workflow below:
## Purpose
You may find that Windows 11 does not allow you to install it with a local account. This is a documented case of Microsoft attempting to push Microsoft accounts, and can be bypassed by following the workflow below:
## Initial Boot to Windows 11 Installer
- Begin installing the OS as normal, selecting the region/language
@@ -18,4 +19,7 @@ tags:
- Set up the local administrator account as normal and finish the OS installation process
!!! warning "Disconnect Internet"
To ensure a clean installation devoid of additional issues, make sure to disconnect the physical/virtual network from the device before proceeding to install Windows 11 as normal. This time, you will not be prompted to login with a Microsoft account.
To ensure a clean installation devoid of additional issues, make sure to disconnect the physical/virtual network from the device before proceeding to install Windows 11 as normal. This time, you will not be prompted to login with a Microsoft account.
## Related Documentation
- [Related Windows and Linux Documentation](<../../../../reference/Windows and Linux/index.md>) — Find the connected deployments, procedures, and references for this subject.
+17
View File
@@ -0,0 +1,17 @@
---
tags:
- Windows and Linux
- Workflows
- Documentation
---
# Windows and Linux
## Purpose
Find workflows for windows and linux. Follow the subject guide to choose the relevant environment and connect this material to the other document types.
## Includes
- Linux
- Windows
## Follow the Subject
[Windows and Linux](<../../reference/Windows and Linux/index.md>) explains the relationships and offers starting points for the documented tasks.
+20 -30
View File
@@ -1,42 +1,32 @@
---
tags:
- Workflows
- Index
- Documentation
---
# Workflows
## Purpose
Runbooks for maintenance, troubleshooting, backups, and day-2 operations.
Find a procedure for maintaining, migrating, or recovering an existing system. Check the target environment and starting state before following its commands.
## Includes
- Backup and DR workflows
- Routine maintenance tasks
- Troubleshooting runbooks
- Applications
- Automation
- Backup and Recovery
- Containers
- Identity and Certificates
- Networking and Access
- Virtualization and Storage
- Windows and Linux
## New Document Template
````markdown
# <Document Title>
## Purpose
<what this runbook exists to solve>
## Start with a Subject
- [Applications](<../reference/Applications/index.md>) — Find applications by the service they provide, then continue to their deployment, authentication, data, and maintenance documentation.
- [Automation](<../reference/Automation/index.md>) — Connect source control, automation controllers, managed hosts, and configuration delivery. Use the documented execution environment and authentication method for each workflow.
- [Backup and Recovery](<../reference/Backup and Recovery/index.md>) — Find backup concepts, repository maintenance, and recovery dependencies. Select the procedure for the affected backup system and distinguish a backup restore from a replica or snapshot operation.
- [Containers](<../reference/Containers/index.md>) — Prepare the Docker or Kubernetes environment used by application deployments, then follow the operating procedures for building, moving, and exposing workloads.
- [Identity and Certificates](<../reference/Identity and Certificates/index.md>) — Connect directory services, certificate trust, single sign-on, and application authentication. Start with the identity system involved, then follow the integration or maintenance procedure.
- [Networking and Access](<../reference/Networking and Access/index.md>) — Find the DNS, proxy, VPN, and remote-access instructions that connect users and services. Use the address plans to identify the intended network before changing connectivity.
- [Virtualization and Storage](<../reference/Virtualization and Storage/index.md>) — Follow the relationship between hypervisors, shared storage, guest disks, and recovery procedures. Select the documented storage design before choosing a maintenance command.
- [Windows and Linux](<../reference/Windows and Linux/index.md>) — Find workstation and server operating-system setup, updates, and repairs. Storage, networking, and identity tasks are linked to their subject guides when they cross operating-system boundaries.
!!! warning "Risk"
- <irreversible actions or data impact>
## Procedure
```sh
# Commands or steps (grouped and annotated)
```
## Validation
- <command + expected result>
## Troubleshooting
### Symptoms
- <what you see>
### Resolution
```sh
# Fix steps
```
````
## Find Related Knowledge
[The subject guides](<../reference/index.md>) connect these workflows to the other document roles.
@@ -1,45 +0,0 @@
---
tags:
- Veeam
- Backup
- Disaster Recovery
---
**Purpose**:
The purpose of this document is to explain the core concepts / terminology of things seen in Veeam Backup & Replication from a relatively high-level. It's more of a quick-reference guide than a formal education.
## Backup Jobs
Backup jobs take many forms, but the most common are explained in more detail below. Note that this is not an exhaustive list of the different kinds of backup jobs, just the ones I am currently most familiar with.
- **Backup**: This is the simplest of the backup job options. A "Backup" backup job will take a backup of a workstation, server, File Server, specific local files and folders on a device, or a GuestVM running in a hypervisor such as Hyper-V, VMWare ESXi, or ProxmoxVE.
- **Backup Copy**:
- This is when you make a copy of backup data stored on the Veeam server, and send it somewhere else, such as an off-site "Service Provider" such as Veeam partners.
- You can also send backup copies to local drives, SMB network shares, NFS shares, File Servers, pretty much anywhere you can send normal backups, but with the key difference being the data is originating from the Veeam backup server itself instead of the original server/VM.
- **SureBackup**: This is where things get a little more complex. SureBackup is where you effectively "Verify" your backups by spinning them up inside of a lab environment. While they are spun up, they are checked to see if they fully boot, they can have antivirus scans, ransomware scans, custom scripts executed, and validate the integrity of the backups. The general core components are listed below:
- **Virtual Lab**: The virtual lab is a virtual machine environment that you set up for Veeam to leverage to spin up backups on a hypervisor that you configure, such as a remote Hyper-V server in the same building, or perhaps if you have Hyper-V locally installed on the same server as Veeam itself, you would configure the virtual lab's hypervisor to point to `127.0.0.1` or `localhost`.
- The virtual lab will have its own unique virtual networking for the VMs to communicate on, so they don't conflict with the production servers/VMs.
- **Application Groups**: Application groups are defined groups of devices that need to be running when the backups are being validated. For example, in my homelab, I have an application group named `Domain Controllers`, and I put `LAB-DC-01` and `LAB-DC-02` into that application group. I use this as the application group associated with the Virtual Lab because most of my services are authenticated with Active Directory, and if the DCs were missing during backup verification, a variety of issues would ensue. When the Backup Verification Lab (Virtual Lab) is launched on the targeted hypervisor, it spins up the application group devices from backups first, ensuring they are running and functional, before the virtual lab starts verifying backup objects designated in the "Linked Jobs", seen in the next section.
- **Linked Jobs**: These are the "Backup Jobs" you want to verify in in the virtual lab mentioned above. If you have a large backup job with a bunch of machines you don't want verified, you can configure "Exclusions" in the SureBackup job settings to exclude those objects/devices from verification.
## Replication Jobs
As the name states, Veeam Backup & Replication can also handle replicating Servers/VMs from either their original locations or from a recent backup and push them into a hypervisor for rapid failover/failback functionality. Very useful for workloads that need to be spun up nearly immediately due to strict RTO requirements. There are some additional notes regarding replication seen below.
!!! warning "Orchestrate Replication & Failover via Veeam, not the Hypervisor"
You want to coordinate anything replication-wise directly in Veeam Backup & Replication, not directly on the hypervisor itself. While you can do this, it is not only slower, but does not give you the option to failback replicas back into production if you spin up a replica directly on its hypervisor.
- **Replication Restore Points**: Similar to backups, replicas can have multiple restore points associated with them, so you have more than one option when spinning up a replica in a hypervisor.
- **Planned Failover**: A planned failover is when you are scheduling the hypervisor to be offline and simply don't have enough resources to live-migrate it to another cluster host, or you might not even have a virtualization cluster to work with in the first place. In cases like this, a "Planned Failover" tells Veeam to make a fresh replica right now, then shuts down the production VM on its hypervisor, and spins up the replica on the replica server. (If you installed Hyper-V on the Veeam server, it would spin up the replica on the backup server itself).
- A "Planned Failover" allows you to perform a "**Failback to Production**" when the failover event has concluded. This means that while the production VM was offline and the replica took over the production load, any changes made such as new files added, applications installed, etc will be replicated back to the production VM when the replica is "Failed back to Production". **This is the ideal choice in most circumstances**.
- **Failover Now**: Failover now means that the production hypervisor is likely completely dead, and may need to be re-built, or you simply dont need to replicate changes back to production hypervisor after the failover event has concluded, such as on a low-priority print server. Any changes made while the replica is operational will be completely lost when the production VM is turned back on again or a restore is pushed back onto a new hypervisor.
## Backup Infrastructure
### Backup Repository
A backup repository is simply a destination to send the backups or backup copies. It can be anything from direct attached storage to a SMB file share on a NAS, or even off-site storage like Backblaze B2 or Amazon S3.
- If you use object storage like Backblaze B2 or Amazon S3, you can configure an "Immutability Period" for backups that are sent to these destinations, meaning if your backup server was hit by ransomware or a malicious actor, neither they nor you could delete the backups in the off-site storage such as Backblaze B2 until the immutability period had passed, such as 7 days, 30 days, or however long you configured.
- You can adjust the immutability period after-the-fact, but backups that have already been pushed to a backup repository will be immutable for the time period configured when they were originally uploaded, and attempts to delete them will tell you when you are allowed to delete them. You won't be able to delete them even from Amazon or Backblaze's own internal tools / websites during this immutability period.
### Backup Proxy
A backup "proxy" simply refers to a machine that is running the "**Veeam Backup Transport**" agent on it. The Veeam Backup & Replication server installs a proxy onto itself, but it also deploys proxies onto workstations, servers, and hypervisors. These proxies are how the "Veeam Backup & Replication Console" interacts with the devices and performs backups and restores.
### Service Provider
Service Providers are not the same as cloud storage providers such as Backblaze B2, Amazon S3, etc. Service Providers are Veeam "partners" who manage, maintain, and deploy Veeam backup appliances at client environments, as well as providing support to clients within the Veeam ecosystem. You can also use Service Providers as a cloud backup destination in Veeam Backup & Replication for off-site backups.
## Misc Terminology
- **Unstructured Data**: This refers to a device such as a windows or linux server that you can use WinRM or SSH to access, and want to backup specific files and folders without backing up the entire device / VM. This is useful in cases where you cannot install a Veeam Agent or the operating system is unsupported by Veeam, or if the device is not operating under a hypervisor, such as a bare-metal server.
- When you add a device to Veeam's "Inventory" via the "Unstructured Data" section, if you want to perform backups on the device, you will have to make a special backup job under "**Backups > File Server**", because Veeam will treat the unstructured data as a file server.
@@ -1,24 +0,0 @@
---
tags:
- Veeam
- Backup
- Disaster Recovery
---
**Purpose**:
This is meant as a high-level generally-speaking best practice retention policy in most use-cases. This document will generally be pretty bare-bones, but the general idea is the following advanced GFS retention period is generally configured on backup copy jobs, specifically ones that have off-site backups, but can also be used for local backup repositories.
Navigate to Jobs > Backup (or Backup Copy) > (Find a Backup Job) > Right-Click > Edit > Storage (or Target) > "**Keep Certain Full Backups for Archival Purposes**: Checked" > Click on the "**Configure**" button.
Optional: Click the "**Save as Default**" button before clicking the "**OK**" button to make this default behavior for new backup jobs.
| **Description** | **Status** | **Value** |
| :--- | :--- | :--- |
| Keep Weekly Full Backups | Enabled | 4 |
| Keep Monthly Full Backups | Enabled | 3 |
| Keep Yearly Full Backups | Enabled | 1 (`3 - 7 for Medical HIPAA`) |
!!! note "7 Daily Backups Assumption"
This document assumes that you at (least) keep 7 daily backups in the normal backup schedule. Meaning **7 daily, 4 weekly, 3 monthly, and 1 yearly** backup is maintained at all times.
**7 daily, 4 weekly, 3 monthly, and 1 yearly**
@@ -1,21 +0,0 @@
---
tags:
- iLO
- Hardware
- Licensing
---
!!! info "Assumptions of Usage"
It should go without saying, using one of these keys does not entitle you to support by Hewlett-Packard Enterprise. These are meant for homelab environments where licensing / auditing does not matter.
| **iLO Version** | **License Key** |
| :--- | :--- |
| iLO Standard Trial | `34T6L-4C9PX-X8D9C-GYD26-8SQWM` |
| iLO 1 Advanced | `247RH-ZPJ8S-7B17D-FCE55-DDD17` |
| iLO 2 / 3 / 4 Advanced | `35DPH-SVSXJ-HGBJN-C7N5R-2SS4W` |
| iLO 2 / 3 / 4 / 5 Advanced | `35SCR-RYLML-CBK7N-TD3B9-GGBW2` |
!!! warning "Do not Use in Production Work Environments"
In (rare) cases, these keys can be used as a temporary solution when working in a work environment, then promptly removed after the work is performed. Leaving them installed on a server could lead to legal consequences if Hewlett-Packard Enterprise asked for it while providing support, and it was using one of these keys, it could fail a software licensing audit.
`REMOVE THE KEY AFTER USAGE`
@@ -1,58 +0,0 @@
---
tags:
- Fedora
- Linux
- Workstation
---
**Purpose**:
This document serves as a general guideline for my workstation deployment process when working with Fedora Workstation 41 and up. This document will constantly evolve over time based on my needs.
## Automate Initial Configurations
```sh
# Set Hostname
sudo hostnamectl set-hostname lab-desktop-01
# Setup Automatic Drive Mounting
echo "/dev/disk/by-uuid/B865-7BDB /mnt/500GB_WINDOWS_OS auto nosuid,nodev,nofail,x-gvfs-show 0 0
/dev/disk/by-uuid/C006EBA006EB95A6 /mnt/640GB_HDD_STORAGE auto nosuid,nodev,nofail,x-gvfs-show 0 0
/dev/disk/by-uuid/24C82CFEC82CCFBA /mnt/1TB_SSD_STORAGE auto nosuid,nodev,nofail,x-gvfs-show 0 0
/dev/disk/by-uuid/D64E9F534E9F2AEF /mnt/120GB_SSD_STORAGE auto nosuid,nodev,nofail,x-gvfs-show 0 0
/dev/disk/by-uuid/16D05248D0522E6D /mnt/2TB_SSD_STORAGE auto nosuid,nodev,nofail,x-gvfs-show 0 0" | sudo tee -a /etc/fstab
# Install Software
sudo yum update -y
sudo yum install -y steam firefox
sudo dnf install -y @xfce-desktop-environment
# Reboot Workstation
sudo reboot
```
!!! warning "Read-Only NTFS Disks (When Using Dual-Boot)"
If you want to dual boot, you need to ensure that the Windows side does not have "Fast Boot" enabled. You can locate the Fast Boot setting by locating the "Change what the power button does" settings, and unchecking the "Fast Boot" checkbox, then shutting down.
The problem with Fast Boot is that it effectively leaves the shared disks between Windows and Linux in a locked read-only state, which makes installing Steam games and software impossible.
## Manually Address Remaining Things
At this point, we need to do some manual work, since not everything can be handled by the terminal.
### Install Software (Software Manager)
Now we need to install a few things:
- NVIDIA Graphics Drivers Control Panel
- Discord Canary
- Betterbird
- Visual Studio Code
- Signal Desktop
- Solaar # (Logitech Unifying Software equivalant in Linux)
### Import XFCE Panel Configuration
At this point, we want to restore our custom taskbar / panels in XFCE, so the easiest way to do that is to import the configuration backup located in Nextcloud.
Backups are located here: https://cloud.bunny-lab.io/f/792649
### Configure Window Snapping
By default, XFCE has a really small threshold for telling windows to "snap" to the sides of the screens, such as a half:half arrangement. This can be adjusted by navigating to "**Applications Menu > Settings > Settings Manager > Windows Manager Tweaks > Placement**"
Once you have reached this window, you will see a slider from "**Small**" to "**Large**". Slide the slider all the way to the right, facing "**Large**". Now windows will snap to the sides of the screen successfully.
@@ -1,76 +0,0 @@
---
tags:
- Fedora
- Linux
- Desktop Environment
- Workstation
---
## Purpose
You may find that you need to install an XFCE desktop environment or something into Fedora Server, if this is the case, for installing something like Rustdesk remote access, you can follow the steps below.
### Install & Configure XFCE
We need to install XFCE and configure it to be the default environment when the server turns on.
```sh
sudo dnf install @xfce-desktop-environment -y
sudo systemctl set-default graphical.target
sudo reboot
```
#### Install Rustdesk:
We need to install Rustdesk into the server.
```sh
curl -L -o /tmp/rustdesk_installer.rpm https://github.com/rustdesk/rustdesk/releases/download/1.4.0/rustdesk-1.4.0-0.x86_64.rpm
cd /tmp
sudo yum install rustdesk_installer.rpm -y
```
!!! info "Configure Rustdesk"
You need to use a tool like "MobaXTerm" or "PuTTy" to leverage X11-Forwarding to allow you to run `rustdesk` in a GUI on your local workstation. From there, you need to configure the relay server information (if you are using a self-hosted Relay). This is also where you would set up a permanent password to the server and document the device ID number.
Be sure to check the box for "**Enable remote configuration modification**" when setting up Rustdesk.
### Configure Automatic Login
For Rustdesk specifically, we have to configure XFCE to automatically login via SDDM then immediately lock the computer once it's logged in, so the XFCE session is running, allowing Rustdesk to connect to it.
**Create SDDM Config File**:
```sh
sudo mkdir -p /etc/sddm.conf.d/
sudo nano /etc/sddm.conf.d/autologin.conf
```
```ini title="/etc/sddm.conf.d/autologin.conf"
[Autologin]
User=nicole
Session=xfce.desktop
```
!!! note "Determining Session Strings"
If you're unsure of the correct session string, check what's available by typing `ls /usr/share/xsessions/`. You will be looking for something like `xfce.desktop`
### Configure Lock on Initial Login
At this point, its not the most secure thing to just leave a server logged-in upon boot, so the following steps will instantly lock the server after logging in, allowing the XFCE session to persist so Rustdesk can attach to it for remote management of the server.
!!! warning "Not Functional Yet"
I have tried implementing the below, but it seems to just ignore it and stay logged-in without locking the device. This needs to be troubleshot further.
```sh
mkdir -p ~/.config/autostart
nano ~/.config/autostart/xfce-lock.desktop
```
```ini title="~/.config/autostart/xfce-lock.desktop"
[Desktop Entry]
Type=Application
Exec=xfce4-screensaver-command -l
Hidden=false
NoDisplay=false
X-GNOME-Autostart-enabled=true
Name=Auto Lock
Comment=Lock the screen on login
```
Lastly, test that everything is working by rebooting the server.
```sh
sudo reboot
```
@@ -1,47 +0,0 @@
---
tags:
- UPS
- APC
- Power
---
**Purpose**: When an APC battery backup's battery dies, you can manually replace the cells and 'refurbish' the battery. The following diagram is how you rewire the cells.
!!! warning "Work in Progress"
This document is still being written
## Wiring Diagram
``` mermaid
graph TB
%% Define cells and connections
Cell1["Cell 1<br>Black (Negative) to Black (Negative)"] -.-> AndersonNeg["Anderson Connector Negative<br>(Black)"]
Cell1 -->|"Red (Positive) to Black (Negative)"| Cell2["Cell 2<br>Red (Positive) to Black (Negative)"]
Cell2 -->|"Red (Positive) to Black (Negative)"| Cell3["Cell 3<br>Red (Positive) to Black (Negative)"]
Cell3 -->|"Red (Positive) to Fuse"| Fuse["30A Fuse"]
Fuse -->|"Red (Positive) to Black (Negative)"| Cell4["Cell 4<br>Red (Positive) to Black (Negative)"]
Cell4 -->|"Red (Positive) to Anderson Connector Positive"| AndersonPos["Anderson Connector Positive<br>(Red)"]
%% Define styles
classDef battery fill:#f2f2f2,stroke:#000,stroke-width:2px;
class Cell1,Cell2,Cell3,Cell4 battery;
classDef fuse fill:#ffcc00,stroke:#000,stroke-width:2px;
class Fuse fuse;
classDef anderson fill:#00ccff,stroke:#000,stroke-width:2px;
class AndersonPos,AndersonNeg anderson;
classDef positive fill:#ff0000,stroke:#000,stroke-width:2px;
class AndersonPos positive;
classDef negative fill:#000000,stroke:#fff,stroke-width:2px;
class AndersonNeg negative;
%% Define line colors for clarity
linkStyle 0 stroke:#000,stroke-width:2px;
linkStyle 1 stroke:#ff0000,stroke-width:2px;
linkStyle 2 stroke:#ff0000,stroke-width:2px;
linkStyle 3 stroke:#ff0000,stroke-width:2px;
linkStyle 4 stroke:#ff0000,stroke-width:2px;
linkStyle 5 stroke:#ff0000,stroke-width:2px;
```
@@ -1,13 +0,0 @@
---
tags:
- UPS
- Backup
- Power
---
| **Battery Backup** | **Status** | **Connected Device(s)** | **Estimated Runtime** | **Shutdown Threshold** | **UPS Web Management** |
| :--- | :--- | :--- | :--- | :--- | :---: |
| Outer-Left `#1` | ![](https://status.bunny-lab.io/api/v1/endpoints/battery-backups_outer-left-1-(virt-node-01--10-port-10gbe-network-switch--pfsense-firewall)/uptimes/7d/badge.svg) | - VIRT-NODE-01<br>- 10-Port 10GbE Network Switch<br>- pfSense Firewall | 10 Minutes | 3 Minutes Remaining | [:fontawesome-solid-car-battery: Manage](http://192.168.3.4:3052){ .md-button } |
| Inner-Left `#2` | ![](https://status.bunny-lab.io/api/v1/endpoints/battery-backups_inner-left-2-(bunny-node-02--24-port-1gbe-network-switch)/uptimes/7d/badge.svg) | - BUNNY-NODE-02<br>- 24-Port 1GbE Network Switch | 12 Minutes | 3 Minutes Remaining | [:fontawesome-solid-car-battery: Manage](http://192.168.3.5:3052){ .md-button } |
| Inner-Right `#3` | ![](https://status.bunny-lab.io/api/v1/endpoints/battery-backups_inner-right-3-(moon-storage-01--wireless-ap)/uptimes/7d/badge.svg) | - MOON-STORAGE-01<br>- Wireless AP | 16 Minutes | 3 Minutes Remaining | [:fontawesome-solid-car-battery: Manage](http://192.168.3.3:3052){ .md-button } |
| Outer-Right `#4` | ![](https://status.bunny-lab.io/api/v1/endpoints/battery-backups_outer-right-4-(lab-draas-01--lab-pool-01--8-port-1gbe-network-switch--internet-modem--poe-surveillance-cameras)/uptimes/7d/badge.svg) | - LAB-DRAAS-01<br>- LAB-POOL-01<br>- 8-Port 1GbE Network Switch<br>- Internet Modem<br>- PoE Surveillance Cameras | 13 Minutes | 3 Minutes Remaining | [:fontawesome-solid-car-battery: Manage](http://192.168.3.33:3052){ .md-button } |
@@ -1,39 +0,0 @@
---
tags:
- Windows
- VSS
- Backup
---
## Purpose
There are times when you may need to delete shadow copies (Volume Shadow Copies) from a drive, commonly to free up disk space. While this is usually straightforward, you may encounter scenarios where shadow copies cannot be deleted through normal means. The following methods provide ways to forcibly remove all shadow copies from a specific volume.
!!! warning
The examples below will **permanently delete all shadow copies** on the specified drive. The examples use drive `D:` > Adjust the drive letter as needed.
## Method 1: Delete Shadow Copies Using `vssadmin`
The `vssadmin` utility is the standard tool for managing shadow copies. It is typically safe and handles deletions gracefully.
However, some antivirus or endpoint protection software may block its execution due to its similarity to behavior used by ransomware. If `vssadmin` fails, use the `diskshadow` method described below.
```cmd
vssadmin delete shadows /for=D: /all /quiet
```
* `/for=D:` specifies the target volume.
* `/all` removes all shadow copies on that volume.
* `/quiet` suppresses confirmation prompts.
## Method 2: Delete Shadow Copies Using `diskshadow`
`diskshadow` is a more direct and lower-level tool than `vssadmin`. It should be used as a fallback option if `vssadmin` fails or is blocked.
```cmd
diskshadow
set context persistent nowriters
delete shadows volume D:
exit
```
Explanation:
* `set context persistent nowriters` ensures the command does not involve writer components (e.g., for backups), reducing the chance of interference.
* `delete shadows volume D:` removes all persistent shadow copies for volume `D:`.
@@ -1,35 +0,0 @@
---
tags:
- Ansible
- WinRM
- Automation
---
# WinRM (Kerberos)
**Name**: "Kerberos WinRM"
```jsx title="Input Configuration"
fields:
- id: username
type: string
label: Username
- id: password
type: string
label: Password
secret: true
- id: krb_realm
type: string
label: Kerberos Realm (Domain)
required:
- username
- password
- krb_realm
```
```jsx title="Injector Configuration"
extra_vars:
ansible_user: '{{ username }}'
ansible_password: '{{ password }}'
ansible_winrm_transport: kerberos
ansible_winrm_kerberos_realm: '{{ krb_realm }}'
```
@@ -1,40 +0,0 @@
---
sidebar_position: 1
tags:
- Ansible
- Automation
---
# AWX Credential Types
When interacting with devices via Ansible Playbooks, you need to provide the playbook with credentials to connect to the device with. Examples are domain credentials for Windows devices, and local sudo user credentials for Linux.
## Windows-based Credentials
### NTLM
NTLM-based authentication is not exactly the most secure method of remotely running playbooks on Windows devices, but it is still encrypted using SSL certificates created by the device itself when provisioned correctly to enable WinRM functionality.
```jsx title="(NTLM) nicole.rappe@MOONGATE.LOCAL"
Credential Type: Machine
Username: nicole.rappe@MOONGATE.LOCAL
Password: <Encrypted>
Privilege Escalation Method: runas
Privilege Escalation Username: nicole.rappe@MOONGATE.LOCAL
```
### Kerberos
Kerberos-based authentication is generally considered the most secure method of authentication with Windows devices, but can be trickier to set up since it requires additional setup inside of AWX in the cluster for it to function properly. At this time, there is no working Kerberos documentation.
```jsx title="(Kerberos WinRM) nicole.rappe"
Credential Type: Kerberos WinRM
Username: nicole.rappe
Password: <Encrypted>
Kerberos Realm (Domain): MOONGATE.LOCAL
```
## Linux-based Credentials
```jsx title="(LINUX) nicole"
Credential Type: Machine
Username: nicole
Password: <Encrypted>
Privilege Escalation Method: sudo
Privilege Escalation Username: root
```
:::note
`WinRM / Kerberos` based credentials do not currently work as-expected. At this time, use either `Linux` or `NTLM` based credentials.
:::
@@ -1,41 +0,0 @@
---
tags:
- Ansible
- Automation
---
# Host Inventories
When you are deploying playbooks, you target hosts that exist in "Inventories". These inventories consist of a list of hosts and their corresponding IP addresses, as well as any host-specific variables that may be necessary to declare to run the playbook. You can see an example inventory file below.
Keep in mind the "Group Variables" section varies based on your environment. NTLM is considered insecure, but may be necessary when you are interacting with Windows servers that are not domain-joined. Otherwise you want to use Kerberos authentication. This is outlined more in the [AWX Kerberos Implementation](../awx/awx-kerberos-implementation.md#job-template-inventory-examples) documentation.
!!! note "Inventory Data Relationships"
An inventory file consists of hosts, groups, and variables. A host belongs to a group, and a group can have variables configured for it. If you run a playbook / job template against a host, it will assign the variables associated to the group that host belongs to (if any) during runtime.
```ini title="https://git.bunny-lab.io/GitOps/awx.bunny-lab.io/src/branch/main/inventories/homelab.ini"
# Networking
pfsense-example ansible_host=192.168.3.1
# Servers
example01 ansible_host=192.168.3.2
example02 ansible_host=192.168.3.3
example03 ansible_host=example03.domain.com # FQDN is required for Ansible in Windows Domain-Joined Kerberos environments.
example04 ansible_host=example04.domain.com # FQDN is required for Ansible in Windows Domain-Joined Kerberos environments.
# Group Definitions
[linuxServers]
example01
example02
[domainControllers]
example03
example04
[domainControllers:vars]
ansible_connection=winrm
ansible_winrm_kerberos_delegation=false
ansible_port=5986
ansible_winrm_transport=ntlm
ansible_winrm_server_cert_validation=ignore
```
@@ -1,62 +0,0 @@
---
tags:
- Ansible
- Automation
---
!!! warning "DOCUMENT UNDER CONSTRUCTION"
This document is a "scaffold" document. It is missing significant portions of several sections and should not be read with any scrutiny until it is more feature-complete down-the-road. Come back later and I should have added more to this document hopefully by then.
**Purpose**:
This is an indexed list of Ansible Playbooks / Workflows that I have developed to deploy and manage various aspects of my lab environment. The list is not dynamically updated, so it may sometimes be out-of-date.
## Linux Playbooks
### Deployments
Deployment playbooks are meant to be playbooks (or a series of playbooks forming a "Workflow Job Template") that deploy a server or piece of software.
- Authentik
- [1-Authentik-Bootstrapper.yml](https://git.bunny-lab.io/GitOps/awx.bunny-lab.io/src/branch/main/playbooks/Linux/Deployments/Authentik/1-Authentik-Bootstrapper.yml)
- [2-Deploy-Cluster.yml](https://git.bunny-lab.io/GitOps/awx.bunny-lab.io/src/branch/main/playbooks/Linux/Deployments/Authentik/2-Deploy-Cluster.yml)
- [3-Deploy-Authentik.yml](https://git.bunny-lab.io/GitOps/awx.bunny-lab.io/src/branch/main/playbooks/Linux/Deployments/Authentik/3-Deploy-Authentik.yml)
- [Check_Cluster_Nodes.yml](https://git.bunny-lab.io/GitOps/awx.bunny-lab.io/src/branch/main/playbooks/Linux/Deployments/Authentik/Check_Cluster_Nodes.yml)
- [Check_Cluster_Pods.yml](https://git.bunny-lab.io/GitOps/awx.bunny-lab.io/src/branch/main/playbooks/Linux/Deployments/Authentik/Check_Cluster_Pods.yml)
- Immich
- [Full_Deployment.yml](https://git.bunny-lab.io/GitOps/awx.bunny-lab.io/src/branch/main/playbooks/Linux/Deployments/Immich/Full_Deployment.yml)
- Keycloak
- [Deploy-Keycloak.yml](https://git.bunny-lab.io/GitOps/awx.bunny-lab.io/src/branch/main/playbooks/Linux/Deployments/Keycloak/Deploy-Keycloak.yml)
- Portainer
- [Deploy-Portainer.yml](https://git.bunny-lab.io/GitOps/awx.bunny-lab.io/src/branch/main/playbooks/Linux/Deployments/Portainer/Deploy-Portainer.yml)
- PrivacyIDEA
- [privacyIDEA.yml](https://git.bunny-lab.io/GitOps/awx.bunny-lab.io/src/branch/main/playbooks/Linux/Deployments/privacyIDEA.yml)
- Rancher RKE2 Kubernetes Cluster
- [PLACEHOLDER]()
- [PLACEHOLDER]()
- [PLACEHOLDER]()
- [PLACEHOLDER]()
- [PLACEHOLDER]()
### Kerberos
This playbook is designed to be chain-loaded before any playbooks that involve interacting with Active Directory Domain-Joined Windows Devices. It establishes a connection with Active Directory using domain credentials, sets up a keytab file (among other things), and makes it so the execution environment that the subsequent jobs are running in are able to run against windows devices. This ensures the connection is encrypted the entire time the playbooks are running instead of using lower-security authentication methods like NTLM, which don't even always work in most circumstances. You can find more information in the [Kerberos Authentication](../awx/awx-kerberos-implementation.md#kerberos-implementation) section of the AWX documentation. `It does require additional setup prior to running the playbook.`
- [Establish_Kerberos_Connection.yml](https://git.bunny-lab.io/GitOps/awx.bunny-lab.io/src/branch/main/playbooks/Linux/Establish_Kerberos_Connection.yml)
!!! warning "Ansible w/ Kerberos is **not** for beginners"
I advise against jumping into the deep-end with setting up Kerberos authentication for your playbooks until you have made yourself more comfortable with how Kubernetes works, or at the very least, you need to read the linked documentation above very closely to ensure nothing goes wrong during the setup.
### Security
Security playbooks do things like secure devices with additional auditing functionality, login notifications, enforcing SSH certificate-based authentication, things of that sort.
- Install SSH Public Key Authentication
- [PLACEHOLDER]()
- SSH Login Notifications
- [PLACEHOLDER]()
## Windows Playbooks
### Deployments
Deployment playbooks are meant to be playbooks (or a series of playbooks forming a "Workflow Job Template") that deploy a server or piece of software.
- Hyper-V - Deploy GuestVM
- [PLACEHOLDER]()
- Query Active Directory Domain Computers
- [PLACEHOLDER]()
- Install BGInfo
- [PLACEHOLDER]()
@@ -1,22 +0,0 @@
---
tags:
- Ansible
- Automation
---
# AWX Projects
When you want to run playbooks on host devices in your inventory files, you need to host the playbooks in a "Project". Projects can be as simple as a connection to Gitea/Github to store playbooks in a repository.
```jsx title="Ansible Playbooks (Gitea)"
Name: Bunny Lab
Source Control Type: Git
Source Control URL: https://git.bunny-lab.io/GitOps/awx.bunny-lab.io.git
Source Control Credential: Bunny Lab (Gitea)
```
```jsx title="Resources > Credentials > Bunny Lab (Gitea)"
Name: Bunny Lab (Gitea)
Credential Type: Source Control
Username: nicole.rappe
Password: <Encrypted> #If you use MFA on Gitea/Github, use an App Password instead for the project.
```
@@ -1,27 +0,0 @@
---
tags:
- Ansible
- Automation
---
# Templates
Templates are basically pre-constructed groups of devices, playbooks, and credentials that perform a specific kind of task against a predefined group of hosts or device inventory.
```jsx title="Deploy Hyper-V VM"
Name: Deploy Hyper-V VM
Inventory: (NTLM) MOON-HOST-01
Playbook: playbooks/Windows/Hyper-V/Deploy-VM.yml
Credentials: (NTLM) nicole.rappe@MOONGATE.local
Execution Environment: AWX EE (latest)
Project: Ansible Playbooks (Gitea)
Variables:
---
random_number: "{{ lookup('password', '/dev/null chars=digits length=4') }}"
random_letters: "{{ lookup('password', '/dev/null chars=ascii_uppercase length=4') }}"
vm_name: "NEXUS-TEST-{{ random_number }}{{ random_letters }}"
vm_memory: "8589934592" #Measured in Bytes (e.g. 8GB)
vm_storage: "68719476736" #Measured in Bytes (e.g. 64GB)
iso_path: "C:\\ubuntu-22.04-live-server-amd64.iso"
vm_folder: "C:\\Virtual Machines\\{{ vm_name_fact }}"
```