The essentials

Quick reference

One focused task per row. Jump to the related section for complete, working examples.

UseSyntaxExamples
Check the cluster featureGet-WindowsFeature -Name Failover-Clustering, ` RSAT-Clustering-PowerShellView examples
Load cluster cmdletsImport-Module FailoverClusters; Get-Command -Module FailoverClusters | Select-Object NameView examples
Identify the clusterGet-Cluster -Name 'AppCluster01' | Format-List Name, Domain, QuarantineDuration, QuarantineThresholdView examples
Inventory nodesGet-ClusterNode -Cluster 'AppCluster01' | Select-Object Name, State, NodeWeight, DynamicWeight, DrainStatusView examples
Inventory clustered rolesGet-ClusterGroup -Cluster 'AppCluster01' | Select-Object Name, State, OwnerNode, GroupType, PriorityView examples
Inspect role resourcesGet-ClusterGroup -Cluster 'AppCluster01' -Name 'App Service' | Get-ClusterResourceView examples
Inspect cluster networksGet-ClusterNetwork -Cluster 'AppCluster01' | Select-Object Name, State, Role, Address, AddressMask, MetricView examples
Inspect quorum configurationGet-ClusterQuorum -Cluster 'AppCluster01' | Format-List Cluster, QuorumResource, QuorumTypeView examples
List validation testsTest-Cluster -Cluster 'AppCluster01' -ListView examples
Preview cluster validationTest-Cluster -Cluster 'AppCluster01' -Include ` 'Inventory','Network','System Configuration' -WhatIfView examples
Inspect preferred ownersGet-ClusterGroup -Cluster 'AppCluster01' -Name 'App Service' | Get-ClusterOwnerNodeView examples
Preview draining a nodeSuspend-ClusterNode -Cluster 'AppCluster01' -Name ` 'APPNODE01' -Drain -TargetNode 'APPNODE02' -Wait ` -WhatIfView examples
Move one clustered roleMove-ClusterGroup -Cluster 'AppCluster01' -Name ` 'App Service' -Node 'APPNODE02' -Wait 300View examples
Resume without failbackResume-ClusterNode -Cluster 'AppCluster01' -Name ` 'APPNODE01' -Failback NoFailbackView examples
Read recent cluster eventsGet-WinEvent -FilterHashtable ` @{LogName='Microsoft-Windows-FailoverClustering/Operational'; ` StartTime=(Get-Date).AddHours(-2)}View examples
Collect focused cluster logsGet-ClusterLog -Cluster 'AppCluster01' -Node 'APPNODE01' ` -TimeSpan 30 -UseLocalTime -Destination ` 'C:/ClusterEvidence'View examples

Failover clustering coordinates workload ownership, health detection, recovery, and quorum across Windows Server nodes; it does not make an application interruption-proof by itself. Establish a validated, supported design first, then resolve the exact cluster, node, and role before every change. Run mutations elevated under an approved maintenance plan, preserve quorum and management access, and verify workload health from the client path after any movement.

Step by step

Detailed examples

01

Establish supported scope and administrative prerequisites

Failover Clustering is a Windows Server feature supported on Standard and Datacenter editions, while particular clustered roles can add edition, licensing, hardware, storage, or networking requirements. Install Failover-Clustering on every node and RSAT-Clustering-PowerShell on each management system; feature installation itself normally does not require a restart, but servicing or dependent roles might. Production nodes should run the same Windows Server version and supported hardware should pass cluster validation. Traditional clusters normally use nodes in one Active Directory domain and require rights to create or use the cluster name object; workgroup and AD-detached clusters are specialized designs with different constraints, not drop-in exceptions. Use an elevated Windows PowerShell session with local administrative or delegated cluster rights. The -Cluster parameter performs cluster-aware remote management, but individual commands can impose extra remoting constraints; notably Get-ClusterLog cannot run remotely without CredSSP, so collect it locally or through an approved management path.

Verify feature, module, edition, and elevation context
$identity = [Security.Principal.WindowsIdentity]::GetCurrent()
$principal = [Security.Principal.WindowsPrincipal]::new($identity)
[pscustomobject]@{
    ComputerName = $env:COMPUTERNAME
    Edition = (Get-ComputerInfo -Property WindowsProductName).WindowsProductName
    Elevated = $principal.IsInRole([Security.Principal.WindowsBuiltInRole]::Administrator)
}
Get-WindowsFeature -Name Failover-Clustering, RSAT-Clustering-PowerShell
Import-Module FailoverClusters
Get-Command -Module FailoverClusters | Select-Object Name

Note: Install-WindowsFeature changes every named server and can trigger a pending restart through dependencies or later patching. Check InstallState and RestartNeeded under a maintenance plan instead of installing features from discovery scripts.

Back to quick reference ↑
02

Inventory ownership, resources, and communication paths

Cluster health is a relationship among nodes, roles, resources, networks, and workload-specific dependencies. Query the named cluster rather than relying on the local default, record node and role ownership before maintenance, and drill into resources when a role is degraded. Network Role controls whether a network carries cluster traffic, client traffic, both, or neither; changing it can sever heartbeats or client access. Read-only cmdlets do not prove application health, so pair the inventory with service-specific probes from outside the cluster.

Capture a read-only topology baseline
$clusterName = 'AppCluster01'
Get-Cluster -Name $clusterName | Format-List Name, Domain, QuarantineDuration, QuarantineThreshold
Get-ClusterNode -Cluster $clusterName | Select-Object Name, State, NodeWeight, DynamicWeight, DrainStatus
Get-ClusterGroup -Cluster $clusterName | Select-Object Name, State, OwnerNode, GroupType, Priority
Get-ClusterGroup -Cluster $clusterName -Name 'App Service' | Get-ClusterResource | Select-Object Name, ResourceType, State, OwnerGroup, OwnerNode
Get-ClusterNetwork -Cluster $clusterName | Select-Object Name, State, Role, Address, AddressMask, Metric
Back to quick reference ↑
03

Protect quorum and schedule validation deliberately

Quorum prevents two partitions from accepting writes at the same time; it is not simply a node-count setting. Inspect the existing witness, vote state, site topology, and failure model before any quorum change. Set-ClusterQuorum has no WhatIf parameter, so a witness change needs an approved design, verified credentials or share permissions, and a recovery plan. Test-Cluster can run before or after creation, but tests are not universally passive: storage validation can take eligible storage offline, -Force can take specified clustered disks or pools offline, and some CSV tests temporarily create a local CliTest2 account. Use -List and -WhatIf for planning, then schedule the real supported validation set during a suitable window and review its report.

Review quorum and prepare a non-storage validation scope
$clusterName = 'AppCluster01'
Get-ClusterQuorum -Cluster $clusterName | Format-List Cluster, QuorumResource, QuorumType
Get-ClusterNode -Cluster $clusterName | Select-Object Name, State, NodeWeight, DynamicWeight
Test-Cluster -Cluster $clusterName -List
Test-Cluster -Cluster $clusterName -Include 'Inventory','Network','System Configuration' -WhatIf

Note: WhatIf confirms command intent, not hardware support, workload compatibility, witness reachability, report success, or the runtime impact of the eventual validation. Never add -Force to exploratory storage validation.

Back to quick reference ↑
04

Drain, maintain, and resume one node at a time

Before maintenance, confirm quorum headroom, destination capacity, possible and preferred owners, replication health, and application-specific failover behavior. Suspend-ClusterNode -Drain uses cluster placement logic and supports WhatIf; -ForceDrain may stop workloads that cannot move safely and should not be a routine shortcut. Move-ClusterGroup and Resume-ClusterNode do not support WhatIf, and a successful move can still interrupt connections or expose application faults. Use explicit role and node names, monitor completion, validate the workload externally, and resume with NoFailback unless an approved policy calls for movement back. Operating-system or firmware servicing can require a reboot even though draining and resuming do not.

Preview a node drain and show the approved execution sequence
$clusterName = 'AppCluster01'
$sourceNode = 'APPNODE01'
$targetNode = 'APPNODE02'
$roleName = 'App Service'
Get-ClusterGroup -Cluster $clusterName -Name $roleName | Get-ClusterOwnerNode
Suspend-ClusterNode -Cluster $clusterName -Name $sourceNode -Drain -TargetNode $targetNode -Wait -WhatIf
# After approval: Move-ClusterGroup -Cluster $clusterName -Name $roleName -Node $targetNode -Wait 300
# After maintenance and workload checks: Resume-ClusterNode -Cluster $clusterName -Name $sourceNode -Failback NoFailback

Note: The commented Move-ClusterGroup and Resume-ClusterNode calls deliberately omit WhatIf because those cmdlets do not implement it. Execute them only through a reviewed runbook with current capacity and quorum evidence.

Back to quick reference ↑
05

Collect time-bounded evidence before changing state

Correlate the FailoverClustering operational log, cluster log, resource states, application logs, storage and network telemetry, and exact UTC or local timestamps before restarting resources. Get-ClusterLog creates files and can be expensive across all nodes, so constrain the node and timespan and use a protected destination with adequate space. It uses GMT by default; -UseLocalTime improves correlation only when the incident record states the timezone. Because remote collection has CredSSP restrictions, prefer running it on a cluster node or a secured administrative session rather than weakening authentication. Log collection is evidence, not a repair.

Collect recent events and one focused node log
$since = (Get-Date).AddHours(-2)
Get-WinEvent -FilterHashtable @{
    LogName = 'Microsoft-Windows-FailoverClustering/Operational'
    StartTime = $since
} | Select-Object TimeCreated, Id, LevelDisplayName, Message
Get-ClusterLog -Cluster 'AppCluster01' -Node 'APPNODE01' -TimeSpan 30 -UseLocalTime -Destination 'C:\ClusterEvidence'

Note: The destination is a placeholder. Create and ACL an incident-specific directory first, and avoid collecting secrets or unrelated tenant data into broadly accessible paths.

Back to quick reference ↑

Sources and further reading

References

Authoritative documentation used to verify and expand this cheat sheet.

  1. MicrosoftCreate a failover clusterlearn.microsoft.com
  2. MicrosoftFailoverClusters Modulelearn.microsoft.com
  3. MicrosoftTest-Clusterlearn.microsoft.com
  4. MicrosoftSuspend-ClusterNodelearn.microsoft.com
  5. MicrosoftWhat is a failover cluster quorum witness in Windows Server?learn.microsoft.com

Help us improve

Found a typo or missing example?

Tell us what would make this cheat sheet clearer, more complete, or more useful.

Share feedback