PowerShell · TROUBLESHOOTING

PowerShell Server Health Check: Practical Windows Server Script Guide

Learn how to build a PowerShell health check for CPU, memory, disk space, services, networking and Windows Server operational status.

Practical Runbook Technical Troubleshooting
Practical Runbook 7 Steps Issues → Solutions → Recommendations
Have a question about this runbook? Post your issue to the TechRunbook Community and get help from other IT professionals.
Ask the Community →
! Issue

Learn how to build a PowerShell health check for CPU, memory, disk space, services, networking and Windows Server operational status.

✓ Solution

01 Core System Info

★ Recommendations

06 Certificate Expiration

01

Core System Info

Start with a baseline so any alert has context. Uptime matters — a server that just rebooted looks different from one that's been up for 400 days with the same disk usage.

$computer = $env:COMPUTERNAME
$os = Get-CimInstance Win32_OperatingSystem
$uptime = (Get-Date) - $os.LastBootUpTime

[PSCustomObject]@{
    ComputerName = $computer
    OSVersion    = $os.Caption
    LastBoot     = $os.LastBootUpTime
    UptimeDays   = [math]::Round($uptime.TotalDays, 1)
}
02

Disk Space

The most common cause of a 3am page. Flag anything under a threshold rather than just reporting raw numbers — you want the script to tell you what needs attention, not make you calculate it.

Get-CimInstance Win32_LogicalDisk -Filter "DriveType=3" | ForEach-Object {
    $freePct = [math]::Round(($_.FreeSpace / $_.Size) * 100, 1)
    [PSCustomObject]@{
        Drive        = $_.DeviceID
        FreeGB       = [math]::Round($_.FreeSpace / 1GB, 1)
        SizeGB       = [math]::Round($_.Size / 1GB, 1)
        FreePercent  = $freePct
        Status       = if ($freePct -lt 10) { "CRITICAL" } elseif ($freePct -lt 20) { "WARNING" } else { "OK" }
    }
}
03

Memory and CPU Pressure

Point-in-time CPU readings are noisy; average over a few samples for something you can actually act on.

$cpu = (Get-Counter '\Processor(_Total)\% Processor Time' -SampleInterval 2 -MaxSamples 3).CounterSamples |
    Measure-Object -Property CookedValue -Average

$mem = Get-CimInstance Win32_OperatingSystem
$memUsedPct = [math]::Round((($mem.TotalVisibleMemorySize - $mem.FreePhysicalMemory) / $mem.TotalVisibleMemorySize) * 100, 1)

[PSCustomObject]@{
    CPUAveragePercent = [math]::Round($cpu.Average, 1)
    MemoryUsedPercent = $memUsedPct
}
04

Critical Services

Don't check every service on the box — maintain an explicit list of what actually matters for this server's role. A generic "any stopped service" check produces so much noise that people stop reading the output.

$criticalServices = @("DNS", "DHCPServer", "NTDS", "Netlogon", "W32Time")

foreach ($svc in $criticalServices) {
    $s = Get-Service -Name $svc -ErrorAction SilentlyContinue
    if ($s) {
        [PSCustomObject]@{
            Service = $svc
            Status  = $s.Status
            Flag    = if ($s.Status -ne 'Running') { "CRITICAL" } else { "OK" }
        }
    }
}
05

Recent Critical and Error Events

Pulling the last 24 hours of Error-level events from the System and Application logs catches problems that haven't caused an outage yet but are heading that way.

$since = (Get-Date).AddHours(-24)
Get-WinEvent -FilterHashtable @{LogName='System','Application'; Level=1,2; StartTime=$since} -ErrorAction SilentlyContinue |
    Select-Object TimeCreated, LogName, Id, LevelDisplayName, Message -First 20

Level 1 is Critical, Level 2 is Error. Limiting to the last 20 keeps the output readable — for a real dashboard, group by Event ID and count occurrences instead of listing every instance.

06

Certificate Expiration

Expired certificates cause outages that are entirely preventable with a 30-second check. Worth including if the server hosts anything over HTTPS or LDAPS.

Get-ChildItem Cert:\LocalMachine\My | Where-Object { $_.NotAfter -lt (Get-Date).AddDays(30) } |
    Select-Object Subject, NotAfter, Thumbprint

Common pitfalls

  1. Alerting on every stopped service instead of an explicit critical-services list — creates alert fatigue fast.
  2. Single-sample CPU readings — always misleading; average over a few seconds.
  3. No threshold tiers — "disk is at 15% free" needs a different response than "disk is at 2% free"; don't flatten that into one alert level.
  4. Running as a low-privilege account — several of these cmdlets need local admin; scheduled tasks should run under an account with the right rights.
Was this runbook helpful?