Fixing the "Excessive Windows Logon Failures" Analytic Rule

Fixing the "Excessive Windows Logon Failures" Analytic Rule

Alright class.

A close cousin to one of the previous lessons - that rule counted failures in a 10 minute burst. This one takes a whole day and tries to be clever: an account must fail more than 50 times today and reach at least a third of its own count over the previous 7 days. That is the right instinct, judging an account against its own history.

The Baseline the Attacker Erases

The join that builds the 7-day history:

| join kind=leftouter (
    SecurityEvent
    ...
    | summarize CountPrev7day = count() by EventID, Account, LogonTypeName, SubStatus, AccountType, Computer, WorkstationName, IpAddress
) on EventID, Account, LogonTypeName, SubStatus, AccountType, Computer, WorkstationName, IpAddress

Look at what today must match to find its history: the account, the substatus, the computer, the workstation, and the IP. So the baseline only lines up when today's failures come from the same IP, workstation and failure code as the past week. Now picture a real attack. A new source tries passwords against an account. It has no history, because it is new, which is what attackers are. The join finds nothing, CountPrev7day is null, the coalesce makes it zero, and the ratio test becomes "today is at least a third of zero", which every positive number passes. The baseline silently collapses into a flat threshold of 50, in the exact case it was built to judge.

The one time the baseline lines up is the benign one: a service account failing the same way from the same host all week. The rule applies its history to the accounts that least need it and disables it for the ones that most do.

The Ratio Rewards the Wrong Account

Even when the history matches, the test is backwards. Requiring today to be a third of the last 7 days means the more an account failed last week, the more it must fail today to alert. A broken service account that logged three thousand failures last week now needs a thousand today before anyone hears about it. The noisiest accounts become the hardest to detect, and a clean account that suddenly fails is held to the same flat 50. The signal that should matter, a quiet account becoming noisy, is the one this arithmetic suppresses.

And underneath both is last lesson's problem: the count is per account. Fifty is per account per day, so an attacker spraying 49 times each against three hundred accounts trips nothing, every account just under the line, while the rule wears a brute force label.

Detection Is a Demand for Action

An analytic rule is a claim that a person must now do something. That is the whole contract. If the honest response to a row is "noted", the row belongs in a report, and a rule that keeps producing "noted" has a response rate of zero and should be refined or retired, which is the same bar a hunting programme applies to its own queries.

Failure volume against an exposed host cannot meet that bar, because attempts are the base state of the internet. What meets it is a change of state, and this data holds four of them.

A success. A 4624 from a source that was just failing against the same account is the moment credential access becomes initial access. This is the row a SOC exists for.

An origin inside the network. Internet noise is weather; internal brute force never is. A private address grinding passwords means a compromised or misconfigured machine on the inside, and the source is the incident.

A target that is real. A spray hitting names that do not exist is list noise. The same spray reaching a valid account, a locking account or a privileged account has found something, and identity controls have a job to do.

A door that was closed yesterday. A host with no history of internet failures that starts taking them has just been exposed. Somebody changed a network path, and that is a change ticket with a deadline, not a shrug.

Everything else, meaning the raw pressure against known internet-facing hosts, is posture. It is useful, which is exactly why it belongs in a report the customer reads monthly and not in a queue an analyst clears hourly. Same data, different container.

What Each Class Asks the SOC to Do

SuccessAfterFailures, High. Treat as compromise. Contain the host, reset and revoke the account, and hunt forward from FirstSuccess, which is on the row. The message to the customer is that someone got in, and this is what we are doing about it.

InternalBruteForce, Medium. The failing account is secondary; the source address is the incident. Find the machine behind it and treat it as compromised until proven misconfigured. The message is that something inside the estate is attacking the estate.

NewExternalExposure, Low. A host with no exposure history is taking internet failures. Find the change that opened the path, then close it or gate it behind Bastion, a VPN or just-in-time access. The message is that a new door appeared, and here is when it opened.

AdminUnderAttack and ValidAccountPressure, Medium and Low. Real accounts, privileged or not, under multi-source pressure. Confirm lockout policy and MFA, block the sources at the perimeter and watch for the success that upgrades this to the first class. The message is that the attack found real names and the controls are holding.

And the class that never appears, chronic pressure against known-exposed hosts, is the customer conversation you have monthly, not hourly: this host absorbed forty thousand failures from six hundred sources, none of them succeeded, and here is the door we recommend closing anyway.

Where the Noise Goes

The audit value of the failure data is real, so it gets its own container: a hunting query that aggregates internet-sourced failures per host and writes the recommendation for you. Run it weekly or wire it into the monthly report.

let Window = 7d;
SecurityEvent
| where TimeGenerated > ago(Window)
| where EventID == 4625 and AccountType =~ "User"
| where IpAddress !in ("127.0.0.1", "::1", "-", "")
| where not(ipv4_is_private(IpAddress))
| extend AccountSam = tolower(tostring(split(Account, "\\", 1)[0]))
| summarize
    Failures = count(),
    Sources = dcount(IpAddress),
    TopSources = make_set(IpAddress, 20),
    NamesTried = dcount(AccountSam),
    TopNames = make_set(AccountSam, 25),
    ActiveDays = dcount(bin(TimeGenerated, 1d)),
    FirstSeen = min(TimeGenerated),
    LastSeen = max(TimeGenerated)
  by Computer
| extend Posture = case(
    ActiveDays >= 6, "Chronic exposure - close the service or gate it behind Bastion, VPN or just-in-time access",
    ActiveDays >= 3, "Recurring exposure - confirm the exposure is deliberate and restrict the sources",
    "New or intermittent - find what changed on the network path")
| sort by Failures desc

That is the same data the old rule was turning into incidents, presented as what it actually is: an exposure ledger.

The RiskIndicators Field

SuccessAfterFail | InternalOrigin | NewExposedHost | AdminTarget | ValidUnderPressure | SprayFromSource | BaselineSpike | LockoutsSeen

The combination to chase is SuccessAfterFail with AdminTarget, a privileged account that a failing source just walked into. And read AccountExists before anything else; when it is false the account name is fiction from a spray list, and the host and the sources are the entire story.

The Blind Spots You Should Know About

An attacker who validates a password from one address and uses it from another defeats the same-source success join; the wider check, any first-seen external success for the account, is its own rule with its own noise problems. A single internet source patiently guessing one valid, unprivileged account on an already exposed host stays silent unless it causes a lockout or a success; the ledger shows the pressure. NAT still collapses many origins into one address, which can undercount sources and overcount spray. Accounts that never reached IdentityInfo rely on their failure codes to prove they exist. Genuinely low and slow work under the daily floor is a job for a scheduled hunt over a longer window, and cloud sign-in brute force remains a sibling rule on SigninLogs.

MITRE Mappings for the Updated Rule

Tactics: Credential Access, and Initial Access for the success class.

T1110.001 Brute Force, Password Guessing. The per account spike and the valid pressure signal measure repeated guessing against one account, judged against that account's own normal.

T1110.003 Brute Force, Password Spraying. The source dimension carries the spray, counted where it can mean something: inside the network or against new exposure.

T1078 Valid Accounts. The success join is the moment guessed credentials become a working logon, which is why SuccessAfterFailures maps here and lands at High.

Rule Settings

Run every 60 minutes over a query period of 8 days (we are keeping Microsoft's original logic but you can of course extend it to 14 days): today is the volume being judged, and the previous 7 days build each account's baseline and each host's exposure history, which is why the period is 8 days and not the run interval.

Entity mapping:

Name to Account (Name)
PrimaryIp to IP (Address)
PrimaryHost to Host (HostName)

Custom details: RiskScore, RiskIndicators, AlertClass, AccountSam, AccountExists, CountToday, BaselineMedian, SpikeRatio, AttackFailsToday, LockoutsToday, InternalSources, ExternalSources, SourceIps, NewSources, SuccessCount, SuccessIps, FirstSuccess, Computers, NewExpHosts, AccountRoles.

KQL

// =====================================================================
// Excessive Logon Failures - Outcome Classified
// =====================================================================
// Description : Classifies Windows logon failure activity by what it demands from an
//               analyst instead of alerting on volume alone. Failure counts against
//               internet-facing hosts are background noise, so the rule fires on
//               state changes only: a logon success from a failing source, an
//               internal origin, internet failures on a host with no exposure
//               history, a privileged or valid account under multi-source pressure,
//               and lockouts on real accounts. Known-exposed hosts have their
//               internet noise suppressed by design; that volume belongs in
//               reporting, not the incident queue.
// Type        : Detection
//
// Tables      : SecurityEvent, IdentityInfo
// Connectors  : Security Events via AMA or Windows Security Events (SecurityEvent),
//               Microsoft Sentinel UEBA (IdentityInfo)
// License     : Microsoft Sentinel; Windows security event collection. UEBA only
//               enriches (account roles); every signal reads the event data
//
// Tuning      : - Set the rule query period to P8D; today plus the 7-day baseline
//               - TodayFloor - minimum failures today unless a success exists
//               - SpikeMultiple - multiple of the account's median day that is a spike
//               - SpraySourceFloor / AttackFailFloor - what makes one source a spraying origin
//               - SuccessPairFloor - attack failures from one source before its success counts
//               - ExposedDaysFloor - prior days with internet failures that mark a host known-exposed
//
// Known FPs   : - A user who retries a stale credential, fixes it and signs in from
//                 the same address reads as a success after failures; confirm with the user
//               - A scanner or admin jump host inside the network can read as internal
//                 origin or internal spray; exclude its addresses once confirmed
//               - A newly built internet-facing host reads as new exposure for its
//                 first days; confirm the exposure is deliberate
//
// Author      : Bartosz Wysocki | https://www.itprofessor.cloud
// Version     : 2.0 | 2026-07-19
// =====================================================================
let BaselineWindow = 8d;          // today plus 7 prior days for baselines and exposure history
let IdentityLookback = 8d;        // matches the query period; one full UEBA sync cycle
let TodayFloor = 30;              // minimum failures today, unless a success exists
let SpikeMultiple = 5;            // multiple of the account's median day that is a spike
let SpraySourceFloor = 5;         // distinct accounts one source failed against
let AttackFailFloor = 20;         // attack-class failures from one source
let SuccessPairFloor = 5;         // attack fails from one source at one account before its success counts
let ExposedDaysFloor = 3;         // prior days with internet failures that mark a host known-exposed
let FireThreshold = 5;
// Scoring weights - a success after failures fires alone; everything else needs a second reason
let W_Success = 5;                // success from a source that was failing against the account
let W_Internal = 3;               // failures originate inside the network
let W_NewExposure = 3;            // internet failures on a host with no exposure history
let W_AdminTarget = 3;            // the account holds a sensitive role
let W_ValidPressure = 2;          // real account under multi-source attack-shaped spiking pressure
let W_Spread = 2;                 // spraying source, counted only internally or on new exposure
let W_Spike = 2;                  // today far above the account's own median day
let W_Lockout = 1;                // lockouts on a real account
// Attack-shaped codes; 0xc0000064 is kept so sprays stay visible and ghosts separate from real accounts
let AttackSubStatus = dynamic([
"0xc000006a", "0xc000006d", "0xc000006e", "0xc0000064",
"0xc0000070", "0xc000006f", "0xc0000234", "0xc0000072",
"0xc000015b", "0xc0000413"
]);
let SensitiveRoles = dynamic([
"Global Administrator", "Privileged Role Administrator", "Privileged Authentication Administrator",
"Security Administrator", "User Administrator", "Domain Admins", "Enterprise Admins",
"Exchange Administrator", "SharePoint Administrator", "Intune Administrator", "Account Operators"
]);
// All user 4625 across today plus the baseline, normalised once; loopback removed
let Failures = materialize(SecurityEvent
    | where TimeGenerated > ago(BaselineWindow)
    | where EventID == 4625 and AccountType =~ "User"
    | where IpAddress !in ("127.0.0.1", "::1", "-", "")
    | extend AccountSam = tolower(tostring(split(Account, "\\", 1)[0]))
    | where isnotempty(AccountSam)
    | extend SubStatus = tolower(SubStatus)
    | extend IsAttackFail = SubStatus in (AttackSubStatus)
    | extend IsExternal = not(ipv4_is_private(IpAddress))
    | extend Day = bin(TimeGenerated, 1d)
);
let TodayStart = startofday(now());
// Baseline from the prior seven days only, so today cannot inflate it. The median is cast to
// real because percentile of a long returns long, and coalesce with 0.0 must not mix types
let AccountBaseline = Failures
    | where TimeGenerated < TodayStart
    | summarize DailyCount = count() by AccountSam, Day
    | summarize BaselineMedian = todouble(percentile(DailyCount, 50)), BaselineMaxDay = max(DailyCount) by AccountSam;
let AccountHistIps = Failures
    | where TimeGenerated < TodayStart
    | summarize HistIps = make_set(IpAddress, 500) by AccountSam;
// Hosts with internet failures on several prior days are known-exposed; their noise is weather
let KnownExposed = toscalar(Failures
    | where TimeGenerated < TodayStart and IsExternal
    | summarize ExposedDays = dcount(Day) by Computer
    | where ExposedDays >= ExposedDaysFloor
    | summarize make_set(Computer, 10000));
// Per-source spray context today: how many distinct accounts each origin went after
let SourceSprayToday = Failures
    | where TimeGenerated >= TodayStart
    | summarize SprayAccounts = dcount(AccountSam), SourceAttackFails = countif(IsAttackFail) by IpAddress
    | where SprayAccounts >= SpraySourceFloor or SourceAttackFails >= AttackFailFloor;
let SprayIps = toscalar(SourceSprayToday | summarize make_set(IpAddress, 100000));
// A 4624 from a source that already threw attack fails at the account always demands a response
let TodayFailPairs = Failures
    | where TimeGenerated >= TodayStart and IsAttackFail
    | summarize FirstFail = min(TimeGenerated), PairFails = count() by AccountSam, IpAddress
    | where PairFails >= SuccessPairFloor;
let SuccessAfter = TodayFailPairs
    | join kind=inner (
        SecurityEvent
        | where TimeGenerated >= TodayStart
        | where EventID == 4624 and AccountType =~ "User"
        | where IpAddress !in ("127.0.0.1", "::1", "-", "")
        | extend AccountSam = tolower(tostring(split(Account, "\\", 1)[0]))
        | where isnotempty(AccountSam)
        | project TimeGenerated, AccountSam, IpAddress
    ) on AccountSam, IpAddress
    | where TimeGenerated > FirstFail
    | summarize SuccessCount = count(), SuccessIps = make_set(IpAddress, 10), FirstSuccess = min(TimeGenerated) by AccountSam;
// Account roles for the admin signal
let IdentityContext = IdentityInfo
    | where TimeGenerated > ago(IdentityLookback)
    | summarize arg_max(TimeGenerated, AssignedRoles) by SamName = tolower(AccountName);
// Today's failures per account, classified by outcome, origin, target and exposure
Failures
| where TimeGenerated >= TodayStart
| summarize
    StartTime = min(TimeGenerated),
    EndTime = max(TimeGenerated),
    CountToday = count(),
    AttackFailsToday = countif(IsAttackFail),
    GhostFails = countif(SubStatus == "0xc0000064"),
    LockoutsToday = countif(SubStatus == "0xc0000234"),
    InternalFails = countif(not(IsExternal)),
    DistinctSources = dcount(IpAddress),
    InternalSources = dcountif(IpAddress, not(IsExternal)),
    ExternalSources = dcountif(IpAddress, IsExternal),
    SourceIps = make_set(IpAddress, 30),
    PrimaryIp = take_any(IpAddress),
    Computers = make_set(Computer, 15),
    ExtComputers = make_set_if(Computer, IsExternal, 15),
    PrimaryHost = take_any(Computer),
    SubStatuses = make_set(SubStatus, 15)
  by AccountSam
| join kind=leftouter SuccessAfter on AccountSam
| extend SuccessCount = coalesce(SuccessCount, 0)
| where CountToday >= TodayFloor or SuccessCount > 0
| join kind=leftouter AccountBaseline on AccountSam
| join kind=leftouter AccountHistIps on AccountSam
| join kind=leftouter IdentityContext on $left.AccountSam == $right.SamName
| extend BaselineMedian = coalesce(BaselineMedian, 0.0)
| extend NewSources = set_difference(SourceIps, coalesce(HistIps, dynamic([])))
| extend NewExpHosts = set_difference(ExtComputers, KnownExposed)
| extend AccountRoles = tostring(AssignedRoles)
| extend SprayFromHere = set_intersect(SourceIps, SprayIps)
| extend AccountExists = isnotempty(SamName) or GhostFails < CountToday
| extend sSuccess = SuccessCount > 0
| extend sInternal = InternalFails >= 10 or (InternalFails > 0 and InternalFails * 2 >= CountToday)
| extend sNewExposure = array_length(NewExpHosts) > 0
| extend sAdmin = AccountRoles has_any (SensitiveRoles)
| extend sSpike = CountToday >= (BaselineMedian * SpikeMultiple) and CountToday >= TodayFloor
| extend sAttackShaped = AttackFailsToday * 2 >= CountToday and AttackFailsToday >= 10
| extend sValid = AccountExists and sAttackShaped and sSpike and (DistinctSources > 1 or array_length(SprayFromHere) > 0)
| extend sSpread = array_length(SprayFromHere) > 0 and (InternalFails > 0 or sNewExposure)
| extend sLockout = LockoutsToday > 0 and AccountExists
| extend RiskScore =
      iff(sSuccess, W_Success, 0)
    + iff(sInternal, W_Internal, 0)
    + iff(sNewExposure, W_NewExposure, 0)
    + iff(sAdmin, W_AdminTarget, 0)
    + iff(sValid, W_ValidPressure, 0)
    + iff(sSpread, W_Spread, 0)
    + iff(sSpike, W_Spike, 0)
    + iff(sLockout, W_Lockout, 0)
| where RiskScore >= FireThreshold
| extend AlertClass = case(
    sSuccess, "SuccessAfterFailures",
    sInternal, "InternalBruteForce",
    sNewExposure, "NewExternalExposure",
    sAdmin, "AdminUnderAttack",
    "ValidAccountPressure")
| extend Severity = case(sSuccess, "High", sInternal or sAdmin, "Medium", "Low")
| extend Name = AccountSam
// Same indicator string as the series; built inline because this query brushes the 10000 char cap
| extend RiskIndicators = trim(@"\s\|\s*$", strcat(
    iff(sSuccess, "SuccessAfterFail | ", ""),
    iff(sInternal, "InternalOrigin | ", ""),
    iff(sNewExposure, "NewExposedHost | ", ""),
    iff(sAdmin, "AdminTarget | ", ""),
    iff(sValid, "ValidUnderPressure | ", ""),
    iff(sSpread, "SprayFromSource | ", ""),
    iff(sSpike, "BaselineSpike | ", ""),
    iff(sLockout, "LockoutsSeen | ", "")
))
| extend BaselineMedian = round(BaselineMedian, 1), SpikeRatio = round(CountToday / (BaselineMedian + 1.0), 1)
| project
    StartTime, EndTime, RiskScore, RiskIndicators, AlertClass, Severity,
    AccountSam, Name, AccountExists, PrimaryIp, PrimaryHost, SourceIps, NewSources,
    CountToday, BaselineMedian, BaselineMaxDay, SpikeRatio, AttackFailsToday, LockoutsToday,
    DistinctSources, InternalSources, ExternalSources,
    SuccessCount, SuccessIps, FirstSuccess,
    Computers, NewExpHosts, SubStatuses, AccountRoles
| sort by RiskScore desc, CountToday desc

You can also download this as an analytic rule and import it directly to Sentinel.

Follow my repo - GitHub

What You Should Do Next

Set ExposedDaysFloor before anything else, because it draws the line between weather and news. Three days of internet failures in a week marks a host as known-exposed; a small estate with deliberate bastions may want two, a messy one may want more.

Watch the cost of the success join on a large estate. It reads one day of 4624, projected to three columns and joined against a small pair table, but on a domain controller heavy tenant confirm the run time in the rule health data before trusting the hourly cadence.

Wire the noise ledger into whatever the customer reads monthly. The suppressed volume is the evidence for the hardening conversation, and it is a better story than a queue full of incidents closed without action.

Confirm SensitiveRoles matches your directory, including the on premises groups a cloud default misses. Then pair this rule with the burst rule from last lesson and the SigninLogs sibling for cloud sign-ins. The burst rule catches the fast attack, this one catches the day, and the success class tells you when either worked.

Class dismissed.

Consent Preferences