SIEM and XDR in Practice: How Analysts Work the Data Layer
Editorial Note: The incidents in this article are composite examples and do not refer to real organizations. Vendor names are used to illustrate technical structures and capabilities, not to recommend products. Capability descriptions reflect public information available from late 2025 into early 2026 and may change with product versions. This article has no sponsorship or financial relationship with any vendor mentioned.
From Alert to Decision: How the SIEM and XDR Layer Actually Works
Two Incidents, One Layer
The first one happened at a mid-sized company whose estate was mostly Microsoft 365 and Azure. The attacker arrived with valid credentials, and every step of the login path looked legitimate: a plausible location, a known client, a clean MFA pass. The EDR reported nothing, because no malicious process was ever created. The SIEM, meanwhile, held forty-some alerts, none of which cleared the bar for an incident on its own: a PowerShell invocation at an odd hour, an unusual enumeration of admin roles, a single outbound connection to an unfamiliar domain.
The first obstacle the analyst hit was naming. The same host appeared as a NetBIOS name in the EDR, as a fully qualified domain name in the SIEM, and as a resource ID in the cloud audit logs, with no automatic mapping between the three. Finding out which business system the machine belonged to meant opening the CMDB; checking for unpatched high-severity vulnerabilities meant opening the vulnerability platform; identifying the owner meant asking IT. Forty-five minutes and five browser tabs later, the timeline still had not assembled itself. By the time it did, privilege escalation was already complete.
The second incident was quieter. A lateral movement campaign ran inside a manufacturer's internal network for two weeks and was finally exposed by an anomalous access pattern on a backup server. The post-incident review did not find badly written rules. The log sources in scope simply never included authentication events from the jump host in question; several machines in the cluster had clock skew of a few minutes, so cross-host event ordering was wrong at the source; and endpoint sensor coverage sat at roughly seventy percent, with the remaining thirty percent presenting as "all clear" in the alerting interface. The line that stood out in the review was blunt: we did not fail to detect it, that data never arrived.
Both failures sit in the same place, between telemetry and human decision. Prophet's 2025 survey of 282 security leaders (predominantly US-based enterprises) describes the load at that layer: organizations receive an average of 960 alerts per day, large enterprises more than 3,000, drawn from roughly 30 different tools; an average alert investigation takes 70 minutes, while an average of 56 minutes elapse before anyone starts; 40% of alerts are never investigated, and 61% of respondents admitted to having dismissed an alert that later proved critical. CrowdStrike's 2026 Global Threat Report (covering 2025 data) compresses the available window: average eCrime breakout time of 29 minutes, a 65% increase in speed from 2024, with a fastest recorded breakout of 27 seconds, and 82% of detections involving no malware. Verizon's 2026 Data Breach Investigations Report adds a distribution correction worth noting: exploitation (31%) overtook credential abuse (13%) as the leading intrusion vector for the first time, so attributing the second incident purely to stolen credentials would misread the actual mix.
The regulatory clock runs through the same layer. Under the EU General Data Protection Regulation, a personal data breach must be notified to the supervisory authority within 72 hours of becoming known, where the breach is subject to notification; the US SEC cybersecurity disclosure rules require a material cybersecurity incident to be disclosed within four business days after the registrant determines that the incident is material. The two clocks therefore do not start at exactly the same point, although both make timely incident discovery and assessment operationally important.
What follows is an attempt to describe what that layer is made of, why it is assembled the way it usually is, and how it actually gets used day to day.
SIEM and XDR: Why Most Teams End Up Running Both
The two acronyms are often discussed as substitutes, but they optimize for different things. XDR derives its depth from high-signal domains: endpoints, identity, email, and cloud workloads. It sees process trees, token usage, and who clicked which link in which message, and it generally carries executable actions with it, such as isolating a host, revoking a session, or disabling an account. SIEM derives its breadth from the infrastructure domain: network devices, server audit logs, database activity, cloud control-plane events, plus the long retention windows that make historical pivoting and forensic reconstruction possible.
Running only one can leave a structural gap, depending on the products and telemetry already in place. An XDR-centric organization may have weaker coverage of server operating systems, network infrastructure, or third-party SaaS if those domains are outside the platform's integrations, while retention may be insufficient for some forensic or audit requirements. An SIEM-centric organization may have broad event coverage but still require additional engineering, correlation, or response tooling to turn those events into high-fidelity investigations and actions. For many enterprises, the practical pattern is therefore a combination of XDR and SIEM capabilities: XDR for high-fidelity detection and response in covered domains, and SIEM for broader correlation, retention, infrastructure telemetry, and reporting.
Analyst firms disagree about the endgame. Some industry forecasts point toward increasing convergence between XDR, SIEM, SOAR, and broader security operations platforms. That is a roadmap judgment rather than a settled fact, and it is not a recommendation for any particular product. The more practical question for a buyer is which side their identity system and cloud footprint already lean toward, and what lock-in costs a more consolidated purchase would carry.
Stack Patterns, and the Four Principles Behind Them
Real deployments cluster into a handful of recognizable shapes. The table below describes structural characteristics and the conditions they tend to fit; it is not a ranking.
| Pattern | Typical composition | Where it tends to fit | Why it is assembled this way |
|---|---|---|---|
| Microsoft-native | Microsoft Defender XDR (Endpoint / Identity / Office 365 / Cloud Apps / Cloud) plus Microsoft Sentinel, presented in a single Defender portal | Organizations where Microsoft 365 and Azure dominate | One identity system and one telemetry standard, so cross-domain join keys align by default and integration effort is lowest |
| Splunk-heavy | Splunk Enterprise / Enterprise Security plus one EDR plus Splunk SOAR plus custom lookup enrichment | Large enterprises and MSSPs with dedicated SOC and SPL engineering capacity | High ceiling on query expressiveness and custom analytics, suited to complex hunting; the cost is ingestion-based pricing plus tuning effort |
| Elastic self-managed or hybrid | Elastic Agent → ECS normalization → Elastic Security (detection engine, rules, ML) plus object-storage cold tier | Strong in-house engineering, cost-sensitive, needs flexible long-horizon queries | Open data model, strong retrieval performance, comparatively controllable licensing; collection coverage and rule tuning are self-owned |
| Endpoint-anchored | CrowdStrike Falcon sensors plus Falcon Next-Gen SIEM plus Fusion SOAR | Organizations already deployed at scale on that sensor | Endpoint telemetry already contains users, processes, and network connections, so less data is exported and re-charged; index-agnostic architecture supports long-retention queries |
| Platform-consolidated | Cortex XSIAM and similar platforms (SIEM + XDR + SOAR + attack surface management) | Mid-to-large enterprises that want fewer consoles and automated grouping | Trades platform consolidation for automatic alert-to-incident grouping, at the price of deeper vendor coupling |
| Cloud-native / data-lake-first | Cribl / Fluent Bit / Logstash pipeline → tiered writes (hot tier into SIEM, cold tier into object storage with federated query) plus a cloud-native SIEM | Very high log volume, multi-cloud, cost-sensitive | Filtering, field trimming, and routing happen in the pipeline, so expensive storage is reserved for data with detection value |
| Managed / MSSP | Managed SIEM / MDR platforms such as Taegis, Exabeam, Rapid7, or Arctic Wolf | Teams without in-house 24×7 coverage | On-call duty and tuning are outsourced, internal staff focus on decisions and response; asset-based pricing makes cost more predictable |
Four principles sit underneath these shapes.
Normalization sets the ceiling on cross-domain correlation. ECS (Elastic Common Schema) and OCSF (Open Cybersecurity Schema Framework) solve the same problem: giving "host," "user," and "network connection" aligned field names across sources. When fields are not aligned, the three host names from the first incident are three separate entities, and any cross-source query degrades into manual comparison. This work is routinely underestimated because it produces no visible alerts.
The cost model determines the architecture. Products billed on ingested volume force the architectural center of gravity into the pipeline: field trimming, event filtering, sampling, tiered routing, otherwise the bill arrives before the detection capability does. Products billed per asset or per platform license tolerate more aggressive full ingestion. Modeling three-year total cost of ownership is more informative than reading a first-year quote, and the lock-in cost of a consolidated purchase deserves equal attention, since migration expenses surface at renewal rather than at signing.
Retention should be tiered by detection value and legal requirement. Hot tiers that support interactive query commonly run 7 to 90 days, with cold data in object storage retrieved on demand. SANS's 2025 State of the SOC Survey (sponsored by Elastic) observed a widespread pattern: organizations store more data than ever and frequently pipe everything into the SIEM without a corresponding management and analysis plan, which ends up reducing visibility rather than increasing it. Retaining everything and being able to find something are two different properties.
The depth-versus-breadth split needs to be decided at design time. Which domains XDR owns for detection and action, which domains SIEM owns for correlation and retention, and who deduplicates at the boundary: if these are not written into the design document, incidents produce the predictable outcome of each system reporting half the picture with nobody owning the merge.
Selection sets the capability ceiling. Actual outcomes depend on how the layer is used day to day, and in most mature teams' post-incident reviews that section appears more often than any product comparison.
The Craft: What Separates a Working Deployment from a Dashboard
Data quality comes before detection rules. Before any rule ships, a source inventory is a precondition: source system, collection method, field completeness, whether timestamps are normalized to UTC, whether NTP is synchronized, and whether outages and packet loss are monitored. Clock skew misaligns the entire timeline, and that misalignment is usually discovered during a review, at which point remediation means re-parsing raw data. Sensor coverage is the other quiet gap: if sensors cover seventy percent of machines, the other thirty percent present as "no anomalies," and that blind spot never surfaces as an alert.
Entity alignment has to be designed ahead of time rather than improvised mid-investigation. The pivot keys should be fixed: host (hostname plus a unique agent ID), user (UPN or objectSID, avoiding display names, which can repeat and be renamed), IP (resolvable through NAT, proxy, and VPN egress, otherwise behavior on a shared egress IP cannot be attributed to an individual), file hash, and process GUID. Joining those keys against the CMDB, the identity directory, and vulnerability data as enrichment, so that asset ownership, business criticality, and known vulnerabilities appear in the alert context itself, is what removes the tab-switching tax. Without enrichment, query cost is transferred to whoever is on shift, and mean time to triage grows accordingly.
Rule governance is a software engineering problem. Writing rules in Sigma, versioning them in Git, validating syntax and field dependencies in CI, and compiling to SPL, KQL, or EQL is already common practice in detection engineering teams. Each rule should carry an owner, an ATT&CK technique mapping, a historical false-positive rate, test cases with both should-hit and should-not-hit sample logs, and a review and retirement date. New rules should run in detection-only or shadow mode for two to four weeks before alerting, because a false-positive rate measured on too few samples carries no statistical meaning: two weeks might yield three hits, a number that proves neither efficacy nor safety. A second high-leverage move is reviewing the alert-volume ranking periodically, since a small number of rules typically account for most of the volume, and tuning those pays off far more than adding rules.
Suppression and allow-listing need expiry dates and approvers. Grouping, deduplication, and suppression are necessary for noise control, but without lifecycle management a system accumulates, over three years, a pile of exceptions nobody dares to delete, and attackers can use exactly those exceptions. Storing expiry dates in rule metadata and routing expired exceptions into a review queue is a cheap control.
Query habits: narrow to wide, unique to general. Start the time window at 15 minutes and expand to 24 hours, 7 days, 30 days; locate first with the most distinctive field (a hash, a rare process name, an anomalous UPN), then expand outward to find siblings; avoid unbounded full scans; save recurring queries as macros or saved searches and share them across the team. These habits do not change detection capability, but they change the duration of a single investigation, and duration is a meaningful variable inside a 29-minute breakout window.
Metrics should be operational. Mean time to triage, mean time to respond, false-positive rate, rule hit distribution, ATT&CK coverage, and alerts per analyst per shift, since past a certain threshold misses become a probability problem rather than an attitude problem. The metric most often skipped, and among the most valuable, is reviewing incidents that were known to have occurred but never generated an alert. That is one of the few reliable ways to find detection gaps, because an alerting system cannot report what it did not see. Gurucul and Cybersecurity Insiders' 2025 Pulse of the AI SOC (739 security leaders) found that 67% of respondents lack visibility into user access behavior and lateral movement, a class of gap that typically surfaces only after the fact.
A short list of recurring anti-patterns: connecting security appliances while leaving identity and business logs unconnected; using the SIEM as a compliance archive that is stored but neither queryable nor detectable; copying community rules wholesale and stacking duplicate alerts; discovering data gaps only after an incident; and reading "no alerts" as "no attack." The last one deserves its own flag, because it appears in executive reporting far more often than it is technically justified.
A meaningful share of these practices is exactly the part generative AI can currently take over, and only that part.
Generative AI at This Layer: What Is Real, and What It Requires
Alert triage and investigation assistance is the area already operating at scale. As of public information from late 2025 into early 2026, Microsoft offers Security Copilot on the Sentinel and Defender side; CrowdStrike offers Charlotte AI on the Falcon side, orchestrable into workflows through Fusion SOAR; Google SecOps integrates Gemini-based capabilities; Elastic provides an AI Assistant; and Palo Alto's Cortex XSIAM provides automated alert grouping and closure. The common effect is compressing the mechanical work of reading an alert, gathering context, and writing a summary, so analysts move from processing every alert to working a pre-filtered set of higher-value items plus spot checks. A self-built path is equally viable: an LLM with RAG over internal runbooks, asset inventories, historical tickets, and prior cases, plus read-only function calling or MCP into the SIEM and threat intelligence platforms, emitting structured output (verdict, confidence, cited evidence, suggested next step), with write operations kept behind human confirmation.
Detection engineering offers the highest leverage and is the most commonly undervalued. Translating natural language into Sigma, SPL, KQL, or EQL while asking the model to explain each rule's intent and plausible evasion paths; deduplicating rules, detecting redundancy, checking field dependencies, and enumerating false-positive scenarios; generating should-hit and should-not-hit regression samples per rule to close the detection-as-code testing loop; and mapping the existing rule set to ATT&CK to find blank tactics. These tasks share a property that makes them safe to adopt early: the output is verifiable by a human quickly, and the cost of an error is contained.
Pipelines and schema work is a safe high-frequency use. Parser and grok generation, field mapping to ECS or OCSF, attribution of sudden log-volume changes, and routing or sampling recommendations. Iteration is fast and verification is direct, which makes this a reasonable first batch of use cases for a team adopting AI.
Enrichment, intelligence, and documentation compound over time. Multi-source IOC enrichment and deduplication summaries, extracting TTPs from sandbox reports and intelligence text and mapping them to ATT&CK, drafting incident response reports and timelines, writing technical notifications for management and customers, generating and versioning runbooks and SOPs, and producing shift handover summaries. CrowdStrike's 2026 report records an 89% increase in attacker AI activity in 2025, including AI-generated phishing and forged content, which makes spending defensive AI capacity against comparable workload volumes easier to justify in a budget conversation.
| Capability | Current maturity | What must be verified before go-live |
|---|---|---|
| Alert triage and investigation summaries | Commercial products and self-built practice both exist | Build an evaluation set from historical real incidents; measure precision, recall, and misclassification; regress continuously |
| Detection rule generation and review | Usable with human gatekeeping | Executability checks per rule, false-positive sample testing, field dependency validation |
| Parsers and schema mapping | Relatively mature | Sample-compare parsed output against raw logs |
| Intelligence enrichment and report drafting | Relatively mature | Fact-checking and source citation to prevent summary drift |
| Automated response actions | Handle with care | Read-only tools may run automatically; isolation, blocking, and deletion require human confirmation plus full audit logging |
The engineering constraints before go-live matter more than model choice. Conclusions must be attached to raw log records or ticket IDs, and unevidenced assertions are unacceptable. Permission tiering belongs in the tooling layer rather than in a process document, with read-only and destructive operations on separate paths. Data governance requires masking, prohibition on uploading credentials and personal identifiers, and confirmation of enterprise training terms or private deployment. Indirect prompt injection is a live threat at this layer rather than a theoretical one: an attacker can place instructions inside log fields, email bodies, or URL content to steer an agent into taking action, and CrowdStrike's 2026 report documented more than 90 organizations targeted with malicious prompt injection against generative AI tooling. All external content therefore has to be treated as untrusted, a constraint that matters most precisely in logs, which are external input by nature.
The gap between adoption and depth of use is worth reading carefully. Gurucul's 2025 survey found 87% of respondents deploying, piloting, or evaluating AI, while only 31% had it in core detection and response workflows; the same survey reported 76% alert fatigue, 73% burnout and staffing shortages, and 64% still operating primarily manually. IBM's 2025 Cost of a Data Breach Report supplies the other side: organizations using security AI and automation extensively averaged roughly USD 1.9 million lower breach costs than those that did not, while the global average breach cost was USD 4.44 million, down 9% year over year. The mean time to identify and contain a breach was 241 days, the lowest in nine years. Read together, the numbers say the benefit is real and that realizing it requires process change rather than procurement.
Two boundaries belong in any rollout plan. Context quality sets the ceiling: without normalized data, asset and identity enrichment, and a case history, the model produces structurally complete but factually unreliable conclusions. Capability maintenance is the second: if judgment is outsourced to a model long-term, a team can find itself without independent analytical capacity when the model errs or the vendor changes. In SANS's 2025 survey, AI/ML tooling had the lowest user satisfaction among tool categories while EDR was the most trusted, a ranking that suggests trust accumulates over time. A low-cost countermeasure is running periodic exercises in which the AI is switched off and the work still has to get done, treated as a capability check rather than a verdict on the tooling.
What Would Have Been Different
Returning to the two incidents, the attribution lands somewhere other than where intuition points. The first lacked context and entity alignment, while its rules were not wrong. The second lacked collection scope, clock synchronization, and sensor coverage, and had little to do with analyst effort. Neither required a product change; both required work at the data layer.
The order of that work is governed by three constraints that cannot be compressed. New rules need two to four weeks in shadow mode before their false-positive rate is statistically meaningful. Sensor deployment follows IT rollout cycles that the security team cannot unilaterally accelerate. Onboarding identity and cloud telemetry typically passes security, IT, and privacy review, whose duration is independent of team skill. Within those constraints, a realistic cadence looks like this: in the first 30 days, complete the source inventory, clock normalization, coverage audit, join key design, and cost and retention tiering; between days 30 and 60, onboard identity and cloud telemetry, complete asset and vulnerability enrichment, stand up 20 to 30 rules starting from high-value ATT&CK techniques in shadow mode, and define alert grouping and suppression policy; between days 60 and 90, move rules-as-code into Git, build the metrics dashboard, introduce AI triage with an evaluation set and a human confirmation loop, and establish the review process for incidents that produced no alert.
That cadence yields a usable baseline. It does not yield mature detection coverage, a tuned behavioral baseline, or analysts who can judge under pressure, all three of which require accumulated incident volume and cannot be compressed. Improving on that 40% of alerts never investigated starts somewhere plainer than better tooling: it starts with every surviving alert carrying context a human can verify.
Further Reading
explore more from our archive:
-
Advanced Streamer Toolkit: Data Funnels, Sponsor Pipelines & Automation Workflows
-
Live Streaming Workflow for Beginners: Streaming Tool Stack & Audio Chain Setup
References
-
NIST — National Institute of Standards and Technology (2006). Guide to Computer Security Log Management (NIST SP 800-92). https://csrc.nist.gov/pubs/sp/800/92/final
-
Nelson, A., Rekhi, S., Scarfone, K., & Souppaya, M. (2025). Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile (NIST SP 800-61 Rev. 3). https://csrc.nist.gov/pubs/sp/800/61/r3/final
-
CrowdStrike (2026). 2026 Global Threat Report. https://www.crowdstrike.com/en-us/global-threat-report/
-
IBM (2025). Cost of a Data Breach Report 2025: The AI Oversight Gap. https://www.ibm.com/think/x-force/2025-cost-of-a-data-breach-navigating-ai
