This page explains the method behind the poster, step by step, with examples. The poster makes a claim. This page is the full explanation. A glossary of terms is at the bottom.
THE PROBLEM: WHY PROMPTS FAIL AS A KNOWLEDGE STORE
Most AI agents today are built the same way. The team writes everything the experts know into one long system prompt. Then they connect a model to some tools, and ship it. This works in a demo. It fails in production. Here is why.
A system prompt is plain text. Text has no structure that a machine can check. Ask these four questions about any behavior your agent shows:
1. Which rule produced this behavior? (a prompt has no rule IDs)2. Which test proves it still works? (you cannot unit-test a sentence)3. Which change made it different? (no diff, no history)4. How often does it apply? (no coverage number)
For knowledge stored in a prompt, the answer to all four is: nobody knows. When the agent is wrong, you cannot point to the line that caused it. And when the model is updated, the new model may read the same prompt differently. All of your knowledge changes behavior at once, silently. This is called drift.
Here is another way to see the problem. A prompt is source code, and the model is an interpreter that re-reads that source on every single call. Nothing is checked ahead of time. When the interpreter changes — a model update — the same source can behave differently. Compiled knowledge does not have this problem. That is the whole idea in the name: compile, don't prompt.
The problem is not a bad model. The problem is a skipped step: knowledge engineering — the work of converting what experts know into a form a machine runs the same way every time. Facts that should be three lines of code are sitting in the same paragraph as the few facts that truly need a model.
THE CORE IDEA: THE EXPLICITNESS LADDER
Knowledge can be stored in different forms. Some forms are precise and testable. Some forms are vague. Picture a ladder — or better, a sieve. Knowledge enters at the top, as words: things an expert knows, things written in a runbook. Each rung below is an attempt to compile that knowledge into a more exact form. Whatever a rung can catch, it keeps. Whatever no rung can catch falls through to the bottom — and stays as plain text in the prompt, where the model must interpret it on every call.
1 · Tacit knowledge— knowledge inside an expert's head. Not written anywhere. The expert makes the right call in seconds but cannot fully explain how.
2 · Runbook— knowledge written down as instructions. This is the source: prose, ready to be compiled. A human can follow it, but nothing checks it.
3 · Rule— knowledge as deterministic code. Same input, same output, every time. Each rule has an ID, a reason, and a test.Example: "if the file path contains /node_modules/, close the finding — it is third-party code."
4 · Score— knowledge as simple math. For signals that are real but not yes/no. Weights you can read and audit — not a black box.Example: "output is not encoded: add 0.25 to the risk score."
5 · Guardrail— knowledge as a hard limit in code around the model: which tools it may call, what shape its answer must have. The model cannot break a limit that code enforces.Example: "the model's final answer must match this JSON schema, or it is rejected."
6 · Prose in a prompt— the fall-through. Not a storage choice — a remainder. This is the knowledge that could not be compiled by any rung above, so the model must interpret it, fresh, on every call. The goal is to make this rung as small as possible.
The whole method is one move, repeated: take each piece of knowledge and push it down through the rungs, as far as it honestly compiles. Most knowledge that feels like "AI judgment" is actually a simple pattern — a rule catches it. Some knowledge is a real risk estimate — a score catches it. Only a small remainder falls through every rung. We call that remainder the residue. The residue is the model's job — and the prompt should contain the residue, and nothing else.
To decide where a piece of knowledge belongs, ask these questions in order:
Is it a fixed pattern? (a path, a name, a type)
→ RULE
Is it a risk signal you can give a number to?
→ SCORE
Does it limit what the model is allowed to do?
→ GUARDRAIL
Does it truly require reasoning over evidence?
→ keep in the prompt
None of the above?
→ think harder
THE FIVE STAGES, IN DETAIL
The case study on the poster is a production system that triages SAST findings (static code-analysis alerts) at a large company. The examples below are generic — replace them with your own domain.
1EXTERNALIZEcollect past decisions as data
Before you write any agent code, collect a corpus: a set of past decisions your experts already made, exported as data. You do not need to create this data. It already exists. Every time an expert closes an alert, marks a finding as a false positive, or rejects a duplicate report, a tool records that decision. Most teams never look at these records again. That is the waste this stage fixes.
One recorded decision looks something like this:
# one record from a scanner database ("disposition" = the decision made)
finding_id: 48213
type: Unchecked_Return_Value
file: src/util/logger.c
sink: fprintf
disposition: Not Exploitable
reason: "return value intentionally ignored, logging call"
decided_by: senior analyst decided_on: 2024-11-02
Your source is whatever system records outcomes: the scanner, the SIEM, the ticket tracker, the bug-bounty platform. Export the resolved items together with the reason they were resolved. A few hundred records is enough to start.
Why this stage is cheap: this is not a data-labeling project. The labels already exist — your experts created them while doing their normal job. In the poster's case study, the corpus was 3,993 findings that analysts had marked "Not Exploitable" over several years. The data was sitting in the scanner database the whole time. Nobody had ever looked at it as one dataset.
2MINEturn repeated patterns into rules
Now group the corpus by its most structural fields: file path, finding type, function name, category. Large groups of identical decisions will appear immediately. Each large group is a candidate for a deterministic rule — a small piece of code that makes the same decision automatically, forever.
A rule is not just an if-statement. Every rule carries four parts:
RULEF-002: third-party dependency code
match: file path contains "/node_modules/"action: resolve as Not Exploitable
rationale:"dependency code — we do not own or patch it here"test: finding in "app/node_modules/lodash/x.js" → resolved ✓
override: a project can disable this rule with a recorded reason
The four parts are what make the system auditable. If an analyst disagrees with a decision, they do not argue with the AI. They look up the rule ID that made the decision, read its rationale, and if the rule is wrong, the rule gets fixed — and its test proves the fix. A sentence inside a prompt can never offer this.
Expect a surprise here. In the case study, 12 rules resolved 67% of the entire backlog — before any AI model ran at all. The first three rules covered the biggest groups. This is the highest-value stage of the method, and it is the stage most teams skip. Mine your history before you write model code.
3STRATIFYscore the rest · give ambiguity a name
After the rules run, some findings remain. Most of them still do not need a model. They need a score: a simple risk number built by adding and subtracting fixed weights. Simple math matters here — anyone can read exactly why an item received its number, and anyone can dispute a weight. A machine-learning classifier at this stage would hide the reasoning you just worked to expose.
A worked example:
# scoring one injection finding (0.0 = safe … 1.0 = dangerous)
start with a neutral base 0.40
the output is written without encoding +0.25 → 0.65# routing by score:
score below the band → close automatically (clearly safe)
score above the band → send to a human (clearly risky)
score inside the band → THE DEAD ZONE (the math cannot decide)
The important design choice: do not use a single cut-off point. A single threshold (for example, 0.5) forces a yes/no decision exactly on the items where the score is least certain. Instead, define a band in the middle of the scale and give it a name: the dead zone. The dead zone is the system saying, out loud, "I cannot decide this with math." Everything outside the band resolves automatically, for free. Only the dead-zone items go to the model.
The dead zone makes ambiguity a named, visible output instead of a hidden guess. It is also what makes the model affordable: the model sees only the small fraction of items that genuinely need it.
4CONSTRAINa tightly limited agent, for the residue only
Now, and only now, the model runs — on perhaps 10–20% of the original volume. But not as a chatbot. If you paste a finding into a chat and ask "is this a false positive?", you get a careful paragraph that commits to nothing. That is useless for automation. Instead, the model runs as a decision worker: it may call a small set of tools to gather evidence (read the data flow, read the source code around the finding, check whether a similar finding was decided before), and it must end by calling one final tool that records a decision. The decision must match a fixed schema — a required shape. Free-text verdicts are rejected.
# the only accepted final answer — anything else is rejected by code
{
"verdict": "not_exploitable" | "exploitable" | "escalate_to_human",
"confidence": 0.0–1.0, // must be OUTSIDE the dead zone"evidence": ["which tool calls support this"]
}
Around the model, hard limits are enforced by code — never requested politely in the prompt:
Dead-zone rejection. If the model reports a confidence that lands back inside the dead zone, the answer is rejected and the model must continue working. It cannot escape into "maybe." A turn limit. If the model uses too many steps, the item is automatically sent to a human. The agent cannot loop forever. Uncertainty becomes an action. "Send to a human" is a valid verdict. "I am not sure — you should check" is not. Unsure means escalate, explicitly. An injection guard. The code being analyzed may contain text written by an attacker, possibly text that tries to give the model instructions. A guard inspects the input before the model reads it. A full trace. Every step the model takes is written to an append-only log.
The trace is what makes the agent testable. You can replay the same evidence later and check that the same verdict comes back. That means you can build a regression test suite for the agent — and when you switch to a new model, you can measure exactly what changed instead of guessing.
5INSTRUMENT & LOOPmeasure everything · corrections become rules
Measure every phase: how many items each stage resolved, how long the model took, how often it hit the dead zone, and how often a human overrode its verdict. Stamp the model version onto every decision, so that a model upgrade is a measurable event, not a mystery. Write each verdict and its reason back into the tool the analysts already use, so they can see why, not only what.
Then close the loop. This is the most important part of the whole method:
1. An analyst disagrees with a verdict and overrides it.
2. The override is saved as new ground truth — a fresh, expert-made label. Free.
3. The same override happens again on similar items.
4. That repeated pattern is written as a new Stage-2 rule.
5. That whole class of item never reaches the model again.
Every correction pushes knowledge down the ladder. Run the loop for a year, and more of the backlog resolves deterministically than on the day you launched. The system becomes cheaper, more predictable, and more auditable over time — the opposite of a prompt-based agent, which drifts.
QUESTIONS PEOPLE ACTUALLY ASK
Isn't this just feature engineering with extra steps?
The value is the ladder, and the default it sets: every piece of knowledge must prove it needs a model. Only the residue earns one. Most agent projects work in the opposite direction — everything goes into the prompt by default. Naming the five stages is what makes the process repeatable by other people, in other domains, instead of one engineer's private habit.
How long does the mining stage take?
Less time than you expect. The biggest groups become visible as soon as you sort the corpus by structural fields, and the first few rules usually clear the largest share in a day. A full rule set is days of work, not weeks. The corpus does the work — you are confirming patterns that already exist, not inventing them.
What happens when the model is wrong inside the dead zone?
The wrong verdict is visible, because analysts review escalations. The analyst overrides it. The override is saved as ground truth. If the same mistake repeats, it becomes a new deterministic rule, and that class of item never reaches the model again. A wrong verdict is a data point, not a hidden failure. The loop turns mistakes into rules.
How do you defend against prompt injection in the analyzed code?
Two layers. First, a guard inspects the untrusted input before the model reads it, and blocks or flags suspicious content. Second, the tools limit the damage: the agent can only call its registered evidence tools — no network access, no code execution, no file writes. The worst possible outcome is one wrong triage decision, and the override loop exists to catch exactly that.
Why local models? A large cloud model would make the agent stage easy.
Two reasons. Compliance: in a regulated company, source code may not leave the network, so a cloud API is not an option. Architecture: the model only ever sees the residue — a small fraction of the volume, already filtered down to genuinely ambiguous cases, with evidence tools attached. Three stages of work made the model's job small. That is what lets a local model handle it. The method makes local models viable — not the other way around.
How do you know the rules don't hide real vulnerabilities (false negatives)?
Every rule is tested against a held-back part of the corpus before deployment, and carries a measured coverage and error number. In production, telemetry tracks the override rate per rule ID. A rule that starts making mistakes shows up as a rising override rate tied to its ID — visible and fixable. A wrong sentence in a prompt shows up as nothing at all.
THE WORKSHEET
Use this checklist to plan the method for your own domain, before writing code. Progress is saved in your browser only.
0 / 16
1EXTERNALIZE
Collect past expert decisions as a dataset.
2MINE
Turn repeated patterns into deterministic rules.
3STRATIFY
Score what remains. Give ambiguity a name.
4CONSTRAIN
Build a tightly limited agent, for the dead zone only.
5INSTRUMENT & LOOP
Measure everything. Corrections become rules.
One question tells you whether this method fits your problem: do your experts make repeated decisions, and does a system record the outcomes? If yes — you have a corpus. Compile it.
GLOSSARY
SAST
Static Application Security Testing — a tool that scans source code for possible security problems, without running the code.
finding
One alert produced by a scanner. Many findings are not real problems.
false positive
A finding that looks like a problem but is not one.
disposition
The recorded decision made about a finding (for example: "Not Exploitable").
compile
To convert knowledge into a fixed, testable form (a rule or a score) before the system runs. The opposite of asking a model to interpret prose while the system runs.
corpus
A collection of past decisions, exported and used as data.
deterministic
Producing the same output for the same input, every time. Code is deterministic; a model is not.
residue
The small set of items left over after rules and scores have done their work — the only items the model sees.
dead zone
The middle band of the score range where the math cannot decide. Named on purpose, so ambiguity is visible.
guardrail
A limit enforced by code around the model — allowed tools, required output shape, maximum steps.
schema
A required data shape. An answer that does not match the schema is rejected automatically.
trace
An append-only log of every step the agent took while reaching a decision.
ground truth
A correct answer confirmed by a human expert, usable for testing and for new rules.
telemetry
Measurements a system records about itself: counts, timings, rates.
drift
Slow, silent change in behavior over time — for example, when a model update reinterprets an unchanged prompt.