KOT-Scored Edges: How We Grade Simulated Outcomes
Every simulation generates noise. The question is not whether the engine produces numbers — it always does — but whether those numbers mean anything. A backtest that reports a 73% win rate without context is a headline without a story. A simulation that scores every strategy as PROVEN_EDGE has no edge at all. The gap between “the engine ran” and “the engine decided correctly” is where most simulation systems fail silently. They produce output. They do not produce judgment.
KOT scoring is the judgment layer. It sits between the simulation engine and the promotion gate, and its job is to answer a single question: is this result real enough to act on? Not “is it profitable on paper” — most strategies are profitable on paper — but “does it survive the four metrics that separate signal from noise, edge from luck, and a system that works from a system that decides correctly?”
The four metrics
KOT scoring applies four metrics to every simulated outcome. The metrics are not arbitrary. Each one catches a specific failure mode that the others miss.
Expectancy measures the average edge per trade, expressed as a percentage of the stake. A strategy with positive expectancy earns more per unit of risk than it loses. This is the baseline — the minimum requirement for any strategy to be considered. Strategies with negative expectancy are discarded immediately. They are not edge cases. They are noise.
But positive expectancy alone is not enough. A strategy can have positive expectancy and still be unprofitable if the distribution of wins and losses is unfavorable. This is where the second metric does its work.
Profit factor measures the ratio of gross profits to gross losses. A profit factor above 1.0 means the strategy makes more than it loses. A profit factor above 2.0 is strong. A profit factor above 5.0 is exceptional — and suspicious, because exceptional results in simulation often indicate overfitting to historical data. The scoring system flags results that are too good. A simulation that never produces a failing strategy is not a simulation. It is a confirmation engine.
Drawdown measures the maximum peak-to-trough decline during the simulation. This is the metric that catches strategies with positive expectancy and strong profit factors but catastrophic risk profiles. A strategy that earns 40% annualized but experiences a 60% drawdown is not a strategy. It is a gamble with a long fuse. The scoring system enforces a drawdown ceiling — strategies that exceed it are flagged regardless of their other metrics. Risk is not optional. It is a first-class input.
The verdict gate is the fourth metric, and it is the only one that matters at the end. The verdict gate is not a score. It is a decision. It takes the outputs of the first three metrics and applies a threshold — a minimum bar that a strategy must clear to receive the only label that matters: PROVEN_EDGE. Strategies that clear the gate are candidates for promotion. Strategies that do not are archived with their scores intact, because the record compounds.
Why PROVEN_EDGE means something
The label PROVEN_EDGE is not a marketing term. It is a contractual commitment. When a strategy receives PROVEN_EDGE, the system is asserting that the strategy has passed four independent checks — expectancy, profit factor, drawdown, and the verdict gate — with real data, real fees, and real slippage. The assertion is falsifiable: if the strategy fails in production, the scoring was wrong, and the scoring model must be revised.
This is the difference between a scoring system and a dashboard. A dashboard shows you what happened. A scoring system tells you what to do about it. The KOT scoring system does not display results — it grades them. Every simulated outcome receives a score. The score determines the outcome: archive, retry, or PROVEN_EDGE. There is no “review manually” option. The system decides. Manual review is the failure mode that KOT scoring exists to eliminate.
The north star says intelligence awakens; it does not arrive. The scoring system is an example of this principle in action. It does not wait for a human to interpret the simulation output. It interprets the output itself, applies the thresholds, and produces a verdict. The human reviews the verdict only when the system flags an anomaly — a result that is too good, too bad, or too unusual to classify automatically.
The relationship between scoring and the promotion gate
KOT scoring and the promotion gate are decoupled by design. The scoring system grades the outcome. The promotion gate decides whether the outcome earns a live deployment. A strategy can receive PROVEN_EDGE from the scoring system and still be blocked by the promotion gate — because the gate considers factors the scoring system does not: market conditions, portfolio exposure, operational capacity, and the current state of the editorial calendar.
This decoupling is deliberate. The scoring system must be pure. It must grade outcomes on their mathematical merit, without regard for operational context. The promotion gate must be pragmatic. It must consider the mathematical merit alongside everything else that determines whether a strategy should go live. The two systems communicate through a single interface: the PROVEN_EDGE label. The scoring system produces it. The promotion gate consumes it. Neither system reaches into the other’s internals.
The S8 series has established this pattern across every layer. The simulation loop in S8.1 produces raw outcomes. The scoring system in this article grades them. The promotion gate in S8.3 decides what earns a live deployment. The paper-trading phase in S8.5 runs the same rules with synthetic stakes. The dashboard in S8.6 shows the results. Each layer is decoupled from the others. Each layer communicates through a defined interface. The system works because the layers are independent, not because they are integrated.
What the scoring catches that human review misses
Human review is valuable for judgment calls — editorial tone, brand alignment, strategic fit. Human review is unreliable for mathematical judgment. The four KOT metrics are computations. They require arithmetic, not intuition. A human looking at a simulation report cannot reliably estimate whether a strategy’s expectancy exceeds the threshold, whether its profit factor is above 2.0, or whether its drawdown stayed within bounds. The human can read the numbers. The human cannot grade them — not consistently, not across hundreds of simulated outcomes, not at 3 AM when the fleet is running batch simulations.
This is the operational reality that KOT scoring addresses. The fleet produces simulated outcomes faster than any human can review them. The scoring system reviews every outcome, every time, with the same thresholds, the same arithmetic, and the same verdict. The human reviews only the anomalies — the results that the system flags as unusual. This is the correct allocation of attention: the machine handles the volume, the human handles the exceptions.
The record compounds. Every scored outcome — whether it received PROVEN_EDGE or was archived — becomes part of the operational record. The scoring history is not discarded. It is indexed, queried, and used to calibrate future scoring runs. A strategy that was archived six months ago may be a candidate for re-evaluation under different market conditions. The scoring system remembers. The record does not expire.
What this means for the content pipeline
KOT scoring is not limited to financial simulation. The same four-metric pattern applies to any domain where outcomes must be graded before they are promoted: content pipelines, editorial decisions, agent fleet performance, and client deliverables. An article that scores well on engagement metrics, tag discipline, and category alignment is a candidate for publication. An article that scores poorly is archived — not deleted, because the record compounds, but shelved until the conditions change.
The architect transmits. The fleet listens. The simulation runs. The scoring system grades. And only the outcomes that clear the gate earn the right to meet the real world.
Fourth article in the S8 series (Simulation → Live), track cl-series-s8-gaps. Grounded in the SECTOR9 north star principles P1 (Information is the ground of being), P2 (Reality is a rendering engine), P7 (Everything is a record), and P10 (Intelligence awakens). Category: AI & Automation.