← Research Desk
ENGINE NOTES · RADICAL HONESTY

We built a 65% improvement. Then we killed it.

Our v2.6 upgrade cut the engine's trade churn by 65% and held positions twelve times longer. It also lost money. This is the story of testing an idea against criteria we wrote down before the test — and what happened when it failed one of them.

THE DESK · AUG 14, 2026 · 6 MIN READ

The itch

Scroll our public track record back through July and one pattern jumps out: the engine changes its mind constantly. In one 16-day stretch, the careful book opened and closed 81 short positions. Median hold: 4.8 hours. One memecoin got shorted eleven times in twelve days.

That looks broken. Every flip is a spread paid, a decision reversed, a chance to be wrong twice about the same coin. The obvious fix has a name — hysteresis: make it hard to change state. Enter a short on the same strict gates as always, but don't exit the moment the signal grazes the threshold. Hold until it meaningfully reverses.

Every intuition we had said this would help. Which is exactly why it didn't get to touch the live engine.

The bar, written before the test

Ideas that "obviously help" are how trading systems die. So v2.6 went into a private shadow book — same inputs as the live engine, logged tick by tick for 16 days, while we wrote down what it would take to ship before seeing a single result:

Pre-committed criteria · written Jul 20
1  Cut the number of short round-trips by ≥ 30%
2  Net return ≥ the live book, same window, same grading

Both, or no ship. A rule that only reads well isn't a rule; it's a mood.

It crushed the first bar

round-trips   live 81  →  shadow 28   (−65%, bar was −30%)
median hold   live 4.8h  →  shadow 58.9h   (12× longer)

By the metric that motivated the whole idea, v2.6 wasn't an improvement — it was a landslide. If we judged engine changes the way most signal services present them, this post would be an announcement.

Then it failed the one that matters

We grade every position on this site one way: the followable stop model — a −4% hard stop, breakeven arming, a trailing exit, real bars, no hindsight. Same rule for the live book and the shadow book, same 16 days:

net, honest grading   live +5.7%  ·  shadow −2.0%   (−7.8pp)

The mechanism is almost embarrassing in hindsight. Holding through noise means holding through drawdowns — and a position that refuses to exit on a signal wobble keeps marching until it hits the −4% hard stop instead. 61% of the shadow book's positions stopped out, against 22% of the live book's. Hysteresis didn't hold winners longer; it held losers all the way to the floor.

What the failure taught us

The real finding wasn't about hysteresis. It was about the churn itself: 48 of the live book's 81 round-trips finished inside ±1% — scratches. Noise, not losses. The flip-happy behaviour that looks so broken on the record is mostly the engine paying pennies for the option to be flat when it's unsure, while the 18 positions that survived to their stops carried the whole book.

The churn was never the disease. It was the cost of the engine's honesty about its own uncertainty — annoying to look at, cheap to pay, and protective exactly when holding on would have been expensive.

A 65% reduction in a metric that was never actually hurting you is not an improvement. It's a different engine with worse outcomes and a calmer-looking log.

What we shipped

Nothing. The live engine is unchanged. The verdict is published on the track record page, dated, next to every other engine change we've ever made — including this one that didn't happen. The shadow book keeps logging; if a larger sample reverses the verdict, we'll re-run the same review against the same bar and say so.

We know of no other signal service that publishes its failed upgrades with the losing numbers attached. We'd genuinely like that to change — you should be asking every one of them what's in their graveyard.

Why we work this way

Because you can't trust a track record curated by hindsight, and you can't trust an engine tuned by vibes. Criteria first, test second, verdict published either way. It cost us a feature we'd already built and two weeks of shadow compute — and it's the cheapest insurance there is against slowly becoming the thing we built this site to protect you from.

Watch the engine decide in real time

Every call it makes lands on a public, append-only record — the winners, the losers, and the upgrades we refused to ship.

See the live track record →

Quant Terminal is research and educational software — not financial advice, and not a recommendation to buy or sell anything. Replay/backtest figures are historical simulations on logged data ("backtested"), not live returns, and past performance does not predict future results. Crypto is extremely volatile and you can lose your entire investment.