We built a 65% improvement. Then we killed it.
Our v2.6 upgrade cut the engine's trade churn by 65% and held positions twelve times longer. It also lost money. This is the story of testing an idea against criteria we wrote down before the test — and what happened when it failed one of them.
The itch
Scroll our public track record back through July and one pattern jumps out: the engine changes its mind constantly. In one 16-day stretch, the careful book opened and closed 81 short positions. Median hold: 4.8 hours. One memecoin got shorted eleven times in twelve days.
That looks broken. Every flip is a spread paid, a decision reversed, a chance to be wrong twice about the same coin. The obvious fix has a name — hysteresis: make it hard to change state. Enter a short on the same strict gates as always, but don't exit the moment the signal grazes the threshold. Hold until it meaningfully reverses.
Every intuition we had said this would help. Which is exactly why it didn't get to touch the live engine.
The bar, written before the test
Ideas that "obviously help" are how trading systems die. So v2.6 went into a private shadow book — same inputs as the live engine, logged tick by tick for 16 days, while we wrote down what it would take to ship before seeing a single result:
2 Net return ≥ the live book, same window, same grading
Both, or no ship. A rule that only reads well isn't a rule; it's a mood.
It crushed the first bar
median hold live 4.8h → shadow 58.9h (12× longer)
By the metric that motivated the whole idea, v2.6 wasn't an improvement — it was a landslide. If we judged engine changes the way most signal services present them, this post would be an announcement.
Then it failed the one that matters
We grade every position on this site one way: the followable stop model — a −4% hard stop, breakeven arming, a trailing exit, real bars, no hindsight. Same rule for the live book and the shadow book, same 16 days:
The mechanism is almost embarrassing in hindsight. Holding through noise means holding through drawdowns — and a position that refuses to exit on a signal wobble keeps marching until it hits the −4% hard stop instead. 61% of the shadow book's positions stopped out, against 22% of the live book's. Hysteresis didn't hold winners longer; it held losers all the way to the floor.
What the failure taught us
The real finding wasn't about hysteresis. It was about the churn itself: 48 of the live book's 81 round-trips finished inside ±1% — scratches. Noise, not losses. The flip-happy behaviour that looks so broken on the record is mostly the engine paying pennies for the option to be flat when it's unsure, while the 18 positions that survived to their stops carried the whole book.
A 65% reduction in a metric that was never actually hurting you is not an improvement. It's a different engine with worse outcomes and a calmer-looking log.
What we shipped
Nothing. The live engine is unchanged. The verdict is published on the track record page, dated, next to every other engine change we've ever made — including this one that didn't happen. The shadow book keeps logging; if a larger sample reverses the verdict, we'll re-run the same review against the same bar and say so.
We know of no other signal service that publishes its failed upgrades with the losing numbers attached. We'd genuinely like that to change — you should be asking every one of them what's in their graveyard.
Why we work this way
Because you can't trust a track record curated by hindsight, and you can't trust an engine tuned by vibes. Criteria first, test second, verdict published either way. It cost us a feature we'd already built and two weeks of shadow compute — and it's the cheapest insurance there is against slowly becoming the thing we built this site to protect you from.
Watch the engine decide in real time
Every call it makes lands on a public, append-only record — the winners, the losers, and the upgrades we refused to ship.
See the live track record →Quant Terminal is research and educational software — not financial advice, and not a recommendation to buy or sell anything. Replay/backtest figures are historical simulations on logged data ("backtested"), not live returns, and past performance does not predict future results. Crypto is extremely volatile and you can lose your entire investment.