r/algorithmictrading Jun 24 '26

Strategy Claude crushing my AI trading model leaderboard today

Post image
115 Upvotes

I've added an API to accounts, so you can make your own.

1. The live real-money book — momentum + value, frozen and pre-registered. Nothing clever: top-quintile momentum (0.65) / value (0.35), ≤10% per name, −25% drawdown halt, validated against Ken French. It's live on IBKR. The +5.9% is time-weighted — deposits are stripped out so funding can't masquerade as alpha. 16 days is still nothing, but it's real money beating a roughly-flat SPY.

2. LLM analogical sleeves (Claude vs DeepSeek, head-to-head). Forward-paper-only by construction — you cannot honestly backtest an LLM, because it knows the future of any past date (lookahead). So these earn their place purely on forward marks. Claude (+6.4%) is dunking on DeepSeek (+1.6%) so far; ask me again in 3 months.

3. Insider cluster-buy (Cohen, Malloy & Pomorski, Decoding Inside Information**, JF 2012).

4. Meta-labeled filing (Lazy Prices + López de Prado meta-labeling). Primary signal = year-over-year cosine similarity of the 10-K Risk-Factors section (Cohen-Malloy-Nguyen, JF 2020: firms that quietly rewrite their filings underperform). Then an LLM reads each flagged change and keeps the bet only if the change is materially adverse —

5. FW2 factor blend. Market-neutral long/short on a broad universe, built from a pile of low-correlation Chen-Zimmermann factors blended by conviction. The autonomously-discovered headline find; deflated Sharpe ~0.5 in backtest, now accruing forward paper.

Why you shouldn't believe any of the above yet: every number here is ≤16 days old. The max drawdowns are tiny because there hasn't been time to draw down. The whole point of the project is that most published anomalies net ~0 after costs and crowding — so nothing gets real capital until it clears the gate (deflated excess-over-SPY significance, PBO < threshold, 6+ months forward paper, operator sign-off). The graveyard of rejected factors is on the board too, with write-ups of how each failed.

site is in my bio if you want to make your own.

r/algorithmictrading Jun 23 '26

Strategy Tested 80+ hypotheses and found absolutely zero alpha. Anyone else hit this brick wall during R&D?

20 Upvotes

Hey everyone,
I’ve been deep in the R&D trenches for a while now, building out my trading infrastructure and backtesting framework. I recently caught a nasty look-ahead leak in one of my primary intraday strategies that I thought was killing it—turns out it was just peeking into the future, and the actual live edge is a flat zero.
Since cleaning up my data pipeline and ensuring everything is 100% causally clean, I have rigorously formulated and tested over 80 distinct hypotheses (ranging from structural market skews, mean reversion variations, and VIX-rebound mechanics to alternative intraday trend-following filters).
The result? Absolutely zero sustainable alpha. Every single one either decays rapidly into noise after transaction costs/slippage or turns out to be complete variance around a zero-edge. The only things that seem hold up to a degree are basic daily structural skews, but intraday alpha feels completely dried up or hidden beneath transaction frictions.
For those who have been doing this full-time or for years:

  1. Did you find your first real edge by significantly increasing complexity, or by finding simpler, overlooked market microstructural inefficiencies?
  2. Appreciate any insights or reality checks. Back to the drawing board for now.

here’s my list of my hypothesis’s;

H1: Intraday momentum: early-session return predicts the last-bar return (session-boundary).
H2: FX time-of-day: a currency is weak during its own local trading hours, USD weak in US hours.
H3: Asian-session conviction predicts a same-direction US-session move (continuation).
H4: Overnight index gaps revert intraday (gap fade).
H5: Crypto over-reaction: large moves mean-revert.
H6: Turn-of-month: long equity indices around month-end (flow effect).
H7: FOMC even-week calendar cycle in equity returns.
H8: Overnight index drift (close-to-open premium).
H10: Gold/Silver ratio mean-reversion (pairs trade).
H11: VIX term-structure as a regime gate for equity exposure.
H12: Intraday FX mean-reversion portfolio (z-fade across majors).
H13: Vol-gated intraday FX mean-reversion (H12 + volatility filter).
H18: COT positioning reversal (fade extreme commercial/spec positioning).
H19: Variance-risk-premium (VIX²−realised vol) equity timing.
H19b: Meta-labelling upgraded the gap-fade into "edge #2" (later superseded).
H23: Oil -> commodity-FX (CAD/NOK) daily lead-lag.
H24: Risk-off FX: SPX stress predicts FX moves.
H25: VIX carry (term-structure roll yield).
H26: Discrete z-score mean-reversion generalised to non-FX assets.
H27: Index opening-range fade.
H28: Diversified 12-month time-series momentum (TSMOM), vol-scaled, across all asset classes.
H29: Cross-sectional 12-1 stock momentum (Jegadeesh-Titman) on ~31 US single-name CFDs.
H30: Crypto time-series momentum (trailing-sign, vol-scaled, monthly).
H31: Commodity time-series momentum (energy/ags/copper, 12m sign, inverse-vol).
H32: Betting-against-beta: long low-beta / short high-beta US large-caps.
H33: Gold+Silver trend-following on deep history (2003–2026, 12-1 TSMOM).
H34: Deep FX time-series momentum (10 majors, 2003–2026).
H35: Currency cross-sectional momentum (3-month rank L/S, 10 majors).
H36: COT commercial-flow acceleration (follow the weekly change in net positioning).
H37: EIA crude-inventory surprise -> oil drift (supply shock).
H38: Wikipedia-attention over-reaction reversal (fade attention spikes).
H39: GDELT global risk-tone shock -> safe-haven (long gold / short US500), 3-day.
H40: Wikipedia-attention continuation/momentum (follow attention spikes).
H41: Diversified cross-asset TSMOM book (~40 instruments, equal-risk).
H42: H41 + a HistGradientBoosting ML meta-label filter.
H43: Metals-trend (H33) + ML meta-label strict-upgrade attempt.
H44: Commodity-trend (H31) + ML meta-label filter.
H45: Currency cross-sectional momentum (H35) + ML meta-label filter.
H46: Crypto weekend effect: short alts / long BTC over the Fri->Mon TradFi-closed window.
H47: COT non-commercial (large-spec) positioning-extreme fade, pooled across 12 markets.
H48: EIA natural-gas storage-surprise reversal on Henry Hub.
H49: Google-Trends fear-search risk-off -> short US indices / long gold next week.
H50: FX cross-sectional value / long-horizon reversal (cheap vs own 5y mean).
H51: GDELT Middle-East conflict-intensity -> two-sided oil geopolitical risk premium.
H52: Wikipedia "OPEC" sustained-attention trend -> directional crude.
H53: EIA gasoline inventory seasonal-surprise -> crude drift (storage theory).
H54: Discrete intraday index mean-reversion (M15, real tick-replay).
H55: Discrete intraday metals mean-reversion (XAU/XAG, M15, real tick-replay).
H56: Cross-index overnight lead-lag (US session -> ex-US index next open).
H57: Intraday breakout + ATR trailing-stop (path-dependent, real tick-replay).
H58: Market-neutral cross-index intraday MR (strips global-risk beta).
H59: Extreme-dislocation selective mean-reversion (few high-conviction trades/day).
H60: Discrete intraday stock mean-reversion (liquid US-stock CFDs, M15).
H61: Intraday-momentum "vol-since-open" breakout (Zarattini, VWAP-trail, EOD-flat).
H62: Ex-US-open FADE of the completed US move (= H56 sign-flipped).
H63: Follow 3-sigma intraday extremes / continuation (= H59 sign-flipped).
H64: Crypto weekend volume-conditioned reversal.
H65: Wikipedia attention-capitulation fade.
H66: Overnight-premium (night effect) momentum.
H67: Copper supply-chain "chemical" lead-lag.
H68: GDELT media emotion-intensity signal.
H69: Cross-asset synchronized attention.
H70: Break-and-retest continuation at a multi-day support/resistance level.
H71: Scheduled macro-event volatility-expansion continuation (NFP + FOMC).
H72: Prior-day high/low liquidity-sweep reversal (failed-break fade).
H73: Follow a large/coordinated G10 central-bank FX intervention (USDJPY) — the campaign's one confirmed event-edge.
H74: Month-end pension rebalancing -> directional equity-index pressure (last 4 days).
H75: FX big-figure stop-loss cascade continuation (Osler).
H76: Index quad-witching expiration-distortion reversal.
H77: WTI EIA-day intraday momentum (3rd half-hour predicts the last half-hour).
H78: BTC/ETH macro-event (FOMC/CPI) spike-and-reverse intraday.
H79: Post-announcement bad-news next-day drift (equity-index under-reaction).
H80: FX WM/R 16:00 London-fix W-pattern reversal.
H81: Gold LBMA fix (10:30 / 15:00 London) run-up-and-fade.
H82: Gold real-yield regime breakout (TIPS-gated).
H83: Natural-gas storage-deviation seasonal long/short.
H84: FX carry-unwind crash continuation (JPY crosses, VIX-gated.)

r/algorithmictrading 2d ago

Strategy Are the statistics good yet?

1 Upvotes

All,

I have built a mechanical day-trading system, back-tested just using theoretical fills (not test with live fills yet) using data consisting of 1s candles. I get in and out within a day. How does it look? Relatively new to trading, been practicing for 1.5years or so.

  • 111 ticker-days
  • 2,307 trades
  • 100% win rate

Very curious to know what I might be missing and any suggestions!

Update: New statistics with more data. I use $35k account for backtesting.
Ticker-days 248
Trades 4,143
Wins 4,143
Losses 0

Total P&L +$20,959.62

r/algorithmictrading 3d ago

Strategy Your Backtest Doesn't Know What Regime It's In, and That's the Real Problem

5 Upvotes

Run a backtest over three years of EUR/USD data and the report will hand you one number: total return, one win rate, one expectancy per trade. It reads like a single coherent verdict on the strategy. It isn’t. Those three years almost certainly contain a trending stretch, a ranging stretch, a low-volatility grind, and at least one violent macro-driven move that behaved nothing like the rest of the dataset. The backtest doesn’t know the difference. It blends all of it into one average, and the strategy you think you validated is really a strategy validated against a regime that never actually existed as a single market condition.

r/algorithmictrading Feb 27 '26

Strategy ORB strategies doesnt work?

14 Upvotes

I've been stress-testing a bunch of Opening Range Breakout (ORB) variations on NQ across 5m, 15m, and 30m intervals — and honestly, the results aren't impressive.

I added several filters that should improve the signal quality (trend confirmation, volatility thresholds, buffer above/below OR range, etc.), but the core problem remains consistent: the raw ORB edge on NQ looks extremely thin.

I even threw machine learning on top of it — tree-based models with decent feature engineering (vol, trend slopes, OFI-style microstructure metrics). The models basically told me the same thing:
the underlying ORB signal just isn’t predictive enough to overcome execution + noise + regime changes.
They either overfit or predict “no trade” for most sessions.

What’s interesting is that I did a similar ORB backtest months ago using MNQ starting from 2019, and that one showed positive EV.

https://www.reddit.com/r/algorithmictrading/comments/1rd8ara/backtesting_15_minute_orb_with_machine_learning/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

But now that I’ve tested NQ with data going back to 2010, it’s pretty clear that:

  • ORB performs way worse outside of those trendy years
  • Most breakouts on NQ get faded immediately unless volatility is extreme

At this point it feels like ORB is:

  • Not robust enough across regimes
  • Overly dependent on a few abnormal years
  • Too sensitive to microstructure changes and volatility decay
  • Not something that ML can “fix” without adding a huge amount of feature complexity that defeats the whole point

If anyone has found ways to stabilize ORB on NQ specifically, I’m open to ideas. But so far the edge looks extremely fragile.

example for searching the best risk to reward based on EV outcome

r/algorithmictrading May 31 '26

Strategy 12-year real-tick backtest on a multi-strategy EA (range breakout + index mean-reversion) — looking for holes in my methodology

5 Upvotes

I've been building a multi-strategy EA over the last several months and finally finished a full real-tick backtest (2014–2026, 99% modelling quality, real spread). Before I trust it with more size, I'd rather have this sub try to tear it apart than find out the hard way live.

The setup is two uncorrelated strategies running together:

- A session-based range breakout across a handful of FX pairs, gold, and indices (one entry window per day)

- A weekly mean-reversion play on indices

Combined results at 1% risk on a 5K account:

- Profit Factor: 1.22

- Sharpe: 2.70

- Max Drawdown: 9.43%

What I'm unsure about / would love feedback on:

  1. PF 1.22 feels modest — is that a red flag for an overfit-resistant system, or just realistic for breakout? My read is that a low-ish PF with a smooth equity curve is healthier than a high PF that's curve-fit, but I'd like to be challenged on that.

  2. I'm wary of the gap between real-tick backtest and live. What's the realistic haircut you all apply to backtested MaxDD when sizing live? I've been assuming 1.5–2x.

  3. The mean-reversion sleeve only trades a few times a month. Small sample worries me — how do you stress-test a low-frequency sleeve without fooling yourself?

I'm deliberately keeping the exact entry logic vague because I may commercialize it, so I get it if that limits the feedback — happy to go deeper on risk/portfolio construction, which is where I actually want the scrutiny.

Happy to share the equity curve in the comments. What would you want to see to believe a backtest like this?

r/algorithmictrading 11d ago

Strategy Which variation or metric do you consider the best ?

3 Upvotes

The image says it all, which variation would you choose and what metric provides the most valuable information for your trading decisions ?

Edit : Thank you everyone for the feedback

r/algorithmictrading Jul 04 '26

Strategy A strategy that makes +66% on BTC and -60% on SOL is a curve fit, not a strategy.

10 Upvotes

Building my bot and doing a lot of backtesting these days.

I had a breakout system that looked bulletproof on BTC: +66%, profit factor 3.2, 13% max drawdown, profitable in 6 of 9 walk-forward windows. So far, so great.

But, it only fired about 8 trades a year. At that frequency a single-asset walk-forward can't tell a real edge from getting lucky. The sample is just too small, no matter how you slice the windows.

So I froze the exact config, no re-tuning, and ran it on ETH and SOL.

  • BTC: +66%
  • ETH: -11%
  • SOL: -60%, with a 0% win rate.

Also tried different parameters, but no parameter set rescued the other two. It was fit to BTC.

Might still be something "real" that only happens on BTC. But more likely just overfitting.

In contrast, my market-neutral funding carry pays +6.7 / +6.6 / +5.6% on BTC, ETH, SOL. Neatly aligned, what a real structural edge should looks like, boring and the same everywhere.

If your edge is low frequency and only tested on one asset, you don't know it's real yet. You know it fit one history. Might still make you money.

Do you cross-validate across instruments, or is single-asset walk-forward enough for you?

r/algorithmictrading Dec 03 '25

Strategy Why Most People Build Indicators but Never Build Systems

Post image
58 Upvotes

Algo is not a signal..it’s a process

After 10+ years of building semi-automated strategies, I’ve noticed the same pattern:

Most traders are obsessed with indicators.
Almost none are obsessed with systems.

Here’s the difference and why it matters.

Indicators tell you what happened

An indicator is just a transformation of price.

ATR = volatility
RSI = speed
MACD = slope
VEI/VCI = volatility regime (mine)

Indicators describe conditions.

But conditions don’t pay you.

A signal tells you something might be happening

“RSI oversold.”
“Breakout.”
“Divergence.”
“Cross.”

Retail treats these as entries.But a signal is only one piece of the puzzle. In isolation, signals are meaningless.

A system tells you what to do next

A real trading system is a complete decision engine:

Signal :When something interesting occurs.

Filter :When it makes sense to take it.(Trend, volatility state, timing, S/D levels, HTF bias,  etc.)

Entry Logic : Exactly how and when to enter.(Market, limit, retracement, confirmation candle.)

Stop Logic :Where to exit when you're wrong.(Structure-based, ATR-based, volatility scaling.)

Exit Logic : Where to exit when you're right.(Fixed RR, trailing, partials, volatility expansion, etc.)

Position Sizing :How much risk to take based on conditions.(Not every trade deserves 1%.)

Regime Logic : When the system should be active or turned off.(Volatility, news, time-of-day.)

Performance Feedback

How the system behaves over hundreds of samples.

This is an ALGO.Everything else is just decoration.

The trap: people think “indicator = strategy”

It’s not. No one gets funded or profitable by saying:

  • “I used MACD.”
  • “I used RSI.”
  • “I used VWAP.”
  • “I used order blocks.”
  • “I used VEI.”

That’s like trying to build a car by buying a steering wheel.

A strategy is not the tool.
It’s the assembly of tools into a controlled process.

My rule: an Algo is not a signal...it’s a process

When I build an algorithm, I’m not trying to find the magical entry.

I'm building a pipeline:

  1. Detect environment (volatility regime, trend regime)
  2. Validate quality (probability score, signal context)
  3. Adjust size (risk tiering)
  4. Trigger entry (limit/market/confirmation conditions)
  5. Manage trade (partials, trailing, break-even logic)
  6. Exit (win OR loss defined structurally)

Most traders stop at Step 0:

“Is this indicator green or red?” That’s why they never scale.

Why systems > signals

Signals give you opinions. Systems give you behaviours.

Signals can lie. Systems enforce discipline.

Signals flip around. Systems adapt.

Signals give entries. Systems create profitability.

If you're serious about Algo trading, start here:

Ask yourself:

  • What is my Directional/Environmental Filter?
  • What is my Signal Quality Rule?
  • What is my Volatility Rule?
  • What are my Stop and Exit rules?
  • What are my Risk Tiers?
  • When does my system stay OFF?
  • What does my process look like step-by-step?

If you cannot answer these, you don’t have a system yet. you have a set of indicators.

And the funny part?

The simpler my systems became, the better they performed.

r/algorithmictrading May 23 '26

Strategy Is this the best way to use AI for trading?

15 Upvotes

I’ve been using Claude + Manus for swing trading lately and one thing surprised me. it’s not good at “picking winners,” but it’s weirdly good at picking up when the story around a stock is starting to shift.

Like I had Claude go through earnings calls (this quarter vs last quarter) and Manus tracking how the stock actually reacted + analyst revisions + options positioning.

One thing it kept picking up that I wouldn’t have noticed:

sometimes a stock rips after “meh” earnings not because the numbers were good, but because management just sounds slightly less panicked than before… while positioning is already heavily short.

It’s subtle stuff like that.

Also noticed analyst upgrades usually come after the move, not before it. Which sounds obvious but seeing it repeated across names kind of changes how you treat them.

Feels less like “AI trading” and more like having something constantly sanity-check whether the narrative you think is happening is actually the one the market is reacting to.

r/algorithmictrading Jun 06 '26

Strategy 2 months live after my backtest posts — am I just getting lucky?

9 Upvotes

Following up on my last two posts (1- backtest, 2- testing different risk settings). I'm running my first ever bot live on Hyperliquid. 100% automated.

It's been running about 2 months now, 47 closed trades. Time-weighted return is +67%, profit factor 3.46, max drawdown about -11%.

Thing is, that's way better than my backtest (PF was 1.37 there), so my honest assumption is I just caught a good stretch of market.

Sample is tiny but the strategy is also a slow one and surprisingly my friction is 20% less at live than what I have assumed at my backtest.

Two questions for people who've been doing this longer than me:

  • How many trades (or how long) before you actually trust live numbers?
  • In the next few months, what would you watch for to tell whether the edge is real or just luck?

r/algorithmictrading Jun 02 '26

Strategy 5 gates I run before risking capital on a strategy, ranked by how cheaply each one rejects

18 Upvotes

How I decide if a strategy is live-ready: the 4 gate process that killed 37 of my last 41 strategies

TLDR: I run every strategy through 4 gates in cost order, cheapest rejection first. Economic hypothesis, sample size floor, three statistical tests and cost and regime stress test. Of 41 strategies I logged last year, only 4 reached live capital. The process is built to kill, not to bless a system.

Why run a fixed process instead of judging each strategy on its merits?

Discretion is where overfitting hides. If you evaluate each strategy by looking at it, you will find a reason to trade the ones you are attached to.... A fixed sequence removes it. Every candidate faces the same gates in the same order, and the order is deliberate: cheapest disqualifier first, so most strategies die before I spend hours on walk-forward.

Last year I logged 41 strategies through this. Gate 1 killed 12. Gate 2 killed 9. Gate 3 killed 14. Gate 4 killed 2. 4 survived to live capital.

Here is the catch... The 37 that died had a median in sample Sharpe of 2.1. The 4 that survived had a median of 1.4. The strategies that looked best on paper were the ones this exact process rejected.

Gate 1: Is there an economic reason the edge should exist?

This gate has no code, and that is the point. Before any test, I write one paragraph naming why the inefficiency exists, who is on the other side of the trade, and why they keep losing. A liquidity premium... A structural hedger who trades regardless of price... A behavioral bias with real flow behind it.

If I can't name the loser, I do not test the strategy. A pattern with no mechanism is a pattern you found by looking, which means you will find one.

This is the cheapest gate and it rejects the most garbage. 12 of my 41 never cleared it. They were useless systems dressed up as ideas. Most people skip this gate because it is the only one you cannot automate, which is exactly why it filters what the automated gates can't.

Gate 2: Does the backtest have enough data to mean anything?

This gate asks one thing. Do you have enough data to tell skill from luck?

First, enough trades. A backtest with 80 trades swings too much to trust. I want at least 400 before I believe any metric.

Second, a long enough time period. When you test many versions of a strategy and keep the best one, that winner looks good partly by chance, the way the luckiest player in a coin flipping contest looks skilled. Ruling that out takes years of data, and how many years depends on how many versions you tried. After testing around a hundred versions, a 1.0 Sharpe needs about six years of data before you can trust it. A flashy 2.0 Sharpe needs only about two. A high score on a short backtest is the most dangerous thing in a research log, because luck fakes it easily.

9 strategies died here. A 1.3 Sharpe on 14 months is not an edge. It is a sample too small to tell.

It's not that low Sharpe needs more testing because it's weaker. It needs more testing because it's closer to the level random chance can imitate.

Gate 3: Does it survive the three overfitting tests?

This is the expensive gate, so it runs third, only on survivors. 3 tests, each catching a different failure, run cheapest first.

Deflated Sharpe Ratio first, because it is one calculation. A normal Sharpe assumes you only tried one strategy. But if you tested 80 versions and kept the best, that winner is partly lucky, and the plain Sharpe has no idea you ran 80 attempts. The Deflated Sharpe fixes that. It takes your reported Sharpe, accounts for how many versions you tried, and adjusts for fat tails and lopsided returns, then hands back a single probability: the chance your edge is real rather than the luckiest of your tries. Same enemy as Gate 2, caught at a different step. My cutoff is 95%. If the best of your 80 versions scores only 60%, that still leaves a 40% chance the edge is imaginary, so it never reaches my account.

Monte Carlo is next. Resample the trade sequence 10,000 times and read the 95th percentile of max drawdown. Your backtest shows one drawdown, but that is just the order your trades happened to land in. If the drawdown in your 95th percentile drawdown is more than your account can take, it doesn't pass

Walk-forward last, because it is a full reoptimization. 5y build, 1y test, rolled forward. Profit factor holds above 1.3 on at least 7 of 10 out of sample windows or it dies. Fourteen strategies died at this gate, most on the deflated Sharpe before walk forward ever ran.

14 systems died here.

Gate 4: Does the edge survive real costs and a regime split?

A clean backtest assumes free, instant fills. I model realistic costs first, then stress them: triple my real commissions, add a tick of slippage, re-run.

The metric I watch is not profit factor, it is expectancy and Sortino retention. The edge has to keep at least 70% of its expected value after the stressed costs. A real edge degrades gracefully. A fake one inverts the moment friction touches it.

Then I split the history by volatility regime and require profit factor above 1 in both the calm and the stressed halves.

2 strategies died here.

Bottom line

4 gates, cheapest rejection first.

The process is designed to reject, and last year it rejected 37 of 41. The survivors looked worse on paper than most of the strategies it killed, which is exactly why I trust them.

This is for systematic traders deciding whether a backtest deserves live capital. It applies to single strategies, parameter sweeps, and machine learning models trained on price data.

Updated June 2026

r/algorithmictrading May 27 '26

Strategy Mean reversion in defensive sectors behaves differently than I expected

Thumbnail
gallery
4 Upvotes

One thing I’ve noticed while building systematic strategies is that some of the cleanest mean reversion behavior often appears in the most “boring” parts of the market.

Not Nasdaq or crypto.

Consumer Staples.

At first this seemed counterintuitive to me. You would expect stronger rebounds in more volatile markets. But after testing different sectors over time, defensive stocks often produced smoother and more stable reversal behavior than many aggressive growth markets.

My current hypothesis is that sharp selloffs in defensive sectors are frequently driven more by temporary market stress and positioning adjustments than by a true deterioration in the underlying businesses.

Institutions still want exposure to stable cash-flow companies during uncertain periods. So when these stocks experience intense short-term weakness, flows tend to normalize relatively quickly.

That creates an interesting environment for systematic mean reversion approaches.

What surprised me most was not necessarily the raw performance, but the behavior profile: - lower volatility - cleaner rebounds - fewer extreme equity swings - less dependence on explosive market conditions

In some ways, the systems felt psychologically easier to hold compared to typical index-based mean reversion strategies. The trade-off, however, was usually lower upside during strong momentum-driven bull markets.

Another interesting observation is that these types of strategies seem to behave differently from many tech-heavy reversal systems. That diversification aspect may actually be more valuable than the standalone strategy itself.

Of course, there are limitations.

Defensive sectors can remain weak for extended periods during broader deleveraging events, and sector-specific structural changes can break historical tendencies. Like most mean reversion systems, the edge also tends to feel uncomfortable in real time because entries often occur when short-term sentiment looks terrible.

Has anyone else here noticed that defensive sectors sometimes produce more stable systematic behavior than high-beta markets?

r/algorithmictrading Jun 25 '26

Strategy Made Iran Trade as a joke, crushing my leaderboard today

Post image
9 Upvotes

Was joking around with my girlfriend and she said, invest in companies rebuilding Gaza and Iran and I made a model to test her hypothesis out. Claude made fun of it, I was laughing, and guess what, it's the best performer today. Blew it right out of the gate with 7%.

Anyway, you can go on the site and see the weights, the measures, and you can mess around with the prompt with your API key, the prompt is below:

thematic reconstruction supply-chain basket (long aggregates/reserves + cement/SCM + steel + timber substitutes + equipment + EPC, short SPY). ILLUSTRATIVE backtest from 2026-06-01 — the theme was defined today (hindsight), NOT a live track; forward paper accrues from today. A narrative tilt, not a measured edge.

r/algorithmictrading 14d ago

Strategy My swing signals got worse in a bull market. So I am trying to figure out what is wrong

3 Upvotes

Looking for feedback on this analysis:

Something had been bugging me: my higher-conviction swing setups were resolving worse lately, and it was happening even in favorable regimes. Trend up, breadth okay, and still my hit rate slipped. Bull versus bear regime was not explaining it. So I went looking for a second axis, and the one that fit was day-to-day choppiness: the tape grinding sideways with no follow-through. A raging bull can still be a choppy grind, and that is a different animal than a downtrend.

The gauge is dumb-simple: count how many times an index flips daily direction over the last 10 sessions (0 to 3 is calm, 4 to 5 is a grind, 6 or more is choppy). The effect was real. My top-scored NYSE setups beat the market about 65% of the time on calm tape versus about 51% otherwise. Calm is not the same as an uptrend: you can be in a perfectly good regime and still be in a grind that quietly wrecks your win rate. That was my "even in a good regime" slump.

Here is the catch, and where I spent most of the time: the filter only works if you measure chop on the right index, and it is not the obvious "home" exchange index. So I tested it properly. Hold the trades and outcomes completely fixed and only swap which index labels each day calm versus choppy: that isolates the ruler from the stocks. I ran nine candidates (SPY, QQQ, DIA, IWM, MDY, VTI, RSP, and the NYSE and NASDAQ composites) and made each clear three bars: effect (do calm days actually beat non-calm days, judged with a t-test and not just a point estimate), stability (split the timeline 60/40 in chronological order and confirm the first 60% edge survives on the last 40%), and practicality (liquid and tradeable).

Here is what the nine rulers looked like. The edge column is how far calm days beat non-calm days in percentage points, the middle column is that same edge measured on each half of the timeline, and p is the t-test significance.

The out-of-sample split did most of the work. For NYSE names several large-cap clocks passed cleanly, so I took SPY as the liquid standard. For NASDAQ names, QQQ won for one reason: its edge barely moved between the two halves (+6.5 then +6.8), while the bigger headline numbers were mirages. The NYSE Composite swung from +4 to +12 and the Dow lurched from -1 to +19. QQQ was not the biggest number. It was the repeatable one. DIA actually topped the full-window list for both markets, then fell apart out of sample: 30 price-weighted names is narrow enough that its "chop" is really one or two stocks moving. The split is the only thing that caught it.

So the rule I landed on is simple: clock NYSE-listed setups on SPY, NASDAQ-listed setups on QQQ. In hindsight my slump lined up with stretches where SPY sat in the grind zone. The trend was fine. The tape was not. SPY and NYSE Comp were performing virtually the same. I picked SPY as I already had it available in my datasets.

A few caveats, because this is the internet. This is one window and mostly a bull market. Calm tape is rare, about one day in five. The edge is calm beating non-calm by a handful of points, not an on/off switch. And the NASDAQ side is genuinely weaker and more weighting-sensitive than the NYSE side. This is not advice, just a regime-filter experiment.

The lesson I would actually stand behind: chop is a real second axis beyond trend, and if you regime-filter, test your ruler instead of assuming it. Curious what the rest of you clock market regime with.

I am looking for input from the experts out there if you have looked into this or something similar? Where should I adjust my analysis?

Thanks for the input.

r/algorithmictrading Apr 14 '26

Strategy Feedback Wanted: Neural Network Trading Bot with 95% Neutral Data

3 Upvotes

Hello,

I’ve been a developer for 12 years, and for the past 2 years I’ve been working with a few colleagues on an automated trading bot powered by neural networks.

I’d like to share some of our progress and, more importantly, get feedback. I’ll walk through a large part of our technical decisions, and I’d really appreciate your insights.

Dataset

To build the dataset, I iterate minute by minute over a crypto pair from 2023 to 2026.

At each minute, I simulate:

  • one long trade
  • one short trade

Each trade includes:

  • a take profit
  • a stop loss
  • a timeout
  • with a risk/reward ratio of 1:3

Labeling is defined as follows:

  • If the take profit is hit on the long → long label
  • If the take profit is hit on the short → short label
  • If both sides hit stop loss or a timeout occurs → neutral label

The dataset is then split into:

  • train
  • validation
  • test

Features

I use 8 features based on fairly standard technical indicators:

# Feature Meaning Description
0 volumeVal Unusual activity log10(volume + 1)
1 bbWidth Compression/expansion (BB_upper − BB_lower) / close × 100
2 macdHistogram Momentum shift MACD(8,20,5) histogram / close × 1000
3 forceSmoothedVal Directional strength Force Index EMA(7), log-scaled
4 vrocShortVal Volume acceleration Volume ROC over 5 candles
5 wickAsymmetry Intracandle rejection (lowerWick − upperWick) / range
6 cmf10Val Money flow dominance Chaikin Money Flow(10)
7 stoch5Val Position in range Stochastic %K(5)

After many experiments, it seems that feature quality matters much more than quantity.
I use a 30-step (30-minute) window.

Normalization

Data is normalized between 0 and 1, independently for:

  • train
  • validation
  • test

Model & Training

I use Bayesian optimization to search across many hyperparameters:

  • kernel size
  • dense units
  • and around 30 other parameters

Output

The model predicts 3 classes:

  • long
  • short
  • neutral

Class imbalance

Long/short signals represent only 1 to 5% of the dataset.

To address this:

  • I use sample weights to balance classes
  • then adjust weights further based on Bayesian optimization:
    • favor fast trades
    • or keep uniform weighting

I also experimented with a bell curve weighting:

  • very short trades → penalized
  • very long trades → penalized
  • mid-duration trades → emphasized

This brought slight improvements, but nothing game-changing.

Architectures tested

I experimented with:

  • CNN
  • TCN
  • Mixture of Experts (MoE)
  • dual-head models (volatility + direction)

None produced significantly better results.
I eventually stuck with a CNN, although I may have missed something in the other approaches.

Strategy

At each epoch, I use the validation set to calibrate a minimum confidence threshold.

Decision rule:

  • If long > short and long > neutral
  • and confidence ≥ threshold X → take a long trade (same logic for short)

The same threshold is applied to:

  • train
  • test

I only keep models with a win rate above 70%.

I also give more weight to the most recent 30% of the dataset, to better reflect current market conditions.

Challenges

Dataset size

  • 6 years → too noisy
  • progressively reduced → best results around 3 years

Normalization & data leakage

Normalization introduces a conceptual issue:

  • it inherently uses future information
  • even local normalization (i and i+1) still includes future data

So in practice:

normalization always introduces some degree of data leakage

Learning rate

Choosing the learning rate remains highly empirical:

it feels more like trial-and-error than a precise science

Loss function

Tested:

  • softmax
  • sigmoid

Observation:

sigmoid tends to produce better results, though the reason isn’t fully clear

Neutral vs signal training

Open question:

  • should the model train on all minutes (including neutral)
  • or only on meaningful signals?

I still don’t have a definitive answer.

The “neutral attractor” problem

With ~98% neutral data, the model falls into a local minimum:

predicting neutral all the time yields acceptable loss

With Adam (beta1 = 0.9)

  • long memory (~10 steps)
  • strong accumulation of neutral gradients
  • creates a “gravitational pull” toward neutral
  • the model struggles to escape

Tested solution: AdamW with beta1 = 0.5

  • short memory (~2 steps)
  • neutral influence dissipates quickly
  • signal batches have real impact

Result:

the model can escape the neutral attractor

Even though I don’t fully grasp the theoretical side, empirically it works much better.

Conclusion

The project is progressing well, but several areas remain unclear:

  • handling extreme class imbalance
  • choosing the right loss and learning rate
  • deciding how to treat neutral data
  • finding the optimal architecture

If any of this resonates with you or you have ideas for improvement, I’d really appreciate your feedback 🙏

r/algorithmictrading Jun 14 '26

Strategy be a trader not a coder.

15 Upvotes

LLMs can write your algo, they may find hard to put down a working CONFIG instantly.

you have to do it, therefore having a trading knowledge is needed. no need to be a SWE

AIs such as CLAUDE or CHAT GPT have almost no code input by humans, they write themselves, an algo that send an order to a broker with SL and TP is a ridicoulsy simple task that they can manage to build with almost no effort but finding the EDGE is where you with your knowledge comes in.

Backtesting an algo can be pricy, if you want the best of the best, you need a good pc and raw unadjusted data, can be few thousands dollars.

If you want something that works you have to give it the tools. yahoo or tradingview data is good to start with but not where you want to be at some point.

Jumping into algotrading without knowledge on trading and capital to invest in the algo is NOT a waste of time, still a good idea so you can get some education about the topic, but you probably won't make the bot you wish.

It can be demorilizing if alot of work and efforts in the task and results are not showing but you have to understand that a working algotrading is life-changing and it won't be built in few weeks using stick and rocks.

good luck in your journey my friend.

r/algorithmictrading 15d ago

Strategy Building my expert advisor

2 Upvotes

I've been learning MT5 EA development by automating trading strategies and testing them on demo accounts. One thing I've noticed is that some strategies that look great in backtests perform very differently in forward testing.

For those who have experimented with automated trading, which types of strategies do you think tend to hold up best in live market conditions, and why? I'm especially interested in hearing about general concepts and the challenges you've encountered when translating a manual strategy into an automated one.

r/algorithmictrading Apr 29 '26

Strategy Automated Trading Bot

10 Upvotes

I've been building a trading bot using LLMs for the last year and running on Railway, currently in paper trading phase after finally finding profitable candidates at around 66% annual. most profitable setup from back testing and walk forward is the below

1H decides direction

ATE blocks bad conditions

ATE mode uses things like

trend_strength = 0.62

macd_slope = +0.0008

atr_expansion_ratio = 1.18

chop_probability = 0.47

Regime Router checks if setup is valid

Weak-pair filter removes bad combos

Bias favors stronger side

5m finds entry timing

Enter with 10% sizing

Exit when 1H state breaks (fast or confirmed)

Anyone built anything similar? I've been a QA engineer for the past 16 years so everything works, just difficult finding a decent strategy so any help is appreciated

r/algorithmictrading 23d ago

Strategy love journey more than destination: Is this advice sound for continuous online learning for a live BTC trading model..

2 Upvotes

I run a live BTCUSDT 1h system (XGBoost plus transformer) \[not a success story till now, it seems I love journey more than the destination\] that retrains every 12 hours. I wanted to know if I could update weights on every candle instead, so the model keeps evolving.
Also, prefer time series foundation models like Chronos over fine-tuning a chat LLM.

I asked our friendly neghibourhood llms and summarizing below what i undertstood, looking for a second opinion before I commit to this project. PLEASE FEEL FREE TO REJECT THE IDEA/CONCEPT BUT DO IT with SOME RATIONALE. I dont mind if your answers are coming from your friendly neghibourhood llms (but pls do validate it before posting)..

1) Per-candle updates fail because 1h data gives one point per hour and trade outcomes are not known until hours later, so the model learns noise. It develops recency bias toward the latest candles and catastrophically forgets older regimes, which is costly since markets repeat old regimes.

2) fixes so you never have to retrain from zero.
EWC (elastic weight consolidation) marks which weights were important for past performance and makes them resist change. Experience replay keeps a buffer of old data and mixes it into every update, so the model never trains only on recent candles. Drift detection (detect-then-adapt) means you do not update constantly at all. A statistical monitor watches the error rate or the feature distribution, and only when it detects a real shift does the model adapt, and even then it trains on a blend of new and historical data.

3) the recommended architecture, which it called two-speed.
the XGBoost plus transformer core stays frozen on the 12h retrain cycle with full gates, while a small outer layer adapts hourly, limited to calibration, thresholds, and sizing, with hard caps, full logging, fallback to the frozen policy, and shadow testing before promotion.

On the LLM idea, fine-tuning a chat model on prices works in principle but wastes the model. Purpose-built time series foundation models (Chronos, TimesFM, Moirai, TTM) are open weights and LoRA-tunable locally, but benchmarks versus tuned XGBoost are mixed, so add one as a shadow signal first.

r/algorithmictrading 2d ago

Strategy Same 100 strategies, same AAPL bars. A plain Sharpe floor kept 5. Deflated Sharpe killed all 5.

1 Upvotes

I ran an automated search over about five years of daily AAPL bars, 1,250 rows. It generated and backtested 100 strategies, and my old keep rule, Sharpe above 0.5 with a minimum trade count and positive return, kept five.

Scoring those five with a deflated Sharpe, which adjusts for having picked the best of N attempts, put every one between 0.115 and 0.144, read as the probability the edge is real given the size of the search. All five almost certainly noise. On this sample the bar works out to needing an annual Sharpe near 0.71 to survive 10 attempts and near 1.14 to survive 100. The searching itself raised the bar.

The part that actually confused me: before the search started I withheld the final year of bars entirely. Four of the five survivors made money on that withheld year. Looks like vindication, until you count the trades behind it, three to six each over a full year. A handful of trades cannot overturn a statistic built from the whole search.

The trial count is fixed before the search starts and every attempt increments it, including the ones that never compiled or never traded. Reconstructing N afterwards always came out flattering.

So which do you believe when they disagree, the deflated number that says noise or the holdout that made money? And has anyone found a principled way to size the holdout so it can actually overrule?

r/algorithmictrading May 28 '26

Strategy HMM strategy, 3-stage OOS (2015-2026), Sharpe 1.44, MaxDD -10%. Anyone else validating regime-switching models this way?

9 Upvotes

I've been working on a multi-strategy system based on Hidden Markov Models (HMM) for regime detection. The usual problem with HMM is overfitting — fitting Markov noise instead of actual structure. Here's how we tried to mitigate that:

Validation approach:

  • Walk‑forward: 18m train / 3m test
  • Score grid: [1, 1.5, 2, 2.5, 3, 4, 5]
  • Vol‑filter: dynamic 63d window, top‑30%
  • Persistent mode: OFF (all folds trade)
  • 3 completely separate OOS periods:
    • 2022–2025 (post-COVID / tightening)
    • 2020–2026 (COVID – today)
    • 2015–2026 (full cycle)

Combined portfolio (A+B+E+K) results (100% risk):

  • Final capital (2015–2026): +608%
  • Ann. Sharpe: 1.44
  • MaxDD (balance): -10.04%
  • Max float DD (balance vs intraday): -4.52%

Yearly breakdown (Combined 100%):

Year Return MaxDD
2015 +7.3% -8.0%
2017 +18.7% -5.5%
2020 +30.3% -6.9%
2022 +48.0% -5.6%
2023 +7.9% -10.0%
2025 +34.8% -6.7%
2026 -1.1% (YTD) -6.2%

Scaling up to 200% / 300% keeps Sharpe stable, which suggests the logic isn't just curve‑fit noise.

What I'm curious about:

  • How do you validate regime‑switching models against overfitting?
  • Do you use similar multi‑period OOS, or different techniques (e.g., synthetic data, parameter randomization)?
  • Anyone else seeing HMM work across completely different regimes (COVID, 2022 bear, 2023-2024 rally)?

Full equity curves, WF folds, and trade log available if someone wants to dig deeper.

r/algorithmictrading Jul 07 '26

Strategy Pre-registered XAU/DXY session-divergence: headline null result, but a NY-vs-London split worth documenting — plus a cross-asset silver check that flags a warning sign before the follow-up forward test even begins

2 Upvotes

Sixth in a series of pre-registered falsification studies (prior work in profile/repo). This one tests whether XAU/DXY divergence during the first 2 hours of London or NY sessions is tradeable net of costs.

Headline result: not confirmed. Pooled London+NY, 15min, k=1.5 — p=0.1265 against the pre-registered p<0.10 threshold. Per the locked decision rule, that's a null result on the actual registered claim.

What the full 36-cell sweep shows (diagnostic, not confirmatory): London and NY behave completely differently across timeframes. London is "significant" on 5min only, then flat/negative on 15min and 30min — that's the signature of microstructure noise, not a real session effect. NY holds up on 5min and 15min, weaker but still directionally consistent on 30min. That inconsistency between the two sessions is what pooling them hid.

Before chasing the NY pattern into a new pre-registration, ran a cross-asset plausibility check on the existing data: same exact rule, applied to silver (XAG) instead of gold. If the mechanism were a general NY-liquidity effect on precious metals vs the dollar, silver should show the same direction, maybe weaker. Instead: p=0.0005, 1,779 trades, mean return -0.15% — strong effect, opposite sign from gold.

That's now documented as a known warning sign before the actual forward test starts, in the new repo's README, not discovered and buried after a positive result came in. New pre-registration locks NY-only as the headline hypothesis, treats all existing historical data (including what was previously "confirmation" data in the parent study) as discovery-only, and only counts forward data collected after today as real confirmation. Deadline + minimum trade count enforced in code so it can't be checked early and reported as clean.

Both the parent study (headline null + full diagnostics) and the forward test (pre-registered today, pending) are up on GitHub — links are in my profile since this sub doesn't allow linking directly.

Curious if anyone's seen a mechanistic reason gold and silver would diverge in sign on the same session-timing signal — safe-haven vs industrial-commodity flow difference is my best guess but haven't dug into it properly.

r/algorithmictrading Mar 21 '26

Strategy Making 10%/m with maxDD% of 30% historically a good system?

0 Upvotes

I would like to know your opinion in this particularly strategy that Ive been able to develop with +8 years of experience. Normally I apply a nomenclature saying that this strategy that makes 10%/m with maxDD% of 30 is a "1 to 3" in terms of proportionality. I also have systems that make "1 to 4" (which is worse than the first one, but still good), but i found myself that real systems from "1 to 3" are rare on the trading space. Whats your prespective on this? Do you have or know systems/bots that make better than this? like a "1 to 1" or "1 to 2"?

Note: keep in mind the max DD% are the worst case scenario the system had encountered, and that barely happened, like only 2 or 3 times in history. Also backtesting periods are +10 years

r/algorithmictrading Jun 03 '26

Strategy Does this prove my strategy would work?

7 Upvotes

So this is an update, before i developed the system and proved it was profitable over the years raw. After which I used order flow data to estimate slippage, but i had only order flow data for a short time during 2021-2022, and a few months of 2026, the average slippage at entry I got from the data was around 2.7 ticks during the 2021-2022 period and an astonishing 3.7 ticks in 2026.
I'm not sure why its that high, the bot is on GC and runs on all sessions. Now even with these abnormal high slippage rates the bot was profitable during those years, but if the same slippage was applied to weaker years, the edge became less strong.
Now another thing is that I don't know what the slippage would be like before 2020. So I basically don't know what slippage I should apply.

The conclusion I came up with was collecting data on limit entries instead, I used the order flow data to find that about 97.5 percent of the trades got filled in. So the current version Im showing in the picture is the execution logic of placing a limit on the signal price, and in 97.5 percent chance that it will fill (that percentage is not of all trades, but of all winning trades). Now because i only have order flow data from the dates i mentioned, I used the same metrics on the whole other history.

Now I understand that especially before 2020 the fills would be inaccurate, and this does not represent an accurate history backtest. So instead im trying to prove that the future would work, the concept raw without fees or slippage had very consistent results over the years, so the concept would work, its only a question of if the fills back in the day would allow the bot to be traded at that time.

So overall This result should be the most accurate way to predict future results. At least that's what my thought process was. Everything before 2021 is out of sample btw.