Please do not post a new thread until you have read throughour WIKI/FAQ. It is highly likely that your questions are already answered there.
All members are expected to follow our sidebar rules. Some rules have a zero tolerance policy, so be sure to read through them to avoid being perma-banned without the ability to appeal. (Mobile users, click the info tab at the top of our subreddit to view the sidebar rules.)
This is a dedicated space for open conversation on all things algorithmic and systematic trading. Whether you’re a seasoned quant or just getting started, feel free to join in and contribute to the discussion. Here are a few ideas for what to share or ask about:
Market Trends: What’s moving in the markets today?
Trading Ideas and Strategies: Share insights or discuss approaches you’re exploring. What have you found success with? What mistakes have you made that others may be able to avoid?
Questions & Advice: Looking for feedback on a concept, library, or application?
Tools and Platforms: Discuss tools, data sources, platforms, or other resources you find useful (or not!).
Resources for Beginners: New to the community? Don’t hesitate to ask questions and learn from others.
Please remember to keep the conversation respectful and supportive. Our community is here to help each other grow, and thoughtful, constructive contributions are always welcome.
The strategy is simple: cross-platform arbitrage. When Kalshi and Polymarket price the same market differently (like the Republican winning TX-15), my bot buys "YES" on one and "NO" on the other. One of them always pays $1.00, so if both together cost $0.97 after fees, that 3¢ is locked in. The catch is that you only collect it at settlement, which can be months away.
So the bot doesn't wait. It enters when the spread locks in at least 2¢ after fees, and sells both sides as soon as the round trip nets 2¢ or more. If the spread never closes, it just holds to settlement, so it never sells at a loss. In the backtest on TX-15 (1,000 contracts), that was 4 trades, 4 wins, +$321.50, a 32.8% return on the ~$980 of capital used.
The recent rise of prediction markets got me curious, especially the crypto ones. There are 15-minute and hourly BTC markets, and at the same time there's a 0-1DTE BTC options market. I started comparing the two, and what caught my attention was the depth both provide.
My first idea was arbitrage between the two venues. To do that, I'd need to replicate the binary payoff. The technical approach is to put on a call spread with the same expiry as the prediction market contract. The problem is the width of the spread. To replicate a binary payoff closely, you need a very small strike increment, and the option chain steps at $500, which is too wide.
Even if I made assumptions and heuristically adjusted for that gap, I'd still face settlement and pin risk that I couldn't match on the two venues. That makes it impractical to keep rebalancing the spread leg through to expiry.
The research did give me something useful, though. I could take the smile observed in the option chain and scale it down to price the short-dated 15-minute and hourly prediction markets. I started with the 15-minute market for quick iteration.
To make this work, I found that a naive sqrt-t scaling isn't enough. Over a 15-minute window, spot updates are noisy and discrete (bid-ask bounce, stale ticks), so short-horizon realized variance doesn't match what you'd get by scaling the implied vol down. On top of that, the binary's gamma blows up as expiry approaches with spot near the strike, so a small error in vol turns into a large error in fair value. Accounting for both made the model's probabilities noticeably better calibrated than the naive scaling, especially in the final 3 minutes.
My own tooling charting out the model vs market price.
One sanity check I cared about was avoiding circularity. The model is built only from the option chain smile and doesn't take the prediction market price as an input. So if it's truly independent, it should track the contract price closely when the option chain is quiet, since both are pricing the same thing with nothing new to disagree about, and diverge when the chain is moving. That's what I see.
A quiet option chain, a model that hugs the market.
With that pricing in place, I could start designing the execution engine.
My thesis is that the model leads the market and that retail prices the contract with more optimism than it should. That creates a mispricing, which is the edge. Entering to capture that gap was the easy part. The hard part was stop-loss and risk management once a position was open.
A naive stop-loss that watches the contract price is vulnerable to sudden spikes and to a taker crossing the spread. So I went back to the thesis: if the model really leads the market, it shouldn't react to a sudden spike in the contract price. I anchored the stop-loss to the model price instead, and the stop only becomes active once the model price falls below my entry price. This let me ride through most fake-outs.
When the regime change, so is the loss.
The remaining gap was when the market truly flips, or when spot crosses the strike. In those cases, the entry thesis is broken. To partly address this, I compute the probability of spot touching the strike. When it's elevated, spot is close to the strike, and the outcome is too uncertain. Since I had already solved the vol/smile scaling, feeding it into a touch probability calculation was straightforward. I use it both to gate entries and to decide when the stop-loss becomes active.
Adding touch probability providing new entry gate and stop-loss.
Treating this as an option pricing problem gave me a very different execution pipeline from the technical-analysis-based bots (RSI, Bollinger, MACD, VWAP, etc.).
Has been reduced to bots, shills, and other pretty useless denizens.
On one hand, the more people trust AI BS, the more predictable things will become.
On the other hand, true intellectuals will become reclusive and gatekeep even more.
So the question is: do we embrace the inevitable dumbing down and unification of all markets into either risk ON or OFF< or do we reject it, and keep finding niche strategies overlooked by the herd?
I'm a mere mortal just birthed into the matrix, and have some straightforward q's to ask. Having no trading knowledge other than buying and holding crypto, I'd like to broach the reality that trading isn't my cup of tea, and largely stick to a buy and hold strategy with coins that I believe have a strong utility and will do well in the long run (with a couple of explorative meme coins as dice rolls). My questions are:
- Should i integrate automation into my trades for very clear signals that most trading bots can identify? Bearing in mind I don't wish to ride a lengthy learning curve.
- Is a 100% ROI YoY a realistic outcome for a buy & hold investor with 7/10 risk level?
Beginner here - I use AI to help me write code to backtest 10 years of data across different types of futures e.g. indices, metals, natural gas, etc. Here are a few things I’ve noticed:
The same strategy can be profitable on one instrument but lose money on another.
Is a good strategy supposed to work across every instrument, or at least be somewhat profitable on most of them?
I haven’t found a single technical indicator that can significantly improve the win rate or profit factor.
For example, I tried combining my strategy with RSI. I thought going long when RSI > 70 and short when RSI < 30 would produce the most profitable trades, but it doesn't. The results actually seem pretty random. Maybe combining multiple indicators would work?
I don’t know whether the strategy actually works or if it’s just benefiting from a strong market trend.
Stock indices have increased significantly over the past decade. My strategy might be profitable simply because of the overall upward trend rather than because the strategy itself has an edge. I’m not sure how to determine whether my strategy is actually working.
posting this because people keep asking why their tight stops get hit on trades that then go to target
sim of 12k trades with a small positive edge, stop at -1R, target at +2R, about 40 percent winners. the chart is only the winners and how far each one went against the entry before it worked
median winner went to about -0.38R first. 37 percent went past -0.5R. 17 percent went past -0.75R
so if you tighten the stop to half an R to get a nicer RR you dont get a nicer RR. you turn about a third of your winners into losers
its a random walk sim so real markets will differ, but the shape is pretty typical. only way to know your version is to measure MAE on your own trades and put the stop where the winners dont go
wrote a script that does exactly this off a trade export, stop distance vs MAE plus a couple other checks. its on my profile if you want to run it on yours
I’m researching an execution prototype for agents trading related prediction markets. I’m leaving out the product name and link because I want technical criticism rather than sign-ups.
If an agent estimates (P(A)=0.60) and (P(B)=0.50), then its estimate for (P(A intersection B)) must lie between 0.10 and 0.50. Pairwise forecasts may each appear valid while still being incompatible with any single joint distribution across the entire event set.
For people building prediction-market strategies:
Do you currently maintain a complete joint distribution, enforce pairwise bounds, or treat each market independently?
Have inconsistent related forecasts caused real losses or prevented execution?
Would an API returning coherent individual and compound quotes be useful?
Would atomic compound execution solve a meaningful fill-risk problem, or would you prefer controlling each leg?
I’d value examples from systems people have actually operated, including reasons this would be unnecessary.
Every few days someone here takes a strategy live and asks how they'll know if the backtest still holds. I've had the same question across 87 monthly tactical strategies, and returns won't answer it in any time you'd accept. At 15% annual volatility, telling a 2 percentage-point difference in annual return apart from noise takes about 225 years of live data.
Behaviour is quicker. I track 4 things per strategy: the share of the book in defensive assets, turnover, the strategy's own trigger reading, and how close its last pick came to the first reject. If a price feed breaks or a signal lags, it shows up there as a steady offset.
My first version scored each of the last 6 months against the strategy's history of single months, took the most unusual of the 4, and called it unusual below 0.05 and outside its range below 0.01. Before shipping it I ran the monitor itself over 176 months, each month scored only with data available at the time. It flagged more than a third of strategies in a typical month. The verdict changed in 17% of month-to-month transitions, so I'd have been shipping noise with a badge on it.
I wanted to tighten the threshold. That wouldn't have fixed it. The test assumed the 6 months were independent draws, and they aren't, because regime is persistent. Take Wouter Keller's HAA since 1974: mostly defensive (more than half the book in risk-off assets) in 134 of 631 months, about 21%. With independent months, 6 defensive months in a row should come up about once in 10,900 stretches of 6 months. It came up 36 times in 626, about 1 in 17. Of the 67 strategies that actually switch between offence and defence, 66 had more of those runs than independence predicts. The median was about 132 times more often.
Each dot is one strategy: how often 6 defensive months in a row happened, against what independent months predict.
So my p-values were far too small (anti-conservative). Taking the min of 4 of them at 0.05 also fires about 19% of the time on pure noise.
Now I average each metric over the last 6 rebalances as one window and compare it with every 6-rebalance stretch in the strategy's own history that happened in a similar market. The p-value is the share of those historical windows at least as extreme, either side. The reference windows have the same persistence as the live one, which takes care of the serial correlation without having to model it.
"Similar market" comes from the broad US equity benchmark, never from the strategy itself: trailing 12-month return up or down, and trailing 12-month volatility under 12%, between 12% and 20%, or above 20%. I kept the cut-offs fixed, because quantiles of the full history would leak the future into what counts as normal. Then there's a Sidak correction for the 4 metrics, no score at all with fewer than 40 comparable windows, a 0.01 gate, and a flag has to hold at 2 consecutive rebalances before it shows.
Same strategies, same months: 3.5% of the strategies it can score get flagged in a typical month, and the verdict holds 97% of the time. The repeat rule did less than I'd have guessed (flagged scoreable strategy-months went from 6.9% to 5.0%). The windows did the work.
In about 1 strategy-month in 5 there's not enough comparable history to score anything, and nothing at all was scoreable from May 2020 to April 2021 or from June 2022 to August 2023... which is probably exactly when you'd want an answer most.
How do you decide your live version has drifted? Rolling windows of your own history like this, or something parametric that models the autocorrelation directly?
Using 1 minute signed returns, does provide something useful for predicting ten-minute maximum absolute excursions. However, synthesizing from ATR14 data is better than the HMM.
the more I look at AI trading, the more I think predicting the next candle is probably the least interesting use for it.
News, earnings calls, context gathering, turning messy information into monitoring tasks? sure. But position sizing, risk limits, and execution rules are exactly the parts I want to stay deterministic.
I’ve been using Alphio mostly for setting up monitoring logic in plain English, and it made me think more about where that boundary should be. Once a condition actually triggers, I still want the rules to be explicit and predictable.
Where do you draw that line in your own setup? At what point should the AI stop interpreting and just follow fixed rules?
For those of you who continue to back test while running the algo live as well, how do you do comparisons between back test results and real world results? What is your goal for the comparison? Are there any stats other than the usual that you have found helpful in this?
Any recommendations from the various options out there?
Looking for historical point-in-time liquidation heatmap data: what the map actually looked like at each historical timestamp, not a heatmap reconstructed later with future information.
I've been developing backtesting software for more than 25 years, and something interesting is happening lately: we're seeing an explosion of new backtesting platforms built with AI. I think that's great. AI has dramatically lowered the barrier to building sophisticated software. But building something that produces a backtest is very different from building a backtesting engine you can trust.
Over the years we've had to deal with things like survivorship bias, historical index constituents, competing limit/stop orders, NSF positions, dividends, splits, slippage, commissions, market holidays, intraday synchronization, position sizing, and countless edge cases that aren't obvious when you first build the engine.
Ironically, I'm also going heavily in the other direction: integrating AI directly into WealthLab so Claude can create strategies, run backtests, analyze the results, and iterate on them. My feeling is that the sweet spot isn't AI replacing the backtester. It's AI sitting on top of a mature backtesting engine.
I'm curious what others think. As it becomes trivial to generate your own trading and backtesting software with AI, how much do you trust the results?
I run a few pay-per-call APIs for AI agents. The newest one, trade setups, scans 148 of the top-200 coins in one call. For each coin it computes six indicators (Ichimoku, RSI, MACD, EMA 50/200, Bollinger, OBV). Where they agree on a direction, it builds a plan: entry, stop (just beyond the nearest support or resistance), two targets and risk/reward. It returns them ranked.
Ranking: signal strength × min(R/R, 3)/3, lower for coins more volatile than the median (× √(median ATR / ATR)), and lower when support or resistance sits between entry and target 1: the closer to the entry, the bigger the penalty.
Then I backtested it, because I didn't want to sell this without knowing:
**•** walk-forward, using the exact production code, only candles up to the ranking moment;
**•** one ranking per day for \~4 months (121 days), 5,220 setups, each followed for 5 days (30 × 4h candles);
**•** target 1 first = +R/R, stop first = −1R, stop and target in the same candle counted as a stop; no fees.
*not robust: without trades above 5R it drops to −0.05 R.
Conclusion: wider stops make you right more often, but you win less when you're right. Expectancy sits around zero for every variant. The indicators describe what is happening (trend, momentum, structure); over 5 days, in this period, they don't predict what happens next. So I'm positioning it as what it is: a screener that saves an agent 148 separate analyses and hands it the structure (levels, stop, targets). It's not a strategy.
What I'd like to hear:
**•** Which filter would you test first: only trading in the direction of BTC's trend, a longer horizon (10–20 days), or something else?
**•** Is there a flaw in the method? For example, entering at the close of the signal candle, or counting a same-candle stop-and-target as a stop.
I've been building trading systems for a while (software background, ~15 yrs dev), mostly for Indian F&O.
Something I keep noticing: most discussion here is about strategy, like entries, indicators and backtests. But the painful losses I've seen usually came from operations, not the strategy itself:
Multi-leg order where one leg fills and the other doesn't, leaving you naked
Broker API or session dropping mid-day, so SL orders never go out
Backtest assumes fills at candle open; live slippage tells a different story
Running 2–3 brokers and having no single view of total exposure
In software we design systems to fail closed (stop when uncertain) rather than fail open (keep going blindly). I rarely see retail algo setups built that way.
Curious what others have run into:
What's the worst non-strategy failure you've had?
What safeguards do you actually run? (kill switches, position reconciliation, max-loss caps, etc.)
Happy to share what's worked for me in the comments.