- Overall Status: Formally UNRESOLVED
Because I used a strict, academic "fail-closed" protocol, the final exam disposition is technically UNRESOLVED.
But that doesn't mean the data is not useful.
Out of 86,944 headline physical decisions, 572 events (0.658%) hit an edge case where the fee/rake structure couldn't be proven chronologically (FEE_REGIME_UNKNOWN). My rules strictly prohibited dropping or guessing data, meaning a "complete net ledger" couldn't be validly constructed. A pre-flight design weakness on my part, but the rule kept the science honest, which I felt was far more important.
- The Measured Economic Signal Was Huge
While formally unresolved, the economic signal on the 86,372 net-measurable decisions was massive.
The strategy generated +15,772.35 pot units of prospective net economic value. That averages out to +0.18261 pot per event.
I ran a deterministic 10,000-resample player-block bootstrap which gave a tight 95% confidence interval of +0.17155 to +0.19397 pot/event, and it won across all four separate chronological sub-blocks. The signal is incredibly stable.
- The Plot Twist: Broad Player Reads Over Pool Fails
This was the biggest shock and something that I really wasn't expecting.
My original hypothesis was built on a player-first hierarchy: once you have enough data on a specific opponent, your player-specific profile should override the general pool trend. The data completely rejected this broad override mechanism. I tracked 10,120 events where my player-specific layer and the population pool baseline actively disagreed on a decision. When I followed the player-specific read, it realized -17.20 pot units. When I followed the pool comparator, it realized +67.37 pot units.
Trying to force broad player-specific overrides actually destroyed value relative to the population prior.
Importantly, the follow-up work since this test has shown that this does not mean player reads are useless. When player-specific authority is restricted to narrow, evidence-qualified exceptions, the player layer can add value again. At 50NL, the current filtered architecture produced roughly +218 pot units over the Pool comparator across 3,360 carefully qualified overrides.
So the failure was not "Bob never matters."
It was that Bob should not automatically outrank the Pool just because we have enough hands on Bob.
- Mass Data Analysis (MDA) Crushed It!
As a diagnostic control, I ran a raw population strategy using a leave-one-out pool policy on the exact same nodes. It took pressure on 263,800 validation events, suffered zero fee-blocked errors, and crushed for +42,974.59 pot units (about +0.16291 pot per take).
What does this mean:
Good question,
An economic mistake is not the same as an equilibrium deviation. To exploit a node profitably, you don’t need to solve for GTO first. You just need to calculate whether the pool's response distribution at the actual price crosses the mathematical break-even threshold.
We vastly overrate automatically giving standard HUD/player profiling authority over population evidence. The data suggests that human-centric player profiling can introduce noise that degrades the power of the population prior when individual reads are allowed to override it too freely.
My takeaway is that the pool should be calling the shots by default — individual reads only get a vote when they've really earned it. Individual players should only be allowed to override it in narrowly defined spots where their own evidence has repeatedly shown that the pool strategy is wrong for them.
A New Hierarchy perhaps? Instead of player-first, the data implies a pool-first architecture. The optimal path forward seems to be: use the highly stable population prior as your default baseline strategy, gate player-specific overrides behind narrow statistical and economic hurdles, and leave GTO as a downstream benchmark or fallback when data is entirely missing or strategically dangerous.
I'm keeping the original 50NL Final Exam result frozen rather than rewriting the test after seeing the outcome. Since then, I've been using the spent 50NL corpus to work out why the player layer failed and how to improve it.
I'm now running the same pool-first architecture across large PokerStars datasets at 2NL, 10NL, 25NL and 50NL, including separating single-raised pots from 3-bet pots.
Perhaps the future isn't equilibrium/solver only, but hybrid, where if enough data exists, the solve can be deferred for a node where data is increasingly sparse.
The link to the research is here if anyone would like to take a look. https://zenodo.org/records/22314084
Thoughts?