The point of calling yourself a quant platform is that you can be wrong in public and prove you fixed it. So we handed our daily pick engine to an adversarial reviewer with one instruction: tear it apart as a prediction system. It did. This is the audit and the repair, with the numbers, because a track record nobody can check is just marketing.

The blunt version

What was broken

Three findings mattered. First, a factor called species-demand fired on nine of nine recent picks, because it was a level proxy: Charizard, Pikachu, and the Eeveelutions are permanently in the top demand percentile, so tagging a pick with high species demand said nothing and selected nothing, while quietly driving the ranking. A factor attached to nearly every pick is not a signal. Second, none of the factors had ever been shown to predict forward returns; the only live scoreboard read uniformly negative while the engine kept publishing. Third, picks with 12-to-36-month theses were being graded on marks one to seven days old, against a 44-card benchmark where a single card repricing could move the whole index, and the benchmark level was re-derived each time rather than frozen at the pick's entry.

The founding miss made it concrete: Base Set Charizard compounded 68 percent over fourteen weeks in clean monthly steps, sitting in the engine's scan universe the entire time, and was never recommended. Two structural blind spots hid it. The growth factors read week-over-week deltas and only fired on sharp jumps, so a card grinding higher in small weekly increments never tripped them, even one compounding to Charizard's 68 percent; and the moment its run crossed a flat plus-25-percent cap, it was blocked as a momentum chase. The guard built to avoid buying tops was filtering out the single best sustained trend on the board.

Each fix backed by the backtest

What we changed

We ran a 28,288-observation backtest (the full study is in Does Momentum Work in Pokémon Cards?) and rebuilt the engine on what it measured. The flat overextension cap is now era-conditional: it blocked cards that went on to beat their cohort by 4.4 points, and only modern parabolas above plus-50-percent actually reverse. A new dead-tape gate removes the study's two worst cohorts outright, the flat tape (-4.0 points) and the bottom-quintile-volatility tape (-5.3 points), which was worth more than any new buy signal. The compounder shape is now a scored factor with its measured thresholds. And species-demand was demoted from a counted factor to a context tag: it can no longer clear the bar or drive the ranking, which the engine bar must now earn through timing signals. The day we shipped it, the count of candidates clearing the bar fell from thirty to ten as the phantom factor stopped inflating it.

9 of 9picks the level-proxy factor fired onbefore the fix
+4.4 ptsalpha the flat cap was blockingnow era-conditional
-5.3 ptsthe dead-tape cohort we now gate outbottom-quintile volatility
30 to 10candidates clearing the bar, day of the fixphantom factor removed
The instrumentation

How we keep ourselves honest now

Fixing the factors is not enough; the point is to keep them under test. Every pick and every factor is now scored on a spread-adjusted net basis, buying at the ask and exiting net of fees, so an edge that does not survive the round trip is visible rather than flattering. The benchmark level is frozen into the ledger at entry, so alpha can never be revised by a later index rebuild. Every candidate that clears the bar, not just the four we publish, is recorded and scored as a shadow cohort, which is the only way to learn whether our ranking and our concentration guards add value or cost it. And the entire decision layer now sits behind a unit-test suite, so a threshold cannot drift silently. The all-era backtest re-runs monthly and will supersede today's numbers on its own.

The audit was an independent adversarial review of the daily and weekly pick engines as prediction systems; the fixes are calibrated on the 28,288-observation point-in-time backtest documented in the linked study (recorded price curves, no lookahead, cohort-relative alpha). Every threshold cited is in the engine's tested decision module. Spread assumptions: entries marked up 8 percent to the ask, exits net of 13 percent fees. The shadow cohort, frozen index anchor, net scoring, and monthly all-era re-measurement are live as of publication; factor attribution and the shadow-cohort edge are visible on the daily buys page and accrue in public from here.