A day after publishing Does Momentum Work in Pokemon Cards, we are retracting its central claim. That piece measured 28,288 point-in-time observations and concluded that a card up 25 to 50 percent over the prior 90 days beat its cohort by 4.7 points over the next 90. We have since rebuilt the study on a much larger sample, and the effect reverses. This piece explains what the bigger sample says, why the first one was wrong, and what we changed in the product because of it.
The curve is a U, and the bottom of it is where most cards sit
Forward 90-day return by trailing 90-day return, 65,685 observations
| Trailing 90d | Observations | Forward 90d | Win rate |
|---|---|---|---|
| down more than 10% | 7,349 | +7.36% | 40.5% |
| down 3 to 10% | 6,890 | -3.67% | 31.6% |
| flat, within 3% | 15,935 | -2.00% | 26.5% |
| up 3 to 10% | 10,394 | -1.18% | 34.6% |
| up 10 to 25% | 12,905 | +0.23% | 39.7% |
| up 25 to 50% | 7,832 | +0.65% | 42.4% |
| up more than 50% | 4,379 | +1.66% | 40.4% |
Read the middle of that table first, because it is where the market actually lives. The three buckets from down 3 percent to up 10 percent hold 33,219 of the 65,685 observations, and all three are negative. The flat bucket alone is 15,935 observations with a 26.5 percent win rate, the lowest on the table by a wide margin. A card that has done nothing for three months is not resting. It is the single worst prior condition we can measure.
Then read the ends. The rising buckets are positive but small, between two tenths of a point and 1.7 points, and they rise gently with momentum rather than peaking in the 25 to 50 band the first study identified. The falling end is not small: cards down more than 10 percent returned 7.36 percent over the following 90 days, roughly four times the best momentum bucket. An independent pass over 57,971 observations, computing returns relative to the same-window cohort and bootstrapping by card, put the trailing-return decile spread at negative 9.7 points, which is the same conclusion stated as a slope.
The first study measured a survivor pool and called it the market
The original backtest ran on cards with deep recorded price curves, roughly 3,298 of them. That sounds like a sampling detail and it is actually the whole result. To have a deep curve in our data, a card has to have been continuously priced and continuously interesting for the better part of two years. Cards that went quiet, got delisted, or drifted into the bulk bins do not qualify. The sample was therefore selected on precisely the outcome the study was trying to measure, which is the oldest mistake in quantitative finance and one we walked into with a large observation count as cover.
The number 28,288 was doing a lot of unearned work. It sounds like a big sample, and it is a big sample of a small and unrepresentative group. The rebuilt study covers 52,041 cards, roughly sixteen times as many, including all the cards that stopped being interesting. Adding them changes the answer, which is the definition of the bias being real.
The second failure was ours as engineers rather than as analysts. The scheduled job meant to re-measure this every month was reading a price table that is only nine days deep, while requiring 104 days of prior history and 90 days of forward outcome per observation. It therefore produced zero observations on every run and reported that as an empty result rather than as an error. A harness that silently measures nothing is worse than no harness, because it looks like oversight. It now reads the full recorded curves and refuses to return a result below a minimum observation count.
The number that decides whether any of this is tradeable
Selling into eBay costs roughly 13 percent between fees and shipping, and buying at the ask rather than the bid costs several points more. Our net-return model puts breakeven at about 24 percent gross over the holding period. Against that floor, a 7.4 percent forward return in the best bucket does not clear, and the momentum buckets are not close. Of all 65,685 observations, roughly a quarter clear the cost floor at 90 days.
That is the honest summary of both studies taken together. The tradeable version of this finding is not a signal you act on card by card at retail. It is a filter that tells you which conditions are worth paying attention to at all, and a warning that the flat, quiet, comfortable-looking card is the one most likely to disappoint.
Three things we removed or reversed today
First, our pick engine had a hard gate that rejected any card whose 90-day tape was falling more than 2 percent. It was the largest single filter in the engine, rejecting 86 to 95 candidates per scan. On 57,971 observations the cohort it kept returned negative 0.79 percent against positive 2.43 percent for the cohort it discarded. It was systematically throwing away the better half of the candidate set, so it is gone. The reading still appears on every pick as disclosure. It no longer filters.
Second, a price penalty that had been written into the engine and never actually applied is now applied. Cards entering above 100 dollars performed materially worse than cheap ones across the sample, and our published picks had been running a median entry price of 347 dollars, which is the worst band we measure.
Third, our gem-rate factor was rewritten for a related reason, described in The Gem Rate Curve. Both mistakes have the same root: a threshold that looked precise, was calibrated on a narrow slice, and selected a proxy for age rather than the property it named.
Point-in-time study over our recorded daily price curves unioned with the canonical daily price accretion, 52,041 cards, span 729 days, as of 2026-09-08. For each card, evaluation points are taken every 30 days where 104 days of prior history and 90 days of forward outcome both exist, giving 65,685 observations. The trailing signal uses only data dated at or before the evaluation day. Forward return is measured from the evaluation day to a window centred 90 days later. The independent cross-check described above computes returns relative to the mean and median of the same evaluation window and bootstraps confidence intervals clustered by card, so that many observations from one card cannot masquerade as independent evidence. The prior study, retracted here, is left published at its original URL with a link to this piece rather than edited, because a research record that quietly rewrites itself is not a record.


