Pivot Confirmation on Three Unseen Weeks — Precision Holds, the Threshold Does Not
The reversal ladder of the previous article was scored on a held-out coin — but in the same weeks it trained on. Here the models trained on everything up to 14 August 2026 and were scored on the three weeks after, six coins, with the threshold fixed in advance from honest training scores. Precision did not fall: 51% two and a half hours after a pivot, 73% after six hours, on 88–93 reversals. What broke instead was the threshold itself — at two rungs of the ladder the models went completely silent while their AUC stayed above 0.9.
① A stricter test: held-out time, fixed threshold
Study 23 measured everything on train with a held-out coin: the model had not seen that coin, but it had seen those very weeks. Here the test is held-out time: train on all six coins up to 14 August 2026, then score 14 August – 4 September.
The ladder is unchanged — "was there a 2.3% pivot exactly N events ago?" — and each example lies wholly on one side of the split. Four input sets are compared: the class (62 batch features), the context (5 causal numbers on one 400-event window), four windows (the same five on 200/400/800/1,600 events, 20 numbers) and four windows plus class. The threshold is set at 30% recall on honest training scores — each training event scored by a model that had not seen its coin — and applied to validation as is.
Validation is small and the article says so: 112 coin-days and 88–93 reversals per rung. With that sample the error on a precision is several percentage points.

② Precision survives the move to new weeks
At rung 43 (150 minutes after the pivot) four windows plus class give 55.0% on train and 51.0% on the unseen weeks; four windows alone, 49.9% and 49.1%. At rung 103 (six hours) validation is even higher than train: 73.3% against 67.9%. Rung 43 is the working point: the highest recall (55–60%) at about 50% precision. Rung 103 is more precise but catches only 38–53% of pivots.
A second score counts signals that miss the pivot event but point to an event within ±0.1% of its price — a miss in time, not in the entry level. It matters most for the single-window context: at rung 103 its strict precision is 24.3%, with tolerance 56.9% — 35 of its 144 signals sit on the pivot and 47 more on its level. Four windows barely need the tolerance (73.3% either way): one window shows where, four windows also show when. The class on top of four windows changes precision by a few points in either direction and halves recall — within the error.

③ The model separates — the threshold does not travel
The failure that did appear is of a different kind. At rung 23 three of the four input sets gave no signal at all on validation; at rung 3 the four-windows-plus-class set was silent. Their AUC on validation was 0.91–0.99 — the models still rank reversals far above everything else — but the threshold taken from training ended up above the entire validation series. The ranking transfers; the absolute level of the scores does not. And the breakdown jumps between sets and rungs, so no single set can be fixed as "the" model yet.
Per coin at rung 43 the picture is uneven but consistent: BTC hits 8 of 9 pivots, ETH 8 of 11, while XRP — with 31 pivots in three weeks — catches 11 and misses 19. One or two pivots per coin sit at the edge of the data, where the answer "a pivot 103 events ago" simply has not arrived yet.
Two things carry forward. The confirmation ladder keeps its precision out of time, which makes it a real candidate for a trading rule. And its weak point is calibration: a threshold from the past will not always admit anything in the future, so a live bot needs a threshold that adapts, not one frozen at training time. Three weeks and ~90 reversals are still a short check — the numbers agree with train and with each other, but decisions need a longer validation.

Reproduce this study
- Research log (.md, Ukrainian): goal, data, plan, scripts, every confirmed stage and table — enough to rerun the study
- Reproduction kit (.zip): the study's scripts, project rules and base scripts that build every class
If this changed how you read the tape, the natural next step is Volume Is the Fuel — Not the Steering Wheel — We recorded the Binance order book every second for six coins over five months and ran eighteen tests on what volume really does.
Volume Is the Fuel — Not the Steering Wheel
We recorded the Binance order book every second for six coins over five months and ran eighteen tests on what volume really does.
About Market Research Lab — What We Collect and Why
Most market commentary is storytelling.
From Calm to Calm: a Standard for What Counts as a Signal (BTC)
We stopped defining market signals with a stopwatch.
The Wall That Goes Quiet: What Resting Liquidity Predicts
We measured resting limit liquidity sitting on both sides of the book second by second across six coins, and asked the only question that matters:….
Comments
Discussion is powered by GitHub. Enable it by adding secrets/giscus.json (repo IDs from giscus.app).