← Market Flow Research 🌙
Four numbers, 28 cells

Four Numbers per Event and a 28-Cell Answer — Where the Signal Really Came From

The previous networks read each book feature through 13 exponential averages and topped out at a correlation of 0.29 with the leg target. Strip the averages away and the same features, raw per event, correlate at 0.09. So the link was made by the smoothing, not by the snapshot. We then gave a transformer six coins, 345,000 training windows and a new kind of output — 28 cells instead of one number — to stop it hiding in the middle. It stopped hiding. It did not start hitting: 5.4% of real pivot-zone signals, with false calls a day away.

Research note · 30 September 2026 · AI-assisted, human-reviewed

① The snapshot barely knows the leg

Studies 16 and 17 fed their networks a 13 × 4 matrix: four book features, each smoothed by 13 EMAs up to 130 events. The strongest single link with the target there was |r| = 0.294 — volume balance under a 130-event EMA. This study asks how much of that is in the features themselves.

The class here is `210×210×4c`: one event, four numbers, no averaging — book balance, change of balance, volume balance and the price feature, all in [−1, 1]. The target is the same leg of a 2.3% zigzag on a 90-event EMA of price.

The answer on BTC train (23,166 events): the strongest feature is volume balance at −0.092, three times weaker than its smoothed version. The link was made by the smoothing. That makes sense: four numbers describe 210 seconds of the book, while the target lives on a 5-hour EMA of price. Without averaging, feature and target are simply on different time scales. The signs agree with study 16 on every variant: balance and its change positive, volume balance and price feature negative.

Pearson correlation of four raw event features with the raw and the leg-smoothed 2.3% zigzag target on BTC train
Pearson correlation of four raw event features with the raw and the leg-smoothed 2.3% zigzag target on BTC train

② Smooth each leg by its own length

The target twitches between pivots, and earlier studies calmed it with one fixed window for all legs. But 2.3% legs vary enormously: median 802 events, shortest 227, longest 2,431. One window undersmooths long legs and flattens short ones.

So here the window is set per leg, equal to the leg's own length — centred, so there is no phase lag. Looking forward is fine for a target, which is labelled in hindsight anyway; it would be a leak for a feature, and none is smoothed this way. Pivot zones are left untouched and nothing crosses from one leg into the next.

It is the strongest smoothing of the whole ladder — median window 802 events — and the only one that added to the link on train: volume balance −0.0957, 104% of the raw target's −0.0921; no fixed window did better than 98%. The honest caveat is in the same table: over all 39,879 events the gain disappears (83% kept), so this is an observation on train, not a ready target. What it does show is that the obstacle was the uniformity of the window, not its size.

Share of the raw correlation between volume balance and the leg target kept after smoothing the target with fixed windows versus a window equal to each leg's length
Share of the raw correlation between volume balance and the leg target kept after smoothing the target with fixed windows versus a window equal to each leg's length

③ 28 cells instead of one number

The minimum of MSE on a noisy target is the conditional mean, which is why every model so far clustered near zero: saying −0.9 and missing costs more than saying nothing. So the output changed. The range [−1, 1] is cut into 28 cells: the two edge cells are fixed at |y| ≥ 0.80 — literally the pivot zone — and the 26 in between are quantiles of the target, about 3.5% of events each. The label is soft (0.70 on the true cell, 0.15 on each neighbour), the loss is cross-entropy, and a "bottom" or "top" call is simply the most probable cell being 1 or 28.

Training moved to all six coins with gaps filled — 387,990 events, split by date (train up to 10 August, test after), about 345,000 training windows. Four transformers differed only in window length, under the same wall-clock budget.

On test (41,462 events): K = 128 is best with 264 bottom hits and 105 top hits — 369 of 6,783 real edge signals, 5.4%. Longer windows are strictly worse (369 → 289 → 223 → 0), because attention costs K²: in the same 40 minutes K = 1024 saw the data 7 times, K = 128 sixty. The model is not silent — it called "top" 3,374 times against 1,243 real ones — but a false call lands on average 444–890 events (about a day or more) from the nearest real signal of the same cell. Unusable as an entry. Four numbers per event are too few; the next studies change the features, not the length of the window.

Hits of the bottom and top pivot cells on test for transformer windows of 128 to 1024 events
Hits of the bottom and top pivot cells on test for transformer windows of 128 to 1024 events

Reproduce this study

🤖 This research — data collection, analysis, charts and the narrated video — was produced with the assistance of AI, then reviewed by a human. We forecast volatility, not direction, and log every prediction. Research, not financial advice.

Comments

Discussion is powered by GitHub. Enable it by adding secrets/giscus.json (repo IDs from giscus.app).

← All research