ProEA Lab · Honest notes on building & testing a real MT5 system · No income claims · Every number links to its source
Order Flow

Multi-Timeframe Confluence Backtest: Does Timeframe Agreement Predict Anything?

Our MTF matrix reads an 11-indicator bias across five timeframes and prints how aligned the stack is. We ported its exact rules, took 1,131 hourly readings across 50 sessions of gold futures, and tested the folklore that full alignment is an entry edge. Full alignment turned out to be common, sticky — and worth nothing for predicting the next hour. The product never claimed otherwise. The folklore does. Every number reproduces from committed code.

PLProEA LabJul 12, 2026 · 7 min read
Five antique compasses on a dark navigation desk all pointing the same way while a fogged window hides the road ahead — poster titled All Aligned, from the ProEA research series.

"Five of five timeframes agree." It feels like the safest moment in trading.

We measured what that feeling is worth over gold's next hour. Nothing detectable. Agreement is not rare or precious — quite the opposite: the whole stack agrees far more often than most traders would guess, and the next hour looks the same either way.

We sell the tool that displays this agreement, so the receipts matter. Our own matrix calls itself a bias visualizer and refuses to make win-rate claims. This study is the independent check on whether that refusal was wisdom or modesty. It was wisdom.

What we tested

The MTF Confluence Matrix reads a trend bias for each timeframe. At the default SMART setting, eleven classic indicators vote up, down, or abstain; the tool then reduces the default five-timeframe stack (15m, 1h, 4h, daily, weekly) to three numbers: percent aligned, net bias, and a strength-weighted conviction score.

Its own docs call it a bias visualizer. They make no win-rate claims and offer no entry signals. The pitch is simpler: replace flipping through five charts, using only confirmed bars so the higher-timeframe cells never repaint.

The folklore around tools like this makes a stronger claim: only trade when everything lines up. Traders repeat it as if alignment itself were an edge. That claim is testable, so we tested it.

The design

Data. The observation clock and the 15-minute stack use the same frozen 50-session GC=F 5-minute tape as studies #1 and #2: 13,640 bars from 2026-04-29 to 2026-07-10. A 200-period EMA on the 4-hour chart needs months of warmup, so we froze two longer series from the same source: hourly bars back to February 2024 and daily bars back five years. All three files are committed.

The rig. We ported the shipped Pine to TypeScript at defaults and annotated every rule with its source line. The port covers the eleven votes, the abstain rules, the consensus math, the confluence reducer with flats excluded, and the confirmed-bar read; every cell sees only the last completed bar of its timeframe.

Observations. We sampled every hour on the 5-minute clock: 1,131 readings, spaced so the primary forward windows never overlap. Of those, 1,124 had a directional net bias; we excluded 7 ties.

Outcomes. We pre-registered two. Hit asks whether price sits on the net-bias side of the entry one hour after a reading. Return is the one-hour move in ATR units, signed by the bias.

The null. We shuffled the confluence column across observations 2,000 times (fixed seed, committed code) and measured how often shuffled alignment separated outcomes as well as the real column. This is the same design that audited our zone score. Because hourly readings of a slow-moving stack are serially dependent, we added a session-cluster bootstrap as the honesty guard and a four-hour horizon as sensitivity.

Finding 1 — full alignment is common, and sticky

Folklore treats "everything agrees" as rare and precious. On this tape, it was neither.

The stack reached 100% alignment in 28.6% of all hourly readings: 323 of 1,131. At the product's default alert threshold of 80%, it was aligned 58.0% of the time. Aligned runs lasted a median of 3 hours; the longest lasted 57.

Horizontal stacked bar of 1,131 hourly readings: 28.6 percent at full alignment, 29.4 percent at 80 to 99, 41.4 percent at 50 to 79, 0.6 percent ties — with a note that aligned runs last a median of three hours.
How often the five-timeframe stack agreed, hour by hour, across 50 sessions. Full agreement is a nearly-one-in-three state.

Stop on that before the statistics. A condition that holds 29% of the time cannot be a rare edge. Whatever full alignment means, "special" is not it — and no test we ran below changes that arithmetic.

Why so common? All eleven trend-followers vote on five nested timeframes of the same price series. They agree for the same reason five thermometers in one room agree: correlation, not confirmation.

Editorial panel on ivory paper reading: Full alignment is common, not special — with an orange marker underline.
Nearly one hour in three, on this tape. That arithmetic disqualifies the folklore before any statistics run.

Finding 2 — alignment level says nothing about the next hour

Start with the baseline. Following the net bias for one hour hit 48.2% across all directional readings, with a mean signed return of −0.020 ATR. On this window, the stack's direction was a coin flip; the test was whether more alignment beat less.

It didn't:

Reading (H = 1 hour)Bottom tercileTop tercileDiffShuffle p
Hit rate by confluence %47.1%47.5%+0.4 pts0.43
Signed return (ATR) by confluence %−0.229−0.020+0.2090.14
Hit rate by conviction score47.1%46.9%−0.1 pts0.50

The session-cluster bootstrap puts the hit-rate difference at a 95% CI of [−5.2, +10.0] points. That band sprawls across zero and ends the conversation. The four-hour horizon reads the same (+2.3 pts, p=0.27).

The flagship moment was the FULL STACK alert: the edge-triggered crossing into 100%, seen in 67 readings. It hit 44.8% over the next hour, below the 48.2% baseline (p=0.77 against random non-event draws). The panel's most-aligned moment carried no next-hour information at all.

Editorial panel on ivory paper reading: More alignment predicts nothing about the next hour — with an orange marker underline.
On this tape, at these horizons. The table above is the whole case.
Bar chart of next-hour hit rate: 46.6 percent at 50 to 79 percent aligned, 50.8 percent at 80 to 99, 48.0 percent at full alignment, and 44.8 percent for the 67 full-stack alert events, against a 48.2 percent baseline drawn as a dashed line. The three alignment bars hug the baseline; the full-stack alert sits below it.
Next-hour hit rate by alignment level, 1,124 directional readings. Every bar hugs the coin-flip line — the alert bar sits under it.

Finding 3 — confluence is a state, not a forecast

Put the two findings together. The matrix's number acquires its real meaning.

Alignment describes now. It says the trend reads across five timeframes currently cohere — a regime statement, just as "it's raining" describes weather without predicting tomorrow's. That is useful: one glance replaces five charts, and the panel can't repaint what a closed bar already said.

What it is not (on this tape, at these horizons) is a forecast. And this rhymes with what study #2 found from a different angle: there, the impulse features (displacement, volume — things that just happened) carried real signal while the composite state-score diluted it. Here, a pure state-description carries none. Twice now, the same lesson from our own products: information about what price just did beats information about how tidy the current picture looks.

Editorial panel on ivory paper reading: Confluence describes the present, not the future — with an orange marker underline.
The number's real meaning, and the only reading this study supports.

What survives, honestly

The product's own framing passes its audit. Its docs say "no win-rate/profit claims" and call the matrix a visualizer. Nothing we measured contradicts a shipped claim. The folklore fails; the tool doesn't.

The workflow value is untouched by any of this. Reading five timeframes in one panel still beats opening five charts. Confirmed-bar reads can't repaint, and the request-budget guard keeps the tool loading where other MTF scripts crash; neither depends on alignment predicting anything.

The audit loop worked again. The matrix ships readable source, so we could run the test. The scripts and frozen data are committed; owners can rerun the rig on their own market, and a different market or horizon may genuinely answer differently.

Limits — read before quoting

One symbol (gold futures), one 50-session observation window, two horizons. A different market, timeframe stack, or bias method may genuinely answer differently — the committed rig makes that rerun a one-command job.

The rig is a faithful line-annotated port of the shipped Pine at defaults, not the TradingView runtime itself. Its 4-hour and weekly buckets are UTC-aligned rather than exchange-session aligned, and forward windows near a session close straddle the overnight gap; both apply identically across alignment levels, so they move the baseline, not the comparison.

The observations are hourly readings of a slow-moving stack, so they are serially dependent; the flat shuffle leans optimistic, which is why the session-cluster bootstrap is the reading that matters — and it straddles zero.

What we're changing

  1. The matrix's docs get a "what confluence is — and isn't" section linking this study, putting the limit in the manual.
  2. The FULL STACK alert's description will call it a state-change notification, not an action cue.
  3. Study #4 in this series will test transitions rather than states: whether a direction flip, rather than sustained alignment, carries anything. Both audits so far point at change, not condition.

If a rerun on another market overturns any number above, the committed script will report it before we do.

Count how often it's true

Before you treat a condition as an edge, count how often it is simply true. On any chart, mark every hour it holds and divide.

If the answer is "a third of the time," you are looking at wallpaper, not an edge. Wallpaper can be useful and honestly descriptive without predicting anything by itself.

The stack agreeing is information about the present. The folklore sold it as information about the future. Only one of those survived contact with its own source code.

More from ProEA Lab

Trailing vs Static Drawdown: The Rule That Moves While You Trade

Two accounts, same drawdown percentage, completely different games: a static floor never moves, a trailing floor climbs behind every new equity peak — which means your best trading day can quietly shrink your cushion. Here's the arithmetic of both floors, the three firm-specific variants that decide everything, and why the dangerous moment is right after you've been winning.

Aug 10 · ProEA Lab

A brass figure climbing a staircase while the stairs behind it fold upward into a rising floor, closing the gap — poster titled The Floor Follows.

Walk-Forward Testing: Stop Grading Your Own Homework

One backtest over one history rewards whoever searched hardest, not whoever found something real. Walk-forward testing is the honest alternative: optimize on a window the strategy can see, judge it on the window it can't, roll forward, repeat — and only ever score the unseen parts. Here's the protocol on any platform, the one ratio worth recording, and what walk-forward still can't prove.

Aug 10 · ProEA Lab

A brass stencil sliding along a long paper chart strip, exposing only one unmarked section at a time while the rest stays covered — poster titled Judge the Unseen Part.

How to Add a Pine Script to TradingView (Both Routes, 2 Minutes)

TradingView has no Import button — and in the current interface the Pine Editor isn't at the bottom of the chart anymore, so half the guides out there point at a tab that no longer exists. Here are both routes as they actually work today: pasting source code into the editor, and pulling a community script from the Indicators dialog — plus the paste mistake behind most 'compile error' support tickets.

Jul 24 · ProEA Lab

A brass pneumatic-tube capsule holding a rolled paper script being slotted into a receiver panel beside a dark candlestick chart — poster titled Two Routes In.