The Fibonacci Disconnect
What the Academic Studies Measure, and What Traders Actually Do
1. The Academic Position, Stated Fairly
It is tempting to open a page like this by claiming that academic finance dismisses Fibonacci analysis as folklore, and that the dismissal is the product of prejudice against practitioners. That framing is common, it is rhetorically convenient, and it is not accurate. It should be dropped, because building an argument on a misrepresentation of the other side guarantees the argument collapses the moment a reader checks.
The academic record on technical analysis is not monolithic, and it is not uniformly hostile. In their review of the field, Park and Irwin (2004) catalogued 92 modern studies of technical trading rules published between 1988 and 2004: 58 reported positive results, 10 were mixed, and 24 were negative. Whatever one makes of publication bias and transaction costs, that distribution is not the signature of a method science has confidently rejected. It is the signature of a contested question.
On support and resistance specifically, one of the more rigorous tests was conducted not by a chartist but by an economist at the Federal Reserve. Osler (2000), in the Federal Reserve Bank of New York's Economic Policy Review, took the support and resistance levels that six major foreign-exchange firms published to their clients in advance, over 1996–98, and scored them against what price subsequently did. She found the levels were "quite successful" at predicting intraday trend interruptions, and that their predictive power persisted for at least five business days after publication. Levels specified before the fact, tested out of sample, at a central bank: this is the opposite of a field reflexively dismissing the idea that price reacts at pre-identified levels.
So the honest statement of the academic position is narrower and more interesting than the folklore story. It is this: when Fibonacci ratios are tested by the specific procedures academics have used, the ratios do not cluster more than chance. That is a real finding. The question this page examines is what those procedures actually measure — and whether it is the thing traders do.
2. What the Strongest Study Actually Did
The most careful test of Fibonacci ratios in a major equity index is Batchelor and Ramyar's Magic numbers in the Dow (Cass Business School, 2006). It is worth setting out exactly what they did, because the detail is where the disconnect lives, and because they deserve to be represented precisely rather than caricatured.
They took daily data on the Dow Jones Industrial Average across 22,194 trading days, January 1915 to June 2003. They identified turning points using the Pagan–Sossounov business-cycle dating procedure, with two deliberate modifications: they used the daily high series for peak candidates and the daily low series for troughs rather than closing prices — because, in their words, "technical analysts in practice employ charts with daily bars rather than single points" — and they added a censor requiring each peak to exceed its preceding trough. A turning point qualified only if it was the most extreme value in the eight months on either side of it. Cycles shorter than sixteen months, and phases shorter than four months that moved less than 20%, were discarded. This yielded 46 turning points across 88 years, roughly one every twenty-three months.
For each run of four consecutive turning points they computed the ratio of the latest price leg to the preceding one — the retracement ratio — and the ratio of the latest leg to the leg before that — the projection ratio. They asked whether these ratios landed near any of ten values (0.382, 0.500, 0.618, 0.786, 1.000, 1.382, 1.618, 2.000, 2.618, 4.236), within a band of ±0.025, more often than a bootstrapped copy of the Dow would produce by chance. The null was a Politis–Romano stationary block bootstrap, 2,000 replications. Counting retracements and projections, bull and bear phases, and four measurement dimensions (price level, log price, percentage price, and duration), they ran 144 such tests.
Their finding: 15 of the 144 tests cleared the 90th percentile, against a chance expectation of about 14.4. Their own summary is the fair one to quote: "A few significant ratios appear, but no more than would be expected by chance given the large number of tests we conduct." They also noted that the 4.236 ratio had too few occurrences to test meaningfully and discarded it.
Two things about this study deserve emphasis before any criticism. First, it is honestly done and honestly reported; there is no sleight of hand in it. Second, Batchelor and Ramyar level a criticism at practitioners that this research programme accepts in full: that popular technical writing "contains dramatic examples of successful predictions of turning points, with no count of misses or false alarms." That is exactly the error corrected elsewhere in this series, and they are right to name it.
3. We Reproduced It
Criticising a study one has not reproduced is cheap. So we rebuilt the Batchelor–Ramyar detector from their published procedure and ran it on the same index over the same window. The implementation and its regression tests live in the research codebase; what matters here is that it reproduces their result.
At their three stated parameter settings our detector finds 50, 54, and 71 turning points, against their reported 46, 60, and 72. We have not tuned it to close those small gaps: the paper does not specify the order in which its censoring rules apply, and our price series carries about 5% more bars than theirs. Landing within a few turns of their counts, untuned, is what a faithful replication looks like, and stopping there rather than fitting to their exact numbers is the discipline the rest of this series demands.
Running their ratio test against a 2,000-replication stationary bootstrap, we reproduce not only their conclusion but its texture. The single ratio that stands out in our replication is 0.618, in the retracement dimension, at roughly the 99th percentile — and 0.618 is precisely the ratio Batchelor and Ramyar themselves highlight. Yet the count of ratios clearing the 90th percentile is two out of twenty in the price-level dimension we ran, against a chance expectation of two. The one ratio a chartist would most expect to matter does stand out, and the study is nonetheless right that the overall pattern is indistinguishable from chance. Both halves of that sentence are true, and holding them together is the whole point.
4. Why the Design Cannot See What a Trader Sees
The disconnect is not in the arithmetic — every ladder level reproduces exactly under its stated formula. It is in the first dependency, the step every study takes before any Fibonacci level is drawn: the identification of the swing. Change how the swing is identified and everything downstream changes with it. A method defined relative to reference points inherits the properties of whatever chooses those points, and none of the academic tests validate their swing selection against how a trader would draw it.
4.1. The turning points are the wrong scale, and some are structurally invisible
Batchelor and Ramyar's eight-month window is designed to find major macro-cycle turns. In a sustained rally it cannot find intermediate lows at all, because a pullback low in a rising market is never the lowest low of the sixteen months surrounding it. This is not a tuning artefact; it is a mathematical property of the filter.
The consequence is stark on the one rally every Fibonacci analyst studies. Between the October 1923 trough and the September 1929 peak, their detector identifies no turning point whatsoever — six years, a 350% advance, and not one intermediate anchor. Yet inside that span sit pullbacks any trader would have drawn a retracement from: roughly −16.7% in early 1926, −15.0% in late 1928, −13.3% in early 1929, among others in the 12–17% range. None can be dated as a trough. They are not down-weighted; they are absent from the sample, removed twice over — once by the window, and again by the rule that discards phases moving less than 20%. A study that cannot see a 16% pullback is not testing how retracements are used.
4.2. The yardstick is the adjacent leg, not a projected level
The deeper mismatch is what gets measured. Batchelor and Ramyar compute the ratio of each leg to the one immediately before it. The reference point moves every time a new turning point prints. There is no durable anchor and no forward projection.
Consider what this does to 1929. The four-point sequence their detector produces there runs from a March 1923 peak to an October 1923 trough to the September 1929 peak. So the size of a six-year, three-hundred-point advance is measured against a six-month, nineteen-point dip. The resulting ratio is 15.42. The largest ratio they test is 4.236. The 1929 top does not register as a hit or a miss; it falls off the end of the ruler and contributes a silent non-occurrence to every hypothesis. Across the sequences their detector produces, roughly one in five exceeds the largest ratio tested — which is also why they found 4.236 too sparse to evaluate. It is sparse because their yardstick rarely reaches that far.
This is a different object from a trader's ladder, which fixes one anchor — a topping swing — and projects levels forward from it, watching the same anchor across many subsequent moves. Ratios between adjacent macro-legs and levels projected from a fixed anchor are not two tolerances on the same test. They are different measurements of different things.
4.3. Close-and-clock criteria hide the fastest reactions
Many studies — though notably not Batchelor and Ramyar, who used high/low bars for exactly this reason — detect touches on closing prices and require the reaction to complete within a fixed number of bars. Both choices quietly truncate the sample, and they truncate it in a biased direction.
A wick to a level followed by an immediate reversal is, definitionally, the fast rejection a level-trader is looking for. Under a close-based rule it does not register as a failure; it never enters the dataset at all, because the close was never at the level. What survives such a filter is price that lingered at the level — which is consolidation, the condition that more often precedes a break than a bounce. So a close-and-clock design systematically discards clean rejections and retains the setups most likely to fail. If that is right, the studies' near-chance reversal rates are partly a prediction of the trader's model rather than evidence against it: measured on the lingering subsample, a lower bounce rate is exactly what one would expect. Any touch criterion must count intraday wicks, and must fix its tolerance band in advance rather than widening it to manufacture hits — a wider band raises the bounce rate without improving the trade, as Tsinaslanidis, Guijarro and Voukelatos (2022) show directly.
4.4. A break is treated as a failure, not as information
In every academic design a level either holds or it does not, and a break is scored zero. This discards the structure the ladder is built on. In the durable model the levels form a cascade: a break of one rung is not a failure but a transition that updates the distribution over the next. A move that breaks 1.618 and accelerates into 2.20 behaves differently from one that stalls at 1.618, and the framework treats those as different states with different expected continuations. A test that records both as "1.618 did not hold" throws away precisely the information a decision tree over the rungs is designed to use.
4.5. There is no concept of decision utility
Finally, the studies measure the wrong quantity for a trader. "Does price reverse at this level more often than chance" is a well-posed question, and the academic answer to it is sound. But it is not the question capital asks. The tradeable question is conditional and sequential: given a retracement to one level, what is the probability of reaching an extension target before a stop level is hit, and at what reward-to-risk? A level can reverse price only 40% of the time and still be the foundation of a profitable rule if the winners pay several times the losers. None of the cited studies computes anything of this shape, and pooling trending and ranging regimes — in which the ladder predicts different behaviour — drags any unconditional average toward chance by construction.
5. The Decisive Test: Hold the Measurement, Swap the Detector
Every criticism above is, on its own, a degree of freedom — the kind of "they tested it wrong, done our way it works" claim that every failed method makes in its defence. The only way to earn it is to change one thing and measure the effect. So we did.
We took a single measurement — the durable forward-projected ladder, armed when price breaks the 1.272 level of a topping swing, resolved when it either completes at 6.8 or is invalidated by a new high, scored on-ladder within ±5% — and we ran it two ways on the same 25 international indices, against the same coverage-corrected null. The only difference between the two runs is which detector supplies the topping-swing anchor: the durable pivot rule, or Batchelor and Ramyar's own Pagan–Sossounov turning points. Everything else is identical code.
| Anchor detector | Anchors | Engage the ladder | On-ladder rate | Chance null |
|---|---|---|---|---|
| Durable topping-swing | 2,186 | 545 (25%) | 78.2% (CI 74.5–81.4) | 54.4% |
| Pagan–Sossounov (B&R) | 329 | 47 (14%) | 38.3% (CI 25.8–52.6) | 23.3% |
The primary split is not accuracy but applicability. The durable detector produces anchors that engage the ladder a quarter of the time; the Pagan–Sossounov detector, only one time in seven, and on the Dow specifically just four times in thirty-six. The reason is the scale problem from Section 4.1: Pagan–Sossounov's censoring leaves only the great macro up-legs as candidate swings, and the downside ladder projected from a swing spanning an entire bull market is priced so far below the market that price makes a new high before ever reaching the first rung. The anchor decides whether the method can be applied at all, before any question of accuracy arises.
Honesty requires the second half of the table, and it cuts against a triumphant reading. Among the few Pagan–Sossounov anchors that do arm, the on-ladder rate is 38.3% against a 23.3% null — weak, but still positive (p ≈ 0.015), not noise. The structure is faintly visible even through the wrong detector. What the durable detector does is not conjure a signal from nothing; it brings a real but faint regularity into focus, lifting a barely-significant 38% over a low baseline to a strongly-significant 78% over a higher one. That is a more defensible claim than a knockout, and it is what the numbers actually say.
The same lesson appears from the other direction in the broader test programme. Run the ladder on every swing an off-the-shelf zig-zag detector produces — fragmenting each durable anchor into dozens of sub-swings — and adherence falls to about 15% against a 16% null, indistinguishable from chance. Run it on durable anchors and it is 75% against 54%, positive in all 25 markets. The ratios do not change between those two results. Only the swing detection does. That is the disconnect, stated as a measurement rather than a grievance: the swing detector is the load-bearing dependency, and no published Fibonacci study validates the one it uses against the swings a trader draws.
6. What We Hold Ourselves To
The argument of Section 5 imposes a burden on us, not only on the studies. If the detector determines the result, then a detector chosen after seeing the data is worthless — it is the same curve-fitting the studies are rightly warned against. The defence is the one Osler's design models: specify the rule in advance, and test it on data that was not used to build it.
The durable result cited above is pre-registered in that sense. The anchor rule, the 1.272 arming gate, the invalidation and completion conditions, the ladder, and the ±5% tolerance were all fixed before scoring, the detection is mechanical and causal, and the test runs across 25 markets rather than the one it was conceived on. Its audit trail is open: of 1,332 candidate topping swings, 975 never armed — the anti-curve-fitting gate rejecting nearly three-quarters of candidates before they could contribute — leaving 344 resolved games, of which 259 landed on the ladder.
Two gaps in that result are stated rather than hidden. The remap stage that handles trends continuing past the 6.8 terminal is not yet implemented, so completions are logged and set aside rather than extended. And roughly a dozen games that remained open at the end of the data are excluded from scoring; excluding unresolved games is defensible, but the count is disclosed so a reader can judge it. Neither gap is large enough to move the headline, and both are named so that no one has to discover them.
This is the standard the academic studies set and largely meet, and it is the standard this framework has to meet too. The claim is not that the studies are careless or that their authors are biased. It is narrower, and more durable: that a geometric method is only as good as its anchors, that the anchors these studies use are not the ones traders draw, and that when the anchors are drawn the way the method is actually used — mechanically, in advance, and tested out of sample — the regularity the studies could not find becomes measurable.
Conclusion: The Disconnect Is Real, But It Is Not Where the Folklore Says
The gap between the academic literature and the trading floor is genuine, but it is not a gap between rigorous scientists and superstitious chartists. On the evidence, the studies are competent and their negative results reproduce. The gap is that the studies and the traders are not looking at the same thing. The studies test whether the sizes of successive macro-cycle trends relate by Fibonacci ratios, scored on whether price reverses at a level within a fixed window. Traders fix a durable anchor, project a ladder forward, count fast wicks as reactions, read a break as a change of state rather than a failure, and care about reward-to-risk rather than reversal frequency. Almost none of that is what the studies measure.
We have tried to show this the hard way: by reproducing the strongest study faithfully, by keeping its honest negative result intact, and by demonstrating on its own turning points that the result turns on the swing detector rather than on Fibonacci. The conclusion is not that the academics are wrong about what they measured. It is that what they measured is not what traders do, and that the difference is large enough to reverse the finding. A method defined relative to reference points can only be tested against the reference points it is actually used with — and that test has, until now, not been run.