Sharpe Ratio Limitations: Why I Promote on Calmar, Not Sharpe
A candidate in my Titan system posted a Sharpe near 1.4, cleared every statistical gate, and still traced a drawdown so deep and so long that I could never have held it through the trough. Here is why Sharpe is blind to the path and the tail, the survival-metric suite (Sortino, Calmar, CVaR, CDaR) that sees what it misses, and the one rule it bought me: promote on Calmar lift, not Sharpe lift. All figures are illustrative and sanitised.

One of the best-looking strategies I ever built for Titan, my live multi-strategy system, would have quietly bankrupted my nerve. It posted a Sharpe ratio around 1.4. It cleared every statistical gate I had: the deflated-Sharpe check against the whole search pool came back positive, the serially-aware bootstrap put the lower bound comfortably above zero, the causality test passed end to end. On paper it was a clean promote, the kind of result you screenshot. Then I looked at the one thing Sharpe cannot see, the path, and the promotion died on the spot.
Its worst peak-to-trough drawdown ran deep into the double digits and stayed underwater for well over a year before it clawed back to a new high. I want to be precise about why that killed it, because it is not the obvious reason. The strategy was not bad. Over the full sample it made money, and the Sharpe was honest. The problem is that no human, and certainly no risk committee, sits in a hole that deep for that long without switching the thing off somewhere near the bottom. And the moment you switch it off in the trough, you have converted a paper drawdown into a realised loss and forfeited the recovery that made the Sharpe respectable in the first place. The maths said hold. Every instinct I have, and every instinct of anyone I would ever answer to, said sell. When the metric and the holder disagree, the holder wins, because the holder has the off switch.
That is the whole problem with optimising Sharpe, in one war story. Sharpe is blind to the two things that actually end accounts: the path of the drawdown and the tail of the loss. Most people, myself included for too long, spend their search budget maximising a number that is silent on the only questions that decide whether you can stay in the trade.
One number cannot describe a path
Here is the thing to internalise before any formula. Sharpe is a property of the return distribution: mean over standard deviation. It throws away the order of the returns entirely. But order is exactly what a human experiences, what a margin engine reacts to, and what a risk committee votes on.
Chart
Two paths, identical Sharpe
Both curves end at the same value with the same volatility, so their Sharpe is identical near 1.4. One you hold without a thought; the other you abandon in the trough. Sharpe never looked at the shape. Illustrative and sanitised.
Picture two equity curves that end at the same place with the same volatility. One climbs as a steady ramp. The other earns the identical return by cratering forty percent and then heroically recovering. Same mean. Same standard deviation. Identical Sharpe. They are not remotely the same asset to own. One you hold through your holidays without a thought. The other you abandon in month nine of the drawdown, right before it turns, and you never see the recovery that the backtest so proudly averaged over. Sharpe rates them equal because it never looked at the shape. The shape was the entire story.
So the discipline is to stop reporting a scalar and start reporting a vector. Each component answers a different question, a candidate has to clear all of them, and they fail independently, which is the point. I call them the survival metrics, because that is the question they speak to: not "how good is this on average" but "will I still be holding it when the average arrives".
The survival suite: what each number sees that Sharpe cannot
Figure
What each survival metric sees that Sharpe cannot
Four metrics around Sharpe, each aimed at a blind spot, each failing independently.
Sortino
downside-only volatility
Sharpe punishes good surprises, because standard deviation is symmetric. Sortino counts only downside deviation, so a big win stops registering as risk. If Sortino sits below Sharpe, your risk is skewed downward.
Calmar
return per worst drawdown
CAGR over the maximum drawdown. The denominator is the exact quantity that gets a strategy killed mid-trough, which is why it is my primary promotion metric. Sharpe ignores the path entirely.
CVaR
the average of the worst tail
Value-at-Risk reads one percentile; CVaR averages everything beyond it, which is where ruin lives. A tame CVaR means no single day is catastrophic.
CDaR
the average of the deepest drawdowns
Max drawdown shows one hole; CDaR averages the deepest ones. A tame CVaR with an ugly CDaR is the signature of the slow-grind drawdown that is hardest to hold.
Compute all of them for every candidate; refuse to let a strong Sharpe paper over a weak survival number.
Four numbers surround Sharpe in my report, each aimed at a specific blind spot.
Sortino fixes the daftest thing Sharpe does: it punishes your good surprises. Standard deviation is symmetric, so a sharp gain moves it exactly as much as a sharp loss. A strategy whose volatility comes from occasional big wins gets marked down for the crime of making money in lumps. Sortino swaps the denominator for downside deviation, computed only over returns below zero, so upside stops counting as risk. The Sortino-minus-Sharpe gap is itself a diagnostic I read every time: if Sortino sits well above Sharpe, the volatility is mostly upside and that is fine. If Sortino sits below Sharpe, your risk is disproportionately to the downside, and the smooth headline number is hiding a left-skewed return stream. That gap has talked me out of more deployments than almost anything else.
Calmar is the one that matters most, and it is return divided by the worst drawdown: CAGR over the absolute maximum drawdown. It is the most honest single-number summary of livability, because the denominator is the exact quantity that gets a strategy killed mid-trough. My 1.4-Sharpe candidate had a perfectly nice numerator and a denominator that told the truth the Sharpe had buried.
CVaR and CDaR handle the tail. Value-at-Risk and max-drawdown each report a single point: the loss at one percentile, the single deepest hole. Neither tells you anything about what lies beyond that point, which is exactly where ruin lives. CVaR (expected shortfall) averages the whole bad tail of bars instead of reading one threshold off it, and CDaR does the same to the drawdown curve. I read them as a pair. A strategy can have a tame CVaR (no single day is catastrophic) and an ugly CDaR (it bleeds in long correlated runs). That specific combination is the signature of the slow-grind drawdown, the kind that is hardest to hold and easiest to switch off at the worst possible time. It is, in other words, precisely the shape of the candidate that nearly fooled me.
None of these four is exotic. What is load-bearing is that I compute every one of them for every candidate and refuse to let a strong Sharpe paper over a weak survival number.
The trap inside Calmar that made the war story worse
There is a subtle way Calmar itself will lie to you, and it bit me on the same strategy. Calmar needs an annualised return in the numerator, and the convenient way to get one is the arithmetic mean of the bars times the periods per year. That number is wrong for anything that gates real capital, because the arithmetic mean overstates what you actually compound by roughly volatility-squared over two per year: the volatility drag. A minus-fifty-percent bar needs a plus-one-hundred-percent bar to undo it, and the arithmetic mean serenely ignores that.
Let me re-derive the gap freshly, with round illustrative numbers so you feel the direction. Take a volatile candidate that draws down thirty-five percent over a multi-year window. Read its return arithmetically and you might see eighteen percent a year, so the quick Calmar is 0.18 over 0.35, about 0.51. Respectable. It clears a lift gate against an incumbent sitting near 0.40. But the volatility drag means the geometric CAGR, the rate you genuinely compounded, is more like eleven percent, so the true Calmar is 0.11 over 0.35, about 0.31. That is below both the incumbent and the gate. Same strategy, same hole, opposite verdict. The arithmetic mean would have promoted a strategy the geometric truth rejects, and it flatters the volatile candidate every single time, never the reverse. So I keep a fast arithmetic Calmar for eyeballing diagnostics and a geometric one, built from the actual equity curve's terminal value, for anything that decides money. The full derivation and the two deliberately separate implementations are in the chapter; the rule you keep is just: never gate a promotion on an arithmetic Calmar.
The rule it bought me: Calmar lift, not Sharpe lift
Here is the portable artefact, the one thing to carry out of this piece. In Titan, a new strategy is never promoted on its standalone metrics at all. It is promoted on whether it makes the existing portfolio better, and the primary verdict is Calmar lift: the change in the combined book's return-per-drawdown, measured relative to what I already hold. Sharpe lift is demoted to a secondary, no-regression check, gated only at "must not get worse".
The asymmetry is the entire discipline:
- Calmar lift is the primary gate, with a real positive threshold. A candidate has to raise the whole book's return-per-drawdown by a meaningful margin, not just nudge it.
- Sharpe lift is secondary, gated only at zero or better. It may not regress the book's return-per-volatility, but improving it earns nothing on its own.
So a candidate that raises Sharpe while flattening Calmar fails. A candidate that raises Calmar without hurting Sharpe passes. I deliberately put the burden of proof on the metric that maps to getting switched off mid-trough, not the one that maps to looking smooth on a monthly return chart. A strategy that improves your average and worsens your worst is not an improvement. It is a future kill-switch event with good marketing, and I have the war story to prove I nearly bought one.
One honest caveat so you do not over-trust the rule. Calmar lift and Sharpe lift are the framework-level slice of the gate, not the whole thing. Full promotion in my system also has to clear a joint risk-of-ruin bound and a Monte-Carlo constraint on the probability of a max drawdown past a threshold I pre-register. The metric suite proposes; the survival math disposes. But if you take one change back to your own pipeline today, make it this: stop ranking candidates by Sharpe lift and start ranking them by drawdown-adjusted lift. It reorganises everything downstream, because it forces every promotion to answer the question your future self will actually ask at the bottom of the hole.
Every survival number obeys the same lies
A closing warning, because it is the one people skip. Surrounding Sharpe with four more metrics buys you nothing if the metrics are computed on a corrupt equity curve. A Calmar built on a look-ahead curve is exactly as worthless as a look-ahead Sharpe, and arguably more dangerous, because "we also checked drawdown" manufactures false confidence. Wrong annualisation units, dropped flat bars, a same-bar return leak, a full-series normaliser, a point estimate with no error bars: every one of those corrupts Calmar, Sortino, CVaR and CDaR just as thoroughly as it corrupts Sharpe. The suite does not launder a bad backtest. It just measures, faithfully, whatever you feed it. Centralise the whole battery in one audited module so the unsafe version cannot be written, and write down, at the point of use, exactly which question each estimator is allowed to answer.
The full chapter, Beyond Sharpe: the metric suite, is free to read. It works through each metric with the code, the geometric-CAGR trap in full, the promotion gate that combines them, and the honest limitations I document rather than hide. This essay is the argument for the posture; that chapter is how you implement it, and it is part of Building a Production Quant Trading System; the complete book, a living digital copy on Leanpub and a print paperback on Amazon, carries the paid half on sizing, portfolio construction and running the system live. If these field notes are useful, the newsletter is where the next one lands.
Carry one sentence out of here: optimise the metric that maps to getting switched off at the bottom, not the one that maps to looking smooth on the way there.
This is an engineering essay, not investment advice, and it contains no tradable strategy. All figures are illustrative and sanitised, and every war story is about a bug or a loss, never a profit.
Chart
Two paths, identical Sharpe
Both curves end at the same value with the same volatility, so their Sharpe is identical near 1.4. One you hold without a thought; the other you abandon in the trough. Sharpe never looked at the shape. Illustrative and sanitised.
Figure
What each survival metric sees that Sharpe cannot
Four metrics around Sharpe, each aimed at a blind spot, each failing independently.
Sortino
downside-only volatility
Sharpe punishes good surprises, because standard deviation is symmetric. Sortino counts only downside deviation, so a big win stops registering as risk. If Sortino sits below Sharpe, your risk is skewed downward.
Calmar
return per worst drawdown
CAGR over the maximum drawdown. The denominator is the exact quantity that gets a strategy killed mid-trough, which is why it is my primary promotion metric. Sharpe ignores the path entirely.
CVaR
the average of the worst tail
Value-at-Risk reads one percentile; CVaR averages everything beyond it, which is where ruin lives. A tame CVaR means no single day is catastrophic.
CDaR
the average of the deepest drawdowns
Max drawdown shows one hole; CDaR averages the deepest ones. A tame CVaR with an ugly CDaR is the signature of the slow-grind drawdown that is hardest to hold.
Compute all of them for every candidate; refuse to let a strong Sharpe paper over a weak survival number.
Further reading
- The Deflated Sharpe Ratio: Why Your Grid-Search Winner Is Probably Noise
Every survival metric obeys the same lies as Sharpe: a Calmar built on an overfit or leaked curve is just as worthless, and just as flattering.
- Suspicion Over Celebration: Inside "Building a Production Quant Trading System"
Beyond Sharpe, the metric suite, is one chapter of the free half; the book review lays out the full guide.
Related posts
Risk of Ruin Monte Carlo: Resample the Cause, Not the Effect
Most risk-of-ruin Monte Carlo resamples a strategy's realised P&L, which quietly bakes in the good luck you are trying to stress and understates the tail. Here is the correction that made my honest drawdown distribution far fatter than my first one, plus the relative gate a long-only sleeve actually needs. First person, British English, illustrative numbers only.
Walk Forward Optimization Is Not Automatically Out of Sample
Everyone sells walk-forward validation as proof a strategy is out-of-sample. It is not. I ran my own walk-forward pipeline on pure random-walk noise and it still produced a healthy positive stitched Sharpe, because out-of-sample is a property of provenance, not partitioning. Here is the random-walk control that tells you whether your pipeline, or your edge, produced the number.
Look-Ahead Bias in a Backtest: The Corrupt-the-Future Test That Catches It
Look-ahead bias is the default state of careless backtest code, not an exotic edge case, and it survives review because it reads as ordinary pandas. Here is the corrupt-the-future causality test I now gate on: poison every price after a date, then assert nothing computed before it moves. If the past shifts when you poison the future, the strategy is reading ahead.