236 articles 4 sections last published 2026-09-09 independent · no sponsored placements

ETF Analysis

Sharpe, Sortino, or Calmar? Choosing the Right Risk-Adjusted Metric for Your Goal

The three ratios differ only in what they call "risk" in the denominator: Sharpe uses total volatility, Sortino uses downside volatility, Calmar uses...

Three risk-adjusted performance metrics — Sharpe, Sortino, and Calmar — compared for long-term ETF investors

Photo by Dynamic Wang on Unsplash

The short version

  • The three ratios differ only in what they call "risk" in the denominator: Sharpe uses total volatility, Sortino uses downside volatility, Calmar uses maximum drawdown.
  • Because each answers a different question, a fund can rank first on one and third on another — the ranking is a statement about the measurer's goal, not only the fund.
  • Bottom line: match the metric to your horizon and your dominant fear. Accumulators lean Sortino; anyone near a withdrawal date should weight Calmar heavily.
4.58%10Y Treasury (Sharpe risk-free anchor, FRED 2026-07-14)
16.5VIX regime (FRED 2026-07-14)
3Different denominators, one numerator
36+Months a deep drawdown can take to recover

Every investor eventually meets a fund that looks brilliant on one measure and mediocre on another. The usual reflex is to ask which number is "right." That is the wrong question. Sharpe, Sortino, and Calmar are not competing estimates of the same quantity — they are three different definitions of risk wearing the same algebraic clothing. Choosing among them is really a decision about what kind of pain you are trying to avoid.

This article walks through what each ratio actually measures, where each one quietly misleads, and how the choice should track your horizon rather than your enthusiasm. There is no single winner. There is only a metric that fits your goal and two that answer a question you were not asking.

Context: one numerator, three denominators

All three ratios share the same top half — excess return, meaning the portfolio's return above a risk-free benchmark. With the 10-year Treasury near 4.58% and the effective federal funds rate at 3.63% (FRED, asof 2026-07-14 and 2026-06-01), the choice of risk-free rate is not academic trivia in 2026: a higher cash rate mechanically shrinks every excess-return numerator, and a strategy that looked attractive against a 0.5% cash rate can look ordinary against 4.5%.

What separates the three is the denominator — the definition of "risk" you divide by:

MetricRisk in the denominatorQuestion it answersOrigin
SharpeStandard deviation of all returns (up and down)How much total variability did I endure per unit of excess return?Sharpe, 1966 / revised 1994
SortinoDownside deviation (only returns below a target)How much harmful variability did I endure per unit of excess return?Sortino & Price, 1994
CalmarMaximum drawdown over the windowHow much peak-to-trough loss did I risk per unit of annual return?Young, 1991

The table is the whole argument in miniature. Sharpe penalizes a fund for large gains as heavily as for large losses, because standard deviation is symmetric. Sortino fixes that asymmetry by counting only returns below a threshold. Calmar ignores the shape of the distribution entirely and asks a blunt, path-dependent question: what was the worst single hole this thing dug?

Sharpe: the honest generalist with a symmetry problem

The Sharpe ratio is the default for a reason. It is simple, it is comparable across asset classes, and William Sharpe's 1994 revision made the definition rigorous. When you read a factsheet's "risk-adjusted return," it is almost always Sharpe.

Its weakness is philosophical, not computational. Standard deviation treats a +8% month and a −8% month as equally undesirable. No long-term investor actually believes that. Upside dispersion is the entire point of holding equities; punishing it is a category error. The practical consequence: a fund with occasional explosive positive months — many trend-following and factor strategies behave this way — will show a depressed Sharpe that overstates how uncomfortable it really was to hold. Sharpe also assumes returns are roughly normal. Real ETF returns have fat tails, so Sharpe tends to understate the probability of the rare, portfolio-defining loss. The academic literature on downside risk, running back to Sortino's work in the 1990s, exists largely to address this blind spot.

Sharpe, Sortino, and Calmar are not competing estimates of the same quantity — they are three definitions of risk wearing the same algebraic clothing.

Sortino: penalize the pain, not the payoff

Sortino replaces total volatility with downside deviation — the dispersion of returns that fall below a minimum acceptable return (often 0%, sometimes the risk-free rate). Everything above the threshold is treated as harmless, which is closer to how a human actually experiences a portfolio.

For an accumulator running dollar-cost averaging over decades, Sortino is usually the more faithful lens. You are adding cash monthly, you are not selling, and upside volatility is a feature you are being paid to accept. The non-obvious catch is statistical, not conceptual: because downside deviation throws away roughly half the observations, it is estimated from a smaller effective sample. That makes Sortino noisier and easier to overfit than Sharpe. Rank forty funds by Sortino over a three-year window and a meaningful share of the top ranks will be luck — the same data-mining hazard that haunts any metric computed from a short, single-regime history. I initially treated a high trailing Sortino as strong evidence of skill; running it across rolling windows, the ordering reshuffled far more than a Sharpe ranking did, which is exactly what you would expect from an estimator built on half the data.

Calmar: the metric that remembers the path

Calmar is the outlier. It divides annualized return by maximum drawdown — the largest peak-to-trough decline over the measurement window. It says nothing about the shape of month-to-month returns and everything about the deepest hole. That makes it uniquely relevant to a real behavioral fact: investors do not abandon plans because volatility was high. They abandon plans because they watched a balance fall 45% and could not stomach the wait to recover.

Drawdown has a second dimension the ratio hides — duration. A 30% drawdown that recovers in six months and a 30% drawdown that grinds sideways for three years produce the same Calmar denominator but very different lived experiences, and very different outcomes if a withdrawal lands mid-trough. This is precisely where Calmar connects to sequence-of-returns risk: for someone drawing down a portfolio, the worst drawdown and when it arrives dominate the arithmetic mean return entirely. Calmar's weakness is fragility. Maximum drawdown is a single worst-case observation, so the whole ratio can hinge on one week in one crisis. Shorten or lengthen the window to include or exclude 2020 or 2022 and the number can move dramatically — a look-back-window sensitivity that makes cross-fund Calmar comparisons treacherous unless the windows are identical.

How the choice should track your horizon

The metric is downstream of the goal. A young accumulator adding to a broad equity position each month cares about long-run compounding and can treat drawdowns as buying opportunities; total volatility is close to irrelevant and Sortino is the natural default. A soon-to-be-decumulator cares almost entirely about the worst path, because a deep drawdown in the first years of withdrawals can be unrecoverable; Calmar and drawdown duration should dominate. Sharpe sits in the middle as the common language for comparing dissimilar strategies on a single, if imperfect, scale.

A discipline worth adopting: never rank on one ratio alone. Compute all three, over identical windows, and read the disagreements. When a fund's Sharpe and Sortino diverge, its return distribution is asymmetric — worth understanding before you hold it. When Calmar disagrees with both, the fund's risk lives in the tail and the path, not the day-to-day wobble. The disagreement is the signal. If you want to pressure-test these numbers on your own holdings rather than trust a factsheet, the mechanics of a reproducible backtest are covered in this walkthrough on backtesting an ETF strategy, and the broader question of comparing overlapping funds honestly is treated in this framework for measuring ETF overlap.

Scoreboard: which metric wins for which job

CategoryBest fitWhy
Comparing dissimilar strategiesSharpeUniversal, symmetric, widely reported — the common denominator.
Long-horizon accumulation (DCA)SortinoRewards upside, penalizes only harmful volatility.
Approaching or in withdrawalCalmarCaptures the worst path, which sequence risk makes decisive.
Robustness / least overfitSharpeUses the full sample; less noisy than half-sample Sortino or single-point Calmar.

Frequently asked questions

Is a higher Sharpe ratio always better? Higher is better only if variability is the risk you care about and returns are roughly symmetric. For a strategy with large positive outliers, Sharpe understates its appeal; for one with fat left tails, it can overstate its safety. Read it alongside Sortino and Calmar.

What is a "good" Sharpe ratio? There is no universal threshold, and any specific number is regime-dependent. A Sharpe computed in a low-volatility, rising market will flatter a fund; the same fund measured across a crisis will look worse. Compare funds over identical windows rather than against a folk-wisdom cutoff.

Why would Sortino and Sharpe disagree? They disagree when the return distribution is asymmetric. A fund with frequent modest gains and rare sharp losses will have a Sortino close to its Sharpe; one with occasional large upside spikes will show a Sortino meaningfully above its Sharpe, because Sharpe penalizes those spikes and Sortino does not.

Does the risk-free rate really change the ranking? It changes the level of every ratio and can change rankings at the margin, because all three subtract the same risk-free rate in the numerator. With cash near 4.5% (FRED, 2026-07-14), low-return, low-volatility funds see their excess return shrink toward zero, compressing their ratios more than higher-returning peers.

Which single metric should a beginner use? If you must pick one, Sharpe is the most robust and the most comparable across sources. But the more useful habit is to compute all three over the same window and treat their disagreement as information about the fund's risk shape.

What this comparison can and can't tell you

These ratios are backward-looking summaries of a single realized history. They cannot tell you whether the next drawdown will resemble the last one, and none of them captures liquidity risk, fund-closure risk, or the tax-cost drag that erodes an active strategy's realized return. Every figure here is regime-dependent: a window that omits 2020 or 2022 produces very different Sortino and Calmar values than one that includes them. Treat all three as descriptions of what happened, not forecasts of what will.

Scenarios where each metric fits

Reader in their 30s, 401(k)-only, contributing monthly, no plans to sell for decades → weight Sortino, largely ignore short-window Calmar noise. Reader within five years of drawing income from the portfolio → weight Calmar and drawdown duration, because a bad early path is the dominant threat, an idea developed further in the work on protecting a 30-year plan from a pre-retirement crash. Reader comparing a factor ETF against a plain index fund → start with Sharpe for a common scale, then check whether Sortino tells a kinder story about the factor's asymmetric payoff.

Editor's read

If forced to privilege one lens for a long-horizon core, the editor leans on Sortino for the accumulation years and shifts weight to Calmar as a withdrawal date approaches — but never reads either in isolation. The most useful practice is not choosing a favorite ratio; it is computing all three over identical windows and paying attention to where they disagree, because that disagreement is a compact description of a fund's true risk shape. A single number that flatters a strategy is usually hiding the question it declined to ask.

Holdings disclosure: this article discusses a methodology rather than specific funds; the editor holds no position tied to the metrics themselves at the time of writing.

Methodology: macro figures (10-year Treasury, effective federal funds rate, VIX, CPI) are from FRED, with as-of dates cited inline (latest 2026-07-14). Metric definitions follow Sharpe (1966, revised 1994), Sortino & Price (1994), and Young (1991). No fund-level return series was used in this conceptual piece; any application to specific ETFs should recompute all three ratios over identical, clearly stated windows.

Key takeaways

  • The three ratios share a numerator and differ only in how they define risk: total volatility (Sharpe), downside volatility (Sortino), maximum drawdown (Calmar).
  • Sharpe is the most robust and comparable but penalizes upside and assumes near-normal returns; it can understate tail risk.
  • Sortino fits long-horizon accumulators but is noisier because it uses roughly half the sample — easy to overfit on short windows.
  • Calmar is decisive for anyone near withdrawal, because it speaks to sequence risk, but it hinges on a single worst-case observation and is highly window-sensitive.
  • Never rank on one ratio alone; compute all three over identical windows and read the disagreements as information.

This article is for educational purposes and does not constitute personalized financial advice. See our full Disclaimer.