Three classic hedge-fund arbitrage strategies, each simulated from the mathematics that defines it: statistical arbitrage on a mean-reverting spread, volatility arbitrage through variance swaps and dispersion, and merger arbitrage on deal spreads — ending with the short-convexity payoff shape all three share.

Module 04 · Strategies, simulated

The Arbitrage Book

"Arbitrage" almost never means free money. It means a spread that should close, a reason to believe it will, and a position sized to survive the times it doesn't. Three of the biggest strategies in the industry, each reduced to the equation that drives it — and each run until it breaks.

Same rules as the rest of the course: every coloured, underlined term, , — opens a full explanation. Each strategy has a set of dials and a simulation behind it; the last section shows why all three die the same way.
01 — Statistical Arbitrage

Two stocks on a rubber band

Find two assets that move together for a structural reason — same industry, same input costs, two share classes of one company. Their price difference, the , wanders but tends to snap back; the statistical name for that tie is . Model it as an : sell it when it's stretched high, buy it when it's stretched low, and wait for the rubber band. The strategy's entire character comes from one number — the of that snap-back.

d = ( − X)dt + ·dW      = ln2 / θ
Annualised
1.42
net of costs
Round trips / year
12
≈ 21d per trade
Win rate
78%
trades closed in profit
Cost drag
−0.2%/yr
what the spread eats
Two years of spread · shaded = position open · dashed = entry bands
Sharpe vs. entry band — too tight pays the spread, too wide never trades

The half-life is the strategy. It sets everything downstream: how long capital is tied up, how many independent bets you get per year, and therefore — through the fundamental law — your information ratio. A 3-day half-life gives you a hundred-odd shots a year and a Sharpe that survives costs; a 90-day half-life gives you four, and the same edge per trade is no longer a business. Two pairs with identical statistical significance can be, in practice, one strategy and one waste of a seat.

Try this
DoSet the half-life to 3 days, then push costs from 0 to 60 bp.
WatchThe fastest, most profitable-looking version of the strategy is the one costs destroy first — and the Sharpe-vs-band curve slides right, telling you to trade less often.
WhyFast reversion means small moves and many trades. Your gross edge per trade shrinks toward the cost of crossing the spread, and the strategy quietly converts from an alpha engine into a fee-paying machine. This is why stat arb is a technology business: whoever trades cheapest can harvest reversion nobody else can reach.
02 — Divergence Risk

The rubber band that turned out to be string

Mean reversion is an assumption, not a law. Companies get acquired, business models diverge, an index rebalances, a short squeeze arrives — collectively, . When a pair stops reverting, the strategy's own logic makes things worse: as the spread widens the model reads it as a better opportunity, and a disciplined system adds. Here's what that does to the distribution of outcomes.

Median year
+6.1%
the typical outcome
Mean year
+4.2%
dragged down by the tail
Worst 5%
−8.4%
5th percentile
Worst run seen
−31%
of 800 simulated years
Annual return across 800 simulated pairs · note the left tail

Look at the shape, not the average. The median is comfortably positive and most years are boring — that's the point of the strategy and the reason it attracts capital. The damage lives in a thin left tail that a year of data will probably never show you. This is : many small gains funded by rare large losses. It is not a flaw in the implementation. It is the trade — you are being paid to absorb the risk that a relationship ends, and sometimes relationships end.

Try this
DoTurn the stop-loss on at . Then tighten it to 2.5σ.
WatchAt 4σ the worst run more than halves while the median barely moves — nearly free insurance. At 2.5σ the median halves as well.
ThenPut the stop back to 4σ and push the gap to 15σ. The stop still helps, but a large part of the loss now arrives before it can fire.
WhyA stop set far from the entry band only ever triggers on genuinely broken pairs, so it costs almost nothing. Move it close and it starts firing on healthy trades that would have reverted — you are paying for the insurance in median return. And no stop escapes the gap itself: a stop is an instruction to sell at a price, while a gap is the market never trading there. You get filled on the far side of the hole. It is the identical failure that jumps inflict on a delta hedge — any control that has to trade during the move is defeated by a move with no during.
03 — Volatility Arbitrage

Selling insurance on movement

Section 08 of the derivation showed that a delta-hedged option pays you the gap between and volatility. Do that deliberately, at size, and you have the volatility-arbitrage business. The cleanest expression is a , which pays the difference in variance — and that squaring is the whole story.

P&Lshort = · (² − ²)      Nvar = vega notional / (2K)
Average month
+$29k
per $100k vega
Months in profit
86%
of 3,000 simulated
Worst 1% of months
−$4.8M
= 23 average months
Break-even vol
20.0%
above this you lose
P&L vs. realised vol — linear in variance means quadratic in vol
3,000 months of selling variance · the tail is the whole risk

Why the parabola is the point. A variance swap is linear in variance, so it's quadratic in volatility. Vol going from 20% to 25% costs you a little. Vol going from 20% to 60% costs you nine times the squared distance — the loss curve accelerates exactly when everything else in the fund is also going wrong. Sellers of variance in February 2018 learned this in a single afternoon when a short-vol product lost most of its value overnight. The mathematics gave no warning because the mathematics was never wrong; the position was simply short a parabola.

: the same trade, pointed at correlation. An index moves less than its members because they don't move together — so index variance is a bet on correlation. Sell index vol, buy single-name vol, and you're short with no directional view at all.
σindex² = σ²·[ 1/n + (1 − 1/n)· ]
0.35
what the market is paying
Your edge
+0.05
implied minus expected
Fair index vol
16.9%
at your correlation
Verdict
sell index vol
buy the single names

Implied correlation above what you expect means the index is expensive relative to its parts — the classic dispersion trade. Drag correlation to 1.00 and watch the index vol converge on the single-name vol: with everything moving together, diversification is gone and the index is just one big stock.

Try this
DoSet crash months to 0%, note the average and the 1-in-100 month. Now set it back to 2%.
WatchTwo percent of months costs about a third of the average profit — and multiplies the bad month roughly fivefold.
ThenSet crashes back to 0% and drag vol-of-vol to 90%. Average realised vol never changes, yet the strategy goes from profitable to losing money.
WhyBoth effects are the same parabola. Because the payoff is quadratic, spread in the vol distribution costs you even when its average is untouched — losses on the high side outrun gains on the low side. And 2% of months is one month every four years: comfortably longer than most track records, most investor patience, and most bonus cycles. A strategy can look flawless across its entire measured history while its defining risk simply hasn't happened yet, which makes the backtest trap and this page the same lesson told twice.
04 — Merger Arbitrage

Picking up the last two dollars

Company A offers $50 a share for Company B. B trades at $48. That is the market's price for two risks — the deal , and your money is tied up until it does. Buy the target, and you are selling insurance against deal failure. The arithmetic is simple enough to do in your head, which is exactly what makes the risk easy to underestimate.

= (P − D) / (V − D)      annualised = (V − P)/P · 365/
83%
what the price is saying
Annualised if it closes
+17%
gross spread 4.2%
Loss if it breaks
−21%
5.0× the gain
Expected value
+0.4%
at your 85% estimate
The whole payoff, to scale — a small step up, a cliff down
Now hold twenty of them. Deal risk looks idiosyncratic — regulators, financing, a shareholder vote — so a portfolio of unrelated deals should diversify beautifully. It does, right up until the thing that breaks deals is the same thing for all of them: credit freezes, markets fall, buyers walk. Add a and watch the diversification evaporate.
Median year
+16%
Worst 1%
−34%
1-in-100 years
Years losing > 20%
4.1%
of 2,000 simulated
Diversification
weak
vs. independent deals
Annual return · gold = deals break independently, blue = a shared crisis factor
Try this
DoHold 60 deals with the crisis chance at 0%, then turn the crisis back to 8%.
WatchWith independent deals, sixty positions nearly erase the risk. Add one shared factor and sixty deals behave more like three.
WhyDiversification divides idiosyncratic risk by n; it does nothing to a common factor. Your is set by the number of independent risks you hold, not the number of tickers. This is the single most-repeated way funds discover they were less diversified than the position list suggested.
05 — The Family Resemblance

Three strategies, one payoff

Put the three side by side and the differences turn out to be cosmetic. Each collects a small, reliable premium; each occasionally pays out something enormous; each is, at bottom, a way of being paid to carry a risk other people want to hand over.

StrategyYou are paid forWhat ends it
Statistical arbitrageProviding liquidity to whoever needed to trade urgentlyThe relationship was structural until it wasn't — a takeover, an index change, a business that genuinely diverged
Volatility arbitrageInsuring other people against movementOne month where movement arrives all at once, and the loss is quadratic in the size of the move
Merger arbitrageAbsorbing the risk that a deal failsA credit event that breaks many deals at once, exactly when your other positions are also falling
All threeSelling insuranceThe correlated tail — the thing that makes every position fail together

Short convexity is not a mistake, it's a job. Someone has to hold these risks, and the premium is real compensation for real exposure. The failure mode isn't running the strategy — it's mistaking the premium for skill, levering it up because the measured volatility looks low, and discovering that measured volatility never contained the event that defines the position. Every metric on the Desk — Sharpe, VaR, drawdown — is estimated from the sample you happen to have. Short-convexity strategies are precisely the ones whose defining risk is usually missing from that sample.

Where to go next
Size itBet sizing is the defence: the Kelly maths assumes you know the distribution, and these strategies are the ones you know least well. Quarter-Kelly exists for this.
Decompose itFactor models ask the harder question — how much of this premium is a risk factor anyone can harvest, and how much is genuinely yours?
Price itThe hedging lab is where the volatility premium comes from in the first place.