Kyle Wisniewski

Every result here is reproducible: open-source code and 229 automated checks, from theory-anchored quantitative tests to publishing checks. Browse the source on GitHub →

Technical foundations · Mathematical notes

Mathematical foundations for pricing, simulation, and tail risk

Six derivations cover the stochastic calculus, pricing, simulation-error, and tail-risk results the lab relies on.

Revised · citations refer to the editions listed below

Reading time
Approximately 16 minutes for the complete set
Prerequisites
Single-variable calculus, probability, and basic derivative pricing

Decision summary

Use the mathematics to expose assumptions, not conceal them

Interpretation

The six notes connect stochastic calculus, pricing, simulation, and tail risk into one chain. Every result in that chain depends on a stated probability model and an error budget.

Decision context

The derivations show when a pricing identity, simulation estimate, or risk metric applies. They also name the assumption or approximation that limits the conclusion.

Intended analytical use

Quantitative-finance students, analysts moving into model work, technical reviewers, and practitioners who need an auditable conceptual reference.

Prerequisites

The collection assumes calculus, probability, and basic derivative pricing. It is a working reference, not a substitute for the cited treatments.

Contents · six notes

Note 01

Itô's Lemma#

4 min · Prerequisites: multivariable calculus and Brownian motion

Ordinary calculus fails for Brownian motion because \(W_t\) accumulates quadratic variation: over a partition of \([0,t]\), \(\sum (\Delta W)^2 \to t\), not zero. The squared increments of a Brownian path are not negligible — they behave, in the limit, like time itself.1 The heuristic multiplication table is

$$(dW)^2 = dt, \qquad dW\,dt = 0, \qquad (dt)^2 = 0.$$

Take an Itô process \(dX_t = \mu_t\,dt + \sigma_t\,dW_t\) and a smooth function \(f(t, x)\). A second-order Taylor expansion gives

$$df = \frac{\partial f}{\partial t}\,dt + \frac{\partial f}{\partial x}\,dX + \tfrac{1}{2}\frac{\partial^2 f}{\partial x^2}\,(dX)^2 + \cdots$$

In ordinary calculus the \((dX)^2\) term would vanish. Here \((dX)^2 = \sigma_t^2\,(dW)^2 + O(dt^{3/2}) = \sigma_t^2\,dt\), so a second-order term survives into the first-order differential:

$$df = \left(\frac{\partial f}{\partial t} + \mu_t\frac{\partial f}{\partial x} + \tfrac{1}{2}\sigma_t^2 \frac{\partial^2 f}{\partial x^2}\right) dt + \sigma_t \frac{\partial f}{\partial x}\,dW_t.$$

That extra \(\tfrac{1}{2}\sigma^2 f_{xx}\,dt\) — the Itô correction — is the single most consequential term in mathematical finance. It is why convexity has a price, why hedged option books bleed or earn theta, and why the drift of a log-price is not the drift of the price. Whenever a payoff is curved and the underlying is volatile, the correction term is where the money is.

Note 02

Solving Geometric Brownian Motion#

3 min · Prerequisite: Note 01, Itô's Lemma

The SDE \(dS_t = \mu S_t\,dt + \sigma S_t\,dW_t\) is solved by applying Itô's lemma to \(f(S) = \ln S\), for which \(f' = 1/S\) and \(f'' = -1/S^2\):

$$d(\ln S_t) = \frac{1}{S_t}\,dS_t - \frac{1}{2}\frac{1}{S_t^2}(dS_t)^2 = \left(\mu - \tfrac{1}{2}\sigma^2\right)dt + \sigma\,dW_t.$$

The right-hand side no longer involves \(S_t\); it integrates directly:

$$S_t = S_0\,\exp\!\left[\left(\mu - \tfrac{1}{2}\sigma^2\right)t + \sigma W_t\right].$$

Log-prices are Gaussian; prices are lognormal; prices stay positive. Two readings of the \(-\tfrac{1}{2}\sigma^2\) term repay attention. First, \(\mathbb{E}[S_t] = S_0 e^{\mu t}\) — the correction exactly offsets the convexity of the exponential, by design. Second, the median path grows at \(\mu - \tfrac{1}{2}\sigma^2\), strictly less than the mean growth rate. A volatile asset's average outcome is dragged upward by a shrinking minority of enormous paths while the typical path does worse. This "volatility drag" is not a market imperfection; it is arithmetic, and it is the quantitative core of why compounding punishes variance.

Note 03

Risk-Neutral Pricing: Assumptions and Derivation#

4 min · Prerequisites: discounting, replication, and binomial trees

Consider one period and two assets: a bond growing at \(r\), and a stock worth \(S\) that moves to \(uS\) or \(dS\) with \(d < e^{r\Delta t} < u\). To price a claim paying \(V_u\) or \(V_d\), build a portfolio of \(\Delta\) shares and \(B\) in bonds that replicates it in both states:

$$\Delta = \frac{V_u - V_d}{(u-d)S}, \qquad V = \Delta S + B = e^{-r\Delta t}\left[\,p^* V_u + (1-p^*) V_d\,\right], \qquad p^* = \frac{e^{r\Delta t} - d}{u - d}.$$

The real-world probability of the up-move never entered. The price is a discounted expectation under an artificial probability \(p^*\) — the unique one that makes the discounted stock a martingale.2 This is the whole content of risk-neutral pricing: no arbitrage plus replication implies prices are expectations under a measure \(\mathbb{Q}\) constructed for accounting convenience, not belief. In continuous time Girsanov's theorem plays the same role, shifting the drift of \(W_t\) so that \(dS_t = r S_t\,dt + \sigma S_t\,dW_t^{\mathbb{Q}}\), and

$$V_0 = e^{-rT}\,\mathbb{E}^{\mathbb{Q}}\!\left[\,\text{payoff}(S_T)\,\right].$$

The derivation requires that the payoff can be replicated under complete markets with continuous frictionless trading. With jumps, stochastic volatility, or transaction costs, replication is imperfect, \(\mathbb{Q}\) is no longer unique, and the arbitrage-free price can become an interval. The selected measure then depends on market prices for non-replicable risk.

Note 04

Feynman-Kac and the Black-Scholes PDE#

4 min · Prerequisites: Note 01, Note 03, and PDE basics

The Feynman-Kac theorem is the bridge between expectations and differential equations. If

$$V(t,x) = \mathbb{E}\!\left[\left. e^{-r(T-t)}\,\varphi(X_T)\,\right|\, X_t = x\right], \qquad dX_s = a(X_s)\,ds + b(X_s)\,dW_s,$$

then \(V\) solves

$$\frac{\partial V}{\partial t} + a(x)\frac{\partial V}{\partial x} + \tfrac{1}{2}b(x)^2\frac{\partial^2 V}{\partial x^2} - rV = 0, \qquad V(T,x) = \varphi(x).$$

The proof idea is one line of Itô: apply the lemma to \(e^{-rt}V(t, X_t)\); since a conditional expectation of a fixed terminal payoff is a martingale, its \(dt\) term must vanish, and that vanishing is the PDE. Now specialize to the risk-neutral stock \(a = rS\), \(b = \sigma S\):

$$\frac{\partial V}{\partial t} + rS\frac{\partial V}{\partial S} + \tfrac{1}{2}\sigma^2 S^2 \frac{\partial^2 V}{\partial S^2} - rV = 0.$$

This is the Black-Scholes equation, and Feynman-Kac explains why it and the risk-neutral expectation of Note 03 are the same object viewed from opposite sides: the expectation is the PDE's stochastic representation; the PDE is the expectation's infinitesimal description. Solving it with the call payoff boundary condition yields the closed form in the pricing laboratory. The same bridge carries the physics intuition: the equation is a heat equation in disguise (substitute \(x = \ln S\) and rescale time), so option value diffuses — kinks in payoffs get smoothed exactly the way heat smooths a temperature spike.

Note 05

Monte Carlo Error Scales as \(O(1/\sqrt{N})\)#

4 min · Prerequisites: variance, sample means, and the central limit theorem

Estimate \(\theta = \mathbb{E}[f(X)]\) by the sample mean \(\hat\theta_N = \frac{1}{N}\sum_{i=1}^N f(X_i)\) over independent draws. The estimator is unbiased, and its standard error follows from nothing deeper than the variance of a mean:3

$$\operatorname{se}(\hat\theta_N) = \frac{\sigma_f}{\sqrt{N}}, \qquad \sigma_f^2 = \operatorname{Var}[f(X)],$$

with the central limit theorem supplying asymptotic normality and hence confidence intervals. Three consequences structure all practical Monte Carlo work:

1. Computational scaling. One more decimal digit of accuracy costs a factor of 100 in samples. Monte Carlo is the method of choice not because it converges fast but because its rate does not deteriorate with dimension — a 100-dimensional basket option converges at the same \(N^{-1/2}\) as a one-dimensional one, while grid methods collapse under \(O(h^{-d})\) nodes.

2. Variance reduction. Since the rate is fixed, accuracy can improve by shrinking \(\sigma_f\): antithetic variates, control variates (price the arithmetic Asian against the geometric one, which is known in closed form), importance sampling for tail events, and quasi-random sequences that trade independence for low discrepancy and near-\(N^{-1}\) behavior.

3. Tail estimates inherit the worst constants. Estimating \(\mathbb{P}[L > \ell] = p\) by naive simulation has relative standard error \(\approx \sqrt{(1-p)/(pN)}\) — for a \(p = 0.1\%\) event, ten percent relative accuracy needs on the order of \(10^7\) paths. Rare-event estimation without importance sampling is statistically unstable unless the estimator changes, which links this note directly to the next one.

Note 06

Fat tails and the failure of normal VaR#

5 min · Prerequisites: quantiles, moments, and basic risk metrics

Parametric-normal VaR takes a mean and standard deviation and reads the quantile off the Gaussian: \( \mathrm{VaR}_\alpha = -( \mu + z_{1-\alpha}\,\sigma )\). The procedure is exact when returns are normal and can materially underestimate tail risk when they are not, because the Gaussian density dies like \(e^{-x^2/2}\) while empirical return distributions die like a power law, \(\mathbb{P}[|r| > x] \sim x^{-\alpha}\) with tail index \(\alpha\) around 3 to 4 for daily equity returns.4 Every moment of the comparison fails in the same direction:

The counting argument. Daily equity moves of five standard deviations should occur, under normality, about once per 14,000 years. The realized record produces them every few years; October 19, 1987 was, on a Gaussian yardstick, roughly a 20σ event — a probability so small it has no physical interpretation. The model is not slightly wrong in the tail; it is wrong by factors of \(10^{3}\) to \(10^{50}\), depending on the tail threshold.

The moment argument. Excess kurtosis of daily index returns is far above the Gaussian's zero. Since sample variance is dominated by the very observations the Gaussian deems nearly impossible, \(\hat\sigma\) is inflated by past crises while the normal quantile formula simultaneously understates how much worse than \(z\hat\sigma\) the next crisis will be. The errors do not cancel; at high confidence levels the understatement wins.

The structural argument. Volatility clustering means returns are a mixture of distributions — calm-regime and stress-regime — and mixtures of normals with different variances are themselves fat-tailed. So even if each day were conditionally Gaussian, unconditional normal VaR would still be miscalibrated. Worse, the stress regime arrives with correlations lurching toward one, so the portfolio-level tail is fatter than any asset-level analysis suggests.

Available responses include historical-simulation VaR, as used on the risk dashboard; Expected Shortfall, which averages the tail beyond a quantile; and extreme-value methods that fit an asymptotic tail family such as the GPD or GEV. Estimates at very high confidence levels remain extrapolations beyond limited observed data.

Sources

Standing references for these notes#

  1. Øksendal, B. (2003), Stochastic Differential Equations: An Introduction with Applications, 6th ed., Springer, pp. 21–84 (Itô integrals, Itô formula, and SDEs). doi:10.1007/978-3-642-14394-6
  2. Shreve, S. E. (2004), Stochastic Calculus for Finance II: Continuous-Time Models, 1st ed., Springer, chs. 4–6, pp. 131–339 (stochastic calculus, risk-neutral pricing, and PDE connections), ISBN 978-0-387-40101-0. publisher record
  3. Glasserman, P. (2003), Monte Carlo Methods in Financial Engineering, Springer, pp. 1–38 and 185–337 (foundations, variance reduction, and quasi-Monte Carlo). doi:10.1007/978-0-387-21617-1
  4. McNeil, A. J., R. Frey, and P. Embrechts (2015), Quantitative Risk Management: Concepts, Techniques and Tools, revised 2nd ed., Princeton University Press, chs. 2 and 5, ISBN 978-0-691-16627-8. publisher catalog