CVaRPort: a tail-risk calculator for a small portfolio

cvarport takes the historical prices of a few ETFs and the amounts you hold, and answers one question: is my portfolio efficient, and how much tail risk could I remove without giving up expected return? It measures the portfolio you have (mean, volatility, Sharpe ratio, historical VaR and CVaR, and how much of each risk every asset carries), then compares it with the best portfolios that can be built from the same assets under the same constraints, on two efficient frontiers: Markowitz's mean-variance one and the mean-CVaR one of Rockafellar and Uryasev.

It is written in Julia: the optimization problems are stated with JuMP almost as they are written below, and solved by HiGHS (linear programs) and Clarabel (quadratic programs). The figures of this page are drawn in R from the program's own output.

This is a study tool, not investment advice: every number it prints is an estimate from past data (see §9).

What you need

flowchart LR
    D["Returns and losses<br/>(§4.1)"] --> E["Estimating μ and Σ<br/>(§4.2)"]
    E --> MV["Mean-variance QP<br/>(§4.3)"]
    MV --> T["Tangency portfolio<br/>(§4.4)"]
    D --> V["VaR and CVaR<br/>(§4.5)"]
    V --> LP["Mean-CVaR LP<br/>(§4.6)"]
    E --> C["Risk contributions<br/>(§4.7)"]
    V --> C
    MV --> X["The example<br/>(§5)"]
    LP --> X
    C --> X

1. What it does

Given a prices file and a weights file, cvarport prints a report like this one, computed on the synthetic data set of examples/ (five years of daily prices of four made-up assets, §5):

cvarport: 1260 return periods, 4 assets, alpha = 0.95, cap = 0.70, Ledoit-Wolf delta = 0.017
mean and vol annualized (x252), rf = 2.00%; VaR and CVaR are one-period losses

portfolio                      mean      vol  Sharpe    VaR95   CVaR95  |     EQTY     BOND     GOLD     SMCP
current                       7.94%   11.37%   0.522    1.04%    1.51%  |    0.600    0.200    0.100    0.100
min variance                  0.85%    4.95%  -0.233    0.45%    0.65%  |    0.144    0.700    0.156    0.000
tangency                      8.90%   11.15%   0.619    0.99%    1.45%  |    0.700    0.300    0.000    0.000
min CVaR                      1.29%    4.97%  -0.143    0.45%    0.65%  |    0.167    0.700    0.133    0.000
min var, mean >= current      7.94%   10.05%   0.591    0.90%    1.31%  |    0.625    0.375    0.000    0.000
min CVaR, mean >= current     7.94%   10.05%   0.591    0.90%    1.31%  |    0.625    0.375    0.000    0.000

risk contributions of the current portfolio (shares of the total)
asset      weight      vol   CVaR95
EQTY        0.600   81.10%   78.31%
BOND        0.200    1.13%    1.19%
GOLD        0.100    0.90%    1.46%
SMCP        0.100   16.88%   19.05%

with expected return at least the current one, CVaR95 goes from 1.51% to 1.31% (-13.8% relative)

How to read it:

Optionally (--frontier F) it also writes both efficient frontiers to a CSV file, one row per point, for plotting.

2. Usage

Requires Julia 1.10 or later. From the repository root:

julia --project=. -e 'using Pkg; Pkg.instantiate()'    # once: install JuMP, HiGHS, Clarabel

2.1 Quick start, without a network connection

julia --project=. examples/synthetic.jl > examples/prices.csv
julia --project=. bin/cvarport.jl examples/prices.csv examples/weights.txt --rf 0.02 --cap 0.7

The first command generates the synthetic prices; the second prints the report of §1. The first run takes a while longer, because Julia compiles JuMP and the solvers.

2.2 Your own portfolio

cvarport reads two plain-text files.

Prices, a CSV with a header date,T1,...,Tn and one row per date, in ascending order, with no missing values:

date,AAA,BBB,CCC
2021-01-04,61.82,103.66,49.74
2021-01-05,62.05,103.71,50.12
...

Weights, one TICKER amount line per asset you hold, with the same tickers as the header of the prices file. Amounts can be weights or money: they are normalized to sum to one, and tickers left out weigh zero. # starts a comment.

# TICKER amount
AAA  6000
BBB  2500
CCC  1500

scripts/fetch.sh downloads daily, weekly or monthly closes from Stooq and merges them into that format, keeping only the dates common to all tickers:

scripts/fetch.sh -i d TICKER1 TICKER2 TICKER3 > prices.csv    # -i d|w|m, Stooq symbols

Stooq may answer scripts with a browser check instead of data; when that happens, fetch.sh stops with an error, and the files can be downloaded from the Stooq page of each ticker and merged by hand into the format above.

Two choices make the numbers more meaningful:

2.3 Options

julia --project=. bin/cvarport.jl prices.csv weights.txt --rf 0.025 --cap 0.6 --frontier frontier.csv
option default meaning
--alpha A 0.95 CVaR confidence level
--rf R 0 annual risk-free rate (e.g. €STR or a short BOT yield)
--cap C 1 maximum weight of any asset
--periods P 252 return periods per year: 252 daily, 52 weekly, 12 monthly
--no-shrink off sample covariance instead of Ledoit-Wolf
--frontier F none write both frontiers to the CSV file F

The cap must be feasible: with $N$ assets the weights can sum to one only if $N \cdot \text{cap} \ge 1$. If the current weights exceed the cap, the report says so, and the comparison of the last line is between portfolios that live under different constraints.

2.4 From Julia

Everything the command line does is a function of the CVaRPort module, which also exposes the building blocks (§3):

using CVaRPort

_, tickers, P = read_prices("examples/prices.csv")
w = read_weights("examples/weights.txt", tickers)
R = simple_returns(P)

res = analyze(R, w; alpha = 0.95, rf = 0.02, cap = 0.7)
print_report(tickers, res)

Sigma, delta = ledoit_wolf(R)
w_cvar, cvar = min_cvar(R, 0.95, Constraints(cap = 0.7))

examples/analysis.jl is a complete example: it runs the analysis of §1 and writes its results to CSV files instead of printing them.

3. How it works

The program is a straight pipeline: two files in, returns, estimates, optimization problems, a report out. Each stage lives in its own file of src/.

flowchart LR
    PR["prices.csv"] --> RP["read_prices<br/>simple_returns"]
    WT["weights.txt"] --> RW["read_weights"]
    RP -- "R (scenarios × assets)" --> LW["ledoit_wolf<br/>mean"]
    LW -- "μ, Σ" --> QP["min_variance<br/>tangency<br/>(QP, Clarabel)"]
    RP -- "R" --> LP["min_cvar<br/>max_return<br/>(LP, HiGHS)"]
    LW -- "μ" --> LP
    QP -- "weights" --> RK["portfolio_stats<br/>var_cvar<br/>contributions"]
    LP -- "weights" --> RK
    RW -- "w" --> RK
    RK --> AN["analyze"]
    AN --> REP["print_report<br/>(stdout)"]
    AN -. "frontier option" .-> WF["write_frontier<br/>(CSV)"]
  1. Data (src/data.jl). read_prices checks the prices file (ascending dates, positive and complete prices) and returns the price matrix $P$; simple_returns turns it into the returns matrix $R$, one row per period and one column per asset; read_weights matches the weights file against the tickers and normalizes it.
  2. Estimation (src/estimate.jl). The expected returns $\hat\mu$ are column means of $R$; the covariance $\hat\Sigma$ is the Ledoit-Wolf shrinkage estimator (§4.2).
  3. Risk (src/risk.jl). For any weight vector: historical VaR and CVaR, risk contributions, and the annualized statistics of the report (§4.5, §4.7, §4.8).
  4. Optimization (src/optim.jl). The constraints are a value of type Constraints; every problem is a JuMP model built on the same feasible set and handed to the right solver (§6): quadratic programs for variance, linear programs for CVaR and for the maximum return.
  5. Report (src/report.jl). analyze computes the reference portfolios and their statistics into a named tuple; print_report formats it; write_frontier sweeps both frontiers and writes them to CSV.

The command line (bin/cvarport.jl) only parses the options and calls these functions. Two of them have more than one method, chosen by Julia's multiple dispatch on the argument types: var_cvar(L, alpha) works on a vector of losses and var_cvar(R, w, alpha) on returns and weights, while print_report writes to stdout unless it is given an IO (see teoria_tipi §13 on polymorphism across languages).

4. The model

4.1 Returns, losses and constraints

There are $N$ assets and $S + 1$ dates. The return of asset $i$ in period $s$ is the simple return

\[r_{s,i} = \frac{P_{s,i}}{P_{s-1,i}} - 1, \qquad s = 1, \dots, S,\]

and the row $r_s \in \mathbb{R}^N$ is a scenario: the model takes the $S$ observed scenarios, each with probability $1/S$, as the distribution of the returns ("historical simulation"). A portfolio is a weight vector $w$, and its loss in scenario $s$ is

\[L_s(w) = -r_s^\top w ,\]

positive when money is lost. Simple returns, unlike log returns, aggregate exactly across assets: the return of the portfolio is $r_s^\top w$. All portfolios belong to the feasible set

\[\mathcal{W} = \lbrace w \in \mathbb{R}^N : \mathbf{1}^\top w = 1, \ 0 \le w_i \le c \rbrace ,\]

fully invested, long-only, with a cap $c$ on each weight; it is not empty exactly when $N c \ge 1$.

4.2 Estimating mu and Sigma

The expected returns are the sample means $\hat\mu = \frac{1}{S} \sum_s r_s$. They are the weakest input of the whole model: the standard error of an annual mean estimated from $Y$ years of data is about $\sigma / \sqrt{Y}$, which for an equity index (volatility around 16%) over five years is about 7% a year, as large as the mean itself. §5 shows it on the synthetic data.

The covariance is better determined, but the sample covariance

\[C = \frac{1}{S} \sum_{s=1}^{S} x_s x_s^\top , \qquad x_s = r_s - \hat\mu ,\]

is noisy when $S$ is not much larger than $N$, and the optimizers exploit its noise. ledoit_wolf shrinks it towards a multiple of the identity (Ledoit and Wolf, 2004):

\[\hat\Sigma = \delta \, m I + (1 - \delta) \, C , \qquad m = \frac{\operatorname{tr} C}{N} .\]

The intensity $\delta \in [0, 1]$ trades the distance of $C$ from the target against the estimation error of $C$, both measured with the normalized Frobenius norm:

\[d^2 = \frac{1}{N} \lVert C - m I \rVert_F^2 , \qquad \bar{b}^2 = \frac{1}{N S^2} \sum_{s=1}^{S} \lVert x_s x_s^\top - C \rVert_F^2 , \qquad \delta = \frac{\min(\bar{b}^2, d^2)}{d^2} .\]

With many observations $\bar{b}^2$ is small and $\delta \to 0$: on the 1260 days of the example $\delta = 0.017$, and on 20000 observations of three assets (a test) it stays below 0.05. The result is always symmetric positive definite, which the quadratic programs need. --no-shrink uses $C$ instead.

4.3 Mean-variance

The minimum-variance portfolio with expected return at least $\tau$ solves the convex quadratic program

\[\min_{w \in \mathcal{W}} \; w^\top \hat\Sigma w \qquad \text{s.t.} \qquad \hat\mu^\top w \ge \tau .\]

Without the sign and cap constraints, and with an equality on the return, the Lagrange conditions are linear and give the closed form of teoria_lagrange §9.4; with them, the optimum is characterized by the KKT conditions (teoria_lagrange §5) and computed numerically. The tests check the solver against the closed forms on an example where the constraints are not active.

min_variance without a target gives the left end of the frontier. The right end is the portfolio of maximum expected return, a linear program (max_return): with a cap $c$, it fills the assets in decreasing order of $\hat\mu_i$ up to $c$ each. mv_frontier sweeps $\tau$ between the two ends.

4.4 Tangency portfolio

With a risk-free rate $r_f$, the tangency portfolio maximizes the Sharpe ratio

\[\max_{w \in \mathcal{W}} \; \frac{(\hat\mu - r_f \mathbf{1})^\top w}{\sqrt{w^\top \hat\Sigma w}} ,\]

which is not a convex problem as stated. If some feasible portfolio beats $r_f$, the change of variables $y = w / (\hat\mu - r_f \mathbf{1})^\top w$ makes it one. The constraints of $\mathcal{W}$ are homogeneous once the budget is dropped, so they carry over to $y$ as $y \ge 0$ and $y_i \le c \, \mathbf{1}^\top y$; the numerator becomes $1$, and maximizing the ratio is the same as minimizing the denominator:

\[\min_{y} \; y^\top \hat\Sigma y \qquad \text{s.t.} \qquad (\hat\mu - r_f \mathbf{1})^\top y = 1, \quad y \ge 0, \quad y_i \le c \, \mathbf{1}^\top y ,\]

after which $w = y / \mathbf{1}^\top y$. Here $r_f$ is per period, the annual rate divided by the number of periods per year. When no feasible portfolio beats $r_f$, tangency throws an error and the report leaves out the row.

4.5 VaR and CVaR

For a loss $L$ and a confidence level $\alpha$ (say 0.95), the value at risk (VaR) at level $\alpha$ is the $\alpha$-quantile of $L$: the loss exceeded only with probability $1 - \alpha$. It says nothing about how large the losses beyond it are, and it is not convex in the weights, so diversification can increase it. The conditional value at risk is, roughly, the average loss in the worst $(1 - \alpha)$ fraction of the cases, and Rockafellar and Uryasev (2000, 2002) characterized it as the value of a one-dimensional minimization:

\[\operatorname{CVaR}_\alpha(L) = \min_{\zeta \in \mathbb{R}} \; \Big( \zeta + \frac{1}{1 - \alpha} \, \mathbb{E}\big[ \max(L - \zeta, 0) \big] \Big) ,\]

whose minimizer is the VaR. CVaR is convex and coherent (subadditive: a combination of portfolios is never riskier than the combination of their risks), which is what makes it optimizable.

With $S$ equally likely scenarios, the expectation is a mean, and the historical CVaR of the portfolio $w$ is

\[\operatorname{CVaR}_\alpha(w) = \min_{\zeta} \; \Big( \zeta + \frac{1}{(1 - \alpha) S} \sum_{s=1}^{S} \max\big( L_s(w) - \zeta, 0 \big) \Big) .\]

The objective is piecewise linear and convex in $\zeta$, with slope changing at the losses: its minimum is at the $k$-th smallest loss, $k = \lceil \alpha S \rceil$, which is the historical VaR. var_cvar evaluates exactly this expression. When $(1 - \alpha) S$ is not an integer the VaR scenario enters the average with a fractional weight: for the ten losses $1, 2, \dots, 10$ and $\alpha = 0.75$, the worst $2.5$ scenarios give $\operatorname{VaR} = 8$ and $\operatorname{CVaR} = (10 + 9 + 0.5 \cdot 8) / 2.5 = 9.2$ (one of the tests).

4.6 Minimum CVaR as a linear program

The formula of §4.5 is jointly convex in $(w, \zeta)$, so minimizing CVaR over the portfolios is a single minimization over $w$ and $\zeta$ together. Each $\max(\cdot, 0)$ becomes an auxiliary variable $u_s$ with two linear constraints, and the problem is the linear program

\[\min_{w \in \mathcal{W}, \ \zeta, \ u} \; \zeta + \frac{1}{(1 - \alpha) S} \sum_{s=1}^{S} u_s \qquad \text{s.t.} \qquad u_s \ge -r_s^\top w - \zeta, \quad u_s \ge 0, \quad \hat\mu^\top w \ge \tau .\]

At the optimum each $u_s$ is as small as allowed, $u_s = \max(L_s(w) - \zeta, 0)$, so the objective is the historical CVaR of the optimal portfolio and $\zeta$ is its VaR: the optimizer and the report use the same definition, and the tests check that they agree. The program has $N + 1 + S$ variables and grows with the number of scenarios (1265 variables for the example), which a simplex solver handles without effort (teoria_ottimizzazione §1; its dual, read through the Lagrangian, is in teoria_lagrange §8.1). min_cvar without a target gives the minimum-CVaR portfolio, and cvar_frontier sweeps $\tau$ as in §4.3.

4.7 Risk contributions

Volatility and CVaR are positively homogeneous of degree one, $\rho(t w) = t \rho(w)$ for $t > 0$: doubling every position doubles the risk. By Euler's theorem on homogeneous functions, such a $\rho$ is the sum of its marginal contributions,

\[\rho(w) = \sum_{i=1}^{N} w_i \frac{\partial \rho}{\partial w_i}(w) ,\]

and $c_i = w_i \, \partial \rho / \partial w_i$ is the part of the risk due to asset $i$ (Tasche, 2008). For the volatility $\sigma(w) = \sqrt{w^\top \hat\Sigma w}$ the gradient is $\hat\Sigma w / \sigma(w)$, so

\[c_i^{\sigma} = \frac{w_i \, (\hat\Sigma w)_i}{\sigma(w)} .\]

For the historical CVaR, let $s^\star$ be the scenario of the VaR, $\zeta^\star = L_{s^\star}(w)$, $\mathcal{T} = \lbrace s : L_s(w) > \zeta^\star \rbrace$ the scenarios beyond it, and $m = (1 - \alpha) S$. Then

\[\frac{\partial \operatorname{CVaR}_\alpha}{\partial w_i} = \frac{1}{m} \Big( \sum_{s \in \mathcal{T}} (-r_{s,i}) + \big( m - \lvert \mathcal{T} \rvert \big) (-r_{s^\star,i}) \Big) ,\]

the average loss of asset $i$ over the tail scenarios, with the same fractional weight on the VaR scenario as in §4.5. Both decompositions add up exactly to the total (the tests check it to $10^{-12}$), and the report prints the shares $c_i / \sum_j c_j$. A share can be negative: an asset that tends to gain when the portfolio loses reduces the risk.

4.8 Annualization

With $p$ periods per year (--periods), the report annualizes linearly, as is customary for simple returns:

\[\text{mean} = p \, \hat\mu^\top w , \qquad \text{vol} = \sqrt{p} \, \sigma(w) , \qquad \text{Sharpe} = \frac{\text{mean} - r_f}{\text{vol}} .\]

VaR and CVaR stay one-period: scaling them by $\sqrt{p}$ would assume normal, independent returns, which is exactly what the tail measures are meant not to assume.

5. The example

The synthetic data set (examples/synthetic.jl) has four assets over 1260 trading days: an equity index (EQTY), a bond index (BOND), gold (GOLD) and small caps (SMCP, correlated 0.8 with EQTY). The daily returns are drawn from a multivariate Student $t$ with 4 degrees of freedom, so they have fat tails, and the figures below come from the program itself:

flowchart LR
    G["examples/synthetic.jl"] --> PR["examples/prices.csv"]
    PR --> A["examples/analysis.jl"]
    WT["examples/weights.txt"] --> A
    A --> CSV["portfolios.csv<br/>contributions.csv<br/>frontier.csv"]
    CSV --> RS["scripts/*.r"]
    PR --> RS
    RS --> IMG["img/*.png"]

The example also shows the weakness of $\hat\mu$ (§4.2). The returns are generated with annual drifts of 7%, 2%, 4% and 9%, but over these five simulated years the sample means come out at 12.7%, 0.0%, $-6.1\%$ and 9.1%: gold, built to earn 4% a year, lost 6% a year in this sample. Every optimizer below takes these estimates at face value.

5.1 The frontiers

Mean-variance and mean-CVaR efficient frontiers of the synthetic example, in the volatility-mean plane and in the CVaR-mean plane: the two frontiers overlap; the current portfolio lies to the right of them, and an arrow points to the portfolio with the same mean and minimum CVaR

The current portfolio (60% EQTY, 20% BOND, 10% GOLD, 10% SMCP) lies inside both frontiers. Keeping its expected return of 7.94%, the minimum-CVaR portfolio takes the one-day CVaR95 from 1.51% to 1.31% and the volatility from 11.4% to 10.1%, by holding 62.5% EQTY and 37.5% BOND and nothing else: gold goes because of its negative sample mean, small caps because they add volatility to equity without adding return. The tangency portfolio sits where the capital market line from $r_f = 2\%$ touches the mean-variance frontier, at the cap on EQTY.

The two frontiers coincide, and they should. The Student $t$ used here is elliptical: every portfolio return has the same distribution up to its mean and scale, so any risk measure that is translation invariant and positively homogeneous, VaR and CVaR included, is a function of the mean and of the volatility of the portfolio, and ranks portfolios as variance does. For elliptical returns, Markowitz is optimal for CVaR too (Embrechts, McNeil and Straumann, 2002); the residual gap in the figure is sampling noise of the historical CVaR. The frontiers separate when the joint distribution is not elliptical: assets with skewed returns, such as options or credit, or crashes that hit several assets at once more often than their correlation suggests. That is when cvarport's second frontier says something the first one cannot.

5.2 The losses

Histogram of the daily losses of the current portfolio, with the VaR95 at 1.04% and the CVaR95 at 1.51%, the tail beyond the VaR highlighted, and the normal density with the same mean and standard deviation

The orange bars are the worst 5% of the days, whose average (with the fractional weight of §4.5) is the CVaR. Against the normal distribution with the same mean and standard deviation, the losses are more concentrated around zero and have a longer tail: the normal model gives a larger VaR (1.15% against 1.04%) but a smaller CVaR (1.45% against 1.51%), and the sample has 4 losses beyond three standard deviations, where a normal would expect 1.7, the largest at 6.0%. This is why cvarport uses the historical scenarios rather than a normal approximation.

5.3 The risk contributions

Bar chart of weights and of shares of volatility and CVaR for each asset of the current portfolio: EQTY has 60% of the weight and about 80% of both risks, SMCP 10% of the weight and 17 to 19% of the risk, BOND and GOLD about 1% of the risk each

EQTY holds 60% of the money and 78-81% of the risk; SMCP, with 10% of the money, carries 17-19%, because it moves with EQTY and is more volatile. BOND and GOLD together hold 30% of the money and about 2-3% of the risk: they are almost uncorrelated with the equity assets, so their marginal contribution is small. The CVaR shares are close to the volatility shares, as expected from the elliptical model of §5.1; with real data the gap between the two columns is where tail risk hides.

6. Solvers and numerics

7. Project layout

To regenerate the example and the figures, from the repository root:

julia --project=. examples/synthetic.jl > examples/prices.csv
julia --project=. examples/analysis.jl
Rscript scripts/frontiers.r
Rscript scripts/losses.r
Rscript scripts/contributions.r

The R scripts use base graphics only, and the Spectral font of this page through the Cairo device.

8. Tests

julia --project=. -e 'using Pkg; Pkg.test()'

The 38 checks, grouped by test set, cover:

The scenarios of the tests are correlated Student $t$ returns of daily size, generated with a fixed seed.

9. Caveats

References