3-12 month returns and earning momentum is consistently profitable. The best performer are no more riskier than worst performers. Hence, standard risk adjustments tend to increase the return spread between the winner and losers.
The cause is overreaction or underreaction to information. There is reversal over weeks to months and years and 5 years, while momentum at 3-12 months. There is seasonality in January with negative returns and positive for every other month.
Tuesday, July 28, 2015
A New Anomaly: The Cross-Sectional Profitability of Technical Analysis - Han, Yang, Zhou 2013
Momentum portfolio sorted by volatility generates better profits than well-known return based momentum strategies. The correlations are low as well. These excess returns are not explained by market timing, investor sentiment, default and liquidity risk. Similar results hold if the portfolios are sorted based on other proxies of information uncertainty (size, distance to default, credit rating, analyst forecast dispersion, earnings volatility). The more noise-to-signal ratio or the more uncertain the information, the more profitable the technical analysis.
Strategy: Buy or remain long the portfolio today when yesterday's price is above its 10-day MA price, to to invest in risk-free asset otherwise. This is compared against buy-and-hold, for the top decile.
Strategy: Buy or remain long the portfolio today when yesterday's price is above its 10-day MA price, to to invest in risk-free asset otherwise. This is compared against buy-and-hold, for the top decile.
Monday, July 27, 2015
Momentum and Autocorrelation in Stock Returns - Lewellen
Role of size and BM factors on stock momentum. Both are negatively auto-correlated and cross-serially correlated over intermediate horizons. The excess covariance of stocks with each other, and not under-reaction, explains momentum in the portfolios.
Firm specific returns and investors under-reaction and belated overreaction does not explain a significant component of momentum. Size and BM factor based momentum is strong and distinct, showing that momentum can't be attributed solely to firm-specific returns - there must be multiple sources of momentum. Momentum shows up in individual stocks and size quintiles, but vanishes at the market level.
Firm specific returns and investors under-reaction and belated overreaction does not explain a significant component of momentum. Size and BM factor based momentum is strong and distinct, showing that momentum can't be attributed solely to firm-specific returns - there must be multiple sources of momentum. Momentum shows up in individual stocks and size quintiles, but vanishes at the market level.
Sources of Momentum
Profits depend on both auto-correlations and the lead-lag relationship. The portfolio weight of asset $i$ in month $t$ is
$$w_{i,t}=\frac{1}{N}(r_{i,t-1}-r_{m,t-1})$$
where $r_{m,t}$ is the equal-weighted market index returns in month $t$. Assume returns have unconditional mean $\mu=E[r_t]$ and autocovariance matrix $\Omega=E[(r_{t-1}-\mu)(r_t-\mu)^T]$. The portfolio return in month t equals:
$$\pi_t=\sum_i w_{i,t}r_{i,t}=\frac{1}{N}\sum_i (r_{i,t-1}-r_{m,t-1})r_{i,t}.$$
Hence, the expected profit is
$$E[\pi_t] = \frac{1}{N}E\Bigg[\sum_i r_{i,t-1}r_{i,t}\Bigg]-\frac{1}{N}E\Bigg[r_{m,t-1}\sum_i r_{i,t}\Bigg] \
= \frac{1}{N} \sum_i (\rho_i+\mu_i^2)-(\rho_m+\mu_m^2),$$
where $\rho_i$ and $\rho_m$ are the autocovariances of the asset i and the equal-weighted index, respectively. Using that fact that average autocovariance equals $tr(\Omega)/N$ and the autocovariance of the market portfolio equals $\varsigma^T\Omega\varsigma/n^2$, where $\varsigma$ is the vector of ones.
$$E[\pi_t]=\frac{1}{N}tr(\Omega)-\frac{1}{N^2}\varsigma^T\Omega\varsigma+\sigma_{\mu}^2=\frac{N-1}{N^2}tr(\Omega)-\frac{1}{N^2}[\varsigma^T\Omega\varsigma-tr(\Omega)]+\sigma_{\mu}^2.$$
This decomposition says that momentum can arise in three ways:
1) stocks might be positively autocorrelated (first term) - meaning stocks with high returns today are expected to have higher returns tomorrow.
2) Cross-serial correlations might be negative - meaning firm with high return today predicts that other firms will have low returns in the future. This is related to excess covariance among stocks.
3) High unconditional mean stocks.
This decomposition is not unique.
$$\pi_t=\sum_i w_{i,t}r_{i,t}=\frac{1}{N}\sum_i (r_{i,t-1}-r_{m,t-1})r_{i,t}.$$
Hence, the expected profit is
$$E[\pi_t] = \frac{1}{N}E\Bigg[\sum_i r_{i,t-1}r_{i,t}\Bigg]-\frac{1}{N}E\Bigg[r_{m,t-1}\sum_i r_{i,t}\Bigg] \
= \frac{1}{N} \sum_i (\rho_i+\mu_i^2)-(\rho_m+\mu_m^2),$$
where $\rho_i$ and $\rho_m$ are the autocovariances of the asset i and the equal-weighted index, respectively. Using that fact that average autocovariance equals $tr(\Omega)/N$ and the autocovariance of the market portfolio equals $\varsigma^T\Omega\varsigma/n^2$, where $\varsigma$ is the vector of ones.
$$E[\pi_t]=\frac{1}{N}tr(\Omega)-\frac{1}{N^2}\varsigma^T\Omega\varsigma+\sigma_{\mu}^2=\frac{N-1}{N^2}tr(\Omega)-\frac{1}{N^2}[\varsigma^T\Omega\varsigma-tr(\Omega)]+\sigma_{\mu}^2.$$
This decomposition says that momentum can arise in three ways:
1) stocks might be positively autocorrelated (first term) - meaning stocks with high returns today are expected to have higher returns tomorrow.
2) Cross-serial correlations might be negative - meaning firm with high return today predicts that other firms will have low returns in the future. This is related to excess covariance among stocks.
3) High unconditional mean stocks.
This decomposition is not unique.
Saturday, July 25, 2015
Anticipating Correlations - Engle
These are my notes on Robert Engle's book 'Anticipating Correlations - a new paradigm for risk management'. Engle is a celebrated Nobel Laureate for his contributions to the development of GARCH model of volatility.
Ch1: Correlation Economics
The movement in the prices of assets are not independent. If they were it would have been possible to construct a portfolio with negligible volatility. Estimating the correlations for big cross-section is a Herculean task, especially when it is recognized that these correlations var over time. Hence, a forward looking correlation estimation is needed for optimal risk-management, portfolio selection and hedging. The main method developed is dynamic conditional correlations (DCC).
There are high correlations between industry sector stocks but lower otherwise. The correlation between different asset classes is lower. For equity of different countries the data should be non-synced (e.g. by taking average over more than one days) before taking correlations.
Changes in asset prices and correlations reflect changing forecasts of future payments. The effect of a news affects all asset prices to a greater or lesser extent, depending on their correlations. The most important reason why these correlations change over time is because the firms change their line of business. A second important factor is the characteristic of the news change (e.g. change in magnitude of the news).
There are high correlations between industry sector stocks but lower otherwise. The correlation between different asset classes is lower. For equity of different countries the data should be non-synced (e.g. by taking average over more than one days) before taking correlations.
Changes in asset prices and correlations reflect changing forecasts of future payments. The effect of a news affects all asset prices to a greater or lesser extent, depending on their correlations. The most important reason why these correlations change over time is because the firms change their line of business. A second important factor is the characteristic of the news change (e.g. change in magnitude of the news).
Saturday, July 18, 2015
Modeling return dynamics via decomposition
The paper Modeling Financial Returns Dynamics via Decomposition - Anatolyev and Gospodinov (2010) points out that predicting excess returns for stocks is much more difficult than simply predicting the direction of change. They hence decompose returns into direction and magnitude change and jointly model them for sign and magnitude using a copula for interaction. This lets them incorporate important non-linearities.
This paper, instead of trying to identify better predictors, look for better ways of using predictors. This is done by decomposing returns into sign and magnitude. Sign has better predictability. We aim to predict the expected returns and the following decomposition model is proposed:
$$E[r_t|F_{t-1}]=E[|r_t|sign(r_t)|F_{t-1}]=f(|r_t|)\times g(sign(r_t))\times \text{interaction copula}.$$
The magnitude is modeled using multiplicative error model, the sign by dynamic binary choice model and a copula for their interaction. This way we are able to model hidden nonlinearities absent from the regression setup. Magnitude and signs have substantial dependence over time but hardly any for returns! e.g. magnitude is like vol which shows significant dependence. One important aspect of the bivariate analysis is that, in spite of a large unconditional correlation between the multiplicative components, they appear to conditionally very weakly dependent.
This opens avenues for strategies as well (Anatolyev and Gerko 2005). The decomposition model is better than predictive regression which is better than buy-and-hold strategy - both in and out of sample. The decomposition model also produces unbiased forecasts.
Introduction
Valuation ratios (dividend price, earning price), yields on short and long term treasury and corporate bonds appear to posses some predictive power at short horizons for timing the market. New variable with incremental predictive power such as share of equity issue in total new equity and debt issues, consumption-wealth ratio, relative valuations of high and low-beta stocks, estimated factors from large economic datasets can be use (Lettau and Ludvigson 2008 review paper).
This paper, instead of trying to identify better predictors, look for better ways of using predictors. This is done by decomposing returns into sign and magnitude. Sign has better predictability. We aim to predict the expected returns and the following decomposition model is proposed:
$$E[r_t|F_{t-1}]=E[|r_t|sign(r_t)|F_{t-1}]=f(|r_t|)\times g(sign(r_t))\times \text{interaction copula}.$$
The magnitude is modeled using multiplicative error model, the sign by dynamic binary choice model and a copula for their interaction. This way we are able to model hidden nonlinearities absent from the regression setup. Magnitude and signs have substantial dependence over time but hardly any for returns! e.g. magnitude is like vol which shows significant dependence. One important aspect of the bivariate analysis is that, in spite of a large unconditional correlation between the multiplicative components, they appear to conditionally very weakly dependent.
This opens avenues for strategies as well (Anatolyev and Gerko 2005). The decomposition model is better than predictive regression which is better than buy-and-hold strategy - both in and out of sample. The decomposition model also produces unbiased forecasts.
Methodological Framework
The key identity is
$$r_t=c+|r_t-c|sign(r_t-c)=c+|r_t-c|(2\mathbb{I}[r_t>c]-1)$$
and hence,
$$E[r_t|F_{t-1}]=c-E[|r_t-c| | F_{t-1}]+2E[|r_t-c|\mathbb{I}[r_t>c]|F_{t-1}],$$
where $c$ is a user defined constant used to model transaction cost, different dynamics of small or large positive and large negative returns. It would be 0 for modeling recession and expansion using GDP. 3% for modeling output gap, and 2% for forecasting inflation. $F_{t-1}$ is all the information available till time $t-1$, which practically consists of all data like lagged returns, volatility, volume and other predictive variables available at time $t-1$. Toy example, where predictive variables are based on realized volatility $RV_{t-1}$:
a) For direct regression model: $E[r_t]=\alpha+\beta RV_{t-1}$ gives a $R^2$ of 0.39%
b) For decomposition model: $E[|r_t|]=\alpha_{|r|}+\beta_{|r|}RV_{t-1}$ and $Pr[r_t>0]=\alpha_{\mathbb{I}}+\beta_{\mathbb{I}}RV_{t-1}$. Assuming the two components are stochastically independent giving $E[r_t]=\alpha_r+\beta_rRV_{t-1}+\gamma_r RV^2_{t-1}$, showing that nonlienarities are covered in the decomposition model, giving $R^2$ of 0.72%.
c) Further adding $\mathbb{I}[r_{t-1}>0]$ and $RV_{t-1}\mathbb{I}[r_{t-1}>0]$ to the regressor list increases the $R^2$ to 1.21%.
It is important to note that it is the augmentation of the sign component which delivers nonlinear dependence, improving the prediction. The driving force behind the predictive ability of the decomposition model is the predictability in the two components. The interaction term is less significant. This is the main theme of this work.
and hence,
$$E[r_t|F_{t-1}]=c-E[|r_t-c| | F_{t-1}]+2E[|r_t-c|\mathbb{I}[r_t>c]|F_{t-1}],$$
where $c$ is a user defined constant used to model transaction cost, different dynamics of small or large positive and large negative returns. It would be 0 for modeling recession and expansion using GDP. 3% for modeling output gap, and 2% for forecasting inflation. $F_{t-1}$ is all the information available till time $t-1$, which practically consists of all data like lagged returns, volatility, volume and other predictive variables available at time $t-1$. Toy example, where predictive variables are based on realized volatility $RV_{t-1}$:
a) For direct regression model: $E[r_t]=\alpha+\beta RV_{t-1}$ gives a $R^2$ of 0.39%
b) For decomposition model: $E[|r_t|]=\alpha_{|r|}+\beta_{|r|}RV_{t-1}$ and $Pr[r_t>0]=\alpha_{\mathbb{I}}+\beta_{\mathbb{I}}RV_{t-1}$. Assuming the two components are stochastically independent giving $E[r_t]=\alpha_r+\beta_rRV_{t-1}+\gamma_r RV^2_{t-1}$, showing that nonlienarities are covered in the decomposition model, giving $R^2$ of 0.72%.
c) Further adding $\mathbb{I}[r_{t-1}>0]$ and $RV_{t-1}\mathbb{I}[r_{t-1}>0]$ to the regressor list increases the $R^2$ to 1.21%.
It is important to note that it is the augmentation of the sign component which delivers nonlinear dependence, improving the prediction. The driving force behind the predictive ability of the decomposition model is the predictability in the two components. The interaction term is less significant. This is the main theme of this work.
Marginal distributions and Copula model
a) Volatility model: Absolute returns $|r_t-c|$ is a positively valued variable and is modeled using multiplicative error framework of Engle (2002) $$|r_t-c|=\psi_t\eta_t,$$ where $\psi_t=E[|r_t-c||F_{t-1}]$ and $\eta_t$ is a positive multiplicative error with $E[\eta_t|F_{t-1}]=1$ and conditional distribution $\mathbb{D}$. $\psi_t$ can be modeled using lograthimic autoregressive conditional duration (LACD) as
$$ln\psi_t=\omega_v+\beta_vln\psi_{t-1}+\gamma_vln|r_{t-1}-c|+\rho_v\mathbb{I}[r_{t-1}>c]+\pmb{x}^T_{t-1}\pmb{\delta}_v.$$ The second last term allows for regime-specific volatility dependence while the last term represents macroeconomic predictors of volatility. $\mathbb{D}$ can be modeled as constant parameter Weibull distribution (or others distributions with the shape parameter vector $\varsigma$ a function of the past).
b) Direction model: The indicator $\mathbb{I}[r_t>c]$ has a conditional distribution of Bernoulli $\mathbb{B}(p_t)$ with probability mass function $f_{\mathbb{I}[r_t>c]}(v)=p^v_t(1-p_t)^{1-v}, v\in {0,1}$, where $p_t$ denotes the conditional 'success' probability $Pr(r_t>c|F_{t-1})=E[\mathbb{I}[r_t>c]|F_{t-1}]$. Christoffersen and Diebold (2006) show a remarkable result that if data are generated by $r_t=\mu_t+\sigma_t\epsilon_t$, where $\mu_t=E[r_t|F_{t-1}]$, $\sigma^2_t=Var[r_t|F_{t-1}]$, and $\epsilon_t$ is a homoskedastic martingale difference with unit variance (i.e. can be modeled as GARCH process) and distribution function $\mathbb{F}_{\varepsilon}$, then
$$Pr[r_t>c|F_{t-1}]=1-\mathbb{F}_{\epsilon}\left(\frac{c-\mu_t}{\sigma_t}\right).$$
This suggests that time-varying volatility can generate sign predictability as long as $c-\mu_t\ne0$. Furthermore Christoffrsen (2007) derive a Gram-Charlier expansion of this distribution and show that $Pr[r_t>c|F_{t-1}]$ depend on the third and fourth conditional cumulants of the standardized errors $\epsilon_t$. Hence, sign predictability would arise from time variability in second and higher-order moments. This leads us to parametrize $p_t$ as a dynamic logit model:
$$p_t=\frac{e^{\theta_t}}{1+e^{\theta_t}}\quad\text{with}\quad\theta_t=\omega_d+\phi_d\mathbb{I}[r_{t-1}>c]+\pmb{y}^T_{t-1}\pmb{\delta}_d,$$
where the last term denotes macroeconomic variables (valuation ratios, interest rate) and realized measure (variance, bipower vriation, realized third and fourth moment of returns).
$$E[r_t|F_{t-1}] = c - E[|r_t-c||F_{t-1}]+2E[|r_t-c|\mathbb{I}[r_t>c]|F_{t-1}]$$
In terms of inference
\[\hat{r}_t=c-\hat{\psi}_t+2\hat{\xi}_t\]
Under conditional independence or if conditional dependence is weak we have
$$\xi_t=E[|r_t-c||F_{t-1}]E[\mathbb{I}[r_t>c]|F_{t-1}]=\psi_tp_t.$$
so,
$$\hat{r}_t=c+(2\hat{p}_t-1)\hat{\psi}_t.$$
Under the general case of dependence, the copula estimation is essential.
$$ln\psi_t=\omega_v+\beta_vln\psi_{t-1}+\gamma_vln|r_{t-1}-c|+\rho_v\mathbb{I}[r_{t-1}>c]+\pmb{x}^T_{t-1}\pmb{\delta}_v.$$ The second last term allows for regime-specific volatility dependence while the last term represents macroeconomic predictors of volatility. $\mathbb{D}$ can be modeled as constant parameter Weibull distribution (or others distributions with the shape parameter vector $\varsigma$ a function of the past).
$$Pr[r_t>c|F_{t-1}]=1-\mathbb{F}_{\epsilon}\left(\frac{c-\mu_t}{\sigma_t}\right).$$
This suggests that time-varying volatility can generate sign predictability as long as $c-\mu_t\ne0$. Furthermore Christoffrsen (2007) derive a Gram-Charlier expansion of this distribution and show that $Pr[r_t>c|F_{t-1}]$ depend on the third and fourth conditional cumulants of the standardized errors $\epsilon_t$. Hence, sign predictability would arise from time variability in second and higher-order moments. This leads us to parametrize $p_t$ as a dynamic logit model:
$$p_t=\frac{e^{\theta_t}}{1+e^{\theta_t}}\quad\text{with}\quad\theta_t=\omega_d+\phi_d\mathbb{I}[r_{t-1}>c]+\pmb{y}^T_{t-1}\pmb{\delta}_d,$$
where the last term denotes macroeconomic variables (valuation ratios, interest rate) and realized measure (variance, bipower vriation, realized third and fourth moment of returns).
c) Copula model:To construct the bivariate conditional distribution of $R_t=[|r_t-c|, \mathbb{I}[r_t>c]]^T$ copula theory is used. In particular,$$F_{R_t}(u,v)=C(F_{|r_t-c|}(u), F_{\mathbb{I}[r_t>c]}(v))$$where $F$ denotes the CDF and $C(u,v)$ is a copula. Most common choices are Frank, Clayton or Farlie-Gumbel-Morgenstern copulas. Once the three ingredients of the joint distribution of $R_t$, i.e. the volatility model, the direction model and the copula are specified, the parameter vector can be estimated by maximum likelihood.
Conditional mean prediction in decomposition model
The main interest is the mean forecast of returns$$E[r_t|F_{t-1}] = c - E[|r_t-c||F_{t-1}]+2E[|r_t-c|\mathbb{I}[r_t>c]|F_{t-1}]$$
In terms of inference
\[\hat{r}_t=c-\hat{\psi}_t+2\hat{\xi}_t\]
Under conditional independence or if conditional dependence is weak we have
$$\xi_t=E[|r_t-c||F_{t-1}]E[\mathbb{I}[r_t>c]|F_{t-1}]=\psi_tp_t.$$
so,
$$\hat{r}_t=c+(2\hat{p}_t-1)\hat{\psi}_t.$$
Under the general case of dependence, the copula estimation is essential.
Empirical Analysis
TBDSunday, June 28, 2015
Commodity index investing
Commodity Index investing and commodity futures prices - Stoll and Whaley (2009)
Provide a comprehensive evaluation of whether commodity index investing is a disruptive force in commodity futures market in general. Institutional investors are active in it because of low correlation with stocks and bonds using managed futures, ETF, ETN and OTC return swaps. main conclusions are a) commodity index investment is not speculation b) rolls have little futures price impact, and inflows and outflows from index do not cause the prices to change. c) failure of wheat futures to converge to the cash price at expiration has not undermined the futures contract's effectiveness as a risk management tool.Limits to Arbitrage and Commodity Index investment: front-running the Goldman roll - Mou (2011)
Rolling causes price impact. Front running has shown IR of 4.4 from 2000 to 2010. Profitability is positively correlated to size of index investment and amount of arbitrage capital employed. Talks about pre-rolling 10-1 business days before GSCI rolling.Speculators, Index investors, and commodity prices - Greely, Currie (2008)
Index investors take the risk of prices (away from producers) and hence do not bring any information and do not effect the prices. Speculators bring information about supply and demand and effect the prices. These activities in turn lowers the cost of capital to commodity producers.Saturday, June 20, 2015
Developing high frequency equity trading models: Infantino and Itzhaki (2010)
Seconds to minutes horizon. PCA based equity market neutral reversal strategy combined with regime switching gives handsome results.
Ultra high frequency traders (millisecond technology players) make their profits by providing liquidity. They do not attempt to correct the mispricing in high frequency domain (second to minute), due to their shorter holding periods (Jurek, Yang 2007).
The model is a mean-reversion model as described in Khandani and Lo (2007) - 'what happened to the quants?' - to analyze the quant meltdown of August 2007. The weight of security i at date t is given by,
$$w_{i,t}=-\frac{1}{N}(R_{i,t-k}-R_{m,t-k})$$
where $R_{m,t-k}=\frac{1}{n}\Sigma_{i=1}^{N} R_{i,t-k}$. This is a market neutral strategy. Daily re-balancing correspond to $k=1$. These produce huge IRs at daily frequency and even more impressive numbers as the frequency is increased to 60 mins to 5 mins. This assumes every security has a CAPM beta close to 1 (which will be addressed using PCA).
Avellaneda and Lee (2010) describe statistical arbitrage with holding period from seconds to weeks. Pairs trading is the fundamental idea based on the expectation that one stock tracks the other, after controlling for beta in the following relationship, for stock P and Q
$$\frac{dP_t}{P_t}=\alpha dt+\beta \frac{dQ_t}{Q_t}+dX_t,$$
where $X_t$ is the mean reverting process to be traded on. The stock returns can be decomposed to systematic and idiosyncratic components by using PCA giving
$$\frac{dP_t}{P_t}=\alpha dt+\Sigma_{j=1}^{n}\beta_j F_t^{(j)}+dX_t,$$
where $F^{(j)}_t$ represent the risk factors of the market/cluster under consideration.
These ideas will be merged and utilized in a slightly different sense in this paper.
PCA: We use PCA for valuation using OLS for predictive modeling. This statistical in nature as 'identity' of the risk factor is not cared about. At seconds time frame, instead of Debt to Equity ratio, Current ratio and Interest coverage it is the positioning and flow of hedge funds, brokers and asset managers which is much more a driving factor. Orthogonality of PCA also avoids multi-collinearity in OLS. PCA have also been shown to identify market factors without bias to market capitalization. Finally, PCA uses implicitly the variance-co-variance matrix of returns, giving different threshold for each stock reversion, based on different combination of PCs for each of them. This address the basic flaw of having to assume a general threshold for the entire universe, with a CAPM based beta close to one for every security.
Model description: The steps are -
1) Define the stock universe - 50 stocks randomly chosen from S&P500. 1000 will need clustering techniques. Collected top of the book bid-ask quotes on the tick data for each trading day (2009).
2) Intervalize dataset - one-second intervals using the first mid-price quote of the second.
3) Calculate log-returns - calculate log returns on the one-second mid-prices.
4) PCA - For N assets and T time steps, demean and calculate the eigenvectors for the first k eigenvalues (of covariance matrix $\Sigma$) as columns into $\Phi$ and then calculate the dimensionally reduced returns $D$ of principal components.
$$D = [\Phi^T(X-M)^T]^T$$
where $M$ is the mean vector of $X$.
5) Build prediction model - Following Campbell, Lo and MacKinlay (1997) we ran regression on future accumulated log returns with the last sum of H-period dimensionally reduced returns in the principal component space:
$$r_{t+1}+...+r_{t+H}=\beta_1\Sigma_0^H D_{t-i,1}+...+\beta_{k}\Sigma_{0}^H D_{t-i,k}+\eta_{t+H,H}$$
Which can be represented in matrix form as:
$$S = \hat{D_t}B.$$
To form the mean-value signal we add back the mean
$$\hat{S}=S + M_t$$
The base assumption is that the principal components explain the main risk factors that should drive the stock's returns in a systematic way, and the residuals are the noise we will try to get rid of. IF we see that the last H-period accumulated log-returns have been higher than the signal, we assume that the stock is overvalues and thus place a sett order. Thus the final signal is $\hat{S} - \Sigma^{H} r_i$.
Since this is a liquidity providing strategy, trading cost should hurt less, relatively. A lag of 1 second is assumed.
Results : shows a negative Sharpe of -1.84 for 2009 with a drawdown of -65% at an annualized volatility of 16.6%. Positive returns for first quarter and then negative.
Since principal components are the main risk factors, they are the ones who can justify the two regimes. Momentum regime is related to the sprouting of dislocations in the market - measured by the cross sectional volatility of the Principal components, $\sigma_D(t)$. The key observation is: as the short term changes in $\sigma_D$ appeared to be more pronounced (identified by very narrow peaks in the $\sigma_D$ time series), cumulative returns from the basic mean-reversion strategy seemed to decrease (or momentum sets up). Changes in $\sigma_D$ over time are defined by $\psi = d\sigma_D/dt$ and the cumulative returns of the basic strategy by $\rho(t)$. We define the measure $E_H$ as:
$$E_H(t)=\sqrt{\Sigma_{i=0}^H[\psi(t-i)]^2}.$$
There is a pretty consistent negative correlation between $E_H(t-1)$ and $\rho(t)$. This allows for the identification of the strength of mean reversion strategy in next second. The sprouting of principal components dislocation at time t triggers momentum at time t+1. The regime switching strategy would then follow $E_H(t)-E_H(t-1)$ at time t. IF this value is greater than zero, we understand that the dislocation is increasing and we trade on 'momentum', otherwise we stick to 'mean-reversion' behavior. We see that 'momentum' seems to be linked to the 'acceleration' of $\sigma_D$.
Results: After applying the regime switching conditioning the Sharpe is +7.67 (2009) with a max drawdown of -1.45% at 10.03% annualized volatility.
2) Selection of Eigenvectors - Selection of eigenvectors could be better. One can get a time series of number of eigenvectors that maximize the Sharpe at a given time, and then run an auto-regressive model to determine the number of eigenvectors to use in future.
3) NLS - Instead of taking simple sum of the returns, one can do weighted sum using beta functions, changing it from OLS to an NLS problem.
4) Other - Numerical speed can be obtained by using SVD decomposition instead of covariance matrix computations, using Marquardt Levenberg algorithm for NLS and GPU.
Ch 1: Introduction
We want a short term valuation and identify the regime if the market will act in line of against the valuation. With so much noise, we should not expect high precision in our solutions. We only need to be slightly precise to generate decent alpha in a high frequency environment, with approx holding periods on the orders of seconds to minutes. By fundamental law of active management: $IR = IC \sqrt{Breadth}$. where, $IR$ is the information ratio, $IC$ is the Information coefficient (correlation between predicted and real values) and $Breadth$ is the number of independent decisions made on a trading strategy in one year. An $IC$ of 0.05 is huge!Ultra high frequency traders (millisecond technology players) make their profits by providing liquidity. They do not attempt to correct the mispricing in high frequency domain (second to minute), due to their shorter holding periods (Jurek, Yang 2007).
The model is a mean-reversion model as described in Khandani and Lo (2007) - 'what happened to the quants?' - to analyze the quant meltdown of August 2007. The weight of security i at date t is given by,
$$w_{i,t}=-\frac{1}{N}(R_{i,t-k}-R_{m,t-k})$$
where $R_{m,t-k}=\frac{1}{n}\Sigma_{i=1}^{N} R_{i,t-k}$. This is a market neutral strategy. Daily re-balancing correspond to $k=1$. These produce huge IRs at daily frequency and even more impressive numbers as the frequency is increased to 60 mins to 5 mins. This assumes every security has a CAPM beta close to 1 (which will be addressed using PCA).
Avellaneda and Lee (2010) describe statistical arbitrage with holding period from seconds to weeks. Pairs trading is the fundamental idea based on the expectation that one stock tracks the other, after controlling for beta in the following relationship, for stock P and Q
$$\frac{dP_t}{P_t}=\alpha dt+\beta \frac{dQ_t}{Q_t}+dX_t,$$
where $X_t$ is the mean reverting process to be traded on. The stock returns can be decomposed to systematic and idiosyncratic components by using PCA giving
$$\frac{dP_t}{P_t}=\alpha dt+\Sigma_{j=1}^{n}\beta_j F_t^{(j)}+dX_t,$$
where $F^{(j)}_t$ represent the risk factors of the market/cluster under consideration.
These ideas will be merged and utilized in a slightly different sense in this paper.
Ch 2: The model
Log returns and cumulative returns: This is only the predictive part of the step. Regime switch will be tackled in next chapter. We use log returns ($ln(1+r_t)$ assumed to be normal) as compounding of returns is easy and normality holds when compounding returns over a larger period. Further prices have log normal distribution and log returns are close approximation of real returns, i.e. $ln(1+r) \approx r$, for $r<<1$. Also, by using cumulative returns we take advantage of CLT, and build a model to predict cumulative returns, with the cumulative returns of the principal components.PCA: We use PCA for valuation using OLS for predictive modeling. This statistical in nature as 'identity' of the risk factor is not cared about. At seconds time frame, instead of Debt to Equity ratio, Current ratio and Interest coverage it is the positioning and flow of hedge funds, brokers and asset managers which is much more a driving factor. Orthogonality of PCA also avoids multi-collinearity in OLS. PCA have also been shown to identify market factors without bias to market capitalization. Finally, PCA uses implicitly the variance-co-variance matrix of returns, giving different threshold for each stock reversion, based on different combination of PCs for each of them. This address the basic flaw of having to assume a general threshold for the entire universe, with a CAPM based beta close to one for every security.
Model description: The steps are -
1) Define the stock universe - 50 stocks randomly chosen from S&P500. 1000 will need clustering techniques. Collected top of the book bid-ask quotes on the tick data for each trading day (2009).
2) Intervalize dataset - one-second intervals using the first mid-price quote of the second.
3) Calculate log-returns - calculate log returns on the one-second mid-prices.
4) PCA - For N assets and T time steps, demean and calculate the eigenvectors for the first k eigenvalues (of covariance matrix $\Sigma$) as columns into $\Phi$ and then calculate the dimensionally reduced returns $D$ of principal components.
$$D = [\Phi^T(X-M)^T]^T$$
where $M$ is the mean vector of $X$.
5) Build prediction model - Following Campbell, Lo and MacKinlay (1997) we ran regression on future accumulated log returns with the last sum of H-period dimensionally reduced returns in the principal component space:
$$r_{t+1}+...+r_{t+H}=\beta_1\Sigma_0^H D_{t-i,1}+...+\beta_{k}\Sigma_{0}^H D_{t-i,k}+\eta_{t+H,H}$$
Which can be represented in matrix form as:
$$S = \hat{D_t}B.$$
To form the mean-value signal we add back the mean
$$\hat{S}=S + M_t$$
The base assumption is that the principal components explain the main risk factors that should drive the stock's returns in a systematic way, and the residuals are the noise we will try to get rid of. IF we see that the last H-period accumulated log-returns have been higher than the signal, we assume that the stock is overvalues and thus place a sett order. Thus the final signal is $\hat{S} - \Sigma^{H} r_i$.
Since this is a liquidity providing strategy, trading cost should hurt less, relatively. A lag of 1 second is assumed.
Results : shows a negative Sharpe of -1.84 for 2009 with a drawdown of -65% at an annualized volatility of 16.6%. Positive returns for first quarter and then negative.
Ch3: Regime switching model
The mean reversion model itself is not profitable, at all times. Change of market behavior has to be determined beyond the fair value (sentiment?). The two main regimes in which the market work is momentum and mean-reversion. Under momentum regime we expect the returns to further diverge from the theoretical returns. Adaptive market hypothesis is applicable, particularly, to high frequency world, where 'poker' is played and irrationality would be common, Lo and Mueller (2010).Since principal components are the main risk factors, they are the ones who can justify the two regimes. Momentum regime is related to the sprouting of dislocations in the market - measured by the cross sectional volatility of the Principal components, $\sigma_D(t)$. The key observation is: as the short term changes in $\sigma_D$ appeared to be more pronounced (identified by very narrow peaks in the $\sigma_D$ time series), cumulative returns from the basic mean-reversion strategy seemed to decrease (or momentum sets up). Changes in $\sigma_D$ over time are defined by $\psi = d\sigma_D/dt$ and the cumulative returns of the basic strategy by $\rho(t)$. We define the measure $E_H$ as:
$$E_H(t)=\sqrt{\Sigma_{i=0}^H[\psi(t-i)]^2}.$$
There is a pretty consistent negative correlation between $E_H(t-1)$ and $\rho(t)$. This allows for the identification of the strength of mean reversion strategy in next second. The sprouting of principal components dislocation at time t triggers momentum at time t+1. The regime switching strategy would then follow $E_H(t)-E_H(t-1)$ at time t. IF this value is greater than zero, we understand that the dislocation is increasing and we trade on 'momentum', otherwise we stick to 'mean-reversion' behavior. We see that 'momentum' seems to be linked to the 'acceleration' of $\sigma_D$.
Results: After applying the regime switching conditioning the Sharpe is +7.67 (2009) with a max drawdown of -1.45% at 10.03% annualized volatility.
Ch4: Potential improvements
1) Clustering - A different set of stocks may give us very different PCs. To counter that we can cluster the stocks into smaller buckets, each characterized by their PCs.2) Selection of Eigenvectors - Selection of eigenvectors could be better. One can get a time series of number of eigenvectors that maximize the Sharpe at a given time, and then run an auto-regressive model to determine the number of eigenvectors to use in future.
3) NLS - Instead of taking simple sum of the returns, one can do weighted sum using beta functions, changing it from OLS to an NLS problem.
4) Other - Numerical speed can be obtained by using SVD decomposition instead of covariance matrix computations, using Marquardt Levenberg algorithm for NLS and GPU.
Ch5: Conclusions
Alpha source found between ultra high frequency and traditional statistical arbitrage environment.
Subscribe to:
Posts (Atom)