import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import pandas_datareader as pdr
import statsmodels.api as sm
import yfinance as yfHerron Topic 2 - Trading Strategies Based on Technical Analysis
FINA 6333 for Spring 2025
%precision 4
pd.options.display.float_format = '{:.4f}'.format
# %config InlineBackend.figure_format = 'retina'Introduction
This notebook covers trading strategies based on technical analysis in three parts:
- What is technical analysis?
- Why might trading strategies based on technical analysis work (or not work)?
- Implement a simple moving average (SMA) trading strategy
I based this lecture notebook on Welch (2022, chap. 12), Lewinson (2020, chap. 2), and Murphy (1999). If you want to learn technical analysis, Murphy (1999) is the best reference and covers more than we can in a week or semester. The practice notebook will cover several other trading strategies based on technical analysis.
What is technical analysis?
Technical analysis is a methodology that analyzes past market data (e.g., prices and volume, plus open interest in futures and options markets) in an attempt to forecast future price movements. If technical analysis can predict future price movements, the market is not weak-form efficient. Welch (2022, sec. 12.2) provides the three degrees of market efficiency:
The Traditional Classification The traditional definition of market efficiency focuses on information. In the traditional classification, market efficiency comes in one of three primary degrees: weak, semi-strong, and strong.
Weak market efficiency says that all information in past prices is reflected in today’s stock prices so that technical analysis (trading based solely on historical price patterns) cannot be used to beat the market. Put differently, the market is the best technical analyst.
Semistrong market efficiency says that all public information is reflected in today’s stock prices, so that neither fundamental trading (based on underlying firm fundamentals, such as cash flows or discount rates) nor technical analysis can be used to beat the market. Put differently, the market is both the best technical and the best fundamental analyst.
Strong market efficiency says that all information, both public and private, is reflected in today’s stock prices, so that nothing — not even private insider information — can be used to beat the market. Put differently, the market is the best analyst and cannot be beat.
In this traditional classification, all finance professors nowadays believe that most U.S. financial markets are not strong-form efficient: Insider trading may be illegal, but it works. However, there are still arguments as to which markets are only semi-strong-form efficient or even only weak-form efficient.
Welch (2022, sec. 12.2) goes on to provide his own taxonomy of true, firm, mild, and nonbelievers in market efficiency. Chapter 12 summarizes market efficiency, classical finance, behavioral finance, arbitrage, limits to arbitrage, and their consequences for managers and investors. You can read Chapter 12 here. We will focus on technical analysis in this notebook, but Welch (2022) is excellent.
Why might trading strategies based on technical analysis work or not?
…Work?
Technical analysis relies on a few ideas:
- Market prices and volume reflect all relevant information, so we can focus on past prices and volume instead of fundamentals and news.
- Market prices move in trends and patterns driven by market participants.
- These trends and patterns tend to repeat themselves because market participants create them.
…Or Not?
The logic above is reasonable. However, if past market prices reflect all relevant information, they should also reflect any prices trends they predict. Therefore, any patterns should be self-defeating, and market prices should follow a random walk. As well, the signal-to-noise ratio in market prices is high! Still, technical analysis provides an opportunity to learn how to implement and back-test trading strategies in Python.
A Random Walk
In a random walk, the price tomorrow equals the price today plus a tiny drift plus noise. In math terms, a random walk is \[P_{t} = \rho P_{t-1} + m P_{t-1} + \varepsilon_t\] where \(m\) is a small drift term and \(\operatorname{E}[\varepsilon] = 0\). If \(\rho > 1\), prices would quickly increase, and, if \(\rho < 1\), prices would quickly decrease. Let us examine the historical record.
ff = (
pdr.DataReader(
name='F-F_Research_Data_Factors_daily',
data_source='famafrench',
start='1900'
)
[0]
.assign(Mkt=lambda x: x['Mkt-RF'] + x['RF'])
.div(100)
)C:\Users\richa\AppData\Local\Temp\ipykernel_6816\3102596058.py:2: FutureWarning: The argument 'date_parser' is deprecated and will be removed in a future version. Please use 'date_format' instead, or read your data in as 'object' dtype and then call 'to_datetime'.
pdr.DataReader(
We can compound market returns to impute market prices relative to the last day of June 1926.
prices = ff['Mkt'].add(1).cumprod()prices.plot()
plt.title('Imputed Market Prices\nAssuming $1 on the Last Day of June 1926')
plt.ylabel('Imputed Market Price ($)')
plt.semilogy()
plt.show()
We need lagged prices to estimate \(\rho\). We will add 10 lags of \(P\) to help us understand the relation between past and future prices.
prices_w_lags = (
pd.concat(
objs=[prices.shift(t) for t in range(11)],
keys=[f'Lag {t}' for t in range(11)],
names=['Price'],
axis=1,
)
)prices_w_lags.tail()| Price | Lag 0 | Lag 1 | Lag 2 | Lag 3 | Lag 4 | Lag 5 | Lag 6 | Lag 7 | Lag 8 | Lag 9 | Lag 10 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Date | |||||||||||
| 2024-12-24 | 15038.5668 | 14870.9709 | 14778.3109 | 14617.9520 | 14633.0240 | 15110.9845 | 15182.7991 | 15109.2172 | 15114.2049 | 15207.4264 | 15073.7225 |
| 2024-12-26 | 15044.1311 | 15038.5668 | 14870.9709 | 14778.3109 | 14617.9520 | 14633.0240 | 15110.9845 | 15182.7991 | 15109.2172 | 15114.2049 | 15207.4264 |
| 2024-12-27 | 14870.6722 | 15044.1311 | 15038.5668 | 14870.9709 | 14778.3109 | 14617.9520 | 14633.0240 | 15110.9845 | 15182.7991 | 15109.2172 | 15114.2049 |
| 2024-12-30 | 14711.1099 | 14870.6722 | 15044.1311 | 15038.5668 | 14870.9709 | 14778.3109 | 14617.9520 | 14633.0240 | 15110.9845 | 15182.7991 | 15109.2172 |
| 2024-12-31 | 14645.9397 | 14711.1099 | 14870.6722 | 15044.1311 | 15038.5668 | 14870.9709 | 14778.3109 | 14617.9520 | 14633.0240 | 15110.9845 | 15182.7991 |
Now we can plot the correlation of price with its lags.
(
prices_w_lags
.dropna()
.corr()
.loc['Lag 0']
.plot(kind='bar')
)
plt.title('Correlations between Imputed Market Price and Its Lags')
plt.ylabel('Correlation')
plt.show()
But these are pairwise correlations. If we estimate conditional correlations, we see that most of the price information is in the first lag!
y = prices_w_lags.dropna()['Lag 0'] # Pt
X = prices_w_lags.dropna().drop('Lag 0', axis=1).pipe(sm.add_constant) # Pt-1, Pt-2, ..., Pt-10
model = sm.OLS(endog=y, exog=X)
fit = model.fit(cov_type='HAC', cov_kwds={'maxlags': 10})
fit.summary()| Dep. Variable: | Lag 0 | R-squared: | 1.000 |
| Model: | OLS | Adj. R-squared: | 1.000 |
| Method: | Least Squares | F-statistic: | 2.195e+06 |
| Date: | Thu, 03 Apr 2025 | Prob (F-statistic): | 0.00 |
| Time: | 16:42:37 | Log-Likelihood: | -1.2518e+05 |
| No. Observations: | 25891 | AIC: | 2.504e+05 |
| Df Residuals: | 25880 | BIC: | 2.505e+05 |
| Df Model: | 10 | ||
| Covariance Type: | HAC |
| coef | std err | z | P>|z| | [0.025 | 0.975] | |
| const | -0.0144 | 0.144 | -0.100 | 0.920 | -0.297 | 0.268 |
| Lag 1 | 0.9574 | 0.033 | 29.042 | 0.000 | 0.893 | 1.022 |
| Lag 2 | 0.0621 | 0.058 | 1.077 | 0.281 | -0.051 | 0.175 |
| Lag 3 | -0.0315 | 0.037 | -0.860 | 0.390 | -0.103 | 0.040 |
| Lag 4 | -0.0220 | 0.040 | -0.553 | 0.580 | -0.100 | 0.056 |
| Lag 5 | 0.0290 | 0.038 | 0.767 | 0.443 | -0.045 | 0.103 |
| Lag 6 | -0.0498 | 0.042 | -1.181 | 0.237 | -0.132 | 0.033 |
| Lag 7 | 0.1215 | 0.048 | 2.518 | 0.012 | 0.027 | 0.216 |
| Lag 8 | -0.1222 | 0.053 | -2.302 | 0.021 | -0.226 | -0.018 |
| Lag 9 | 0.1331 | 0.046 | 2.908 | 0.004 | 0.043 | 0.223 |
| Lag 10 | -0.0771 | 0.030 | -2.535 | 0.011 | -0.137 | -0.017 |
| Omnibus: | 15929.955 | Durbin-Watson: | 1.994 |
| Prob(Omnibus): | 0.000 | Jarque-Bera (JB): | 4213325.277 |
| Skew: | -1.820 | Prob(JB): | 0.00 |
| Kurtosis: | 65.389 | Cond. No. | 9.90e+03 |
Notes:
[1] Standard Errors are heteroscedasticity and autocorrelation robust (HAC) using 10 lags and without small sample correction
[2] The condition number is large, 9.9e+03. This might indicate that there are
strong multicollinearity or other numerical problems.
plt.bar(
x=fit.params.index[1:],
height=fit.params[1:],
yerr=2*fit.bse[1:]
)
plt.title('Conditional Correlations between Imputed Market Price and Its Lags\nBars Indicate Two Standard Errors')
plt.ylabel('Conditional Correlation')
plt.xlabel('Price')
plt.show()
Signal-to-Noise Ratio
Recall, we can express a random walk as \(P_{t} = \rho P_{t-1} + m P_{t-1} + \varepsilon_t\). Since \(\rho = 1\), we can subtract \(P_{t-1}\) from both sides, then divide by \(P_{t-1}\) on both sides. This transformation expresses a random walk in terms of returns: \(r_{t-1,t} = m + e_t\), where \(\operatorname{E}[e_t] = 0\) and \(\operatorname{SD}[e_t] = s\), so \(\operatorname{E}[r_{t-1, t}] = m\). We can think of the signal-to-noise ratio as \(\frac{m}{s}\). How high is this ratio?
m, s = ff['Mkt'].mean(), ff['Mkt'].std()Here \(m\) is about 4 basis points per day!
m0.0004
However, \(s\) is about 108 basis points per day!
s0.0108
Therefore, the signal-to-noise ratio is less than 0.04! We want this ratio above 2 to reject that a drift (or a strategy) is zero.
m/s0.0397
Recall that means grow linearly with time and standard deviations growth with the square-root of time. So, if we want \(\sqrt{t} \times \frac{m}{s} \geq 2\), we need \(t \geq \left(2 \times \frac{s}{m} \right)^2\) days! Even with market portfolio noise, which is diversified and low, we needat least a decade! During this decade, the true values of \(m\) and \(s\) can change!
(2 * s / m)**2 / 25210.0486
Implement a simple moving average (SMA) trading strategy
The over-simplified goal of technical analysis is to “buy low, and sell high.” The \(n\)-day SMA reduces noise in market prices, removing market fluctuations and providing estimates of “true” prices. While the market price is above the SMA, the SMA rises. While the market price is below the SMA, the SMA falls. So, if we buy the stock as it cross the SMA from below and sell the stock as it crosses the SMA from above, we mechanically buy low and sell high! Here, we will implement a long-only 20-day SMA strategy with Bitcoin:
- Buy when the closing price crosses SMA(20) from below
- Sell when the closing price crosses SMA(20) from above
- No short-selling
We can simplify this strategy to “long if above SMA(20), otherwise neutral”. First, we will need Bitcoin returns data.
btc = (
yf.download(
tickers='BTC-USD',
auto_adjust=False,
progress=False,
multi_level_index=False
)
.assign(Return=lambda x: x['Adj Close'].pct_change())
)btc.head()| Adj Close | Close | High | Low | Open | Volume | Return | |
|---|---|---|---|---|---|---|---|
| Date | |||||||
| 2014-09-17 | 457.3340 | 457.3340 | 468.1740 | 452.4220 | 465.8640 | 21056800 | NaN |
| 2014-09-18 | 424.4400 | 424.4400 | 456.8600 | 413.1040 | 456.8600 | 34483200 | -0.0719 |
| 2014-09-19 | 394.7960 | 394.7960 | 427.8350 | 384.5320 | 424.1030 | 37919700 | -0.0698 |
| 2014-09-20 | 408.9040 | 408.9040 | 423.2960 | 389.8830 | 394.6730 | 36863600 | 0.0357 |
| 2014-09-21 | 398.8210 | 398.8210 | 412.4260 | 393.1810 | 408.0850 | 26580100 | -0.0247 |
Next we:
- Use
.rolling(20).mean()to add aSMA20column containing SMA(20) to ourbtcdata frame - Use
np.select()to add aPositioncolumn containing:1(long) when the adjusted close is greater than SMA(20)0(neutral) when the adjusted close is less than (or equal to) SMA(20)- We use
.shift()to compare yesterday’s closing prices, avoiding a look-ahead bias np.select()tests multiple conditions and provides a default, making it more flexible framework thannp.where()
- Add a
Strategycolumn containing:ReturnifPosition == 10ifPosition == 0- We could earn the risk-free rate instead of 0 percent, but earning 0 percent simplifies this example
btc = (
btc
.assign(
SMA20=lambda x: x['Adj Close'].rolling(20).mean(),
Position=lambda x: np.select(
condlist=[
x['Adj Close'].shift() > x['SMA20'].shift(),
x['Adj Close'].shift() <= x['SMA20'].shift()
],
choicelist=[1, 0],
default=np.nan
),
Strategy=lambda x: x['Position'] * x['Return']
)
)I find it helpful to plot Adj Close, SMA20, and Position for a sort window with one or more crossings.
fig, ax = plt.subplots(2, 1, sharex=True)
df = btc.loc['2023-02-13':'2023-02-28']
df[['Adj Close', 'SMA20']].plot(ax=ax[0], ylabel='BTC-USD ($)')
df[['Position']].plot(ax=ax[1], ylabel='Position', legend=False)
plt.suptitle('Bitcoin SMA(20) Strategy')
plt.show()
We can compare the long-run performance of buy-and-hold and SMA(20).
df = btc[['Return', 'Strategy']].dropna()
(
df
.add(1)
.cumprod()
.rename_axis(columns='Strategy')
.rename(columns={'Return': 'Buy-And-Hold', 'Strategy': 'SMA(20)'})
.plot()
)
plt.semilogy()
plt.ylabel('Value ($)')
plt.title(f'Value of $1 Invested at Close on {df.index[0] - pd.offsets.Day(1):%B %d, %Y}')
plt.show()
In the practice notebook, we will dig deeper on this strategy and others.
df.add(1).prod()Return 248.1171
Strategy 237.1394
dtype: float64