Herron Topic 2 - Trading Strategies Based on Technical Analysis

FINA 6333 for Spring 2025

Author

Richard Herron

import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import pandas_datareader as pdr
import statsmodels.api as sm
import yfinance as yf
%precision 4
pd.options.display.float_format = '{:.4f}'.format
# %config InlineBackend.figure_format = 'retina'

Introduction

This notebook covers trading strategies based on technical analysis in three parts:

  1. What is technical analysis?
  2. Why might trading strategies based on technical analysis work (or not work)?
  3. Implement a simple moving average (SMA) trading strategy

I based this lecture notebook on Welch (2022, chap. 12), Lewinson (2020, chap. 2), and Murphy (1999). If you want to learn technical analysis, Murphy (1999) is the best reference and covers more than we can in a week or semester. The practice notebook will cover several other trading strategies based on technical analysis.

What is technical analysis?

Technical analysis is a methodology that analyzes past market data (e.g., prices and volume, plus open interest in futures and options markets) in an attempt to forecast future price movements. If technical analysis can predict future price movements, the market is not weak-form efficient. Welch (2022, sec. 12.2) provides the three degrees of market efficiency:

The Traditional Classification The traditional definition of market efficiency focuses on information. In the traditional classification, market efficiency comes in one of three primary degrees: weak, semi-strong, and strong.

Weak market efficiency says that all information in past prices is reflected in today’s stock prices so that technical analysis (trading based solely on historical price patterns) cannot be used to beat the market. Put differently, the market is the best technical analyst.

Semistrong market efficiency says that all public information is reflected in today’s stock prices, so that neither fundamental trading (based on underlying firm fundamentals, such as cash flows or discount rates) nor technical analysis can be used to beat the market. Put differently, the market is both the best technical and the best fundamental analyst.

Strong market efficiency says that all information, both public and private, is reflected in today’s stock prices, so that nothing — not even private insider information — can be used to beat the market. Put differently, the market is the best analyst and cannot be beat.

In this traditional classification, all finance professors nowadays believe that most U.S. financial markets are not strong-form efficient: Insider trading may be illegal, but it works. However, there are still arguments as to which markets are only semi-strong-form efficient or even only weak-form efficient.

Welch (2022, sec. 12.2) goes on to provide his own taxonomy of true, firm, mild, and nonbelievers in market efficiency. Chapter 12 summarizes market efficiency, classical finance, behavioral finance, arbitrage, limits to arbitrage, and their consequences for managers and investors. You can read Chapter 12 here. We will focus on technical analysis in this notebook, but Welch (2022) is excellent.

Why might trading strategies based on technical analysis work or not?

…Work?

Technical analysis relies on a few ideas:

  1. Market prices and volume reflect all relevant information, so we can focus on past prices and volume instead of fundamentals and news.
  2. Market prices move in trends and patterns driven by market participants.
  3. These trends and patterns tend to repeat themselves because market participants create them.

…Or Not?

The logic above is reasonable. However, if past market prices reflect all relevant information, they should also reflect any prices trends they predict. Therefore, any patterns should be self-defeating, and market prices should follow a random walk. As well, the signal-to-noise ratio in market prices is high! Still, technical analysis provides an opportunity to learn how to implement and back-test trading strategies in Python.

A Random Walk

In a random walk, the price tomorrow equals the price today plus a tiny drift plus noise. In math terms, a random walk is \[P_{t} = \rho P_{t-1} + m P_{t-1} + \varepsilon_t\] where \(m\) is a small drift term and \(\operatorname{E}[\varepsilon] = 0\). If \(\rho > 1\), prices would quickly increase, and, if \(\rho < 1\), prices would quickly decrease. Let us examine the historical record.

ff = (
    pdr.DataReader(
        name='F-F_Research_Data_Factors_daily',
        data_source='famafrench',
        start='1900'
    )
    [0]
    .assign(Mkt=lambda x: x['Mkt-RF'] + x['RF'])
    .div(100)
)
C:\Users\richa\AppData\Local\Temp\ipykernel_6816\3102596058.py:2: FutureWarning: The argument 'date_parser' is deprecated and will be removed in a future version. Please use 'date_format' instead, or read your data in as 'object' dtype and then call 'to_datetime'.
  pdr.DataReader(

We can compound market returns to impute market prices relative to the last day of June 1926.

prices = ff['Mkt'].add(1).cumprod()
prices.plot()
plt.title('Imputed Market Prices\nAssuming $1 on the Last Day of June 1926')
plt.ylabel('Imputed Market Price ($)')
plt.semilogy()
plt.show()

We need lagged prices to estimate \(\rho\). We will add 10 lags of \(P\) to help us understand the relation between past and future prices.

prices_w_lags = (
    pd.concat(
        objs=[prices.shift(t) for t in range(11)],
        keys=[f'Lag {t}' for t in range(11)],
        names=['Price'],
        axis=1,
    )
)
prices_w_lags.tail()
Price Lag 0 Lag 1 Lag 2 Lag 3 Lag 4 Lag 5 Lag 6 Lag 7 Lag 8 Lag 9 Lag 10
Date
2024-12-24 15038.5668 14870.9709 14778.3109 14617.9520 14633.0240 15110.9845 15182.7991 15109.2172 15114.2049 15207.4264 15073.7225
2024-12-26 15044.1311 15038.5668 14870.9709 14778.3109 14617.9520 14633.0240 15110.9845 15182.7991 15109.2172 15114.2049 15207.4264
2024-12-27 14870.6722 15044.1311 15038.5668 14870.9709 14778.3109 14617.9520 14633.0240 15110.9845 15182.7991 15109.2172 15114.2049
2024-12-30 14711.1099 14870.6722 15044.1311 15038.5668 14870.9709 14778.3109 14617.9520 14633.0240 15110.9845 15182.7991 15109.2172
2024-12-31 14645.9397 14711.1099 14870.6722 15044.1311 15038.5668 14870.9709 14778.3109 14617.9520 14633.0240 15110.9845 15182.7991

Now we can plot the correlation of price with its lags.

(
    prices_w_lags
    .dropna()
    .corr()
    .loc['Lag 0']
    .plot(kind='bar')
)
plt.title('Correlations between Imputed Market Price and Its Lags')
plt.ylabel('Correlation')
plt.show()

But these are pairwise correlations. If we estimate conditional correlations, we see that most of the price information is in the first lag!

y = prices_w_lags.dropna()['Lag 0'] # Pt
X = prices_w_lags.dropna().drop('Lag 0', axis=1).pipe(sm.add_constant) # Pt-1, Pt-2, ..., Pt-10
model = sm.OLS(endog=y, exog=X)
fit = model.fit(cov_type='HAC', cov_kwds={'maxlags': 10})
fit.summary()
OLS Regression Results
Dep. Variable: Lag 0 R-squared: 1.000
Model: OLS Adj. R-squared: 1.000
Method: Least Squares F-statistic: 2.195e+06
Date: Thu, 03 Apr 2025 Prob (F-statistic): 0.00
Time: 16:42:37 Log-Likelihood: -1.2518e+05
No. Observations: 25891 AIC: 2.504e+05
Df Residuals: 25880 BIC: 2.505e+05
Df Model: 10
Covariance Type: HAC
coef std err z P>|z| [0.025 0.975]
const -0.0144 0.144 -0.100 0.920 -0.297 0.268
Lag 1 0.9574 0.033 29.042 0.000 0.893 1.022
Lag 2 0.0621 0.058 1.077 0.281 -0.051 0.175
Lag 3 -0.0315 0.037 -0.860 0.390 -0.103 0.040
Lag 4 -0.0220 0.040 -0.553 0.580 -0.100 0.056
Lag 5 0.0290 0.038 0.767 0.443 -0.045 0.103
Lag 6 -0.0498 0.042 -1.181 0.237 -0.132 0.033
Lag 7 0.1215 0.048 2.518 0.012 0.027 0.216
Lag 8 -0.1222 0.053 -2.302 0.021 -0.226 -0.018
Lag 9 0.1331 0.046 2.908 0.004 0.043 0.223
Lag 10 -0.0771 0.030 -2.535 0.011 -0.137 -0.017
Omnibus: 15929.955 Durbin-Watson: 1.994
Prob(Omnibus): 0.000 Jarque-Bera (JB): 4213325.277
Skew: -1.820 Prob(JB): 0.00
Kurtosis: 65.389 Cond. No. 9.90e+03


Notes:
[1] Standard Errors are heteroscedasticity and autocorrelation robust (HAC) using 10 lags and without small sample correction
[2] The condition number is large, 9.9e+03. This might indicate that there are
strong multicollinearity or other numerical problems.
plt.bar(
    x=fit.params.index[1:],
    height=fit.params[1:],
    yerr=2*fit.bse[1:]
)
plt.title('Conditional Correlations between Imputed Market Price and Its Lags\nBars Indicate Two Standard Errors')
plt.ylabel('Conditional Correlation')
plt.xlabel('Price')
plt.show()

Signal-to-Noise Ratio

Recall, we can express a random walk as \(P_{t} = \rho P_{t-1} + m P_{t-1} + \varepsilon_t\). Since \(\rho = 1\), we can subtract \(P_{t-1}\) from both sides, then divide by \(P_{t-1}\) on both sides. This transformation expresses a random walk in terms of returns: \(r_{t-1,t} = m + e_t\), where \(\operatorname{E}[e_t] = 0\) and \(\operatorname{SD}[e_t] = s\), so \(\operatorname{E}[r_{t-1, t}] = m\). We can think of the signal-to-noise ratio as \(\frac{m}{s}\). How high is this ratio?

m, s = ff['Mkt'].mean(), ff['Mkt'].std()

Here \(m\) is about 4 basis points per day!

m
0.0004

However, \(s\) is about 108 basis points per day!

s
0.0108

Therefore, the signal-to-noise ratio is less than 0.04! We want this ratio above 2 to reject that a drift (or a strategy) is zero.

m/s
0.0397

Recall that means grow linearly with time and standard deviations growth with the square-root of time. So, if we want \(\sqrt{t} \times \frac{m}{s} \geq 2\), we need \(t \geq \left(2 \times \frac{s}{m} \right)^2\) days! Even with market portfolio noise, which is diversified and low, we needat least a decade! During this decade, the true values of \(m\) and \(s\) can change!

(2 * s / m)**2 / 252
10.0486

Implement a simple moving average (SMA) trading strategy

The over-simplified goal of technical analysis is to “buy low, and sell high.” The \(n\)-day SMA reduces noise in market prices, removing market fluctuations and providing estimates of “true” prices. While the market price is above the SMA, the SMA rises. While the market price is below the SMA, the SMA falls. So, if we buy the stock as it cross the SMA from below and sell the stock as it crosses the SMA from above, we mechanically buy low and sell high! Here, we will implement a long-only 20-day SMA strategy with Bitcoin:

  1. Buy when the closing price crosses SMA(20) from below
  2. Sell when the closing price crosses SMA(20) from above
  3. No short-selling

We can simplify this strategy to “long if above SMA(20), otherwise neutral”. First, we will need Bitcoin returns data.

btc = (
    yf.download(
        tickers='BTC-USD',
        auto_adjust=False,
        progress=False,
        multi_level_index=False
    )
    .assign(Return=lambda x: x['Adj Close'].pct_change())
)
btc.head()
Adj Close Close High Low Open Volume Return
Date
2014-09-17 457.3340 457.3340 468.1740 452.4220 465.8640 21056800 NaN
2014-09-18 424.4400 424.4400 456.8600 413.1040 456.8600 34483200 -0.0719
2014-09-19 394.7960 394.7960 427.8350 384.5320 424.1030 37919700 -0.0698
2014-09-20 408.9040 408.9040 423.2960 389.8830 394.6730 36863600 0.0357
2014-09-21 398.8210 398.8210 412.4260 393.1810 408.0850 26580100 -0.0247

Next we:

  1. Use .rolling(20).mean() to add a SMA20 column containing SMA(20) to our btc data frame
  2. Use np.select() to add a Position column containing:
    1. 1 (long) when the adjusted close is greater than SMA(20)
    2. 0 (neutral) when the adjusted close is less than (or equal to) SMA(20)
    3. We use .shift() to compare yesterday’s closing prices, avoiding a look-ahead bias
    4. np.select() tests multiple conditions and provides a default, making it more flexible framework than np.where()
  3. Add a Strategy column containing:
    1. Return if Position == 1
    2. 0 if Position == 0
    3. We could earn the risk-free rate instead of 0 percent, but earning 0 percent simplifies this example
btc = (
    btc
    .assign(
        SMA20=lambda x: x['Adj Close'].rolling(20).mean(),
        Position=lambda x: np.select(
            condlist=[
                x['Adj Close'].shift() > x['SMA20'].shift(), 
                x['Adj Close'].shift() <= x['SMA20'].shift()
            ],
            choicelist=[1, 0],
            default=np.nan
        ),
        Strategy=lambda x: x['Position'] * x['Return']
    )
)

I find it helpful to plot Adj Close, SMA20, and Position for a sort window with one or more crossings.

fig, ax = plt.subplots(2, 1, sharex=True)
df = btc.loc['2023-02-13':'2023-02-28']
df[['Adj Close', 'SMA20']].plot(ax=ax[0], ylabel='BTC-USD ($)')
df[['Position']].plot(ax=ax[1], ylabel='Position', legend=False)
plt.suptitle('Bitcoin SMA(20) Strategy')
plt.show()

We can compare the long-run performance of buy-and-hold and SMA(20).

df = btc[['Return', 'Strategy']].dropna()

(
    df
    .add(1)
    .cumprod()
    .rename_axis(columns='Strategy')
    .rename(columns={'Return': 'Buy-And-Hold', 'Strategy': 'SMA(20)'})
    .plot()
)
plt.semilogy()
plt.ylabel('Value ($)')
plt.title(f'Value of $1 Invested at Close on {df.index[0] - pd.offsets.Day(1):%B %d, %Y}')
plt.show()

In the practice notebook, we will dig deeper on this strategy and others.

df.add(1).prod()
Return     248.1171
Strategy   237.1394
dtype: float64

References

Lewinson, Eryk. 2020. Python for Finance Cookbook: Over 50 Recipes for Applying Modern Python Libraries to Financial Data Analysis. Packt Publishing Ltd.
Murphy, John J. 1999. Technical Analysis of the Financial Markets: A Comprehensive Guide to Trading Methods and Applications. Penguin.
Welch, Ivo. 2022. Corporate Finance. 5th ed. https://book.ivo-welch.info/home/; Ivo Welch.