匡醍量化|大富翁量化

Factor ML Course Syllabus: From Alpha Mining to Live Trading

中文 📅 2024-08-27 👁 views this month —

§ Factor Investing & Machine Learning Strategies

Course Syllabus

Syllabus Notes

1. Introduction

1.1. The Origins of Factor Investing

1.2. Hunting for Alpha

1.3. From CAPM to Multi-Factor Models

1.4. From Factor Analysis to Factor Investing

1.5. From Factor Models to Trading Strategies

1.6. Course Structure

This course is designed for professional quantitative researchers, those transitioning into the field, or professionals from other disciplines determined to explore quantitative research with a rigorous, professional attitude.

Upon completing and mastering this course, you will possess proficient factor analysis skills, master state-of-the-art machine learning strategy construction methods, and become a quantitative researcher with innovative research capabilities and a competitive edge.

The curriculum covers the entire process from factor mining and testing to building machine learning models. If you intend to engage in independent trading, you should supplement your learning with "Quantitative Trading 24 Lessons."


2. Factor Preprocessing Pipeline

2.1. Factor Data Sources

2.2. Factor Generation

2.3. Factor Preprocessing

2.3.1. Outlier Clipping

2.3.2. Missing Values Handling

2.3.3. Distribution Adjustment

2.3.4. Standardization

2.3.5. Neutralization

This chapter and the next form the foundation for factor testing. We will introduce the basic principles and technical implementation details of factor testing using extensive example code, laying a solid groundwork for understanding the Alphalens factor analysis framework.


3. Factor Testing Methods

3.1. Regression Method

3.2. IC Analysis

3.3. Layered Backtest Method

3.4. Code Implementation of Factor Testing

3.4.1. Generating Factors: Further Modularization

3.4.2. Factor Preprocessing: Integrating Real Data

3.4.3. Calculating Forward Returns

3.4.4. Implementing Regression Analysis

3.4.5. Implementing IC Analysis

3.4.6. Implementing Layered Backtest

3.5. Differences and Connections Between the Three Methods

This chapter introduces the principles and implementation code for the regression method, IC method, and layered backtest method. Upon completion, you will be able to implement a simple factor analysis framework yourself, which is highly beneficial for understanding Alphalens.


4. Introduction to Alphalens

4.1. Slope Factor: Definition and Implementation

4.2. How to Calculate Factors and Collect Price Data for Alphalens

4.3. How Alphalens Handles Data Preprocessing

4.4. Factor Analysis and Report Generation

4.5. References

Alphalens highly abstracts the factor testing process, encapsulating the steps discussed in the previous two chapters into two functions, with extensive customization achieved through parameters. We will introduce the input data format required by Alphalens and how it uses parameters to control layering, missing value handling, and forward return calculation behaviors.

By the end of this chapter, you will master the most basic usage of Alphalens.


5. Alphalens Report Analysis

5.1. Return Analysis

5.1.1. Alpha and Beta

5.1.2. Layered Return Mean Chart

5.1.3. Layered Return Violin Plot

5.1.4. Factor-Weighted Long-Short Portfolio Cumulative Return Chart

5.1.5. Layered Drivers of Returns

5.1.6. Robustness of Long-Short Portfolio Returns

5.2. Event Study

5.3. IC Analysis

5.4. Turnover Analysis

5.5. References

Alphalens reports are not self-explanatory. For instance, it doesn't specify the units for Alpha and Beta, nor what a basis point (bps) unit represents; it certainly won't tell you what constitutes a "good" Alpha versus one that is "too good to be true." Some reports calculate metrics differently from what you might expect or have heard.

To accurately understand these reports, we employ three approaches: 1. Reading and debugging the source code. Through this, we discovered that bps stands for one ten-thousandth, defined in the plotting.py file. 2. Using synthetic data, which helps us understand what theoretical reports for the best factors should look like. 3. Searching through GitHub issues and the Quantopian community Archive documentation to find answers in other users' questions.

This is currently the only tutorial on the internet that thoroughly explains Alphalens.


6. Alphalens Advanced Techniques (1)

6.1. Excluding Functional Errors

6.1.1. Outdated Alphalens Versions

6.1.2. MaxLossExceedError

6.1.3. Timezone Issues

6.2. Factor Monotonicity

6.3. Revisiting Price Data Collection

6.4. How to Analyze Factors Above Daily Frequency?

6.5. Deep Dive into Factor Layering

6.5.1. Factors with Determined Trading Signals

6.5.2. Discrete Value Factors

In this chapter, we introduce how to troubleshoot potential errors in Alphalens, both programmatic and logical. We also discuss how to conduct factor analysis at frequencies higher than daily. Many online tutorials don't even realize this presents a problem because they have never performed analysis at this level.

We also delve into Alphalens' layering mechanism, including how to handle cases where factor values are discrete.


7. Alphalens Advanced Techniques (2)

7.1. Refactoring the Factor Testing Process

7.2. Parameter Tuning: Saving Your Factors

7.2.1. Correcting Factor Direction

7.2.2. Filtering Non-Linear Layering

7.2.3. Using Optimal Layering Methods

7.2.4. Grid Search

7.3. Overfitting Detection Methods

7.3.1. Out-of-Sample Testing

7.3.2. Plotting Parameter Plateaus

Using Alphalens for factor testing is like an interview; you must expose the factor's potential as much as possible before evaluating its quality. This chapter introduces how to exhaustively挖掘 (mine) a factor's potential while avoiding the deception of overfitting. Besides out-of-sample testing, we will teach you how to assess the degree of overfitting by plotting parameter plateaus.

Visualization is crucial, especially when collaborating with others.


8. Alpha101 Factor Introduction

8.1. Data and Operators in Alpha101 Factors

8.2. Interpreting Alpha101 Factors

8.3. How to Implement Alpha101 Factors?

The Alpha101 factor library is a factor library published by World Quant in 2015. About 80% of the factors in it are officially used by World Quant (at the time of publication). We will introduce how to read the formulas of Alpha101 factors and implement their operators.

There are already good open-source libraries for implementing the entire factor library, which we will also introduce. This will become one of the treasures in your arsenal.


9. Talib Technical Factors

9.1. Ta-lib Function Grouping

9.2. Warm-up Periods (Unstable Periods)

9.3. Oscillation Indicators

9.3.1. RSI

9.3.2. ADX - Average Directional Movement Index

9.3.3. APO - Absolute Price Oscillator

9.3.4. PPO - Percentage Price Oscillator

9.3.5. Aroon Oscillator

9.3.6. Money Flow Index

9.3.7. Balance of Power

9.3.8. William's R

9.3.9. Stochastic Oscillator

9.4. Volume Indicators

9.4.1. Chaikin A/D Line

9.4.2. OBV

9.5. Volatility Indicators

9.5.1. ATR and NATR - Average True Range

9.6. 8 Types of Moving Averages

9.7. Overlap Studies

9.7.1. Bollinger Bands

9.7.2. Hilbert Trendline and Sine Wave Indicator

9.7.3. Parabolic SAR

9.8. Momentum Indicators

Most Alpha101 factors are volume-price factors. For understandable reasons, they do not repeat classic technical factors that have existed for years, but these factors still possess Alpha potential. In this section, we briefly introduce the talib library and explain the warm-up periods of technical indicators -- a potentially niche topic. The warm-up period is not just NaNs; for example, the warm-up period for RSI is quite long, being 3 times the win parameter.

There are many Talib technical indicators; we will introduce a few from each category, focusing on how to revamp these factors under new technical conditions. Taking RSI as an example, we will discuss intelli-RSI and Connor's RSI. This way, you not only gain new factors but also enhance your ability for innovative research.

Even experienced practitioners might be hearing about some of the factors we introduce for the first time, such as the Hilbert Sine Wave, which is one of the paid technical indicators sold on platforms like TradingView.


10. Other Volume-Price Factors

10.1. Low-Probability Events

10.2. Max Drawdown

10.3. pct_rank

10.4. Volatility

10.5. Z-Score

10.6. Sharpe Ratio

10.7. First-Derivative Factors

10.8. Second-Derivative Factors

10.9. Frequency-Domain Factors

10.10. TSFresh Factor Library

10.11. Behavioral Finance Factors

10.11.1. Integer Barrier Factors

10.11.2. Breakout Failure Factors

10.11.3. Gap Factors

10.11.4. Regret Aversion Theory Factors

Some low-probability factors are easy to construct. Perhaps because of this, they lack names and haven't found their way into academic papers. However, their Alpha is real. For example, the index's largest single-day drop or longest consecutive drop. The underlying principle is probability regression after extreme events.

In short, this is a show-off and innovative chapter. We will introduce second-derivative factors, frequency-domain factors, and behavioral finance factors. For instance, frequency-domain factors use Fast Fourier Transform (FFT) or wavelet transforms to identify the operational cycles of main market participants for prediction. While others are still using wavelets to smooth noise, we have already started using them to explore the patterns of main market participants!


11. Fundamental and Alternative Factors

11.1. Fama-French Five-Factor Model

11.1.1. Market Factor

11.1.2. Size Factor

11.1.3. Value Factor

11.1.4. Profitability Factor

11.1.5. Investment Factor

11.2. Alternative Factors

11.2.1. Social Media Sentiment Factors

11.2.2. Web Traffic Factors

11.2.3. Satellite Image Factors

11.2.4. Patent Application Factors

11.3. Technical Routes for Web Crawling

This part will focus more on concepts. Because alternative factors are either purchased or crawled, we don't want to dwell on crawling techniques.


12. Factor Mining Methods

12.1. Where Do New Factors Come From?

12.2. Online Resources

12.3. Factor Orthogonality Testing

12.4. Discussing the Factor Zoo

This is also a chapter that discusses broad topics, but it is still full of干货 (substance). We will talk about methods for finding resources, such as how to find papers and data. By now, we have introduced hundreds of factors (excluding parameters and cycles), so we also need to see how many factors are truly independent. Therefore, we will introduce orthogonality testing methods.


13. Machine Learning Overview

13.1. Machine Learning Classification

13.1.1. Machine Learning, Deep Learning, Reinforcement Learning

13.1.2. Supervised, Unsupervised, and Reinforcement Learning

13.1.3. Regression and Classification

13.2. Introduction to Machine Learning Models

13.3. Three Elements of Machine Learning

13.4. Basic Workflow of Machine Learning

13.5. Application Scenarios for Machine Learning

A quick introduction to machine learning. Is the world continuous or quantized? This is an ancient philosophical question that also determines the basic models of machine learning -- regression or classification?


14. Core Concepts of Machine Learning

14.1. Bias and Variance

14.2. Overfitting and Regularization Penalties

14.3. Loss Functions, Objective Functions, Metric Functions, and Distance Functions

14.3.1. Loss Functions and Objective Functions

14.3.2. Metric Functions

14.3.3. Distance Functions

This course is an applied course and does not intend to involve too much theory. However, if you understand no principles at all, you can only copy examples without being able to extend them. Therefore, we decided to explain basic concepts closely related to the application layer -- only by understanding these concepts can we know how to choose objective functions, evaluate strategies, prevent overfitting, etc.


15. SKLearn General Toolkit

15.1. Data Preprocessing: preprocessing

15.2. metrics

15.3. Model Interpretation and Visualization

15.4. Built-in Datasets

sklearn is a very powerful machine learning library, winning people's love with its rich models and easy-to-use interface. In this chapter, we first introduce sklearn's general toolkit -- used to handle common problems encountered regardless of which algorithm model we adopt, such as data preprocessing, model evaluation, model interpretation and visualization, and built-in datasets.


16. Model Optimization

16.1. Optimization Overview

16.2. k-fold Cross Validation

16.3.1. Grid Search

16.3.2. Random Search

16.3.3. Bayesian Optimization

16.4. Rolling Forecasting

Machine learning in the quantitative field has its own特殊性 (special characteristics), such as in cross-validation, where we actually need to use a method called Rolling Forecasting (also known as Walk-Forward Optimization).


17. Clustering: Finding Pair Trading Targets

17.1. Overview of Clustering Algorithms

17.2. HDBSCAN Algorithm Principles

17.3. Finding Pair Trading Targets

17.3.1. HDBSCAN Example

17.3.2. Result Evaluation

17.3.3. Pair Selection

In quantitative trading, Pair Trading is an important type of arbitrage strategy, with the prerequisite of finding two targets that can be paired. In this chapter, we will introduce the advanced HDBSCAN clustering method, demonstrate how to implement clustering through it, and then use related methods in statsmodels to perform cointegration pair tests to find targets that can be paired. Finally, we will demonstrate how to assemble all of this into a complete trading strategy.

This will be the first effective machine learning strategy you learn.


18. From Decision Trees to LightGBM

18.1. Decision Trees

18.1.1. Decision Tree Classification

18.1.2. Decision Tree Regression

18.2. LightGBM

18.2.1. Familiarizing with Training Data

18.2.2. Building the First Classifier

18.2.3. Visualizing Feature Importance

18.2.4. Viewing Model Trees

18.2.5. Cross-Validation

18.2.6. Tuning

Limited by the high noise in financial data, end-to-end trading strategies are not yet feasible; also limited by the size of labeled data, deep learning and other AI models are not suitable for constructing trading strategies. Among machine learning models, the best model at present is the Gradient Boosting Decision Tree (GBDT) model. Representative implementations are XGBoost and LightGBM.

Since LightGBM surpasses XGBoost in most tasks in terms of both speed and accuracy, our course will focus on introducing LightGBM.

This chapter will comprehensively introduce the LightGBM model and demonstrate through examples how to use it, how to inspect and visualize the generated models, how to perform cross-validation, and how to tune parameters.


19. Price Prediction Based on LightGBM Regression Model

19.1. Strategy Principles

19.2. Strategy Implementation

19.3. Strategy Optimization Ideas

Asset pricing is one of the core issues in quantitative research. If we can provide a reasonable pricing for assets, we can generate trading signals.

Pricing is a regression problem. Although it is difficult to implement end-to-end price prediction models, we have cleverly designed a regression model that can predict future prices (theoretically self-consistent).

We cannot guarantee that this model will always be effective, and there are many improvement plans we haven't had time to explore. However, starting from this point, you already have a leading advantage in building machine learning trading models.


20. Trading Strategy Based on LightGBM Classification Model

20.1. Strategy Implementation

20.1.1. Top/Bottom Finding Algorithm

20.1.2. Labeling Tools

20.1.2.1. Basic Layout
20.1.2.2. Initialization
20.1.2.3. Component Updates

20.1.3. Building the Model

20.1.3.1. Model Base Class
20.1.3.2. V2
20.1.3.3. V3

20.2. Algorithm Optimization

20.2.1. Sample Balancing

20.2.2. Multi-Cycle and Micro Data

20.2.3. Market Atmosphere

20.2.4. Using ID as a Feature

In this chapter, we will build a trading model based on a LightGBM classification model. In other words, it is not responsible for predicting prices but tells you whether to buy or sell. After completing this chapter, you will surely agree that models should be built this way; the rest is just workload: you need to build the system, label data, construct features, and then train the model.


21. The Future New World

21.1. How to Obtain Free Computing Power

21.2. CNN Price Prediction

21.2.1. How to Provide Data for Training

21.2.2. How to Construct Feature Data

21.2.3. How to Define the Model

21.2.4. Training

21.2.5. Production Deployment

21.2.6. CNN Principles and Performance Optimization

21.3. Transformer

21.4. Reinforcement Learning

21.5. Other Important Intelligent Algorithms

21.5.1. Kalman Filter

21.5.2. Genetic Algorithms

We previously discussed why deep learning is not yet suitable for building quantitative trading models. In the first part of this chapter, we will use a CNN price prediction example to illustrate why. After understanding these limitations, perhaps you can invent a novel model suitable for quantitative trading. This part doesn't teach you portable tools and experiences. However, if you are a research-oriented and innovative person, you will also find this content very valuable.

Reinforcement learning is a direction we are optimistic about, especially in commodity futures and cryptocurrency trading. We will introduce some introductory knowledge and learning resources.

There are two other important intelligent algorithms that are neither machine learning nor deep learning or reinforcement learning, but are indeed commonly used in quantitative finance: Kalman filters and genetic algorithms. However, this part has no code, leaving more exploration space for you.


Notes

1. This syllabus is not a course textbook directory. For example, many chapters in the course have "Extended Reading" sections, which are not displayed here.

2. The course content also includes exercises, which are not displayed here.

3. The course content also includes supplementary materials, such as complete Alpha101 factor implementation code (from data acquisition, factor extraction, factor testing to backtesting) and other example codes, which are not displayed here.