Factor ML Course Syllabus: From Alpha Mining to Live Trading
§ Factor Investing & Machine Learning Strategies
Course Syllabus
1. Introduction
1.1. The Origins of Factor Investing
1.2. Hunting for Alpha
1.3. From CAPM to Multi-Factor Models
1.4. From Factor Analysis to Factor Investing
1.5. From Factor Models to Trading Strategies
1.6. Course Structure
This course is designed for professional quantitative researchers, those transitioning into the field, or professionals from other disciplines determined to explore quantitative research with a rigorous, professional attitude.
Upon completing and mastering this course, you will possess proficient factor analysis skills, master state-of-the-art machine learning strategy construction methods, and become a quantitative researcher with innovative research capabilities and a competitive edge.
The curriculum covers the entire process from factor mining and testing to building machine learning models. If you intend to engage in independent trading, you should supplement your learning with "Quantitative Trading 24 Lessons."
2. Factor Preprocessing Pipeline
2.1. Factor Data Sources
2.2. Factor Generation
2.3. Factor Preprocessing
2.3.1. Outlier Clipping
2.3.2. Missing Values Handling
2.3.3. Distribution Adjustment
2.3.4. Standardization
2.3.5. Neutralization
This chapter and the next form the foundation for factor testing. We will introduce the basic principles and technical implementation details of factor testing using extensive example code, laying a solid groundwork for understanding the Alphalens factor analysis framework.
3. Factor Testing Methods
3.1. Regression Method
3.2. IC Analysis
3.3. Layered Backtest Method
3.4. Code Implementation of Factor Testing
3.4.1. Generating Factors: Further Modularization
3.4.2. Factor Preprocessing: Integrating Real Data
3.4.3. Calculating Forward Returns
3.4.4. Implementing Regression Analysis
3.4.5. Implementing IC Analysis
3.4.6. Implementing Layered Backtest
3.5. Differences and Connections Between the Three Methods
This chapter introduces the principles and implementation code for the regression method, IC method, and layered backtest method. Upon completion, you will be able to implement a simple factor analysis framework yourself, which is highly beneficial for understanding Alphalens.
4. Introduction to Alphalens
4.1. Slope Factor: Definition and Implementation
4.2. How to Calculate Factors and Collect Price Data for Alphalens
4.3. How Alphalens Handles Data Preprocessing
4.4. Factor Analysis and Report Generation
4.5. References
Alphalens highly abstracts the factor testing process, encapsulating the steps discussed in the previous two chapters into two functions, with extensive customization achieved through parameters. We will introduce the input data format required by Alphalens and how it uses parameters to control layering, missing value handling, and forward return calculation behaviors.
By the end of this chapter, you will master the most basic usage of Alphalens.
5. Alphalens Report Analysis
5.1. Return Analysis
5.1.1. Alpha and Beta
5.1.2. Layered Return Mean Chart
5.1.3. Layered Return Violin Plot
5.1.4. Factor-Weighted Long-Short Portfolio Cumulative Return Chart
5.1.5. Layered Drivers of Returns
5.1.6. Robustness of Long-Short Portfolio Returns
5.2. Event Study
5.3. IC Analysis
5.4. Turnover Analysis
5.5. References
Alphalens reports are not self-explanatory. For instance, it doesn't specify the units for Alpha and Beta, nor what a basis point (bps) unit represents; it certainly won't tell you what constitutes a "good" Alpha versus one that is "too good to be true." Some reports calculate metrics differently from what you might expect or have heard.
To accurately understand these reports, we employ three approaches: 1. Reading and debugging the source code. Through this, we discovered that bps stands for one ten-thousandth, defined in the plotting.py file. 2. Using synthetic data, which helps us understand what theoretical reports for the best factors should look like. 3. Searching through GitHub issues and the Quantopian community Archive documentation to find answers in other users' questions.
This is currently the only tutorial on the internet that thoroughly explains Alphalens.
6. Alphalens Advanced Techniques (1)
6.1. Excluding Functional Errors
6.1.1. Outdated Alphalens Versions
6.1.2. MaxLossExceedError
6.1.3. Timezone Issues
6.2. Factor Monotonicity
6.3. Revisiting Price Data Collection
6.4. How to Analyze Factors Above Daily Frequency?
6.5. Deep Dive into Factor Layering
6.5.1. Factors with Determined Trading Signals
6.5.2. Discrete Value Factors
In this chapter, we introduce how to troubleshoot potential errors in Alphalens, both programmatic and logical. We also discuss how to conduct factor analysis at frequencies higher than daily. Many online tutorials don't even realize this presents a problem because they have never performed analysis at this level.
We also delve into Alphalens' layering mechanism, including how to handle cases where factor values are discrete.
7. Alphalens Advanced Techniques (2)
7.1. Refactoring the Factor Testing Process
7.2. Parameter Tuning: Saving Your Factors
7.2.1. Correcting Factor Direction
7.2.2. Filtering Non-Linear Layering
7.2.3. Using Optimal Layering Methods
7.2.4. Grid Search
7.3. Overfitting Detection Methods
7.3.1. Out-of-Sample Testing
7.3.2. Plotting Parameter Plateaus
Using Alphalens for factor testing is like an interview; you must expose the factor's potential as much as possible before evaluating its quality. This chapter introduces how to exhaustively挖掘 (mine) a factor's potential while avoiding the deception of overfitting. Besides out-of-sample testing, we will teach you how to assess the degree of overfitting by plotting parameter plateaus.
Visualization is crucial, especially when collaborating with others.
8. Alpha101 Factor Introduction
8.1. Data and Operators in Alpha101 Factors
8.2. Interpreting Alpha101 Factors
8.3. How to Implement Alpha101 Factors?
The Alpha101 factor library is a factor library published by World Quant in 2015. About 80% of the factors in it are officially used by World Quant (at the time of publication). We will introduce how to read the formulas of Alpha101 factors and implement their operators.
There are already good open-source libraries for implementing the entire factor library, which we will also introduce. This will become one of the treasures in your arsenal.
9. Talib Technical Factors
9.1. Ta-lib Function Grouping
9.2. Warm-up Periods (Unstable Periods)
9.3. Oscillation Indicators
9.3.1. RSI
9.3.2. ADX - Average Directional Movement Index
9.3.3. APO - Absolute Price Oscillator
9.3.4. PPO - Percentage Price Oscillator
9.3.5. Aroon Oscillator
9.3.6. Money Flow Index
9.3.7. Balance of Power
9.3.8. William's R
9.3.9. Stochastic Oscillator
9.4. Volume Indicators
9.4.1. Chaikin A/D Line
9.4.2. OBV
9.5. Volatility Indicators
9.5.1. ATR and NATR - Average True Range
9.6. 8 Types of Moving Averages
9.7. Overlap Studies
9.7.1. Bollinger Bands
9.7.2. Hilbert Trendline and Sine Wave Indicator
9.7.3. Parabolic SAR
9.8. Momentum Indicators
Most Alpha101 factors are volume-price factors. For understandable reasons, they do not repeat classic technical factors that have existed for years, but these factors still possess Alpha potential. In this section, we briefly introduce the talib library and explain the warm-up periods of technical indicators -- a potentially niche topic. The warm-up period is not just NaNs; for example, the warm-up period for RSI is quite long, being 3 times the win parameter.
There are many Talib technical indicators; we will introduce a few from each category, focusing on how to revamp these factors under new technical conditions. Taking RSI as an example, we will discuss intelli-RSI and Connor's RSI. This way, you not only gain new factors but also enhance your ability for innovative research.
Even experienced practitioners might be hearing about some of the factors we introduce for the first time, such as the Hilbert Sine Wave, which is one of the paid technical indicators sold on platforms like TradingView.
10. Other Volume-Price Factors
10.1. Low-Probability Events
10.2. Max Drawdown
10.3. pct_rank
10.4. Volatility
10.5. Z-Score
10.6. Sharpe Ratio
10.7. First-Derivative Factors
10.8. Second-Derivative Factors
10.9. Frequency-Domain Factors
10.10. TSFresh Factor Library
10.11. Behavioral Finance Factors
10.11.1. Integer Barrier Factors
10.11.2. Breakout Failure Factors
10.11.3. Gap Factors
10.11.4. Regret Aversion Theory Factors
Some low-probability factors are easy to construct. Perhaps because of this, they lack names and haven't found their way into academic papers. However, their Alpha is real. For example, the index's largest single-day drop or longest consecutive drop. The underlying principle is probability regression after extreme events.
In short, this is a show-off and innovative chapter. We will introduce second-derivative factors, frequency-domain factors, and behavioral finance factors. For instance, frequency-domain factors use Fast Fourier Transform (FFT) or wavelet transforms to identify the operational cycles of main market participants for prediction. While others are still using wavelets to smooth noise, we have already started using them to explore the patterns of main market participants!
11. Fundamental and Alternative Factors
11.1. Fama-French Five-Factor Model
11.1.1. Market Factor
11.1.2. Size Factor
11.1.3. Value Factor
11.1.4. Profitability Factor
11.1.5. Investment Factor
11.2. Alternative Factors
11.2.1. Social Media Sentiment Factors
11.2.2. Web Traffic Factors
11.2.3. Satellite Image Factors
11.2.4. Patent Application Factors
11.3. Technical Routes for Web Crawling
This part will focus more on concepts. Because alternative factors are either purchased or crawled, we don't want to dwell on crawling techniques.
12. Factor Mining Methods
12.1. Where Do New Factors Come From?
12.2. Online Resources
12.3. Factor Orthogonality Testing
12.4. Discussing the Factor Zoo
This is also a chapter that discusses broad topics, but it is still full of干货 (substance). We will talk about methods for finding resources, such as how to find papers and data. By now, we have introduced hundreds of factors (excluding parameters and cycles), so we also need to see how many factors are truly independent. Therefore, we will introduce orthogonality testing methods.
13. Machine Learning Overview
13.1. Machine Learning Classification
13.1.1. Machine Learning, Deep Learning, Reinforcement Learning
13.1.2. Supervised, Unsupervised, and Reinforcement Learning
13.1.3. Regression and Classification
13.2. Introduction to Machine Learning Models
13.3. Three Elements of Machine Learning
13.4. Basic Workflow of Machine Learning
13.5. Application Scenarios for Machine Learning
A quick introduction to machine learning. Is the world continuous or quantized? This is an ancient philosophical question that also determines the basic models of machine learning -- regression or classification?
14. Core Concepts of Machine Learning
14.1. Bias and Variance
14.2. Overfitting and Regularization Penalties
14.3. Loss Functions, Objective Functions, Metric Functions, and Distance Functions
14.3.1. Loss Functions and Objective Functions
14.3.2. Metric Functions
14.3.3. Distance Functions
This course is an applied course and does not intend to involve too much theory. However, if you understand no principles at all, you can only copy examples without being able to extend them. Therefore, we decided to explain basic concepts closely related to the application layer -- only by understanding these concepts can we know how to choose objective functions, evaluate strategies, prevent overfitting, etc.
15. SKLearn General Toolkit
15.1. Data Preprocessing: preprocessing
15.2. metrics
15.3. Model Interpretation and Visualization
15.4. Built-in Datasets
sklearn is a very powerful machine learning library, winning people's love with its rich models and easy-to-use interface. In this chapter, we first introduce sklearn's general toolkit -- used to handle common problems encountered regardless of which algorithm model we adopt, such as data preprocessing, model evaluation, model interpretation and visualization, and built-in datasets.
16. Model Optimization
16.1. Optimization Overview
16.2. k-fold Cross Validation
16.3. Parameter Search
16.3.1. Grid Search
16.3.2. Random Search
16.3.3. Bayesian Optimization
16.4. Rolling Forecasting
Machine learning in the quantitative field has its own特殊性 (special characteristics), such as in cross-validation, where we actually need to use a method called Rolling Forecasting (also known as Walk-Forward Optimization).
17. Clustering: Finding Pair Trading Targets
17.1. Overview of Clustering Algorithms
17.2. HDBSCAN Algorithm Principles
17.3. Finding Pair Trading Targets
17.3.1. HDBSCAN Example
17.3.2. Result Evaluation
17.3.3. Pair Selection
In quantitative trading, Pair Trading is an important type of arbitrage strategy, with the prerequisite of finding two targets that can be paired. In this chapter, we will introduce the advanced HDBSCAN clustering method, demonstrate how to implement clustering through it, and then use related methods in statsmodels to perform cointegration pair tests to find targets that can be paired. Finally, we will demonstrate how to assemble all of this into a complete trading strategy.
This will be the first effective machine learning strategy you learn.
18. From Decision Trees to LightGBM
18.1. Decision Trees
18.1.1. Decision Tree Classification
18.1.2. Decision Tree Regression
18.2. LightGBM
18.2.1. Familiarizing with Training Data
18.2.2. Building the First Classifier
18.2.3. Visualizing Feature Importance
18.2.4. Viewing Model Trees
18.2.5. Cross-Validation
18.2.6. Tuning
Limited by the high noise in financial data, end-to-end trading strategies are not yet feasible; also limited by the size of labeled data, deep learning and other AI models are not suitable for constructing trading strategies. Among machine learning models, the best model at present is the Gradient Boosting Decision Tree (GBDT) model. Representative implementations are XGBoost and LightGBM.
Since LightGBM surpasses XGBoost in most tasks in terms of both speed and accuracy, our course will focus on introducing LightGBM.
This chapter will comprehensively introduce the LightGBM model and demonstrate through examples how to use it, how to inspect and visualize the generated models, how to perform cross-validation, and how to tune parameters.
19. Price Prediction Based on LightGBM Regression Model
19.1. Strategy Principles
19.2. Strategy Implementation
19.3. Strategy Optimization Ideas
Asset pricing is one of the core issues in quantitative research. If we can provide a reasonable pricing for assets, we can generate trading signals.
Pricing is a regression problem. Although it is difficult to implement end-to-end price prediction models, we have cleverly designed a regression model that can predict future prices (theoretically self-consistent).
We cannot guarantee that this model will always be effective, and there are many improvement plans we haven't had time to explore. However, starting from this point, you already have a leading advantage in building machine learning trading models.
20. Trading Strategy Based on LightGBM Classification Model
20.1. Strategy Implementation
20.1.1. Top/Bottom Finding Algorithm
20.1.2. Labeling Tools
20.1.2.1. Basic Layout
20.1.2.2. Initialization
20.1.2.3. Component Updates
20.1.3. Building the Model
20.1.3.1. Model Base Class
20.1.3.2. V2
20.1.3.3. V3
20.2. Algorithm Optimization
20.2.1. Sample Balancing
20.2.2. Multi-Cycle and Micro Data
20.2.3. Market Atmosphere
20.2.4. Using ID as a Feature
In this chapter, we will build a trading model based on a LightGBM classification model. In other words, it is not responsible for predicting prices but tells you whether to buy or sell. After completing this chapter, you will surely agree that models should be built this way; the rest is just workload: you need to build the system, label data, construct features, and then train the model.
21. The Future New World
21.1. How to Obtain Free Computing Power
21.2. CNN Price Prediction
21.2.1. How to Provide Data for Training
21.2.2. How to Construct Feature Data
21.2.3. How to Define the Model
21.2.4. Training
21.2.5. Production Deployment
21.2.6. CNN Principles and Performance Optimization
21.3. Transformer
21.4. Reinforcement Learning
21.5. Other Important Intelligent Algorithms
21.5.1. Kalman Filter
21.5.2. Genetic Algorithms
We previously discussed why deep learning is not yet suitable for building quantitative trading models. In the first part of this chapter, we will use a CNN price prediction example to illustrate why. After understanding these limitations, perhaps you can invent a novel model suitable for quantitative trading. This part doesn't teach you portable tools and experiences. However, if you are a research-oriented and innovative person, you will also find this content very valuable.
Reinforcement learning is a direction we are optimistic about, especially in commodity futures and cryptocurrency trading. We will introduce some introductory knowledge and learning resources.
There are two other important intelligent algorithms that are neither machine learning nor deep learning or reinforcement learning, but are indeed commonly used in quantitative finance: Kalman filters and genetic algorithms. However, this part has no code, leaving more exploration space for you.
Notes
1. This syllabus is not a course textbook directory. For example, many chapters in the course have "Extended Reading" sections, which are not displayed here.
2. The course content also includes exercises, which are not displayed here.
3. The course content also includes supplementary materials, such as complete Alpha101 factor implementation code (from data acquisition, factor extraction, factor testing to backtesting) and other example codes, which are not displayed here.