NumPy & Pandas Syllabus for Quant Data Processing
§ NUMPY AND PANDAS IN QUANTITATIVE TRADING
Course Syllabus
1. Introduction
2. NumPy Core Syntax (1)
2.1. Basic Data Structures
2.1.1. Creating Arrays
Common methods for creating arrays, along with several built-in arrays frequently used in quantitative finance.
2.1.2. Adding, Deleting, and Modifying Elements
How to append, insert, delete, and modify elements in an array?
2.1.3. Indexing, Reading, and Searching
Introduction to indexing, slicing, searchsorted, etc.
2.1.4. Inspecting Arrays
2.2. Array Operations
2.2.1. Dimensionality Increase
2.2.2. Dimensionality Reduction
2.2.3. Transposition
3. NumPy Core Syntax (2)
3.1. Structured Arrays
3.2. Arithmetic Operations
3.2.1. Logical Operations and Comparisons
3.2.2. Set Operations
3.2.3. Mathematical and Statistical Operations
Matrix operations and statistical functions such as mean, variance, covariance, and percentiles.
3.3. Type Conversion and Typing
In-depth understanding of NumPy data types and their conversions, along with the typing library, to help write robust code.
4. NumPy Core Syntax (3)
4.1. Handling Data with np.nan
Data obtained from third parties may contain np.nan; technical indicator values during the warm-up period are often np.nan as well. This section introduces np.isnan, nanmean, nanmax, and other nan* functions for calculating mean or maximum values when data contains None or np.nan.
4.2. Random Numbers and Sampling
Random number generation and sampling are high-frequency operations in quantitative finance, particularly useful for synthetic data generation.
4.3. I/O Operations
Introduction to reading and saving CSV files and other I/O operations.
4.4. Dates and Times
How to convert time and date formats from market data obtained via other libraries?
4.5. String Operations
How to perform string searches and other operations within NumPy arrays?
5. NumPy Quantitative Scenario Applications
5.1. Continuous Value Statistics
Example: Efficiently finding consecutive price limits (up/down), N-day winning/losing streaks, and calculating streaks in Connor's RSI.
5.2. Cumulative Sum and Intraday Average Price Line
The intraday average price line is crucial for intraday trading. Generally, two attacks on the average price line that fail to break it signal an intraday buy (or sell). How do we calculate this average price line?
5.3. Moving Average Calculation
How to quickly calculate moving averages using NumPy? Introduction to a convolution algorithm.
5.4. Rational Selection of Adaptive Parameters
Often, adaptive parameters are required. How to select them? Percentiles are often a good solution.
5.5. Calculating Maximum Drawdown
With experience, you can identify the major rebound on February 7th. On rebound days, you want to target stocks with the largest declines. How to select them?
5.6. Determining Long-Term Trends for Individual Stocks
Do not buy stocks with long-term bearish trends. The key is how to determine this. This section introduces polynomial regression.
5.7. Function Routines in Alpha101
Alpha101 contains several basic functions upon which factors are built. How to implement them efficiently?
5.8. Finding Similar Candlestick Patterns
Introduction to corrcoef and correlate.
5.9. Asset Portfolio Return and Volatility Example
Start by randomly generating several assets, then calculate their expected returns and volatility. This is one of the high-frequency application scenarios.
6. NumPy High-Performance Programming Practices
6.1. Broadcasting
In-depth look at NumPy's efficient underlying principles.
6.2. Using NumExpr
6.3. Enabling Multithreading
6.4. Using the Bottleneck Library
6.5. Other Alternatives to NumPy
7. Pandas Core Syntax (1)
7.1. Basic Data Structures
7.1.1. Series
7.1.2. Creating DataFrames
7.1.3. Rapid Exploration of DataFrames
Index, info, describe, columns, head, tail, etc.
7.1.4. Merging and Joining DataFrames
7.1.5. Deleting Rows and Columns
7.1.6. Indexing, Reading, and Modifying
Introduction to indexing and data selection in Pandas.
7.1.7. Transposition
7.1.8. Resampling
Intraday real-time minute-level data is surprisingly expensive. Therefore, we need to synthesize it from tick-level data ourselves. This is resampling.
8. Pandas Core Syntax (2)
8.1. Logical Operations and Comparisons
DataFrames contain the features we extract. How to select the top 30 columns with the highest PE and lowest PB?
8.2. Groupby Operations
The factor analysis data table contains industry labels and each company's PE value. How to select the top 5 stocks with the strongest PE in each industry?
8.3. Multi-Index and Advanced Indexing
One of the more difficult concepts in Pandas.
8.4. Window Functions
For calculating sliding window indicators such as moving averages.
8.5. Mathematical and Statistical Operations
Basic quant functions including mean, variance, covariance, percentile, diff, pct_change, rank, etc.
9. Pandas Core Syntax (3)
9.1. Data Preprocessing
What to do during factor analysis preprocessing, such as handling missing values, winsorization, and deduplication?
9.2. I/O Operations
How to read data from CSVs, web pages, databases, Parquet files, etc.?
9.2.1. CSV
Beyond basic operations, we will also introduce how to accelerate CSV reading.
9.2.2. PKL and HDF5
9.2.3. Parquet
9.2.4. HTML and MD
9.2.5. SQL
9.3. Dates and Times
How to convert time and date formats from market data obtained via other libraries?
9.4. String Operations
DataFrames store basic security information, such as names and codes. How to exclude stocks from the STAR Market?
10. Pandas Core Syntax (4)
10.1. Tables and Styling
Giving Pandas rich conditional formatting capabilities similar to Excel.
10.2. Pandas Built-in Plotting Functions
11. Pandas Quantitative Scenario Applications
11.1. Implementing TongdaXin Routines via Rolling Methods
Implementing TongdaXin formula methods such as HHV, LLV, HHVBARS, LAST, etc.
11.2. Filling Missing Adjustment Factors for Minute-Level Data
Introduction to the newly introduced as-of-join feature. A must-know scenario for quantitative finance.
11.3. Preparing Data for Alphalens
The most common DataFrame operations when using Alphalens for factor analysis.
12. Pandas Performance
12.1. Memory Optimization
Compressing memory usage by using category dtype and more compact data types.
12.2. Optimizing Iterations
Use itertuples instead of iterrows, use apply to optimize iterations, and filter before calculating.
12.3. Using NumPy and Numba
12.4. Using eval or query
12.5. Other Alternatives to Pandas
12.5.1. Modin
One line of code to replace Pandas, gaining multi-core and memory-unlimited computational power.
12.5.2. Polars
The fastest table solution.
12.5.3. Dask
Distributed table processing, capable of running on thousands of nodes.