匡醍量化|大富翁量化

NumPy & Pandas Syllabus for Quant Data Processing

中文 📅 2024-08-27 👁 views this month —

§ NUMPY AND PANDAS IN QUANTITATIVE TRADING

Course Syllabus

1. Introduction

2. NumPy Core Syntax (1)

2.1. Basic Data Structures

2.1.1. Creating Arrays

Common methods for creating arrays, along with several built-in arrays frequently used in quantitative finance.

2.1.2. Adding, Deleting, and Modifying Elements

How to append, insert, delete, and modify elements in an array?

2.1.3. Indexing, Reading, and Searching

Introduction to indexing, slicing, searchsorted, etc.

2.1.4. Inspecting Arrays

2.2. Array Operations

2.2.1. Dimensionality Increase

2.2.2. Dimensionality Reduction

2.2.3. Transposition

3. NumPy Core Syntax (2)

3.1. Structured Arrays

3.2. Arithmetic Operations

3.2.1. Logical Operations and Comparisons

3.2.2. Set Operations

3.2.3. Mathematical and Statistical Operations

Matrix operations and statistical functions such as mean, variance, covariance, and percentiles.

3.3. Type Conversion and Typing

In-depth understanding of NumPy data types and their conversions, along with the typing library, to help write robust code.

4. NumPy Core Syntax (3)

4.1. Handling Data with np.nan

Data obtained from third parties may contain np.nan; technical indicator values during the warm-up period are often np.nan as well. This section introduces np.isnan, nanmean, nanmax, and other nan* functions for calculating mean or maximum values when data contains None or np.nan.

4.2. Random Numbers and Sampling

Random number generation and sampling are high-frequency operations in quantitative finance, particularly useful for synthetic data generation.

4.3. I/O Operations

Introduction to reading and saving CSV files and other I/O operations.

4.4. Dates and Times

How to convert time and date formats from market data obtained via other libraries?

4.5. String Operations

How to perform string searches and other operations within NumPy arrays?

5. NumPy Quantitative Scenario Applications

5.1. Continuous Value Statistics

Example: Efficiently finding consecutive price limits (up/down), N-day winning/losing streaks, and calculating streaks in Connor's RSI.

5.2. Cumulative Sum and Intraday Average Price Line

The intraday average price line is crucial for intraday trading. Generally, two attacks on the average price line that fail to break it signal an intraday buy (or sell). How do we calculate this average price line?

5.3. Moving Average Calculation

How to quickly calculate moving averages using NumPy? Introduction to a convolution algorithm.

5.4. Rational Selection of Adaptive Parameters

Often, adaptive parameters are required. How to select them? Percentiles are often a good solution.

5.5. Calculating Maximum Drawdown

With experience, you can identify the major rebound on February 7th. On rebound days, you want to target stocks with the largest declines. How to select them?

Do not buy stocks with long-term bearish trends. The key is how to determine this. This section introduces polynomial regression.

5.7. Function Routines in Alpha101

Alpha101 contains several basic functions upon which factors are built. How to implement them efficiently?

5.8. Finding Similar Candlestick Patterns

Introduction to corrcoef and correlate.

5.9. Asset Portfolio Return and Volatility Example

Start by randomly generating several assets, then calculate their expected returns and volatility. This is one of the high-frequency application scenarios.

6. NumPy High-Performance Programming Practices

6.1. Broadcasting

In-depth look at NumPy's efficient underlying principles.

6.2. Using NumExpr

6.3. Enabling Multithreading

6.4. Using the Bottleneck Library

6.5. Other Alternatives to NumPy

7. Pandas Core Syntax (1)

7.1. Basic Data Structures

7.1.1. Series

7.1.2. Creating DataFrames

7.1.3. Rapid Exploration of DataFrames

Index, info, describe, columns, head, tail, etc.

7.1.4. Merging and Joining DataFrames

7.1.5. Deleting Rows and Columns

7.1.6. Indexing, Reading, and Modifying

Introduction to indexing and data selection in Pandas.

7.1.7. Transposition

7.1.8. Resampling

Intraday real-time minute-level data is surprisingly expensive. Therefore, we need to synthesize it from tick-level data ourselves. This is resampling.

8. Pandas Core Syntax (2)

8.1. Logical Operations and Comparisons

DataFrames contain the features we extract. How to select the top 30 columns with the highest PE and lowest PB?

8.2. Groupby Operations

The factor analysis data table contains industry labels and each company's PE value. How to select the top 5 stocks with the strongest PE in each industry?

8.3. Multi-Index and Advanced Indexing

One of the more difficult concepts in Pandas.

8.4. Window Functions

For calculating sliding window indicators such as moving averages.

8.5. Mathematical and Statistical Operations

Basic quant functions including mean, variance, covariance, percentile, diff, pct_change, rank, etc.

9. Pandas Core Syntax (3)

9.1. Data Preprocessing

What to do during factor analysis preprocessing, such as handling missing values, winsorization, and deduplication?

9.2. I/O Operations

How to read data from CSVs, web pages, databases, Parquet files, etc.?

9.2.1. CSV

Beyond basic operations, we will also introduce how to accelerate CSV reading.

9.2.2. PKL and HDF5

9.2.3. Parquet

9.2.4. HTML and MD

9.2.5. SQL

9.3. Dates and Times

How to convert time and date formats from market data obtained via other libraries?

9.4. String Operations

DataFrames store basic security information, such as names and codes. How to exclude stocks from the STAR Market?

10. Pandas Core Syntax (4)

10.1. Tables and Styling

Giving Pandas rich conditional formatting capabilities similar to Excel.

10.2. Pandas Built-in Plotting Functions

11. Pandas Quantitative Scenario Applications

11.1. Implementing TongdaXin Routines via Rolling Methods

Implementing TongdaXin formula methods such as HHV, LLV, HHVBARS, LAST, etc.

11.2. Filling Missing Adjustment Factors for Minute-Level Data

Introduction to the newly introduced as-of-join feature. A must-know scenario for quantitative finance.

11.3. Preparing Data for Alphalens

The most common DataFrame operations when using Alphalens for factor analysis.

12. Pandas Performance

12.1. Memory Optimization

Compressing memory usage by using category dtype and more compact data types.

12.2. Optimizing Iterations

Use itertuples instead of iterrows, use apply to optimize iterations, and filter before calculating.

12.3. Using NumPy and Numba

12.4. Using eval or query

12.5. Other Alternatives to Pandas

12.5.1. Modin

One line of code to replace Pandas, gaining multi-core and memory-unlimited computational power.

12.5.2. Polars

The fastest table solution.

12.5.3. Dask

Distributed table processing, capable of running on thousands of nodes.