匡醍量化|大富翁量化

Pandas Core Syntax for Quant: Series, Indexing, and Financial Data

中文 📅 2025-03-28 👁 views this month —

Pandas holds a central position in quantitative trading. Many data sources returning data via Python SDKs typically output pandas.DataFrame objects. Factor analysis libraries like Alphalens and performance evaluation tools like empyrical rely heavily on Pandas.

This is hardly surprising. Pandas’ creator, Wes McKinney, was originally a researcher at the asset management firm AQR Capital Management. While handling large-scale financial analysis tasks, he found that Python’s existing tools (such as NumPy) were inefficient for structured data analysis. Consequently, he began developing Pandas in 2008 and open-sourced it in 2009.

1.1. Basic Data Structures

The core data structures in Pandas are Series (similar to a one-dimensional array) and DataFrame (a rectangular data table).

1.2. Series

A Series consists of an array of data along with an associated set of data labels (i.e., an index). The simplest Series can be created from a single array. Series objects are presented interactively, with the index on the left and values on the right (typically, an index ranging from 0 to N-1 is automatically created, where N is the length of the data). Compared to NumPy arrays, you can select individual or grouped values from a Series using index labels. It can also be viewed as an ordered dictionary of fixed length; in scenarios where dictionaries are used, Series can serve as a substitute. For many applications, the most practical feature of Series is its ability to automatically align index labels during arithmetic operations.

from pandas import Series, DataFrame

1.2.1. Creating from Arrays (Default Index)

obj = Series([1, 3, 5, 7]) 
obj

Converting a Python list directly via pd.Series() generates an integer index starting from 0 by default.

1.2.2. Creating from Dictionaries (Custom Index)

Dictionary keys automatically become the index, while values become the data:

obj = Series({"a": 4, "b": 3, "c": 2, "d": 1})
obj

50%


Conversely, a Series can be converted back to a dictionary using the to_dict method:

cprint("转换回字典:{}",obj.to_dict())

1.2.3. Custom Indexing

Specify any immutable object as the index using the index parameter:

obj = Series([90, 85, 92], index=["数学", "英语", "物理"], dtype="float64")
obj

50%

1.2.4. Creating Series with Timestamp Indices

Generate time-series data:

dates = pd.date_range("20230308", periods=4)
s = Series([100, 200, 300, 400], index=dates)

50%

1.2.5. Querying Array Values and Index Objects

You can retrieve the underlying array values and index objects using the array and index attributes of a Series:

obj = Series([1, 3, 5, 7], index=["a","b","c","d"]) 
print("数组值:{}", obj.array)
print("索引对象:{}", obj.index)

Unlike NumPy arrays, Series allows selection of individual or grouped values via index labels:

print(obj["a"])

obj["d"] = 6
print(obj[["a","c","d"]])

print(obj[obj>5])

print(obj * 2)

print(np.exp(obj))

50%


When constructing Series and Pandas objects, any arrays or other label sequences used are converted into index objects:

obj = Series(np.arange(3), index=["a","b","c"])
index = obj.index
print("index:",index)
print("index[1:]",index[1:])

# 注意Index对象是不可变的,因此用户不能对其修改
index[1]="d" # TypeError

Due to the immutability of Index objects, they can be safely shared across multiple data structures:

labels = pd.Index(np.arange(3))
print(labels)

obj = Series([1.5,-2.5,0],index=labels)
print(obj)

Below is a summary of common indexing methods and attributes:

Method/Attribute Description Example
append Concatenate additional index objects to create a new index. new_index = index1.append(index2)
diff Calculate the difference set of indices. diff_index = index1.diff(index2)
intersection Calculate the intersection of indices. common_index = index1.intersection(index2)
union Calculate the union of indices. union_index = index1.union(index2)
isin Return a boolean array indicating whether each value is contained in the passed set. bool_array = index.isin(['a', 'b'])
delete Delete elements at specified positions, returning a new index. new_index = index.delete(0)
drop Drop passed values, returning a new index. new_index = index.drop('a')
insert Insert elements at specified positions, returning a new index. new_index = index.insert(1, 'new_value')
is_monotonic Return True if the index is monotonically increasing or decreasing. is_monotonic = index.is_monotonic
is_unique Return True if the index has no duplicate values. is_unique = index.is_unique

Method/Attribute Description Example
unique Return an array of unique values in the index. unique_values = index.unique()
reindex Rearrange data according to a new index, filling missing values with NaN. new_series = series.reindex(new_index)
reset_index Reset the index to default integer indexing, turning the original index into a column. df_reset = df.reset_index()
set_index Set a specific column as the index. df.set_index('column_name', inplace=True)
sort_values Sort the index values. sorted_index = index.sort_values()
to_series Convert the index to a Series. index_series = index.to_series()
values Return the NumPy array of the index. index_values = index.values
name Get or set the name of the index. index_name = index.name
shape Return the shape (length) of the index. index_shape = index.shape
size Return the length of the index. index_size = index.size

The following code example, relevant to financial quantitative trading, demonstrates how to use Pandas to calculate stock moving averages (MA) and the Relative Strength Index (RSI):

import pandas as pd
import numpy as np

# 模拟股票价格数据
dates = pd.date_range("2025-01-01", periods=100)
prices = pd.Series(np.random.randint(100, 200, size=100), index=dates)

# 计算移动平均线 (MA)
ma_10 = prices.rolling(window=10).mean()
ma_20 = prices.rolling(window=20).mean()

# 计算相对强弱指数 (RSI)
delta = prices.diff()
gain = (delta.where(delta > 0, 0)).rolling(window=14).mean()
loss = (-delta.where(delta < 0, 0)).rolling(window=14).mean()
rs = gain / loss
rsi = 100 - (100 / (1 + rs))

# 输出结果
result = pd.DataFrame({
    "Price": prices,
    "MA_10": ma_10,
    "MA_20": ma_20,
    "RSI": rsi
})

print(result.tail())