Fixing A-Share Data: Limit Prices, ST Status, and Adjustments
Data is like blood; fields are blood types. This chapter dives into the world of "price limits" and "ST" stocks in A-share data, revealing how the 007 assistant uses the Tushare API to perfectly fix missing key fields, enabling quantitative strategies to thrive in real market environments!
Preface
Building on the previous chapter, we have essentially established the basic architecture for scheduled daily data acquisition. However, some fields were incorrect. Today, we focus on fixing and refining these fields...

"007, our scheduled daily data acquisition system is basically in place, but the data fields are not complete enough. Key information such as limit prices, ST status, and adjustment factors are missing." I said to my AI assistant while checking the data in ClickHouse.
"Received 🫡, I will solve this problem for you immediately." 007 responded instantly.
This is the seventh day of our quantitative trading system development. In the past few days, we built a system that can periodically acquire daily data from Tushare and store it in ClickHouse. However, during actual usage, I found that several key fields were missing from the data, which are crucial for quantitative strategy development.
Problem Analysis: Missing Key Fields
In quantitative trading, several fields are critical for strategy formulation and backtesting:
Limit Prices (limit_up and limit_down): The A-share market has price limits, which directly affect the execution of trading strategies. Without knowing the limit prices, one might devise trading plans that cannot be executed.
ST Status (is_st): ST (Special Treatment) stocks have special price limits (±5% instead of ±10%) and higher risk. Many strategies choose to avoid ST stocks.
Adjustment Factor (adjust): For long-term backtesting, the correct adjustment factor is key to ensuring data consistency. Without it, one cannot properly handle the impact of ex-rights and ex-dividend on stock prices.
Reviewing our existing code, we found that these fields were either completely missing or simply filled with default values:
# Add limit and ST information
df['limit_up'] = None
df['limit_down'] = None
df['is_st'] = 0
# Add adjustment factor (simplified here; in reality, it should be obtained from the adj_factor interface)
df['adjust'] = 1.0
This clearly does not meet actual needs. We need to obtain real data from Tushare to populate these fields.
Solution: Using Tushare API to Obtain Complete Fields
After discussion, 007 and I decided to use the following Tushare APIs to obtain the missing fields:
- stk_limit: To obtain limit price data.
- namechange: To obtain stock name change information, used to determine ST status.
- adj_factor: To obtain adjustment factor data.
1. Obtaining Limit Prices
Limit prices are an important reference point for trading strategies. In the A-share market, the price limit for ordinary stocks is ±10%, for ST stocks it is ±5%, and for the STAR Market and ChiNext it is ±20%. We need to obtain this data from Tushare's stk_limit interface.
007 modified the _enrich_daily_data method to add code for obtaining limit prices:
# Obtain limit prices
limit_df = self._call_tushare_api(
api_name='stk_limit',
params={'trade_date': trade_date},
fields='ts_code,trade_date,up_limit,down_limit'
)
# Add limit information
if not limit_df.empty:
# Merge limit information
df = pd.merge(df, limit_df[['ts_code', 'up_limit', 'down_limit']],
on='ts_code', how='left')
else:
# If no limit data is obtained, add empty columns as placeholders
df['up_limit'] = None
df['down_limit'] = None
This code first calls Tushare's stk_limit interface to obtain limit price data for a specified trading day, then merges this data with the original daily data using pd.merge. If the acquisition fails or the data is empty, it adds empty columns as placeholders.
2. Determining ST Status
ST (Special Treatment) refers to the special handling of listed companies with financial problems or other risks. ST stocks have stricter price limits and higher risks, so they usually require special treatment in quantitative strategies.
Determining whether a stock is an ST stock is not as simple as querying a single interface. We need to judge based on stock name change records. If a stock's name contains "ST" and the current date is between the start and end dates of the ST status, then the stock is an ST stock on that date.
007 implemented a complex but efficient ST status judgment logic:
# Obtain ST stock information
namechange_df = self._call_tushare_api(
api_name='namechange',
params={},
fields='ts_code,name,start_date,end_date,change_reason'
)
# Process ST information
st_codes = []
if not namechange_df.empty:
# Filter out stocks with "ST" in their names
st_df = namechange_df[namechange_df['name'].str.contains('ST', na=False)]
# Check if each stock is ST on the target date
for _, row in st_df.iterrows():
ts_code = row['ts_code']
start_date_str = row['start_date']
end_date_str = row['end_date']
# Convert date format
start_date = datetime.datetime.strptime(start_date_str, '%Y%m%d')
target_date = datetime.datetime.strptime(trade_date, '%Y%m%d')
# Handle end date
if pd.isna(end_date_str) or end_date_str is None:
end_date = datetime.datetime.now()
else:
end_date = datetime.datetime.strptime(end_date_str, '%Y%m%d')
# Check if the current date is within the ST date range
if (target_date - start_date).days >= 0 and (target_date - end_date).days <= 0:
st_codes.append(ts_code)
logger.debug(f"Stock {ts_code} is an ST stock on {trade_date}")
# Add ST flag
df['is_st'] = df['ts_code'].isin(st_codes).astype(int)
This code first obtains name change records for all stocks, then filters out records containing "ST" in the name. For each record, it checks if the target date is between the ST start and end dates. If so, it adds the stock to the ST stock list. Finally, it adds the is_st field to the original data, with values of 0 or 1, by checking if each stock is in the ST stock list.
3. Obtaining Adjustment Factors
The adjustment factor is key to handling stock ex-rights and ex-dividend. In the A-share market, dividend distributions and stock splits by listed companies cause price changes, making historical data discontinuous. By using adjustment factors, we can adjust historical prices into a continuous series, facilitating analysis and backtesting.
007 added code to obtain adjustment factors:
# Obtain adjustment factors
adj_factor_df = self._call_tushare_api(
api_name='adj_factor',
params={'trade_date': trade_date},
fields='ts_code,trade_date,adj_factor'
)
# Add adjustment factors
if not adj_factor_df.empty:
# Merge adjustment factor information
df = pd.merge(df, adj_factor_df[['ts_code', 'adj_factor']],
on='ts_code', how='left')
# Fill missing values with 1.0
df['adj_factor'] = df['adj_factor'].fillna(1.0)
else:
# If no adjustment factor data is obtained, add default value
df['adj_factor'] = 1.0
# Rename adj_factor to adjust
df = df.rename(columns={'adj_factor': 'adjust'})
This code calls Tushare's adj_factor interface to obtain adjustment factor data, then merges this data with the original daily data using pd.merge. For missing adjustment factors, it fills in the default value of 1.0 (indicating no adjustment is needed). Finally, it renames the column from adj_factor to adjust to match our database structure.
Special Challenges in Processing Historical Data
For current-day data, we can directly obtain the limit prices, ST status, and adjustment factors for that day. However, for historical data, the situation is more complex because we need to process data over a period of time.
007 designed a more efficient processing method for historical data:
# Convert limit information to a dictionary for easy lookup
limit_info = {}
for _, row in limit_df.iterrows():
ts_code = row['ts_code']
trade_date = row['trade_date']
key = f"{ts_code}_{trade_date}"
limit_info[key] = {
'up_limit': row['up_limit'],
'down_limit': row['down_limit']
}
# Add limit and ST information
df['limit_up'] = None
df['limit_down'] = None
df['is_st'] = 0
for i, row in df.iterrows():
ts_code = row['ts_code']
trade_date = row['trade_date']
key = f"{ts_code}_{trade_date}"
# Add limit information
if key in limit_info:
df.at[i, 'limit_up'] = limit_info[key]['up_limit']
df.at[i, 'limit_down'] = limit_info[key]['down_limit']
# Add ST information
if key in st_info:
df.at[i, 'is_st'] = st_info[key]
This method first converts limit and ST information into dictionaries, with keys being the combination of stock code and trading date, and values being the corresponding data. Then, it iterates through each row of the original data, constructs a key based on the stock code and trading date, and looks up the corresponding data in the dictionary. This method is more flexible than using pd.merge directly and can handle more complex situations.
For adjustment factors, a similar approach is adopted:
# Convert adjustment factor information to a dictionary for easy lookup
adj_factor_info = {}
for _, row in adj_factor_df.iterrows():
ts_code = row['ts_code']
trade_date = row['trade_date']
key = f"{ts_code}_{trade_date}"
adj_factor_info[key] = row['adj_factor']
# Add adjustment factor for each record
df['adjust'] = 1.0 # Default value
for i, row in df.iterrows():
ts_code = row['ts_code']
trade_date = row['trade_date']
key = f"{ts_code}_{trade_date}"
# Add adjustment factor information
if key in adj_factor_info:
df.at[i, 'adjust'] = adj_factor_info[key]
Optimizing API Calls: Using the Tushare Library Directly
During implementation, we discovered an important issue: the initial code attempted to call the Tushare API directly via HTTP requests, which was not only inefficient but also prone to various network issues.
"007, we should use the Tushare library directly for data acquisition, rather than via HTTP requests," I reminded.
"Received 🫡, I will modify the code immediately." 007 responded quickly.
The modified code is as follows:
# Initialize Tushare API
ts.set_token(self.token)
self.pro = ts.pro_api() # No longer passing api_url parameter
This method uses the official Tushare library's API directly, making the code not only more concise but also more reliable, avoiding various problems that HTTP requests might bring.
Testing and Verification
After completing the code modifications, we conducted tests to ensure that the newly added fields could be correctly obtained and stored.
First, we obtained current-day data:
python main.py daily --batch-size 1000
Then, we obtained historical data:
python main.py history --days 7 --batch-size 1000
Finally, we checked the data in ClickHouse:
python main.py info
2025-05-22 17:10:24,352 - day_bar_fetcher - INFO - Tushare API initialization successful
2025-05-22 17:10:24,393 - day_bar_fetcher - INFO - Redis connection successful
2025-05-22 17:10:24,538 - day_bar_fetcher - INFO - ClickHouse connection successful
2025-05-22 17:10:24,554 - day_bar_fetcher - INFO - Table RealTime_DailyLine_DB.day_bar ensured to exist
2025-05-22 17:10:24,555 - day_bar_fetcher - INFO - Scheduler initialization successful
2025-05-22 17:10:24,585 - day_bar_fetcher - INFO - ==================================================
2025-05-22 17:10:24,585 - day_bar_fetcher - INFO - Time range of existing data in ClickHouse: 20250515 - 20250521
2025-05-22 17:10:24,588 - day_bar_fetcher - INFO - Total of 26953 data records in ClickHouse
2025-05-22 17:10:24,588 - day_bar_fetcher - INFO - ==================================================

The test results show that all fields have been correctly obtained and stored. In particular, we can see that key fields such as limit prices, ST status, and adjustment factors have been populated with real data, rather than simple default values.
Results and Reflections
Through this field repair, our scheduled daily data acquisition system has become more complete and practical. Now, the system can obtain and store the following key fields:
- Basic OHLC data: Open, High, Low, Close, Volume, Turnover
- Limit Prices: Up limit price, Down limit price
- ST Status: Whether it is an ST stock
- Adjustment Factor: Used for forward or backward adjustment calculations
These fields provide a solid data foundation for subsequent quantitative strategy development. In particular, with correct limit prices and ST status, we can more accurately simulate the actual trading environment; with adjustment factors, we can properly handle the continuity issues of historical data.
"007, you did a great job! This field repair has made our system more complete," I praised sincerely.
Next Steps
After completing the field repair, our scheduled daily data acquisition system is quite complete. However, on the path of quantitative trading system development, we still have a long way to go. Next, we plan to:
- System Stability Testing: Run the system in a real environment for a long time to ensure its stability and reliability.
- Data Quality Monitoring: Develop monitoring tools to check the completeness and accuracy of data in real time.
- Expand Data Sources: In addition to daily data, consider adding finer-grained data such as minute-level and tick data.
- Strategy Backtesting Module: Develop a strategy backtesting module to conduct strategy backtesting using the acquired data.
These plans will be implemented gradually in the coming days. I believe that with 007's assistance, our quantitative trading system will become increasingly powerful.
The 21-day challenge continues, looking forward to more breakthroughs and achievements!