匡醍量化|大富翁量化 This is a long-form selection guide for the "Data and Storage" topic: stitching together experiences scattered across a dozen older articles into a single decision path. After reading, you should be able to answer two questions: Where does my data come from? Where is my data stored, and what do I use to query it?
| Your Scenario | Data Source | Storage | Query |
|---|---|---|---|
| Beginner / A-share Daily Research | AkShare + BaoStock | Parquet (partitioned by year) | DuckDB |
| Stable Production / Multi-asset Research | Tushare Pro or Broker Source | Parquet + Adjustment Factor Table | DuckDB |
| QMT Live Trading | XtQuant (Market Data + Trading) | Local Incremental Cache (Parquet) | DuckDB |
| Multi-user Collaboration / Service-oriented | Any of the above | Parquet | ClickHouse |
Tushare Pro – Established, standardized fields, large community. Covers daily data, financials, indices, and futures comprehensively. Paid tiers offer significantly better stability and documentation. Suitable as a "stable foundation" for research data. Drawbacks: Points/payment thresholds, rate limits on some interfaces, and average real-time performance (suitable for research, not intraday trading).
AkShare – Free, open-source, with an extensive API coverage (stocks, funds, futures, macroeconomics, alternative data). Zero cost, fast updates, and often the first stop for many. Drawbacks are clear: it is essentially a collection of web scrapers; interfaces may intermittently break when upstream pages change; field naming is inconsistent. Production use requires wrapping it with cleaning, retry logic, and schema validation.
BaoStock – Free, with fewer but stable interfaces. Clean and sufficient for historical daily and financial data. Suitable for scenarios where you "don't want to maintain scrapers and only need clean daily data"; the trade-off is limited variety and dimensions.
XtQuant / QMT – Broker-side market data and trading interfaces. Offers full push market data, Level-1/Level-2 (depending on broker activation), and order placement. For live trading pipelines, it is essentially the only correct solution: high-quality market data and controllable trading. Drawbacks: Requires broker permissions, depends on the client being online, and historical data must be stored locally; additionally, version inconsistencies between broker servers and official documentation are common, so rely on actual returns when troubleshooting.
Conclusion: Start research with AkShare / BaoStock; use Tushare Pro as a fallback for critical data; live trading must go through QMT / XtQuant. When using multiple sources, store them in a unified schema (see Section 4) so downstream scripts don't need to know where the data comes from.
Three candidates: CSV, HDF5, and Parquet.
Practical Advice: Use Parquet for "batch write, batch read" data like daily bars and financials; intermediate products in research scripts can be arbitrary, but unify the warehouse layer to Parquet.
Directory layout is more important than format. Two recommended layouts:
daily/year=2026.parquet # Full market, one file per year: suitable for full-market scans
daily/symbol=600519/year=2026.parquet # Partitioned by asset: suitable for single-asset long history and incremental updates
For daily data volume (millions of rows per year), use the former: fewer files, faster full-market queries. For Tick-level data or scenarios requiring frequent incremental updates, use the latter: modifying one asset touches only one file.
read_parquet() to query files with zero deployment and zero import. Extremely fast for full-market scans on a single machine; it is the veritable "researcher's database."Conclusion: Put metadata in SQLite; use DuckDB to directly query Parquet for research and analysis; consider ClickHouse only for scaling and service-oriented architectures—do not start with a service.
Below are two classes continuously used in the Moonshot series (full implementation available in the Moonshot Practice series' accompanying store.py, 544 lines, production-verified). Comments have been simplified, but the logic remains unchanged.
First, the trading calendar—the "header" of the data warehouse. All incremental updates rely on it to determine missing days:
class CalendarModel:
"""Calendar Model: is_open / prev columns, indexed by date"""
def get_trade_dates(self, start, end):
mask = (df.index >= start) & (df.index <= end) & (df.is_open == 1)
return df[mask].index.tolist()
# Also includes is_trade_date / prev_trade_day / get_next_trade_day /
# floor / ceil / shift / delta, but only the most common ones are listed here
Then, unified storage: using (date, asset) as the unique key, reading/writing Parquet via Polars, with inputs and outputs remaining in pandas. The core function is get_and_fetch—if local data is missing a segment, it calls your provided fetch_data_func to fill that segment. The application layer never needs to worry about data completeness:
store = ParquetUnifiedStorage(store_path, calendar, fetch_data_func=fetch_bars)
barss = store.get_and_fetch(start, end) # If missing, fetch; if fetched, store; if stored, read
dv_store = ParquetUnifiedStorage(store_path, calendar, fetch_data_func=fetch_dv_ttm)
dv_ttm = dv_store.get_and_fetch(start, end)
Three hard rules for the write side (all within append_data):
# 1. Deduplicate by (date, asset), keeping the latest → Incremental updates are idempotent, repeated runs do not produce dirty data
deduped = combined.unique(subset=["date", "asset"], keep="last").sort(["date", "asset"])
# 2. Write sorted by date+asset → Range queries are naturally continuous
# 3. Write to disk with lz4 compression → Daily data volume reads/writes are in milliseconds
sorted_df.collect().write_parquet(self._file_path, compression="lz4")
For queries, use Polars lazy scan directly (zero import, zero service):
lazy_df = pl.scan_parquet(self._file_path)
df = (lazy_df.filter((pl.col("asset").is_in(["600519", "000001"]))
& (pl.col("date") >= start))
.collect().to_pandas())
Three engineering details worth copying verbatim:
CalendarModel, not file existence—suspended trading days and holidays are not misjudged as missing;fetch_data_func is a pluggable parameter. Switching between Tushare / AkShare / XtQuant requires changing only this function, while the semantics of get_and_fetch remain unchanged—this is the implementation of Section 1's "multi-source coexistence, unified storage";The comparative conclusions in this article are based on our daily usage experience with A-share daily and minute-level data. Specific performance numbers are strongly correlated with data scale, field width, disk type, and query patterns—recommend testing the code skeleton in this article with your own data before making a final selection.
24 articles · grouped by level