匡醍量化|大富翁量化

Data & Storage

This is a long-form selection guide for the "Data and Storage" topic: stitching together experiences scattered across a dozen older articles into a single decision path. After reading, you should be able to answer two questions: Where does my data come from? Where is my data stored, and what do I use to query it?

One-Sentence Conclusion

Your Scenario Data Source Storage Query
Beginner / A-share Daily Research AkShare + BaoStock Parquet (partitioned by year) DuckDB
Stable Production / Multi-asset Research Tushare Pro or Broker Source Parquet + Adjustment Factor Table DuckDB
QMT Live Trading XtQuant (Market Data + Trading) Local Incremental Cache (Parquet) DuckDB
Multi-user Collaboration / Service-oriented Any of the above Parquet ClickHouse

1. How to Choose Data Sources

Tushare Pro – Established, standardized fields, large community. Covers daily data, financials, indices, and futures comprehensively. Paid tiers offer significantly better stability and documentation. Suitable as a "stable foundation" for research data. Drawbacks: Points/payment thresholds, rate limits on some interfaces, and average real-time performance (suitable for research, not intraday trading).

AkShare – Free, open-source, with an extensive API coverage (stocks, funds, futures, macroeconomics, alternative data). Zero cost, fast updates, and often the first stop for many. Drawbacks are clear: it is essentially a collection of web scrapers; interfaces may intermittently break when upstream pages change; field naming is inconsistent. Production use requires wrapping it with cleaning, retry logic, and schema validation.

BaoStock – Free, with fewer but stable interfaces. Clean and sufficient for historical daily and financial data. Suitable for scenarios where you "don't want to maintain scrapers and only need clean daily data"; the trade-off is limited variety and dimensions.

XtQuant / QMT – Broker-side market data and trading interfaces. Offers full push market data, Level-1/Level-2 (depending on broker activation), and order placement. For live trading pipelines, it is essentially the only correct solution: high-quality market data and controllable trading. Drawbacks: Requires broker permissions, depends on the client being online, and historical data must be stored locally; additionally, version inconsistencies between broker servers and official documentation are common, so rely on actual returns when troubleshooting.

Conclusion: Start research with AkShare / BaoStock; use Tushare Pro as a fallback for critical data; live trading must go through QMT / XtQuant. When using multiple sources, store them in a unified schema (see Section 4) so downstream scripts don't need to know where the data comes from.

2. How to Choose Storage Formats

Three candidates: CSV, HDF5, and Parquet.

Practical Advice: Use Parquet for "batch write, batch read" data like daily bars and financials; intermediate products in research scripts can be arbitrary, but unify the warehouse layer to Parquet.

Directory layout is more important than format. Two recommended layouts:

daily/year=2026.parquet               # Full market, one file per year: suitable for full-market scans
daily/symbol=600519/year=2026.parquet # Partitioned by asset: suitable for single-asset long history and incremental updates

For daily data volume (millions of rows per year), use the former: fewer files, faster full-market queries. For Tick-level data or scenarios requiring frequent incremental updates, use the latter: modifying one asset touches only one file.

3. How to Choose Query Engines

Conclusion: Put metadata in SQLite; use DuckDB to directly query Parquet for research and analysis; consider ClickHouse only for scaling and service-oriented architectures—do not start with a service.

4. The Storage Layer We Actually Use

Below are two classes continuously used in the Moonshot series (full implementation available in the Moonshot Practice series' accompanying store.py, 544 lines, production-verified). Comments have been simplified, but the logic remains unchanged.

First, the trading calendar—the "header" of the data warehouse. All incremental updates rely on it to determine missing days:

class CalendarModel:
    """Calendar Model: is_open / prev columns, indexed by date"""

    def get_trade_dates(self, start, end):
        mask = (df.index >= start) & (df.index <= end) & (df.is_open == 1)
        return df[mask].index.tolist()

    # Also includes is_trade_date / prev_trade_day / get_next_trade_day /
    # floor / ceil / shift / delta, but only the most common ones are listed here

Then, unified storage: using (date, asset) as the unique key, reading/writing Parquet via Polars, with inputs and outputs remaining in pandas. The core function is get_and_fetch—if local data is missing a segment, it calls your provided fetch_data_func to fill that segment. The application layer never needs to worry about data completeness:

store = ParquetUnifiedStorage(store_path, calendar, fetch_data_func=fetch_bars)
barss = store.get_and_fetch(start, end)   # If missing, fetch; if fetched, store; if stored, read

dv_store = ParquetUnifiedStorage(store_path, calendar, fetch_data_func=fetch_dv_ttm)
dv_ttm = dv_store.get_and_fetch(start, end)

Three hard rules for the write side (all within append_data):

# 1. Deduplicate by (date, asset), keeping the latest → Incremental updates are idempotent, repeated runs do not produce dirty data
deduped = combined.unique(subset=["date", "asset"], keep="last").sort(["date", "asset"])
# 2. Write sorted by date+asset → Range queries are naturally continuous
# 3. Write to disk with lz4 compression → Daily data volume reads/writes are in milliseconds
sorted_df.collect().write_parquet(self._file_path, compression="lz4")

For queries, use Polars lazy scan directly (zero import, zero service):

lazy_df = pl.scan_parquet(self._file_path)
df = (lazy_df.filter((pl.col("asset").is_in(["600519", "000001"]))
                     & (pl.col("date") >= start))
             .collect().to_pandas())

Three engineering details worth copying verbatim:

  1. Calendar First: All "is data missing?" checks go through CalendarModel, not file existence—suspended trading days and holidays are not misjudged as missing;
  2. Decouple Fetching and Storage: fetch_data_func is a pluggable parameter. Switching between Tushare / AkShare / XtQuant requires changing only this function, while the semantics of get_and_fetch remain unchanged—this is the implementation of Section 1's "multi-source coexistence, unified storage";
  3. Unified Adjustment Standard: The warehouse stores "unadjusted price + adjustment factor"; forward/backward adjustment is calculated at query time—storing forward-adjusted prices directly is a classic pitfall, as forward-adjusted prices differ across different lookback windows.

5. Common Pitfalls

6. Further Reading

The comparative conclusions in this article are based on our daily usage experience with A-share daily and minute-level data. Specific performance numbers are strongly correlated with data scale, field width, disk type, and query patterns—recommend testing the code skeleton in this article with your own data before making a final selection.

24 articles · grouped by level

Intermediate (18)

Practitioner (6)

← All topics · Tag cloud · 中文专题