匡醍量化|大富翁量化

Zillionare 1.0: A High-Performance Distributed Quant Framework

中文 📅 2020-11-30 👁 views this month —

Introduction to Zillionare 1.0

Zillionare (Da Fu Weng / "Monopoly") is a distributed, high-speed quantitative trading framework. Built on Python 3.8+, it leverages asynchronous I/O and microservices architecture to achieve the following goals:

  • A distributed, high-performance quantitative computing platform with seamless, scalable compute resources.
  • Localized market data storage with low-latency synchronization, easily supporting real-time triggering of trading signals at minute-level granularity or higher.
  • Support for private deployment to protect proprietary strategies.
  • Virtualized deployment for easier maintenance.
  • Integration of numerous quantitative factors.
  • Machine learning-based trading strategies.
  • Deep learning-based trading strategies.
  • A high-speed backtesting framework.

If you are not viewing this documentation from the Quantide official website, please click the link Quantide to navigate to the official site for reading. We provide additional documentation and tutorials on the website, formatted for better readability.

Why Develop Zillionare?

The author began focusing on quantitative trading in 2019. After testing various commercial and open-source products, the author found that no existing product offered low-latency, accurate data at an affordable cost, particularly for trading strategies based on machine learning and deep learning.

Some products excel in data source organization and reliability but do not provide offline storage. If every data point must be fetched online, issues arise regarding data request response speed and costs.

Slow Data Request Response Times

Info Our strategies often require scanning the entire market to discover trading signals as comprehensively as possible. However, this leads to a massive volume of data requests.

Consider a moving average strategy that utilizes annual line data. When scanning the entire market, a single scan requires fetching 1 million data points from the server (assuming 4,000 targets in the China A-shares market, with a maximum of 250 trading days of annual data per target).

Fetching such large volumes of data involves an astonishing number of network requests and time consumption. For example, using a specific SDK, fetching 1 million data points in a single session typically takes 7–10 hours. Consequently, implementing any real-time or near-real-time trading strategy becomes impossible. Even in non-real-time scenarios, such as database initialization after initial installation, the wait time for some open-source projects can stretch to days (even with breakpoint support), which is unacceptable.

**Costs**
Info Fetching all data from the server imposes a significant load, meaning services cannot be provided for free.

For instance, JoinQuant is an excellent data provider. Users can apply for a one-year free trial, which offers 1 million data points (requests) per day. While this scheme is generous, it is insufficient for actual production environments and even for independent research. The author frequently encounters situations where the quota is exhausted before a single round of unit tests completes.

The fundamental solution to these issues is to store historical K-line data locally and synchronize only the differential data between the local storage and the upstream server in real time.

Some products offer similar localized data storage solutions, but they come with limitations and defects.

For example, some products divide data-fetching APIs into "offline" and "online" categories. Using the offline API allows reading data from local storage, but the data is not real-time and cannot guide real-time trading. To obtain real-time data, one must still use the online API, either by fetching all required data at once (negating the advantage of local storage) or by fetching differential data via the online API and manually merging it with offline data from the local database. Ideally, the product should transparently provide the data the user needs, hiding these technical details internally, allowing users to focus more on strategy development.

**Technical Debt**
Info From a technical perspective, the Python community has progressed significantly. Some products started early and locked their architecture into outdated tech stacks, creating technical debt.

Users of quantitative trading frameworks often need to perform secondary development. During this process, they are inadvertently locked into these outdated technologies, which is unfair to secondary developers.

For instance, some products use synchronous I/O for both reading local data (from databases) and remote data. For applications developed in Python, this is a severe performance design flaw. Python’s limited computational resources are often consumed by I/O wait times.

**Quality Control**
Info The quality control of some commercial projects is difficult to assess. However, if open-source projects do not use CI (Continuous Integration) or provide test coverage reports, their quality control remains hard to evaluate despite being open-source.
Therefore, the author decided to develop a framework dedicated to resolving these issues and introducing machine learning and deep learning into trading strategies.

Of course, due to limitations in personal technical skills and vision, Zillionare certainly has its shortcomings. Whether it meets the author's expectations depends on user feedback.

What Are the Technical Features of Zillionare?

High Performance To ensure Zillionare’s high performance, the author engaged in deep thinking regarding architecture, component selection, algorithms, and data structures, making trade-offs from a product positioning perspective.

Info The Zillionare architecture utilizes multi-process, distributed microservices, message queues, load balancing, and caching (in-memory) databases. For market data synchronization (both real-time and non-real-time), it employs a multi-process collaboration model, significantly increasing network concurrency. In terms of process collaboration, it supports both load balancing modes (via Nginx) and parallel modes (supported by built-in messaging mechanisms).

Network components are chosen to support asynchronous I/O wherever possible. For example, aioredis is used for accessing Redis databases, and asyncpg is used for database access.

After in-depth research into the primary data type—market data—and considering its distribution and processing characteristics, we decided to use Redis for persisting market data and Numpy structured arrays for in-memory representation of market data (other frameworks generally use Pandas). For the rationale behind these choices, please refer to [[TODO]].

Additionally, we continuously optimize from an operational perspective. For instance, due to CDN acceleration and multi-process modes, importing one year of China A-share data (including 30-minute and longer K-lines) in Zillionare takes only 170–240 seconds (depending on CPU cores and network bandwidth). If sufficient CPU cores are available, importing several years of data takes approximately 15–20 minutes (in actual tests, importing 70 months of data took 776 seconds using 32 processes). These data points will continue to be optimized to ensure good performance even on lower-spec computers. Users who have used other products may have a deep appreciation for Zillionare’s speed.

**Ease of Use and Quality Control** Before the 1.0 release, Zillionare’s two most critical early components underwent multiple rewrites to ensure proper decoupling and aggregation of functionalities, guaranteeing API stability and readability.
Info

If you have the opportunity to delve into Zillionare’s development toolchain, you will find that quality control occurs at every stage: from code submission and unit testing to continuous integration.

To foster an active open-source community for Zillionare, the author has published a series of articles introducing Python development toolchains. Please refer to [[TODO]].

For deployment, we support both virtualized deployment and pip installation. Virtualized deployment is fully automated with one click, while the pip installation mode provides an installation wizard to guide users through configuration after installation, making the entire process straightforward.

**Keeping Pace with Community Latest Technologies**
Info Initially, Zillionare used the `cookie-cutter` pyproject template to build the development framework. This template used Makefiles, pip, and requirements files to manage dependencies and build the project, twine for publishing, and tools like tox and pytest for quality control, with Sphinx for documentation generation.

In dependency management, we observed the community migrating toward new standards and tools. Poetry, which fully complies with PEP 517, has become the community’s preferred tool for dependency, build, and version management. Therefore, Zillionare completely replaced its build framework last November.

The documentation tool was also switched to MkDocs to reflect the reality that Markdown formats are preferred by both authors and readers.

Furthermore, we integrated Travis CI and Codecov for continuous integration and pre-commit-hooks to ensure the quality of committed code. Before the 1.0 release, the test coverage reached 93%.

Moving forward, Zillionare will maintain this technical positioning and strive for greater standardization in project releases. We will determine project directions via RFCs, define API interfaces, and ensure interface usability through technical testing coverage.

**AI-First Orientation**
Info

Although Zillionare’s AI functionalities are still being implemented, our design fully reflects an AI-first orientation. For example, our design prioritizes momentum strategies, with less attention paid to text and financial data. This is because AI is not yet fully reliable in text analysis, and trading based on news carries low certainty. Financial data presents similar challenges. Currently, AI performs best in momentum strategies, which is why Zillionare focuses on this area.

Regarding quantitative factors, we do not view them, as in traditional financial engineering, as "golden geese" that can be simply combined to generate returns. Instead, we treat them primarily as data preprocessing and feature extraction inputs for machine learning and deep learning. This also reflects our AI-oriented design philosophy, which we believe is the correct direction.

**Comprehensive and Detailed Documentation**
Info Zillionare provides rich documentation, including beginner guides, tutorials, and detailed API documentation.

Beyond these features, the author strives to explain not just "what" but also "why" in the documentation, sharing more of the thinking behind technical decisions, making it easier for readers to understand and adapt.

# Who Is Zillionare For?
  1. Zillionare is positioned to provide quantitative trading tools and strategies for the China A-share market, with a focus on momentum strategies. Zillionare currently has no plans to integrate futures or digital currencies.
  2. High-frequency arbitrage trading is not Zillionare’s focus area—this field relies more on hardware, network speed, or fundamentally, capital; algorithms and strategies may not be dominant here. Zillionare is positioned to serve the complex China A-share market, bringing open-source AI technology to it. Zillionare’s architecture is designed to handle market data at minute-level granularity or higher. While deep learning on Tick-level data has its merits, this domain should primarily focus on arbitrage trading.
  3. The historical market data provided by Zillionare starts from 2015 and currently includes all stocks with 30-minute and longer K-line data.
  4. Users of quantitative trading frameworks often have needs for secondary development. Zillionare’s advanced and standardized technology stack ensures that user investments remain effective over a long period in the future.
  5. From a tech stack perspective, Zillionare is developed purely in Python. The operating system requirement is Ubuntu > 18. Users unfamiliar with this OS can deploy it via our containerized deployment method.

Compared to other products, Zillionare’s feature set is tightly focused. We prefer to remain focused and concentrate on doing things well.