匡醍量化|大富翁量化

20 Months of Self-Taught Quant: From Retail to Probabilities

中文 📅 2026-04-17 👁 views this month —

20 Months of Self-Taught Quant

Over the past two years, the reported salary thresholds for quantitative roles have been roughly as follows: in Shanghai, a fresh graduate Junior Quant can expect a base salary plus bonus of 600,000–800,000 RMB; in Hong Kong, this figure is around 1.2–1.5 million HKD; and if you manage to join a top-tier hedge fund in New York, the starting salary is typically $300,000–$400,000 USD.

Seeing these numbers, many people’s first reaction is: “I need to learn Python and build a machine learning model to trade stocks.”

That was exactly my initial thought. As a retail investor in China A-shares who had paid a steep tuition, I believed that once I learned quantitative methods and could use code to analyze the market, I would be able to precisely predict tomorrow’s price movements and achieve financial freedom.

It wasn’t until I stumbled through 20 months of self-directed learning—watching my account balance fluctuate wildly from losses to gains and back again—that I gradually realized: the core of quantitative trading is not writing code, nor is it predicting the future. It is more like a process of retraining yourself to understand the market.

Looking back, the truly useful things I learned in those 20 months can be summarized into a few key areas. If you also want to start learning quantitative trading from scratch, this might help you avoid some pitfalls.

Table of Contents:

  • From Guessing Up/Down to Calculating Probabilities
  • Before Writing Strategies, Learn to Doubt Yourself
  • Stocks Are Not Viewed Individually
  • The Optimal Solution Calculated May Not Be Executable Live
  • At the Options Stage, Realizing Volatility Itself Costs Money
  • The Final Bottleneck: Estimation Error

From Guessing Up/Down to Calculating Probabilities

As a retail investor, my favorite activity was analyzing charts and guessing price directions: “This candlestick has a long lower shadow, so it must go up tomorrow!”

This is typical absolute thinking. Later, when I studied quantitative trading, the first thing that woke me up was probability theory: in the market, the word “certain” is the most expensive.

1. Retail Investors Look at Single Points; Quant Looks at “Conditions”

I used to think that because a certain stock was a good company, it should rise. Later, I understood that isolated events have no trading value; what matters is conditional probability:

$$ P(A \mid B)=\frac{P(A \cap B)}{P(B)} $$

I no longer ask, “Will it go up tomorrow?” I only ask the program: “Over the past 10 years, when the broader market fell below its 30-day moving average, but the stock’s net inflow exceeded 100 million RMB (Condition B), what was the probability that it rose the next day (Event A)?” Only by adding strict constraints did I realize that the “patterns” I previously thought I saw were all illusions.

2. The Essence of Being Trapped Is Ignorance of Bayes

What is the worst way for a retail investor to die? Holding on stubbornly. After buying and the price crashes, you look everywhere for bullish news to comfort yourself. Bayes’ formula cured my “stubbornness”:

$$ P(H \mid D)=\frac{P(D \mid H)P(H)}{P(D)} $$

No matter how bullish I was before buying (prior), as long as it broke below my quantitative defense line today (new data), my level of conviction must be adjusted downward (posterior). Stop-losses must be executed. The biggest taboo in trading is stubbornness; when new data emerges, previous judgments must be updated.

3. Why Did I Lose Everything Even When I Got the Big Trend Right?

$$ \mathbb{E}[X]=\sum_i x_iP(X=x_i) \quad \text{with} \quad \mathrm{Var}(X)=\mathbb{E}\left[(X-\mathbb{E}[X])^2\right] $$ I once backtested a strategy with high returns (positive expectation) and went all-in. Then, an extreme market event occurred, causing a 40% drawdown. I panicked and cut my losses. Immediately after I sold, the strategy started making money. I finally understood: expectation determines whether you can make money, while variance determines whether you can survive long enough to make it.

My Full-Time Self-Study Checklist (Approx. 1 Month):

  • Core Reading: Sheldon Ross’s A First Course in Probability (Read chapters 1–7 completely, treating homework problems as interview questions).
  • Practical Task 1: Write a Monte Carlo simulator in Python to calculate the probability of ruin within 100 trades if a strategy has a 55% win rate, a 1:1 payout ratio, and your capital can only withstand 10 consecutive losses.
  • Practical Task 2: Write a Bayesian tracking script to simulate a trend-following strategy that loses for 5 consecutive days, observing how its “posterior validity” drops below the 20% warning line.
import numpy as np

# 测算:胜率55%,赔率1:1,初始资金10万,每次下注1万,100笔交易内的破产概率
def simulate_ruin(win_rate=0.55, initial_capital=10, bet_size=1, n_simulations=10000, n_trades=100):
    ruin_count = 0
    for _ in range(n_simulations):
        capital = initial_capital
        for _ in range(n_trades):
            # 赢了+1万,输了-1万
            capital += bet_size if np.random.rand() < win_rate else -bet_size
            if capital <= 0:
                ruin_count += 1
                break
    return ruin_count / n_simulations

ruin_prob = simulate_ruin()
print(f"正期望策略的破产概率(方差的惩罚):{ruin_prob*100:.2f}%")
# 哪怕期望收益为正,只要仓位管理不对,你依然可能死在半路上。

Before Writing Strategies, Learn to Doubt Yourself

After learning some coding, I started writing strategies frantically. One day, I tuned a parameter set that yielded a perfect 50% annualized return, so excited that I didn’t sleep all night. I launched it for live trading, and lost 15% in half a month.

At that moment, I realized that the human brain is too good at deceiving itself. I had to learn statistics to pour cold water on my enthusiasm.

1. Assume You Are a Liar First

Now, when I get good backtest results, my first reaction is not joy, but hypothesis testing: assume that this 50% return was entirely due to my luck (the null hypothesis $H_0$). What is the probability (p-value) of purely getting lucky enough to produce this result? If the probability is high, it means the strategy is garbage.

2. Parameter Tuning Is Self-Deception

I used to try various parameters: if the 5-day moving average didn’t work, try the 10-day one. I could always find a profitable set. Later, I learned this is called the multiple comparisons problem. If you try 1,000 times, you will inevitably hit a lucky streak. Without Bonferroni correction ($\alpha_{\text{adj}}=\frac{\alpha}{n}$), the so-called “strongest strategy” I selected was merely a perfect fit to a segment of historical noise.

3. Accept That You Can’t Calculate Accurately

$$ \hat{\theta}{MLE} = \arg\max\theta \sum_{i=1}^n \log f(x_i\mid\theta) $$ How does the market actually work? No one knows. Maximum Likelihood Estimation (MLE) taught me pragmatism: take the current batch of data and reverse-engineer a set of parameters that are “most likely to generate this data” for practical use. Don’t pursue perfect prediction; just make it work for the current sample.

My Full-Time Self-Study Checklist (Approx. 1.5 Months):

  • Core Reading: Casella & Berger’s Statistical Inference (Very hardcore, specifically cures various assumptions about statistical indicators).
  • Practical Task 1: Pull the last 10 years of S&P 500 data, write a Rolling Window, and use MLE to estimate the kurtosis and fat-tail characteristics of returns.
  • Practical Task 2: Hand-code a permutation tester: fix future returns, shuffle trading signals, and rerun 5,000 times. See where the original strategy’s Sharpe ratio ranks among these 5,000 random results. If it doesn’t rank in the top 5%, discard it immediately.
import numpy as np

# 用置换检验打碎你的“回测幻觉”
def permutation_test(signal, future_returns, n_permutations=5000):
    strategy_returns = signal * future_returns
    actual_sharpe = np.mean(strategy_returns) / np.std(strategy_returns) * np.sqrt(252)

    random_sharpes = []
    for _ in range(n_permutations):
        # 固定未来收益,只打乱信号与收益的对应关系
        shuffled_signal = np.random.permutation(signal)
        shuffled_returns = shuffled_signal * future_returns
        sharpe = np.mean(shuffled_returns) / np.std(shuffled_returns) * np.sqrt(252)
        random_sharpes.append(sharpe)

    p_value = np.mean(np.array(random_sharpes) >= actual_sharpe)
    return actual_sharpe, p_value

np.random.seed(42)
future_returns = np.random.normal(0.0005, 0.015, 1000)
signal = np.where(np.random.randn(1000) > 0, 1, -1)

sharpe, p_val = permutation_test(signal, future_returns)
print(f"原策略夏普: {sharpe:.2f}")
print(f"随机打乱信号后仍然不弱于它的概率: {p_val:.4f}")
# 如果 p 值大于 0.05,这个策略大概率只是噪音。

Stocks Are Not Viewed Individually

Retail investors look at stocks one by one. But quantitative trading cannot be played this way. When managing dozens of stocks, you must learn linear algebra to see the network behind them.

1. Stocks Are Interconnected

I used to go all-in on 5 stocks, thinking I was diversified, but when the market dropped, they all crashed. This was because their correlation was too high. Later, I learned to use the covariance matrix $\Sigma$. The risk of the entire portfolio is actually a matrix operation:

$$ \sigma_p^2=\mathbf{w}'\Sigma\mathbf{w} $$

I am not buying isolated stocks; I am buying the relationships between them.

2. Filtering Out Useless News

$$ \Sigma=Q\Lambda Q^\top $$ Reading news to trade stocks is too exhausting. PCA (Principal Component Analysis) helped me reduce dimensions. It filters out the chaotic ups and downs of hundreds of stocks and extracts only two or three core axes driving the broader market. Since the covariance matrix is symmetric, orthogonal decomposition is used here. By seeing through the main axes, I no longer need to stare at individual stock noise.

My Full-Time Self-Study Checklist (Approx. 1.5 Months):

  • Core Reading: David C. Lay’s Linear Algebra and Its Applications (This book contains many brilliant discussions on dynamical systems).
  • Required Derivation: On paper, use eigendecomposition to break down the variance formula for multi-asset portfolios $\sigma_p^2=\mathbf{w}'\Sigma\mathbf{w}$ down to the most basic orthogonal vectors.
  • Practical Task: Download daily frequency data for Nasdaq 100 constituent stocks, write a PCA-based Statistical Arbitrage backtest yourself: first extract the first few principal components, then study the residual series after stripping out common factors.
import numpy as np

# 模拟 50 只股票的收益率矩阵(带有一定的系统性相关)
np.random.seed(42)
n_stocks = 50
n_days = 1000
market_factor = np.random.normal(0, 0.01, n_days)
# 股票收益 = 市场因子暴露 + 个股特异性噪音
returns = np.outer(market_factor, np.random.uniform(0.5, 1.5, n_stocks)) + np.random.normal(0, 0.02, (n_days, n_stocks))
centered_returns = returns - returns.mean(axis=0, keepdims=True)

# PCA:协方差矩阵特征分解
cov_matrix = np.cov(centered_returns, rowvar=False)
eigenvalues, eigenvectors = np.linalg.eigh(cov_matrix)

# 对特征值降序排列
idx = np.argsort(eigenvalues)[::-1]
eigenvalues = eigenvalues[idx]
eigenvectors = eigenvectors[:, idx]

variance_explained = eigenvalues / np.sum(eigenvalues)
top_k = 3
factor_loadings = centered_returns @ eigenvectors[:, :top_k]
common_component = factor_loadings @ eigenvectors[:, :top_k].T
residuals = centered_returns - common_component

print(f"前三个主成分一共解释了总方差的: {variance_explained[:top_k].sum()*100:.2f}%")
print(f"去掉公共因子后,残差日波动率均值约为: {residuals.std(axis=0).mean():.4f}")

The Optimal Solution Calculated May Not Be Executable Live

After calculating the perfect portfolio weights, what do you do if you can’t actually buy them in live trading? I started learning calculus and optimization theory to learn how to compromise with reality.

1. Don’t Look Too Far; Focus on the Present

$$ f(x)=f(a)+f'(a)(x-a)+\frac{f''(a)}{2}(x-a)^2+\cdots $$ Predicting the future is too difficult. The lesson Taylor expansion taught me is: just hedge the small step of risk right now (first and second derivatives). Don’t guess the big trend; just manage current survival.

2. Dancing in Shackles

$$ \min_{\mathbf{w}}\mathbf{w}'\Sigma\mathbf{w} \quad \text{s.t.} \quad \mu'\mathbf{w}\ge\bar{\mu},; \mathbf{1}'\mathbf{w}=1 $$ Real trading has fees, slippage, and position limits. Convex optimization solvers do exactly this: under these rigid “shackles,” hard-calculate the least-bad solution that you can actually operate.

My Full-Time Self-Study Checklist (Approx. 1 Month):

  • Core Reading: Boyd & Vandenberghe’s Convex Optimization (Thoroughly understand the KKT conditions; this is the core code for compromising with reality in optimization problems).
  • Derivation Task: Prove the economic meaning of Lagrange Multipliers under portfolio capital constraint (i.e., “shadow price”).
  • Practical Task: Use scipy.optimize to build a long-short portfolio model, adding constraints like “maximum leverage = 1.5” and “turnover per period < 10%,” plus an additional turnover penalty term, then force a solution.
import numpy as np
from scipy.optimize import minimize

# 一个带有现实约束的“脏活”优化器
def constrained_portfolio_optimization(
    expected_returns,
    cov_matrix,
    previous_weights,
    target_return=0.03,
    max_leverage=1.5,
    max_turnover=0.10,
    turnover_penalty=0.01,
):
    n = len(expected_returns)
    
    def objective(weights):
        variance = np.dot(weights.T, np.dot(cov_matrix, weights))
        turnover_cost = turnover_penalty * np.sum(np.abs(weights - previous_weights))
        return variance + turnover_cost
    
    # 限制1:权重和为1(不能无限借钱)
    # 限制2:预期收益必须达到目标
    # 限制3:总杠杆不能超过 max_leverage
    # 限制4:单期换手率不能超过 max_turnover,这里按 0.5 * sum(|w_t - w_{t-1}|) 计算
    constraints = [
        {'type': 'eq', 'fun': lambda w: np.sum(w) - 1},
        {'type': 'ineq', 'fun': lambda w: np.dot(w, expected_returns) - target_return},
        {'type': 'ineq', 'fun': lambda w: max_leverage - np.sum(np.abs(w))},
        {'type': 'ineq', 'fun': lambda w: max_turnover - 0.5 * np.sum(np.abs(w - previous_weights))}
    ]
    
    # 现实的枷锁:单只股票仓位不能超过30%,可以做空但不能低于-10%
    bounds = tuple((-0.1, 0.3) for _ in range(n))
    
    initial_weights = previous_weights.copy()
    result = minimize(objective, initial_weights, method='SLSQP', bounds=bounds, constraints=constraints)
    return result.x

# 市场并不完美,求解器给你的往往也是一个满是“妥协”的解。

At the Options Stage, Realizing Volatility Itself Costs Money

Finally, I touched the deep end of quantitative trading: stochastic calculus.

1. Bumps Themselves Cost Money

Let $Y=f(X,t)$, and $X_t$ follows a diffusion process driven by Brownian motion. Then Itô’s Lemma adds a second-order term. Many people get stuck for the first time when learning stochastic calculus at this step: $(dW_t)^2=dt$.

$$ dY= \left( \frac{\partial f}{\partial t} + \mu\frac{\partial f}{\partial X} + \frac{1}{2}\sigma^2\frac{\partial^2 f}{\partial X^2} \right)dt + \sigma\frac{\partial f}{\partial X}dW_t $$

This extra second-order term tells me: for options positions and dynamic hedging, the price path itself generates P&L. Even if you only look at the start and end points, you might miss the risks that truly matter in between.

2. Abandon Personal Views

$$ \frac{\partial V}{\partial t} + \frac{1}{2}\sigma^2S^2\frac{\partial^2 V}{\partial S^2} + rS\frac{\partial V}{\partial S} -rV=0 $$ In the Black-Scholes equation, $\mu$, representing the expected return rate, disappears. Option pricing does not directly depend on your subjective assumption of the stock’s long-term return rate, but rather depends more on the current price, strike price, maturity, risk-free rate, and volatility. As long as hedging is done well, whether I personally think the market will go up or down tomorrow is not that important. What truly matters is how you handle volatility and risk.

My Full-Time Self-Study Checklist (Approx. 2 Months):

  • Ultimate Bible: Björk’s Arbitrage Theory in Continuous Time (Explains martingale theory and measure change very rigorously).
  • Practical Task 1: Derive option pricing using a Binomial Tree model on paper, and prove that as the step size $\Delta t \to 0$, it converges to the Black-Scholes analytical solution.
  • Practical Task 2: Write a dynamic Delta Hedging Simulator for options. Given a stock price path with jumps, calculate how much friction cost you spend hedging daily, and see how large the final hedging error actually is.
import numpy as np

# 这是一个极其简化的 Delta 对冲模拟器:同时看看对冲误差和摩擦成本
def simulate_delta_hedging(S0=100, K=100, T=1.0, r=0.05, sigma=0.2, steps=252):
    from scipy.stats import norm
    dt = T / steps
    
    # 模拟一条真实的带有跳跃的股价路径
    prices = [S0]
    for _ in range(steps):
        jump = np.random.normal(0, sigma * np.sqrt(dt))
        # 偶尔来个暴跌(模拟黑天鹅)
        if np.random.rand() < 0.01: jump -= 0.05
        prices.append(prices[-1] * np.exp((r - 0.5*sigma**2)*dt + jump))
        
    def get_delta(S, t):
        if t >= T: return 1.0 if S > K else 0.0
        tau = max(T - t, 1e-12)
        d1 = (np.log(S/K) + (r + sigma**2/2)*tau) / (sigma*np.sqrt(tau))
        return norm.cdf(d1)
    
    # 每天调整仓位,计算摩擦成本
    hedging_cost = 0
    current_delta = get_delta(S0, 0)
    cash_account = -current_delta * prices[0]
    for i in range(1, steps):
        new_delta = get_delta(prices[i], i*dt)
        # 假设每次调仓有万分之五的手续费
        delta_change = new_delta - current_delta
        trade_cost = abs(delta_change) * prices[i] * 0.0005
        cash_account -= delta_change * prices[i] + trade_cost
        hedging_cost += trade_cost
        current_delta = new_delta
    
    option_payoff = max(prices[-1] - K, 0)
    hedge_value = current_delta * prices[-1] + cash_account
    hedge_error = hedge_value - option_payoff
    return hedging_cost, hedge_error

cost, hedge_error = simulate_delta_hedging()
print(f"一年下来,摩擦成本大约吃掉了 ${cost:.2f}")
print(f"最终对冲误差约为 ${hedge_error:.2f}")

The Final Bottleneck: Estimation Error

After 20 months, I no longer look at K-line charts, nor do I argue in stock forums about tomorrow’s market direction.

I finally understand that many retail investors transitioning to quantitative trading end up on a dead end because they only learned to write code but did not learn to revere estimation error. They take noisy historical data, calculate a seemingly perfect parameter, and go all-in. Essentially, they are still “gambling.”

The current me acknowledges that historical data is dirty and models are fragile. I don’t pursue predicting tomorrow accurately; I only do high-probability things, strictly hedge risks, diversify positions, and leave the rest to time.

If you are currently walking this path and don’t want to step on these pitfalls again from scratch, you can first check out our Quantitative 24 Lessons to build the overall framework. Once you have the foundation in probability, statistics, and backtesting, it will be easier to connect factor research, feature engineering, and strategy iteration by following our Factor Mining and Machine Learning Strategy Course.

This is probably what I truly learned in these 20 months.