Supervised vs. Reinforcement Learning: AI Trading Strategies
Supervised vs. Reinforcement Learning: Two AI Approaches to Trading
1. A Real-World Analogy
Imagine teaching a child to trade stocks:
Method 1: Supervised Learning
You give them 1,000 candlestick charts.
Each chart is labeled "up tomorrow" or "down tomorrow."
You teach them: "This pattern = up," "That pattern = down."
Then you test them: show new charts and ask for predictions.
Method 2: Reinforcement Learning
You give them $1 million in virtual capital.
You say: "Trade on your own. If you make money, I’ll give you candy. If you lose, I’ll slap your hand."
They don’t memorize specific patterns but gradually learn "when to buy and when to sell."
💡 This is the fundamental difference.
2. Supervised Learning: Exams with Standard Answers
2.1 What is Supervised Learning?
Supervised Learning = Teacher-guided = Standard answers exist
Three Key Elements:
| Element | Description | Financial Example |
|---|---|---|
| Input (X) | Historical stock data | Price, volume, MACD, etc. |
| Label (Y) | Standard answer | Up/down tomorrow |
| Goal | Learn mapping function | Predict Y from X |
2.2 Example in Finance
Training Data:
┌─────────────────────────────────────┬──────────┐
│ Input Features (X) │ Label (Y)│
├─────────────────────────────────────┼──────────┤
│ 20-day return: +15% │ │
│ Volume: 2x normal │ Up │
│ MACD: Golden Cross │ (+1) │
│ RSI: 65 │ │
├─────────────────────────────────────┼──────────┤
│ 20-day return: -10% │ │
│ Volume: Shrinking │ Down │
│ MACD: Death Cross │ (-1) │
│ RSI: 30 │ │
└─────────────────────────────────────┴──────────┘
What the Model Learns:
"Large gains + high volume + MACD golden cross → Up"
"Large drops + low volume + MACD death cross → Down"
2.3 Advantages of Supervised Learning
✅ Simple and Direct — Clear objectives, easy to train
✅ High Interpretability — Understand why the model predicts up/down
✅ High Data Efficiency — Every chart is used for training
✅ Mature Tools Available — XGBoost, LightGBM, and neural networks are well-established
2.4 Challenges in Finance
❌ Challenge 1: Defining Labels is Hard
Consider "moving average bullish alignment":
- March 2020: Fed easing → Surged
- March 2022: Fed tightening → Crashed
- Same pattern, opposite labels, confusing the model
❌ Challenge 2: Severe Overfitting
- The model memorizes every detail of historical data.
- But when the market changes, the patterns break.
- 90% accuracy in backtests, 50% in live trading (no better than flipping a coin).
❌ Challenge 3: Short-term Focus, Ignoring Long-term
- Predicts "up tomorrow" but ignores "down the day after."
- Buy today, gain 1%; sell tomorrow, lose 10%.
- Net result: Loss.
3. Reinforcement Learning: Learning Through Practice
3.1 What is Reinforcement Learning?
Reinforcement Learning = No teacher = Only rewards and penalties
Four Key Elements:
| Element | Description | Financial Example |
|---|---|---|
| State | Current market condition | Price, positions, capital, etc. |
| Action | Decision taken | Buy/sell/hold, how much |
| Reward | Feedback signal | Profit = positive reward, Loss = negative reward |
| Goal | Maximize cumulative long-term reward | Learn optimal trading strategy |
3.2 Example in Finance
Round 1:
State: Cash $1M, Moutai stock $1,800
Action: Buy $500k worth
Result: Stock rises to $1,900 in one week
Reward: +$27k ✅
Round 2:
State: Cash $500k, Positions $500k
Action: Buy another $300k
Result: Stock falls to $1,700 in one week
Reward: -$53k ❌
Round 3:
State: Cash $200k, Positions $800k
Action: Sell $400k
Result: Stock falls to $1,600 in one week
Reward: +$40k ✅ (Reduced loss)
After 1,000 rounds, the model learns:
"Sell in batches after large gains."
"Wait after large drops; don’t rush to bottom-fish."
"Always keep cash; never go all-in."
3.3 Advantages of Reinforcement Learning
✅ Delayed Rewards — Not just about tomorrow, but long-term returns
✅ Adaptability — Strategies adjust automatically as the market changes
✅ No Labels Needed — No manual labeling of up/down required
✅ Exploration Capability — Tries new strategies, discovering hidden patterns
3.4 Challenges in Finance
❌ Challenge 1: Slow Training
- Requires thousands of trades to learn.
- Live trading is impossible; must simulate with historical data.
- Simulation differs from live trading.
❌ Challenge 2: Designing Rewards is Hard
- Profit = +1, Loss = -1?
- But risk control is also critical.
- How to incorporate "avoiding margin calls" into the reward function?
❌ Challenge 3: High Cost of Exploration
- Must try various strategies to find the best one.
- But the cost of trial and error in financial markets is extremely high.
- One major loss can wipe out the account.
4. Core Comparison
4.1 Learning Method Comparison
| Comparison Item | Supervised Learning | Reinforcement Learning |
|---|---|---|
| Learning Style | Teacher-guided (standard answers) | Practice-based (only rewards/penalties) |
| Goal | Prediction accuracy | Maximize profits |
| Time Horizon | Short-term (up/down tomorrow) | Long-term (30-day total return) |
| Data Requirements | Requires labeled data | Only needs historical prices |
| Adaptability | Poor (fails when market changes) | Strong (automatically adjusts strategies) |
| Interpretability | High (knows why it predicts) | Low (black-box strategy) |
| Training Difficulty | Simple | Difficult |
4.2 Trading Metaphors
What is Supervised Learning like?
Memorizing candlestick charts.
"This pattern historically went up, so I buy."
But not knowing *why* it should go up.
What is Reinforcement Learning like?
Gaining experience through live trading.
"Last time I bought here and made money, so I buy again."
"Last time I bought here and lost, so I wait."
Gradually forming one’s own trading discipline.
5. Which is Better for Quantitative Finance?
5.1 Scenarios Suitable for Supervised Learning
| Scenario | Description |
|---|---|
| ✅ Stable Data Patterns | Certain factors remain effective long-term (e.g., P/E, ROE); market styles don’t switch frequently. |
| ✅ Short-term Prediction | Intraday high-frequency trading, arbitrage strategies (calendar spread, cross-market arbitrage). |
| ✅ Clear Labels Exist | Up/down binary classification, return regression. |
5.2 Scenarios Suitable for Reinforcement Learning
| Scenario | Description |
|---|---|
| ✅ Long-term Strategy Optimization | Asset allocation, position sizing, stop-loss/take-profit strategies. |
| ✅ Complex Decision-Making | Dynamic multi-factor weights, multi-asset portfolio optimization, strategies considering transaction costs. |
| ✅ High Market Adaptability Required | Trend following, dynamic hedging. |
5.3 Practical Implementation Advice
For Beginners:
1. Start with supervised learning (XGBoost, LSTM).
2. Understand data, features, and models.
3. Recognize the limitations of supervised learning.
For Advanced Users:
1. Experiment with reinforcement learning (DQN, PPO).
2. Start with simple scenarios (single stock, fixed position sizing).
3. Gradually increase complexity.
For Experts:
1. Combine supervised and reinforcement learning.
2. Use supervised learning for feature extraction.
3. Use reinforcement learning for decision optimization.
6. Summary
One-Sentence Takeaway
| Method | Core Logic |
|---|---|
| Supervised Learning | "This pattern historically went up, so I buy now." |
| Reinforcement Learning | "Buying now leads to long-term profitability, so I buy now." |
Selection Guide
┌─────────────────────────────────────────────┐
│ │
│ What is your scenario? │
│ │
│ 1. Clear labels, stable patterns │
│ → Use supervised learning (XGBoost/LSTM)│
│ │
│ 2. Need long-term optimization, dynamic │
│ adjustment │
│ → Use reinforcement learning (DQN/PPO) │
│ │
│ 3. Need both │
│ → Supervised learning for features + │
│ Reinforcement learning for decisions │
│ │
└─────────────────────────────────────────────┘
💡 After reading this, you should know which method to choose.
Remember: There is no best method, only the most suitable one.