GBDT Regression Secrets: Beyond DeepSeek
Decision trees are a cornerstone of machine learning. Fundamentally, they are algorithms that convert hard-coded if-else logic into models trained on data, achieving functional equivalence to manual coding while offering superior scalability.
for 有房, 年薪 in [("有", "40万"), ("有", "20万"), ("无", "100万")]:
if 有房 == "有" and 年薪 > "30万":
print("见家长!")
else:
print("下次一定")
This approach significantly enhances algorithmic universality: as long as labeled data exists, no manual coding is required to convert inputs into a decision tree model. The more complex the conditions, the more pronounced this advantage becomes. Furthermore, the training process naturally incorporates statistical features of data distribution and includes error tolerance (provided the data labels are correct).
The Single-Cell Organism: Decision Trees
For example, suppose I am an assistant to a "CEO" (a term often used in Chinese romance novels for a dominant, wealthy male lead) and need to schedule whether he should work tomorrow based on his habits. I have collected the following historical data:
data = {
'天气': ['晴', '晴', '晴', '晴', '阴', '阴', '雨', '雨'],
'气温': ['高温', '高温', '舒适', '凉爽', '凉爽', '凉爽', '凉爽', '凉爽'],
'宜工作': [0, 0, 1, 1, 1, 1, 0, 0],
}
df = pd.DataFrame(data)
df
| 天气 | 气温 | 宜工作 | |
|---|---|---|---|
| 0 | 晴 | 高温 | 0 |
| 1 | 晴 | 高温 | 0 |
| 2 | 晴 | 舒适 | 1 |
| 3 | 晴 | 凉爽 | 1 |
| 4 | 阴 | 凉爽 | 1 |
| 5 | 阴 | 凉爽 | 1 |
| 6 | 雨 | 凉爽 | 0 |
| 7 | 雨 | 凉爽 | 0 |
We can use a decision tree to train a model to schedule his business trips. If he enters a passionate romance with a female celebrity, we simply add a new condition: if he studied English the night before, he won’t work. In this case, we only need to update the data.
data = {
'天气': ['晴', '晴', '晴', '晴', '阴', '阴', '雨', '雨'],
'气温': ['高温', '高温', '舒适', '凉爽', '凉爽', '凉爽', '凉爽', '凉爽'],
'学英语':[0, 1, 0, 0, 1, 0, 0, 1],
'宜工作': [0, 0, 1, 1, 1, 1, 0, 0],
}
df = pd.DataFrame(data)
df
The decision tree model below is simple, but it illustrates the entire process of building a decision tree:
import numpy as np
import pandas as pd
from sklearn.tree import DecisionTreeClassifier, plot_tree
import matplotlib.pyplot as plt
# Create sample data
data = {
'天气': ['晴', '晴', '晴', '晴', '阴', '阴', '雨', '雨'],
'气温': ['高温', '高温', '舒适', '凉爽', '凉爽', '凉爽', '凉爽', '凉爽'],
'学英语':[0, 0, 0, 0, 1, 0, 0, 1],
'宜工作': [0, 0, 1, 1, 0, 1, 0, 0],
}
df = pd.DataFrame(data)
df
# Convert categorical variables to numerical values
df['天气'] = df['天气'].map({'晴': 0, '阴': 1, '雨': 2})
df['气温'] = df['气温'].map({'高温': 0, '舒适': 1, '凉爽': 2})
X = df[['天气', '气温', '学英语']]
y = df['宜工作']
# Create and train the decision tree model
clf = DecisionTreeClassifier()
clf.fit(X, y)
# Visualize the decision tree
plt.figure(figsize=(12, 8))
plot_tree(clf, filled=True, feature_names=['天气', '气温', '学英语'], class_names=['诸事不宜', '宜工作'], fontsize=10)
plt.title('Should the CEO Work?')
plt.show()
# Predict whether he should work the next day
weather = "晴"
temp = "高温"
dating=0
sample = pd.DataFrame([(weather, temp, dating)], columns=["天气", "气温", "学英语"])
sample['天气'] = sample['天气'].map({'晴': 0, '阴': 1, '雨': 2})
sample['气温'] = sample['气温'].map({'高温': 0, '舒适': 1, '凉爽': 2})
prediction = clf.predict(sample)
dating_desc = "Did not study English" if dating == 0 else "Studied English last night"
if prediction[0] == 1:
print(weather, temp, dating_desc, "宜工作")
else:
print(weather, temp, dating_desc, "诸事不宜")