scikit-learn 机器学习实战教程
适用:Python 中级,Windows + Miniconda + VS Code。
适用版本:scikit-learn 1.5 / 1.6(当前最新稳定版)。
风格约定:所有示例都用numpy在脚本内构造小数据集,不依赖任何真实大文件,可直接复制运行。
1. 简介与定位:机器学习在化学研究中的角色
在催化与材料科学领域里,机器学习(尤其是 scikit-learn 这类经典算法)通常扮演三类角色:
-
活性 / 性能预测(描述符 → 性能)
你有一批材料样本,每个样本由一组"描述符"(descriptor)刻画,例如电负性、d 带中心、吸附能、配位数、形成能等;对应的标签是实验或 DFT 算出的性能(如过电位、TOF、转化频率、选择性)。目标是从"描述符"回归或分类出"性能"。 -
材料空间探索与聚类
在几百上千个候选材料里,先用 PCA 降维到 2~3 维可视化,再用 KMeans 聚类,找到结构相似的"材料家族",指导后续 DFT 精算的取舍。 -
机器学习势函数(MLIP)的"上游预处理与特征工程"
严格意义上的 MLIP(如训练神经网络势)多用 PyTorch / MACE / NequIP 等,但 scikit-learn 常被用于:标准化原子环境特征、用线性回归或岭回归做"快速基线模型"、用 PCA 压缩描述符维度。
本教程能帮你建立的完整工作流:
原始描述符数据
├── 缺失值处理 (SimpleImputer)
├── 特征标准化 (StandardScaler / MinMaxScaler)
├── 划分训练/测试集 (train_test_split)
├── 特征工程管线 (Pipeline + ColumnTransformer)
├── 建模 (线性/正则化/随机森林/SVM)
├── 评估 (交叉验证 + R²/MAE/RMSE/混淆矩阵)
├── 调参 (GridSearchCV)
└── 保存/加载 (joblib)
2. 安装与版本检查
在 Miniconda 环境里(建议为科研项目单独建环境):
# 新建环境(可选,Python 3.14 若 sklearn 尚未发布对应 wheel 时可用 3.12)
conda create -n ml python=3.12
conda activate ml
# 安装 scikit-learn 及配套科学计算库
conda install scikit-learn numpy pandas matplotlib
# 或使用 pip
pip install scikit-learn numpy pandas matplotlib
版本检查脚本(每个脚本开头建议运行一次,确认 API 一致性):
# version_check.py
import sklearn
import numpy as np
import pandas as pd
print("scikit-learn 版本:", sklearn.__version__) # 期望 1.5.x 或 1.6.x
print("numpy 版本:", np.__version__)
print("pandas 版本:", pd.__version__)
# 检查关键 API 是否存在(1.4+ 新增的 RMSE 函数)
from sklearn.metrics import root_mean_squared_error
print("root_mean_squared_error 可用:", callable(root_mean_squared_error))
提示:
root_mean_squared_error是 1.4 版本引入的官方 RMSE 函数。旧写法
mean_squared_error(y_true, y_pred, squared=False)已弃用,本教程一律使用新函数。
3. sklearn 统一接口与核心概念
scikit-learn 最强大的设计之一:几乎所有模型都遵循同一套接口。你学会一个,就学会了所有。
3.1 三类核心对象
| 对象类型 | 关键方法 | 作用 |
|---|---|---|
| Estimator(估计器/模型) | fit(X, y) | 从数据学习参数 |
| Predictor(预测器) | predict(X) | 输出预测值 |
| Transformer(变换器) | fit(X)、transform(X)、fit_transform(X) | 数据预处理(标准化、降维等) |
几乎所有模型(LinearRegression、RandomForestRegressor、SVC、PCA、KMeans…)都是 Estimator;有监督模型同时是 Predictor;预处理类和 PCA 是 Transformer。
3.2 数据形状约定(极其重要)
X(特征矩阵)必须是二维,形状(n_samples, n_features):n_samples:样本数(材料个数)n_features:特征数(描述符个数)
y(标签/目标)通常是一维,形状(n_samples,)。- 特征矩阵必须是数值型,类别型变量需先编码(如
OneHotEncoder)。
# shape_demo.py
import numpy as np
from sklearn.linear_model import LinearRegression
# 3 个材料样本,每个有 4 个描述符 -> (3, 4)
X = np.array([
[1.9, -1.2, 0.5, 3.0], # 材料 A 的 4 个描述符
[1.7, -1.5, 0.8, 2.8], # 材料 B
[2.1, -0.9, 0.4, 3.3], # 材料 C
])
y = np.array([0.45, 0.52, 0.38]) # 对应的催化活性 (3,)
model = LinearRegression()
model.fit(X, y) # fit 学习参数
y_pred = model.predict(X) # predict 输出预测
print("X 形状:", X.shape, " y 形状:", y.shape)
print("预测值:", y_pred)
常见错误:把单个样本写成 np.array([1.9, -1.2, 0.5, 3.0])(一维),会导致
ValueError: Expected 2D array。修正方法是 X.reshape(1, -1)。
4. 数据准备与预处理
4.1 划分训练集与测试集:train_test_split
# split_demo.py
import numpy as np
from sklearn.model_selection import train_test_split
# 构造 100 个样本、5 个描述符
rng = np.random.default_rng(42)
X = rng.normal(size=(100, 5))
y = rng.normal(size=100)
# test_size=0.2 -> 80% 训练 / 20% 测试;random_state 保证可复现
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
print(X_train.shape, X_test.shape) # (80, 5) (20, 5)
要点:
- 始终设置
random_state,否则每次运行划分结果不同,实验不可复现。 - 若标签是类别且不平衡,可用
stratify=y让训练/测试集保持相同类别比例。
4.2 标准化:StandardScaler 与 MinMaxScaler
描述符量纲差异巨大(如电负性 ~1-3,而吸附能可能 ~ -5 到 5 eV),必须标准化,否则量纲大的特征会主导模型。
# scaler_demo.py
import numpy as np
from sklearn.preprocessing import StandardScaler, MinMaxScaler
X = np.array([
[1.9, -1.2, 0.5],
[1.7, -1.5, 0.8],
[2.1, -0.9, 0.4],
[1.8, -1.1, 0.6],
])
# StandardScaler:减均值、除标准差 -> 每列均值 0、标准差 1
std_scaler = StandardScaler()
X_std = std_scaler.fit_transform(X)
# MinMaxScaler:缩放到 [0, 1] 区间
mm_scaler = MinMaxScaler()
X_mm = mm_scaler.fit_transform(X)
print("标准化后:\n", X_std)
print("缩放后:\n", X_mm)
print("每列均值(应为0):", X_std.mean(axis=0))
选择建议:
- StandardScaler:默认首选,适合大多数线性模型、SVM、PCA(这些模型对"均值 0 方差 1"敏感)。
- MinMaxScaler:当你明确需要输出在固定区间(如后续可视化),或数据有固定上下界时使用。
- 树模型(随机森林、梯度提升)不要求标准化,因为它们是按阈值切分,对单调缩放不敏感。但标准化了也没坏处。
4.3 处理缺失值:SimpleImputer
实验或 DFT 数据常缺失(某个吸附能没算出来)。
# imputer_demo.py
import numpy as np
from sklearn.impute import SimpleImputer
X = np.array([
[1.9, np.nan, 0.5],
[1.7, -1.5, np.nan],
[2.1, -0.9, 0.4],
[1.8, -1.1, 0.6],
])
# strategy: 'mean' 均值 / 'median' 中位数 / 'most_frequent' 众数 / 'constant' 常数
imputer = SimpleImputer(strategy='mean')
X_filled = imputer.fit_transform(X)
print("填补后:\n", X_filled)
要点:SimpleImputer 的"均值"必须只在训练集上拟合(见第 13 节数据泄漏),配合 Pipeline 可以自动保证这一点。
4.4 管线:Pipeline
Pipeline 把"预处理 + 模型"串成一条流水线,保证预处理只对训练集拟合、对测试集只变换,避免手动犯错。
# pipeline_demo.py
import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge
from sklearn.model_selection import cross_val_score
rng = np.random.default_rng(0)
X = rng.normal(size=(200, 5))
y = 2 * X[:, 0] - 1.5 * X[:, 1] + rng.normal(0, 0.1, 200)
# Pipeline 步骤列表:('名称', 对象)
pipe = Pipeline([
('scaler', StandardScaler()), # 第 1 步:标准化
('model', Ridge(alpha=1.0)), # 第 2 步:岭回归
])
# 一步完成:先标准化,再在标准化后的数据上做交叉验证
scores = cross_val_score(pipe, X, y, cv=5, scoring='r2')
print("5 折交叉验证 R² 均值:", scores.mean().round(3))
4.5 列变换:ColumnTransformer
当不同列需要不同预处理时使用。例如你的数据里既有连续描述符(电负性、吸附能),又有类别变量(晶相:fcc/bcc/hcp)。
# coltransformer_demo.py
import numpy as np
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.pipeline import Pipeline
from sklearn.ensemble import RandomForestRegressor
# 构造 DataFrame:连续列 + 类别列
df = pd.DataFrame({
'电负性': [1.9, 1.7, 2.1, 1.8, 2.0],
'd带中心': [-1.2, -1.5, -0.9, -1.1, -1.3],
'晶相': ['fcc', 'bcc', 'hcp', 'fcc', 'bcc'], # 类别列
})
y = np.array([0.45, 0.52, 0.38, 0.47, 0.41])
# 对连续列标准化,对类别列做独热编码
preprocessor = ColumnTransformer([
('num', StandardScaler(), ['电负性', 'd带中心']),
('cat', OneHotEncoder(), ['晶相']),
])
pipe = Pipeline([
('prep', preprocessor),
('model', RandomForestRegressor(n_estimators=100, random_state=42)),
])
pipe.fit(df, y)
print("训练完成。转换后的特征维度:", pipe.named_steps['prep'].fit_transform(df).shape)
5. 监督学习 — 回归
回归用于预测连续数值(过电位、吸附能、活性、选择性、形成能等)。
5.1 线性回归:LinearRegression
最简单、最可解释的基线模型。
# regression_baseline.py
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import r2_score, mean_absolute_error, root_mean_squared_error
rng = np.random.default_rng(42)
n = 300
# 3 个描述符:电负性、d带中心、吸附能
X = rng.normal(size=(n, 3))
# 真值关系:活性 = 0.8*电负性 - 0.5*d带中心 + 0.3*吸附能 + 噪声
y = 0.8 * X[:, 0] - 0.5 * X[:, 1] + 0.3 * X[:, 2] + rng.normal(0, 0.2, n)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
lr = LinearRegression()
lr.fit(X_train, y_train)
y_pred = lr.predict(X_test)
print("系数(特征权重):", lr.coef_.round(3))
print("截距:", round(lr.intercept_, 3))
print("R²:", r2_score(y_test, y_pred).round(3))
print("MAE:", mean_absolute_error(y_test, y_pred).round(3))
print("RMSE:", root_mean_squared_error(y_test, y_pred).round(3))
线性回归适用场景:样本少、特征少、想快速得到一个可解释的基线。系数正负号直接告诉你"哪个描述符促进/抑制活性"。
5.2 正则化回归:Ridge(岭)与 Lasso
当特征很多、样本很少,或特征间高度相关(共线性,如电负性和 d 带中心常强相关)时,普通线性回归会过拟合、系数不稳定。正则化通过惩罚大的系数来缓解。
# regularized_regression.py
import numpy as np
from sklearn.linear_model import LinearRegression, Ridge, Lasso
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
rng = np.random.default_rng(7)
n = 80 # 样本少
p = 30 # 特征多(n < p 的情况普通最小二乘会失败或过拟合)
X = rng.normal(size=(n, p))
# 只有前 5 个描述符真正有贡献,其余 25 个是噪声
true_coef = np.zeros(p)
true_coef[:5] = [1.0, -0.8, 0.5, 0.3, -0.6]
y = X @ true_coef + rng.normal(0, 0.5, n)
# Ridge:L2 正则化,alpha 越大惩罚越强
ridge = Pipeline([('s', StandardScaler()), ('m', Ridge(alpha=1.0))])
# Lasso:L1 正则化,能把不重要特征的系数压到 0(自动特征选择)
lasso = Pipeline([('s', StandardScaler()), ('m', Lasso(alpha=0.1))])
for name, model in [('Ridge', ridge), ('Lasso', lasso)]:
s = cross_val_score(model, X, y, cv=5, scoring='r2')
print(f"{name} 交叉验证 R² 均值: {s.mean():.3f}")
# Lasso 的稀疏性
lasso.fit(X, y)
n_nonzero = np.sum(lasso.named_steps['m'].coef_ != 0)
print("Lasso 非零系数个数(应接近 5):", n_nonzero)
选择建议:
- Ridge:特征多但都可能有微弱贡献时,优先 Ridge(稳定、几乎不会得到 0 系数)。
- Lasso:想要自动筛选出少数关键描述符时用 Lasso(稀疏解)。
- 两者之间的折中是 ElasticNet(L1+L2 混合)。
5.3 随机森林回归:RandomForestRegressor
非线性、对噪声和缺失值鲁棒、能给出特征重要性,是科研里最常用的"黑盒强基线"。
# random_forest_reg.py
import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import train_test_split
from sklearn.metrics import r2_score, mean_absolute_error, root_mean_squared_error
rng = np.random.default_rng(1)
n = 500
X = rng.normal(size=(n, 6))
# 非线性关系:活性与 d带中心(第1列)呈抛物线(火山型曲线)
d_center = X[:, 1]
y = 3.0 - 2.0 * (d_center + 0.5) ** 2 + 0.4 * X[:, 0] + rng.normal(0, 0.3, n)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
rf = RandomForestRegressor(n_estimators=200, random_state=42)
rf.fit(X_train, y_train)
y_pred = rf.predict(X_test)
print("R²:", r2_score(y_test, y_pred).round(3))
print("MAE:", mean_absolute_error(y_test, y_pred).round(3))
print("RMSE:", root_mean_squared_error(y_test, y_pred).round(3))
print("特征重要性:", rf.feature_importances_.round(3))
关键参数:
n_estimators:树的棵数,通常越大越稳(如 100~500),再大收益递减。max_depth:树的深度,None 表示不限制;样本少时限制它能防过拟合。random_state:必须固定以保证可复现。
5.4 支持向量回归:SVR
适合中小样本、特征不多的高维非线性问题;对小数据常优于随机森林,但需要仔细调参和标准化。
# svr_demo.py
import numpy as np
from sklearn.svm import SVR
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.model_selection import train_test_split
from sklearn.metrics import r2_score
rng = np.random.default_rng(3)
n = 120
X = rng.normal(size=(n, 4))
y = np.sin(X[:, 0]) + 0.3 * X[:, 1] ** 2 + rng.normal(0, 0.1, n)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)
# SVR 对特征尺度敏感,必须先标准化
svr = Pipeline([
('scaler', StandardScaler()),
('svr', SVR(kernel='rbf', C=10.0, epsilon=0.1)),
])
svr.fit(X_train, y_train)
y_pred = svr.predict(X_test)
print("SVR R²:", r2_score(y_test, y_pred).round(3))
关键参数:
kernel:'rbf'(默认,最常用)、'linear'、'poly'。C:正则化强度的倒数,越大越倾向拟合训练集(易过拟合)。epsilon:不敏感区间宽度,控制模型对噪声的容忍度。
5.5 回归模型选择速查
| 情况 | 推荐模型 |
|---|---|
| 样本少、要可解释基线 | LinearRegression |
| 特征多/共线性/样本少 | Ridge(稳)或 Lasso(稀疏) |
| 非线性、样本中等偏多、要特征重要性 | RandomForestRegressor |
| 非线性、样本少、特征不多、愿意调参 | SVR |
6. 监督学习 — 分类
分类用于预测离散类别,例如"高活性 / 低活性"、"稳定 / 不稳定"、"选择性好 / 差"。
6.1 逻辑回归:LogisticRegression
名字带"回归"但用于二分类。可解释性好,能输出概率。
# classification_demo.py
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score, confusion_matrix, classification_report
rng = np.random.default_rng(5)
n = 400
X = rng.normal(size=(n, 3))
# 用描述符的线性组合决定类别:活性高(1) / 低(0)
logit = 1.5 * X[:, 0] - 1.2 * X[:, 1] + 0.5 * X[:, 2]
p = 1 / (1 + np.exp(-logit)) # sigmoid 转概率
y = (rng.random(n) < p).astype(int) # 按概率采样得到 0/1 标签
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42, stratify=y)
clf = LogisticRegression(max_iter=1000) # 提高迭代次数避免不收敛警告
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
y_prob = clf.predict_proba(X_test)[:, 1] # 属于"高活性"的概率
print("准确率:", accuracy_score(y_test, y_pred).round(3))
print("混淆矩阵:\n", confusion_matrix(y_test, y_pred))
print(classification_report(y_test, y_pred, target_names=['低活性', '高活性']))
6.2 随机森林分类与支持向量机分类
# rf_svc_classify.py
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
rng = np.random.default_rng(11)
n = 600
X = rng.normal(size=(n, 5))
# 非线性分类边界:两个描述符的平方和决定类别
y = ((X[:, 0] ** 2 + X[:, 1] ** 2) > 1.0).astype(int)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0, stratify=y)
# 随机森林(无需标准化)
rf = RandomForestClassifier(n_estimators=200, random_state=42)
rf.fit(X_train, y_train)
print("RandomForestClassifier 准确率:", accuracy_score(y_test, rf.predict(X_test)).round(3))
# SVC(对尺度敏感,需要标准化;probability=True 才能输出概率)
svc = Pipeline([
('scaler', StandardScaler()),
('svc', SVC(kernel='rbf', C=10.0, probability=True)),
])
svc.fit(X_train, y_train)
print("SVC 准确率:", accuracy_score(y_test, svc.predict(X_test)).round(3))
7. 无监督学习
无监督学习不需要标签 y,用于探索数据结构。
7.1 KMeans 聚类:KMeans
把材料按描述符相似度分成若干组("材料家族")。
# kmeans_demo.py
import numpy as np
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler
rng = np.random.default_rng(0)
# 构造 3 类材料,每类 50 个,围绕不同中心分布
centers = np.array([[-2, -2], [0, 2], [2, -1]])
X = np.vstack([rng.normal(c, 0.5, (50, 2)) for c in centers])
X = StandardScaler().fit_transform(X)
kmeans = KMeans(n_clusters=3, n_init=10, random_state=42)
kmeans.fit(X)
labels = kmeans.labels_ # 每个样本所属簇
centers_found = kmeans.cluster_centers_
print("各簇样本数:", np.bincount(labels))
print("簇中心:\n", centers_found)
print("惯性(簇内平方和,越小越紧凑):", round(kmeans.inertia_, 3))
要点:
n_init:用不同初始中心重复运行的次数,取最优;设n_init=10可避免旧版本的警告。- 簇数
n_clusters需要用肘部法(画惯性 vs. 簇数曲线)或领域知识确定。 - KMeans 对特征尺度敏感,聚类前必须标准化。
7.2 PCA 降维:PCA
把高维描述符压缩到 2~3 维,用于可视化材料空间,也常作为下游模型(如聚类)的预处理。
# pca_demo.py
import numpy as np
from sklearn.decomposition import PCA
from sklearn.preprocessing import StandardScaler
rng = np.random.default_rng(1)
n = 200
# 20 维描述符,但真实信息集中在少数几个方向
X = rng.normal(size=(n, 20))
X[:, 0] = 5 * X[:, 0] # 第 1 维方差放大,主导主成分
X[:, 1] = 3 * X[:, 1]
X = StandardScaler().fit_transform(X)
# 降到 2 维
pca = PCA(n_components=2)
X_pca = pca.fit_transform(X)
print("原始维度:", X.shape, "-> 降维后:", X_pca.shape)
print("各主成分解释方差比例:", pca.explained_variance_ratio_.round(3))
print("累计解释方差比例:", pca.explained_variance_ratio_.cumsum().round(3))
要点:
n_components可以是整数(保留几个成分)或0.0~1.0的浮点数(保留累计解释方差达到该比例所需的最少成分),如n_components=0.95。- PCA 前必须标准化,否则方差大的特征会"霸占"主成分。
- 主成分是特征的线性组合,物理含义需要结合
pca.components_(各原始特征在每个主成分上的载荷)来解读。
8. 模型评估
8.1 交叉验证:cross_val_score
单次 train/test 划分的结果受"运气"影响大。交叉验证把数据分成 k 份,轮流用 k-1 份训练、1 份测试,得到更稳健的评估。
# crossval_demo.py
import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import cross_val_score
rng = np.random.default_rng(2)
X = rng.normal(size=(300, 8))
y = X[:, 0] - 0.5 * X[:, 1] ** 2 + rng.normal(0, 0.5, 300)
model = RandomForestRegressor(n_estimators=100, random_state=42)
# 不同评分指标
for scoring in ['r2', 'neg_mean_absolute_error', 'neg_mean_squared_error']:
scores = cross_val_score(model, X, y, cv=5, scoring=scoring)
print(f"{scoring}: 均值 {scores.mean():.3f}")
# 回归负号说明:sklearn 约定"越大越好",所以 MAE/MSE 取负值
scores_mae = cross_val_score(model, X, y, cv=5, scoring='neg_mean_absolute_error')
print("真实 MAE:", (-scores_mae.mean()).round(3))
8.2 回归评估指标:R²、MAE、RMSE
# regression_metrics.py
import numpy as np
from sklearn.metrics import r2_score, mean_absolute_error, mean_squared_error, root_mean_squared_error
y_true = np.array([0.5, 0.3, 0.8, 0.1, 0.6])
y_pred = np.array([0.55, 0.28, 0.75, 0.15, 0.58])
print("R² (越接近1越好):", r2_score(y_true, y_pred).round(3))
print("MAE (平均绝对误差):", mean_absolute_error(y_true, y_pred).round(3))
print("MSE (均方误差):", mean_squared_error(y_true, y_pred).round(3))
print("RMSE (均方根误差):", root_mean_squared_error(y_true, y_pred).round(3))
指标含义(回归):
- R²:决定系数,1 为完美,0 表示和"用均值预测"一样差,可为负(比均值还差)。
- MAE:平均绝对误差,与目标同量纲,直观、对异常值不敏感。
- RMSE:均方根误差,对大误差(异常值)更敏感。
8.3 分类评估指标:准确率、混淆矩阵
# classification_metrics.py
import numpy as np
from sklearn.metrics import (
accuracy_score, confusion_matrix, classification_report, ConfusionMatrixDisplay
)
y_true = np.array([0, 1, 0, 1, 1, 0, 1, 0, 1, 0])
y_pred = np.array([0, 1, 0, 0, 1, 0, 1, 1, 1, 0])
print("准确率:", accuracy_score(y_true, y_pred))
print("混淆矩阵:\n", confusion_matrix(y_true, y_pred))
print(classification_report(y_true, y_pred))
# 可视化混淆矩阵(需 matplotlib)
import matplotlib.pyplot as plt
ConfusionMatrixDisplay.from_predictions(y_true, y_pred)
plt.show()
要点:当类别不平衡时(如"高活性"材料只占 5%),准确率会误导人——全部预测为"低活性"也能拿到 95% 准确率。此时应看混淆矩阵、精确率/召回率/F1(classification_report 已给出)。
8.4 学习曲线与过拟合/欠拟合
# learning_curve_demo.py
import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import learning_curve
rng = np.random.default_rng(0)
X = rng.normal(size=(500, 3))
y = X[:, 0] - 0.5 * X[:, 1] + rng.normal(0, 0.3, 500)
train_sizes, train_scores, test_scores = learning_curve(
LinearRegression(), X, y, cv=5,
train_sizes=np.linspace(0.1, 1.0, 10),
scoring='r2'
)
train_mean = train_scores.mean(axis=1)
test_mean = test_scores.mean(axis=1)
plt.plot(train_sizes, train_mean, 'o-', label='训练集 R²')
plt.plot(train_sizes, test_mean, 'o-', label='验证集 R²')
plt.xlabel('训练样本数'); plt.ylabel('R²'); plt.legend(); plt.title('学习曲线')
plt.show()
如何判断:
- 过拟合:训练得分高、验证得分低(两条线之间大缺口)。
- 欠拟合:训练和验证得分都低(两条线贴在一起但都不高)。
- 高偏差(欠拟合):加数据没用,应换更强模型或加特征。
- 高方差(过拟合):加数据、加正则化、简化模型有用。
9. 超参数调优:GridSearchCV
超参数(如 alpha、n_estimators、C、max_depth)是在训练前设定的,不能靠 fit 学出来。GridSearchCV 在给定网格里穷举并用交叉验证选出最优组合。
# gridsearch_demo.py
import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import GridSearchCV, train_test_split
from sklearn.metrics import r2_score
rng = np.random.default_rng(9)
n = 400
X = rng.normal(size=(n, 6))
y = 3.0 - 2.0 * (X[:, 1] + 0.5) ** 2 + 0.4 * X[:, 0] + rng.normal(0, 0.3, n)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
param_grid = {
'n_estimators': [50, 100, 200],
'max_depth': [None, 5, 10],
}
grid = GridSearchCV(
RandomForestRegressor(random_state=42),
param_grid,
cv=5,
scoring='neg_mean_squared_error',
n_jobs=-1, # 并行使用所有 CPU 核
)
grid.fit(X_train, y_train)
print("最优参数:", grid.best_params_)
print("最优交叉验证得分(负MSE):", round(grid.best_score_, 3))
best_model = grid.best_estimator_
y_pred = best_model.predict(X_test)
print("测试集 R²:", r2_score(y_test, y_pred).round(3))
要点:
grid.best_estimator_已经在整个训练集上重新拟合过了,可直接用于预测。- 网格要合理,不要过大(组合数 = 各参数取值数的乘积,
3×3=9组合 × 5 折 = 45 次拟合,再乘 n_jobs 加速)。 - 更高级的替代:
RandomizedSearchCV(随机抽样,适合大网格)。
10. 特征重要性与特征选择
10.1 特征重要性
不同模型给出"重要性"的方式不同:
# feature_importance.py
import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.linear_model import Ridge
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
rng = np.random.default_rng(4)
X = rng.normal(size=(500, 5))
# 只有第 0、2、4 个描述符有贡献
y = 2 * X[:, 0] - 1 * X[:, 2] + 0.5 * X[:, 4] + rng.normal(0, 0.2, 500)
feature_names = ['电负性', 'd带中心', '吸附能', '配位数', '形成能']
# 1) 树模型:feature_importances_(相对重要性,和为 1)
rf = RandomForestRegressor(n_estimators=200, random_state=42).fit(X, y)
print("随机森林重要性:")
for name, imp in zip(feature_names, rf.feature_importances_):
print(f" {name}: {imp:.3f}")
# 2) 线性模型:标准化后的系数绝对值(正负号表示方向)
pipe = Pipeline([('s', StandardScaler()), ('m', Ridge(alpha=1.0))]).fit(X, y)
print("岭回归标准化系数:")
for name, c in zip(feature_names, pipe.named_steps['m'].coef_):
print(f" {name}: {c:.3f}")
10.2 特征选择:SelectKBest
用统计检验筛选出与目标相关性最强的 k 个特征,是降维和去噪的常用手段。
# feature_selection.py
import numpy as np
from sklearn.feature_selection import SelectKBest, f_regression
rng = np.random.default_rng(6)
X = rng.normal(size=(300, 10)) # 10 个描述符
y = 2 * X[:, 2] - 1.5 * X[:, 7] + rng.normal(0, 0.3, 300) # 只有第 2、7 个有用
# f_regression:单变量线性相关检验(F 值),保留前 3 个
selector = SelectKBest(score_func=f_regression, k=3)
X_new = selector.fit_transform(X, y)
print("原始特征数:", X.shape[1], "-> 保留特征数:", X_new.shape[1])
print("被保留特征的索引:", selector.get_support(indices=True)) # 应为 [2, 7, 某噪声]
print("各特征 F 分数:", selector.scores_.round(2))
注意:SelectKBest 是单变量筛选,只考虑每个特征与目标的单独关系,不考虑特征间交互。树模型的 feature_importances_ 能捕捉交互,但两者都应视为"参考"而非"真理"。
11. 模型保存与加载:joblib / pickle
训练好的模型(尤其是含 Pipeline 的)应保存下来,避免每次重复训练。
# save_load.py
import numpy as np
import joblib
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.ensemble import RandomForestRegressor
rng = np.random.default_rng(0)
X = rng.normal(size=(300, 5))
y = X[:, 0] - X[:, 1] ** 2 + rng.normal(0, 0.3, 300)
pipe = Pipeline([
('scaler', StandardScaler()),
('model', RandomForestRegressor(n_estimators=200, random_state=42)),
])
pipe.fit(X, y)
# 保存(joblib 对含 numpy 数组的 sklearn 对象更高效)
joblib.dump(pipe, 'catalyst_model.joblib')
# 加载
loaded = joblib.load('catalyst_model.joblib')
# 验证加载后的模型可正常预测
new_sample = np.array([[1.0, -0.5, 0.3, 0.2, -0.1]])
print("预测活性:", loaded.predict(new_sample))
要点:
- 推荐用
joblib而非pickle:对大型 numpy 数组更高效,且是 sklearn 官方推荐。 - 也可以用标准库
pickle(import pickle; pickle.dump(pipe, open('m.pkl','wb'))),但加载时需确保环境(sklearn 版本)一致。 - 版本一致性:保存与加载时 sklearn 版本差异过大可能导致加载失败,记录下训练时的版本号。
12. 科研实战示例:描述符 → 催化活性
下面是一个完整端到端示例,模拟真实的"材料描述符 → 催化活性"研究流程:构造数据、预处理、建模评估、降维聚类。
# catalyst_workflow.py
# 完整工作流:材料描述符 -> 催化活性预测 + 材料空间降维聚类
import numpy as np
import joblib
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.impute import SimpleImputer
from sklearn.ensemble import RandomForestRegressor
from sklearn.decomposition import PCA
from sklearn.cluster import KMeans
from sklearn.metrics import r2_score, mean_absolute_error, root_mean_squared_error
# ---------- 1. 构造模拟数据 ----------
rng = np.random.default_rng(42)
n_materials = 600
# 6 个常见催化描述符
descriptor_names = ['电负性', 'd带中心(eV)', 'CO吸附能(eV)', '配位数', '形成能(eV)', '晶格常数(Å)']
X = rng.normal(size=(n_materials, 6))
# 给每个描述符一个物理量纲(模拟真实量级)
X[:, 0] = X[:, 0] * 0.3 + 1.8 # 电负性 ~ 1.8
X[:, 1] = X[:, 1] * 0.8 - 1.5 # d带中心 ~ -1.5 eV
X[:, 2] = X[:, 2] * 0.6 - 0.3 # CO吸附能
X[:, 3] = rng.integers(4, 13, n_materials) # 配位数 4~12
X[:, 4] = X[:, 4] * 0.5 - 2.0 # 形成能
X[:, 5] = X[:, 5] * 0.2 + 3.8 # 晶格常数
# 人为制造"火山型"关系:活性在 d带中心 = -1.0 eV 附近最高(Sabatier 原理)
d_center = X[:, 1]
activity = 5.0 - 4.0 * (d_center + 1.0) ** 2 + 0.3 * X[:, 0] + rng.normal(0, 0.4, n_materials)
# 人为制造少量缺失值(模拟 DFT 没算完的吸附能)
missing_mask = rng.random(n_materials) < 0.08
X[missing_mask, 2] = np.nan
# ---------- 2. 划分数据 ----------
X_train, X_test, y_train, y_test = train_test_split(
X, activity, test_size=0.2, random_state=42
)
# ---------- 3. 构建管线:填补缺失 -> 标准化 -> 随机森林 ----------
pipe = Pipeline([
('imputer', SimpleImputer(strategy='mean')), # 先填补缺失
('scaler', StandardScaler()), # 再标准化
('model', RandomForestRegressor(random_state=42)),
])
# ---------- 4. 超参数调优 ----------
param_grid = {
'model__n_estimators': [100, 200],
'model__max_depth': [None, 8],
}
grid = GridSearchCV(pipe, param_grid, cv=5,
scoring='neg_mean_squared_error', n_jobs=-1)
grid.fit(X_train, y_train)
# ---------- 5. 评估 ----------
best = grid.best_estimator_
y_pred = best.predict(X_test)
print("=== 催化活性预测结果 ===")
print("最优参数:", grid.best_params_)
print("R²:", r2_score(y_test, y_pred).round(3))
print("MAE:", mean_absolute_error(y_test, y_pred).round(3))
print("RMSE:", root_mean_squared_error(y_test, y_pred).round(3))
# ---------- 6. 特征重要性 ----------
rf_model = best.named_steps['model']
print("\n=== 描述符重要性 ===")
for name, imp in zip(descriptor_names, rf_model.feature_importances_):
print(f" {name}: {imp:.3f}")
# ---------- 7. 材料空间降维与聚类 ----------
# 先在完整数据集上填补 + 标准化(注意:这里用于探索,不涉及泄漏)
X_clean = SimpleImputer(strategy='mean').fit_transform(X)
X_scaled = StandardScaler().fit_transform(X_clean)
pca = PCA(n_components=2)
X_pca = pca.fit_transform(X_scaled)
print("\n=== PCA 降维 ===")
print("前两个主成分解释方差比例:", pca.explained_variance_ratio_.round(3))
kmeans = KMeans(n_clusters=3, n_init=10, random_state=42)
clusters = kmeans.fit_predict(X_scaled)
print("\n=== KMeans 聚类(3 类材料家族)===")
print("各类样本数:", np.bincount(clusters))
# ---------- 8. 保存最终模型 ----------
joblib.dump(best, 'catalyst_activity_model.joblib')
print("\n模型已保存至 catalyst_activity_model.joblib")
运行这个脚本,你能看到:随机森林通常能得到较高的 R²(因为它能捕捉火山型非线性关系),特征重要性会正确地指出 d带中心 是最重要的描述符(因为它与活性有强非线性关系),PCA 把 6 维压缩到 2 维,KMeans 把材料分成 3 个家族。
13. 常见坑(务必避开)
13.1 数据泄漏(Data Leakage)——最严重
先做任何预处理再切分数据是科研代码里最常见的错误。若先对全量数据标准化再切分,测试集的信息(均值/方差)会"泄漏"进训练阶段,导致评估结果虚高。
# 错误示范:标准化用到了全量数据(含测试集)-> 数据泄漏
scaler = StandardScaler()
X_all_scaled = scaler.fit_transform(X_all) # 泄漏!
X_train, X_test = train_test_split(X_all_scaled, y)
# 正确示范:先切分,scaler 只 fit 训练集
X_train, X_test, y_train, y_test = train_test_split(X, y)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train) # 只在训练集 fit
X_test_scaled = scaler.transform(X_test) # 测试集只 transform
# 最稳妥:用 Pipeline 让 sklearn 自动管理上述顺序
规则:任何 fit(包括 StandardScaler.fit、SimpleImputer.fit、PCA.fit)都只能接触训练集。用 Pipeline 是自动防泄漏的最佳手段。
13.2 标准化顺序
- 树模型(随机森林、梯度提升)不需要标准化。
- 线性模型、逻辑回归、SVM、SVR、KMeans、PCA、神经网络都需要标准化。
- 若用
Pipeline,标准化自动发生在交叉验证的每一折内部(正确);手动在循环外标准化则可能泄漏。
13.3 小样本过拟合
计算化学数据常常只有几十个样本。此时:
- 别用太多特征(特征数接近甚至超过样本数时,线性模型会过拟合)。
- 优先用正则化(Ridge/Lasso)或简单模型,而不是复杂模型。
- 交叉验证用
LeaveOneOut或KFold(n_splits=5),并报告均值 ± 标准差而非单次结果。 - 一个评估结果如果没有交叉验证或独立测试集,就不要写进论文。
13.4 随机种子
任何带随机性的模型(随机森林、KMeans)和划分(train_test_split、KFold)都要设 random_state,否则结果不可复现、审稿人或你自己无法复现实验。
13.5 类别不平衡
分类问题里若正负样本严重失衡,准确率无意义。要用 F1、ROC-AUC、精确率/召回率,并考虑 class_weight='balanced'(在 LogisticRegression、RandomForestClassifier、SVC 中都支持)。
13.6 数据形状错误
X 必须是二维 (n_samples, n_features)。单个样本预测时要 reshape(1, -1)。用 sklearn.utils.validation 报的 Expected 2D array 错误基本都源于此。
14. 延伸资源
- 官方文档(首选):https://scikit-learn.org/stable/ —— 每个类的参数说明和示例最权威。
- 用户指南(User Guide):https://scikit-learn.org/stable/user_guide.html —— 系统讲解各算法原理与预处理方法。
- 《Hands-On Machine Learning with Scikit-Learn, Keras & TensorFlow》(Aurélien Géron):机器学习的经典实战书,代码质量高。
- 《Pattern Recognition and Machine Learning》(Bishop):理论深度,适合想懂原理时查。
- scikit-learn 官方示例:
sklearn.datasets和 examples 目录提供了大量可复现案例。
最后提醒:机器学习在催化里是辅助工具,不是替代物理。模型的预测永远要回到 DFT 或实验去验证,尤其是外推到训练集分布之外的材料时,更要警惕"看起来很好但实际不可靠"的过拟合结果。
评论交流
欢迎留下你的想法