Skip to content

线性回归

线性回归是基础的监督学习算法,用于基于一个或多个预测变量(自变量 Independent variables)预测一个连续目标变量(Target variable)。它假设预测变量与目标变量之间存在线性关系(Linear relationship)。

在数学上,这种关系建模为:

对于简单线性回归(一个预测变量):

Y = β₀ + β₁X + ε

对于多元线性回归(多个预测变量):

Y = β₀ + β₁X₁ + β₂X₂ + ... + βₚXₚ + ε

其中:

  • Y:因变量(Dependent variable,目标变量),是我们想要预测的变量。
  • X、X₁、…、Xₚ:自变量(Independent variables,预测变量 Predictors)。
  • β₀:截距(Intercept,当所有 X 都为 0 时 Y 的值)。
  • β₁、…、βₚ:系数(Coefficients,斜率),表示在保持其他预测变量不变的情况下,相应 X 变化一个单位时 Y 的变化量。
  • ε:误差项(Error term),表示实际值与预测值之间的差(残差 Residuals)。

线性回归的目标是找到系数(β₀, β₁, …)的最优值,以最小化误差,通常使用平方误差和(Sum of squared errors - SSE)或残差平方和(Residual Sum of Squares - RSS)来衡量。

  • 正线性关系(Positive Linear Relationship): 随着预测变量的增加,目标变量也倾向于增加(正斜率系数)。

  • 负线性关系(Negative Linear Relationship): 随着预测变量的增加,目标变量倾向于减少(负斜率系数)。

基于预测变量的数量:

  • 简单线性回归(Simple Linear Regression - SLR): 使用单个预测变量(X)预测目标变量(Y)。
  • 多元线性回归(Multiple Linear Regression - MLR): 使用两个或多个预测变量(X₁、X₂、…)预测目标变量(Y)。

SLR 建模一个自变量和一个因变量之间的线性关系。它找到通过数据点的最佳拟合直线。

Python 实现(概念性 - 手动计算)

Section titled “Python 实现(概念性 - 手动计算)”

虽然通常使用像 Scikit-learn 这样的库,但理解底层计算有所帮助。此示例演示了使用最小二乘法公式手动估计系数。

import numpy as np
import matplotlib.pyplot as plt
def estimate_coefficients(x, y):
n = np.size(x)
# Calculate means
mean_x, mean_y = np.mean(x), np.mean(y)
# Calculate cross-deviation and deviation about x
SS_xy = np.sum(y * x) - n * mean_y * mean_x
SS_xx = np.sum(x * x) - n * mean_x * mean_x
# Calculate regression coefficients (slope b1 and intercept b0)
b1 = SS_xy / SS_xx
b0 = mean_y - b1 * mean_x
return (b0, b1)
def plot_regression_line(x, y, b):
plt.figure(figsize=(8, 6))
# Plot actual points as scatter plot
plt.scatter(x, y, color="purple", marker="o", s=30, label='Data Points') # 数据点
# Predict response vector
y_pred = b[0] + b[1] * x
# Plot the regression line
plt.plot(x, y_pred, color="green", label='Regression Line') # 回归线
plt.xlabel('Independent Variable (X)')
plt.ylabel('Dependent Variable (Y)')
plt.title('Simple Linear Regression Example')
plt.legend()
plt.grid(True)
plt.show()
# --- Main execution ---
# Sample data
x = np.array([0, 1, 2, 3, 4, 5, 6, 7, 8, 9])
y = np.array([1, 3, 2, 5, 7, 8, 8, 9, 10, 12]) # 示例数据
# Estimate coefficients
b = estimate_coefficients(x, y)
print(f"Estimated coefficients:\nIntercept (b0) = {b[0]:.4f} \nSlope (b1) = {b[1]:.4f}")
# Plot the regression line
plot_regression_line(x, y, b)
Estimated coefficients:
Intercept (b0) = 1.2364
Slope (b1) = 1.1697

此代码计算样本数据的截距和斜率,并将生成的回归线绘制在数据点上。(该图会显示散点,绿线代表最佳线性拟合。)

使用 Scikit-learn 是标准且更实用的方法。这里我们使用 Diabetes 数据集(仅加载一个特征用于 SLR)。

import matplotlib.pyplot as plt
import numpy as np
from sklearn import datasets, linear_model
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error, r2_score
# Load the diabetes dataset
diabetes = datasets.load_diabetes()
# Use only one feature (e.g., the third feature, BMI)
X = diabetes.data[:, np.newaxis, 2] # np.newaxis keeps it a 2D array
y = diabetes.target
# Split the data into training/testing sets
# Using only 30 samples for testing for demonstration
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=30, random_state=42)
# Create linear regression object
regr = linear_model.LinearRegression()
# Train the model using the training sets
regr.fit(X_train, y_train)
# Make predictions using the testing set
y_pred = regr.predict(X_test)
# The coefficients
print(f'Coefficient (Slope): {regr.coef_[0]:.2f}') # 系数(斜率)
print(f'Intercept: {regr.intercept_:.2f}') # 截距
# The mean squared error
print(f'Mean squared error (MSE): {mean_squared_error(y_test, y_pred):.2f}') # 均方误差
# The coefficient of determination (R^2 score): 1 is perfect prediction
print(f'Coefficient of determination (R^2): {r2_score(y_test, y_pred):.2f}') # 决定系数 (R²)
# Plot outputs
plt.figure(figsize=(8, 6))
plt.scatter(X_test, y_test, color='blue', label='Actual Data') # 实际数据
plt.plot(X_test, y_pred, color='red', linewidth=2, label='Predicted Line') # 预测线
plt.xlabel('Feature (BMI)') # 特征 (BMI)
plt.ylabel('Target (Disease Progression)') # 目标变量(疾病进展)
plt.title('Simple Linear Regression (Diabetes Dataset)')
# plt.xticks(()) # Optional: Hide x-axis ticks if desired
# plt.yticks(()) # Optional: Hide y-axis ticks if desired
plt.legend()
plt.grid(True)
plt.show()
Coefficient (Slope): 999.53
Intercept: 152.00
Mean squared error (MSE): 3260.93
Coefficient of determination (R^2): 0.48

该输出显示了学习到的系数(斜率)、截距、MSE(均方误差,实际值与预测值之间的平均平方差)和 R² 分数(特征能够解释的目标变量的方差比例)。该图将显示测试数据点以及模型学习到的回归线。

MLR 扩展了 SLR,使用多个自变量预测单个因变量。模型在更高维度中拟合一个超平面(Hyperplane),而不是一条直线。

模型方程为:Y = β₀ + β₁X₁ + β₂X₂ + ... + βₚXₚ + ε

每个系数 βᵢ 的解释是,在保持所有其他预测变量不变的情况下,Xᵢ 增加一个单位时 Y 的期望变化量。

我们将使用 Scikit-learn 的 California Housing 数据集,因为 Boston 数据集存在伦理问题且已被弃用。

import matplotlib.pyplot as plt
import numpy as np
from sklearn import datasets, linear_model
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error, r2_score
# Load the California housing dataset
california = datasets.fetch_california_housing()
X = california.data
y = california.target
feature_names = california.feature_names
# Split the dataset into training and testing sets (e.g., 70% train, 30% test)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
# Create linear regression object
mlr = linear_model.LinearRegression()
# Train the model using the training sets
mlr.fit(X_train, y_train)
# Make predictions using the testing set
y_pred_train = mlr.predict(X_train)
y_pred_test = mlr.predict(X_test)
# The coefficients
print('Coefficients:') # 系数
for feature, coef in zip(feature_names, mlr.coef_):
print(f' {feature}: {coef:.4f}')
print(f'Intercept: {mlr.intercept_:.4f}') # 截距
# Evaluate the model
mse_test = mean_squared_error(y_test, y_pred_test) # 测试均方误差 (MSE)
r2_test = r2_score(y_test, y_pred_test) # 测试决定系数 (R²)
print(f'\nTest Mean Squared Error (MSE): {mse_test:.4f}')
print(f'Test Coefficient of Determination (R^2): {r2_test:.4f}')
# --- Residual Plot ---
# Residuals are the difference between actual and predicted values
train_residuals = y_train - y_pred_train # 训练集残差
test_residuals = y_test - y_pred_test # 测试集残差
plt.figure(figsize=(10, 6))
plt.scatter(y_pred_train, train_residuals,
color="steelblue", s=10, label='Train data', alpha=0.5) # 训练数据
plt.scatter(y_pred_test, test_residuals,
color="orange", s=10, label='Test data', alpha=0.5) # 测试数据
plt.hlines(y=0, xmin=np.min(y_pred_train), xmax=np.max(y_pred_train),
color='black', lw=2, linestyle='--')
plt.xlabel('Predicted Values') # 预测值
plt.ylabel('Residuals') # 残差
plt.title('Residual Plot') # 残差图
plt.legend()
plt.grid(True)
plt.show()
Coefficients:
MedInc: 0.4444
HouseAge: 0.0097
AveRooms: -0.1134
AveBedrms: 0.6992
Population: -0.0000
AveOccup: -0.0035
Latitude: -0.4198
Longitude: -0.4336
Intercept: -36.9112
Test Mean Squared Error (MSE): 0.5306
Test Coefficient of Determination (R^2): 0.5969

输出显示了每个特征的系数,表明了在保持其他特征不变的情况下,它对目标变量(房屋中位数价格)的估计影响。R² 表示模型解释了测试集中约 59.7% 的方差。残差图(Residual plot)有助于诊断模型拟合情况;理想情况下,残差应随机分布在零附近,没有明显的模式。

线性回归依赖于几个关键假设,以确保结果有效且可靠:

  • 线性性(Linearity): 预测变量与目标变量之间的关系是线性的。
  • 独立性(Independence): 观测值之间相互独立(对于时间序列数据尤其重要)。残差也应独立(无自相关 Autocorrelation)。
  • 同方差性(Homoscedasticity): 残差的方差在预测变量的各个水平上保持恒定(即,误差的分布是一致的)。
  • 残差的正态性(Normality of Residuals): 残差近似呈正态分布。这对于假设检验和置信区间比对预测准确性更重要。
  • 无完美多重共线性(No Perfect Multicollinearity): 预测变量之间不应存在完全线性相关。高(但不完全)的多重共线性(Multicollinearity)会夸大系数估计的方差,使其不稳定。

违反这些假设可能导致估计有偏或低效,以及推断不正确。诊断图(如上所示的残差图)和统计检验用于检查这些假设。

在处理高维(High dimensionality)或多重共线性问题时,通常首选正则化线性模型,如 Ridge、Lasso 和 Elastic Net。它们在成本函数中添加一个惩罚项(Penalty term),将系数收缩到零,这可以提高模型的泛化能力(Model generalization)并处理多重共线性。