线性回归
回归算法 - 线性回归
Section titled “回归算法 - 线性回归”线性回归介绍
Section titled “线性回归介绍”线性回归是基础的监督学习算法,用于基于一个或多个预测变量(自变量 Independent variables)预测一个连续目标变量(Target variable)。它假设预测变量与目标变量之间存在线性关系(Linear relationship)。
在数学上,这种关系建模为:
对于简单线性回归(一个预测变量):
Y = β₀ + β₁X + ε
对于多元线性回归(多个预测变量):
Y = β₀ + β₁X₁ + β₂X₂ + ... + βₚXₚ + ε
其中:
Y:因变量(Dependent variable,目标变量),是我们想要预测的变量。X、X₁、…、Xₚ:自变量(Independent variables,预测变量 Predictors)。β₀:截距(Intercept,当所有 X 都为 0 时 Y 的值)。β₁、…、βₚ:系数(Coefficients,斜率),表示在保持其他预测变量不变的情况下,相应 X 变化一个单位时 Y 的变化量。ε:误差项(Error term),表示实际值与预测值之间的差(残差 Residuals)。
线性回归的目标是找到系数(β₀, β₁, …)的最优值,以最小化误差,通常使用平方误差和(Sum of squared errors - SSE)或残差平方和(Residual Sum of Squares - RSS)来衡量。
线性关系类型
Section titled “线性关系类型”-
正线性关系(Positive Linear Relationship): 随着预测变量的增加,目标变量也倾向于增加(正斜率系数)。
-
负线性关系(Negative Linear Relationship): 随着预测变量的增加,目标变量倾向于减少(负斜率系数)。
线性回归的类型
Section titled “线性回归的类型”基于预测变量的数量:
- 简单线性回归(Simple Linear Regression - SLR): 使用单个预测变量(
X)预测目标变量(Y)。 - 多元线性回归(Multiple Linear Regression - MLR): 使用两个或多个预测变量(
X₁、X₂、…)预测目标变量(Y)。
简单线性回归(SLR)
Section titled “简单线性回归(SLR)”SLR 建模一个自变量和一个因变量之间的线性关系。它找到通过数据点的最佳拟合直线。
Python 实现(概念性 - 手动计算)
Section titled “Python 实现(概念性 - 手动计算)”虽然通常使用像 Scikit-learn 这样的库,但理解底层计算有所帮助。此示例演示了使用最小二乘法公式手动估计系数。
import numpy as npimport matplotlib.pyplot as plt
def estimate_coefficients(x, y): n = np.size(x)
# Calculate means mean_x, mean_y = np.mean(x), np.mean(y)
# Calculate cross-deviation and deviation about x SS_xy = np.sum(y * x) - n * mean_y * mean_x SS_xx = np.sum(x * x) - n * mean_x * mean_x
# Calculate regression coefficients (slope b1 and intercept b0) b1 = SS_xy / SS_xx b0 = mean_y - b1 * mean_x
return (b0, b1)
def plot_regression_line(x, y, b): plt.figure(figsize=(8, 6)) # Plot actual points as scatter plot plt.scatter(x, y, color="purple", marker="o", s=30, label='Data Points') # 数据点
# Predict response vector y_pred = b[0] + b[1] * x
# Plot the regression line plt.plot(x, y_pred, color="green", label='Regression Line') # 回归线
plt.xlabel('Independent Variable (X)') plt.ylabel('Dependent Variable (Y)') plt.title('Simple Linear Regression Example') plt.legend() plt.grid(True) plt.show()
# --- Main execution ---# Sample datax = np.array([0, 1, 2, 3, 4, 5, 6, 7, 8, 9])y = np.array([1, 3, 2, 5, 7, 8, 8, 9, 10, 12]) # 示例数据
# Estimate coefficientsb = estimate_coefficients(x, y)print(f"Estimated coefficients:\nIntercept (b0) = {b[0]:.4f} \nSlope (b1) = {b[1]:.4f}")
# Plot the regression lineplot_regression_line(x, y, b)输出(示例)
Section titled “输出(示例)”Estimated coefficients:Intercept (b0) = 1.2364Slope (b1) = 1.1697此代码计算样本数据的截距和斜率,并将生成的回归线绘制在数据点上。(该图会显示散点,绿线代表最佳线性拟合。)
Python 实现 (Scikit-learn)
Section titled “Python 实现 (Scikit-learn)”使用 Scikit-learn 是标准且更实用的方法。这里我们使用 Diabetes 数据集(仅加载一个特征用于 SLR)。
import matplotlib.pyplot as pltimport numpy as npfrom sklearn import datasets, linear_modelfrom sklearn.model_selection import train_test_splitfrom sklearn.metrics import mean_squared_error, r2_score
# Load the diabetes datasetdiabetes = datasets.load_diabetes()
# Use only one feature (e.g., the third feature, BMI)X = diabetes.data[:, np.newaxis, 2] # np.newaxis keeps it a 2D arrayy = diabetes.target
# Split the data into training/testing sets# Using only 30 samples for testing for demonstrationX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=30, random_state=42)
# Create linear regression objectregr = linear_model.LinearRegression()
# Train the model using the training setsregr.fit(X_train, y_train)
# Make predictions using the testing sety_pred = regr.predict(X_test)
# The coefficientsprint(f'Coefficient (Slope): {regr.coef_[0]:.2f}') # 系数(斜率)print(f'Intercept: {regr.intercept_:.2f}') # 截距# The mean squared errorprint(f'Mean squared error (MSE): {mean_squared_error(y_test, y_pred):.2f}') # 均方误差# The coefficient of determination (R^2 score): 1 is perfect predictionprint(f'Coefficient of determination (R^2): {r2_score(y_test, y_pred):.2f}') # 决定系数 (R²)
# Plot outputsplt.figure(figsize=(8, 6))plt.scatter(X_test, y_test, color='blue', label='Actual Data') # 实际数据plt.plot(X_test, y_pred, color='red', linewidth=2, label='Predicted Line') # 预测线plt.xlabel('Feature (BMI)') # 特征 (BMI)plt.ylabel('Target (Disease Progression)') # 目标变量(疾病进展)plt.title('Simple Linear Regression (Diabetes Dataset)')# plt.xticks(()) # Optional: Hide x-axis ticks if desired# plt.yticks(()) # Optional: Hide y-axis ticks if desiredplt.legend()plt.grid(True)plt.show()输出(示例)
Section titled “输出(示例)”Coefficient (Slope): 999.53Intercept: 152.00Mean squared error (MSE): 3260.93Coefficient of determination (R^2): 0.48该输出显示了学习到的系数(斜率)、截距、MSE(均方误差,实际值与预测值之间的平均平方差)和 R² 分数(特征能够解释的目标变量的方差比例)。该图将显示测试数据点以及模型学习到的回归线。
多元线性回归(MLR)
Section titled “多元线性回归(MLR)”MLR 扩展了 SLR,使用多个自变量预测单个因变量。模型在更高维度中拟合一个超平面(Hyperplane),而不是一条直线。
模型方程为:Y = β₀ + β₁X₁ + β₂X₂ + ... + βₚXₚ + ε
每个系数 βᵢ 的解释是,在保持所有其他预测变量不变的情况下,Xᵢ 增加一个单位时 Y 的期望变化量。
Python 实现 (Scikit-learn)
Section titled “Python 实现 (Scikit-learn)”我们将使用 Scikit-learn 的 California Housing 数据集,因为 Boston 数据集存在伦理问题且已被弃用。
import matplotlib.pyplot as pltimport numpy as npfrom sklearn import datasets, linear_modelfrom sklearn.model_selection import train_test_splitfrom sklearn.metrics import mean_squared_error, r2_score
# Load the California housing datasetcalifornia = datasets.fetch_california_housing()X = california.datay = california.targetfeature_names = california.feature_names
# Split the dataset into training and testing sets (e.g., 70% train, 30% test)X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
# Create linear regression objectmlr = linear_model.LinearRegression()
# Train the model using the training setsmlr.fit(X_train, y_train)
# Make predictions using the testing sety_pred_train = mlr.predict(X_train)y_pred_test = mlr.predict(X_test)
# The coefficientsprint('Coefficients:') # 系数for feature, coef in zip(feature_names, mlr.coef_): print(f' {feature}: {coef:.4f}')print(f'Intercept: {mlr.intercept_:.4f}') # 截距
# Evaluate the modelmse_test = mean_squared_error(y_test, y_pred_test) # 测试均方误差 (MSE)r2_test = r2_score(y_test, y_pred_test) # 测试决定系数 (R²)print(f'\nTest Mean Squared Error (MSE): {mse_test:.4f}')print(f'Test Coefficient of Determination (R^2): {r2_test:.4f}')
# --- Residual Plot ---# Residuals are the difference between actual and predicted valuestrain_residuals = y_train - y_pred_train # 训练集残差test_residuals = y_test - y_pred_test # 测试集残差
plt.figure(figsize=(10, 6))plt.scatter(y_pred_train, train_residuals, color="steelblue", s=10, label='Train data', alpha=0.5) # 训练数据plt.scatter(y_pred_test, test_residuals, color="orange", s=10, label='Test data', alpha=0.5) # 测试数据
plt.hlines(y=0, xmin=np.min(y_pred_train), xmax=np.max(y_pred_train), color='black', lw=2, linestyle='--')
plt.xlabel('Predicted Values') # 预测值plt.ylabel('Residuals') # 残差plt.title('Residual Plot') # 残差图plt.legend()plt.grid(True)plt.show()输出(示例)
Section titled “输出(示例)”Coefficients: MedInc: 0.4444 HouseAge: 0.0097 AveRooms: -0.1134 AveBedrms: 0.6992 Population: -0.0000 AveOccup: -0.0035 Latitude: -0.4198 Longitude: -0.4336Intercept: -36.9112
Test Mean Squared Error (MSE): 0.5306Test Coefficient of Determination (R^2): 0.5969输出显示了每个特征的系数,表明了在保持其他特征不变的情况下,它对目标变量(房屋中位数价格)的估计影响。R² 表示模型解释了测试集中约 59.7% 的方差。残差图(Residual plot)有助于诊断模型拟合情况;理想情况下,残差应随机分布在零附近,没有明显的模式。
线性回归的假设
Section titled “线性回归的假设”线性回归依赖于几个关键假设,以确保结果有效且可靠:
- 线性性(Linearity): 预测变量与目标变量之间的关系是线性的。
- 独立性(Independence): 观测值之间相互独立(对于时间序列数据尤其重要)。残差也应独立(无自相关 Autocorrelation)。
- 同方差性(Homoscedasticity): 残差的方差在预测变量的各个水平上保持恒定(即,误差的分布是一致的)。
- 残差的正态性(Normality of Residuals): 残差近似呈正态分布。这对于假设检验和置信区间比对预测准确性更重要。
- 无完美多重共线性(No Perfect Multicollinearity): 预测变量之间不应存在完全线性相关。高(但不完全)的多重共线性(Multicollinearity)会夸大系数估计的方差,使其不稳定。
违反这些假设可能导致估计有偏或低效,以及推断不正确。诊断图(如上所示的残差图)和统计检验用于检查这些假设。
在处理高维(High dimensionality)或多重共线性问题时,通常首选正则化线性模型,如 Ridge、Lasso 和 Elastic Net。它们在成本函数中添加一个惩罚项(Penalty term),将系数收缩到零,这可以提高模型的泛化能力(Model generalization)并处理多重共线性。
- Scikit-learn 线性模型文档:https://scikit-learn.org/stable/modules/linear_model.html
- StatQuest:线性回归解释:https://statquest.org/video-index/(搜索“Linear Regression”)