Skip to content

R - 非线性最小二乘

虽然线性回归功能强大,但许多现实世界中的现象并不遵循直线关系。它们更适合用非线性模型来描述,例如生物学中的指数增长曲线或经济学中的幂律。非线性最小二乘法(NLS)是一种将非线性模型拟合到数据集的方法。其目标与线性回归相似,都是找到使观测数据与模型预测值之间平方差之和最小的参数值。

在 R 中,用于非线性最小二乘的标准函数是 nls()。该过程涉及定义模型公式、提供数据,以及——最重要的是——为算法提供合理的参数初始估计值。然后,该函数会迭代地精炼这些估计值以找到最佳拟合。

nls() 的基本语法是:

nls(formula, data, start)
  • formula:一个非线性模型公式,包含变量和你希望估计的参数(例如,y ~ a * x^b)。
  • data:一个包含公式中所用变量的数据框。
  • start:一个命名列表或向量,包含参数的初始猜测值(例如,list(a = 1, b = 2))。

让我们将模型 y = b1*x^2 + b2 拟合到一些示例数据。一个关键步骤是选择好的初始值。我们通常可以通过绘制数据来获得线索。

# 加载用于数据操作和绘图的现代包
library(tidyverse)
# 1. 准备数据
xvalues <- c(1.6, 2.1, 2, 2.23, 3.71, 3.25, 3.4, 3.86, 1.19, 2.21)
yvalues <- c(5.19, 7.43, 6.94, 8.11, 18.75, 14.88, 16.06, 19.12, 3.21, 7.58)
# 创建一个 tibble(一种现代数据框)
data_to_fit <- tibble(x = xvalues, y = yvalues)
# 2. 可视化数据以猜测初始值
# 曲线看起来像一个向上开口的抛物线(x^2 项),并向上平移。
# 我们猜测 b1 ≈ 1 且 y 截距 b2 ≈ 2 或 3。
ggplot(data_to_fit, aes(x = x, y = y)) +
geom_point(color = "blue", size = 3) +
labs(title = "Data and Initial Guesses", x = "X Values", y = "Y Values")
# 3. 定义初始值并拟合模型
start_values <- list(b1 = 1, b2 = 3)
# 使用 nls() 查找最佳拟合参数
nls_model <- nls(y ~ b1 * x^2 + b2, data = data_to_fit, start = start_values)
# 4. 检查模型摘要
# summary() 比直接打印模型提供的信息更丰富
print(summary(nls_model))
# 5. 获取参数的置信区间
print(confint(nls_model))

摘要提供了关于模型拟合的详细信息:

Formula: y ~ b1 * x^2 + b2
Parameters:
Estimate Std. Error t value Pr(>|t|)
_
b1 1.19235 0.05203 22.917 1.48e-08 ***
b2 1.99692 0.44858 4.452 0.00193 **
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Residual standard error: 0.3677 on 8 degrees of freedom
Number of iterations to convergence: 4
Achieved convergence tolerance: 4.887e-07
Waiting for profiling to be done...
2.5% 97.5%
b1 1.072228 1.312471
b2 0.962804 3.031045

评估模型的最佳方法是在原始数据上绘制拟合曲线。我们可以使用 predict() 函数从模型生成点。

# 使用模型的拟合值扩展原始数据
fit_data <- broom::augment(nls_model)
# 绘制原始数据点和拟合的非线性最小二乘曲线
ggplot(data = data_to_fit, aes(x = x, y = y)) +
geom_point(color = "blue", size = 3, alpha = 0.7) +
geom_line(data = fit_data, aes(y = .fitted), color = "red", size = 1) +
labs(
title = "Nonlinear Least Squares Fit",
subtitle = "y = 1.19*x^2 + 2.00",
x = "X Values",
y = "Y Values"
) +
theme_minimal()

使用 nls() 时一个常见问题是未能收敛。这通常意味着以下两种情况之一:

  1. 初始值选择不当:参数的初始猜测值与最优值相差太远。尝试绘制你的数据并进行更明智的猜测。
  2. 模型不合适:你所选的模型公式可能与数据的基础结构存在根本性的拟合不良。考虑使用不同的模型。