R - 非线性最小二乘
R - 非线性最小二乘法
Section titled “R - 非线性最小二乘法”虽然线性回归功能强大,但许多现实世界中的现象并不遵循直线关系。它们更适合用非线性模型来描述,例如生物学中的指数增长曲线或经济学中的幂律。非线性最小二乘法(NLS)是一种将非线性模型拟合到数据集的方法。其目标与线性回归相似,都是找到使观测数据与模型预测值之间平方差之和最小的参数值。
在 R 中,用于非线性最小二乘的标准函数是 nls()。该过程涉及定义模型公式、提供数据,以及——最重要的是——为算法提供合理的参数初始估计值。然后,该函数会迭代地精炼这些估计值以找到最佳拟合。
nls() 的基本语法是:
nls(formula, data, start)formula:一个非线性模型公式,包含变量和你希望估计的参数(例如,y ~ a * x^b)。data:一个包含公式中所用变量的数据框。start:一个命名列表或向量,包含参数的初始猜测值(例如,list(a = 1, b = 2))。
示例:拟合二次模型
Section titled “示例:拟合二次模型”让我们将模型 y = b1*x^2 + b2 拟合到一些示例数据。一个关键步骤是选择好的初始值。我们通常可以通过绘制数据来获得线索。
# 加载用于数据操作和绘图的现代包library(tidyverse)
# 1. 准备数据xvalues <- c(1.6, 2.1, 2, 2.23, 3.71, 3.25, 3.4, 3.86, 1.19, 2.21)yvalues <- c(5.19, 7.43, 6.94, 8.11, 18.75, 14.88, 16.06, 19.12, 3.21, 7.58)
# 创建一个 tibble(一种现代数据框)data_to_fit <- tibble(x = xvalues, y = yvalues)
# 2. 可视化数据以猜测初始值# 曲线看起来像一个向上开口的抛物线(x^2 项),并向上平移。# 我们猜测 b1 ≈ 1 且 y 截距 b2 ≈ 2 或 3。ggplot(data_to_fit, aes(x = x, y = y)) + geom_point(color = "blue", size = 3) + labs(title = "Data and Initial Guesses", x = "X Values", y = "Y Values")
# 3. 定义初始值并拟合模型start_values <- list(b1 = 1, b2 = 3)
# 使用 nls() 查找最佳拟合参数nls_model <- nls(y ~ b1 * x^2 + b2, data = data_to_fit, start = start_values)
# 4. 检查模型摘要# summary() 比直接打印模型提供的信息更丰富print(summary(nls_model))
# 5. 获取参数的置信区间print(confint(nls_model))摘要提供了关于模型拟合的详细信息:
Formula: y ~ b1 * x^2 + b2
Parameters: Estimate Std. Error t value Pr(>|t|)_b1 1.19235 0.05203 22.917 1.48e-08 ***b2 1.99692 0.44858 4.452 0.00193 **---Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Residual standard error: 0.3677 on 8 degrees of freedom
Number of iterations to convergence: 4Achieved convergence tolerance: 4.887e-07
Waiting for profiling to be done... 2.5% 97.5%b1 1.072228 1.312471b2 0.962804 3.031045使用 ggplot2 可视化拟合结果
Section titled “使用 ggplot2 可视化拟合结果”评估模型的最佳方法是在原始数据上绘制拟合曲线。我们可以使用 predict() 函数从模型生成点。
# 使用模型的拟合值扩展原始数据fit_data <- broom::augment(nls_model)
# 绘制原始数据点和拟合的非线性最小二乘曲线ggplot(data = data_to_fit, aes(x = x, y = y)) + geom_point(color = "blue", size = 3, alpha = 0.7) + geom_line(data = fit_data, aes(y = .fitted), color = "red", size = 1) + labs( title = "Nonlinear Least Squares Fit", subtitle = "y = 1.19*x^2 + 2.00", x = "X Values", y = "Y Values" ) + theme_minimal()常见挑战:收敛性
Section titled “常见挑战:收敛性”使用 nls() 时一个常见问题是未能收敛。这通常意味着以下两种情况之一:
- 初始值选择不当:参数的初始猜测值与最优值相差太远。尝试绘制你的数据并进行更明智的猜测。
- 模型不合适:你所选的模型公式可能与数据的基础结构存在根本性的拟合不良。考虑使用不同的模型。