Skip to content

梯度下降优化

梯度下降(Gradient Descent)是一种基础的优化算法,用于寻找函数的最小值。在机器学习中,它被广泛用于通过迭代调整模型的参数(权重 weights 和偏置 biases)来最小化模型的损失函数(loss function)。

其核心思想是朝着与当前点梯度(或近似梯度)相反的方向迈步,因为这是最陡峭的下降方向。这些步骤的大小由学习率(learning rate)决定。

让我们在 TensorFlow 2.x 中实现一个简单的梯度下降优化,以最小化函数 f(x) = (log(x))^2。这个函数的最小值出现在 x = 1 处,此时 f(1) = (log(1))^2 = 0。

首先,我们导入 TensorFlow 并定义我们想要优化的变量 x。我们还定义了损失函数和优化器(optimizer)。

import tensorflow as tf
# Define the variable to optimize. Initialize it to a value, e.g., 2.0.
# 定义要优化的变量。将其初始化为一个值,例如 2.0。
# We use tf.Variable because its value needs to change during optimization.
# 我们使用 tf.Variable 是因为其值需要在优化过程中改变。
x = tf.Variable(2.0, name='x', dtype=tf.float32)
# Define the loss function: f(x) = (log(x))^2
# 定义损失函数: f(x) = (log(x))^2
# Note: log(x) is typically natural logarithm (ln(x)) in TensorFlow.
# 注意:在 TensorFlow 中,log(x) 通常指自然对数 (ln(x))。
# Ensure x remains positive for log(x) to be defined.
# 确保 x 保持正值以使 log(x) 有定义。
def loss_function():
# tf.math.log computes the natural logarithm.
# tf.math.log 计算自然对数。
# Adding a small epsilon can improve stability if x might approach zero,
# 添加一个小的 epsilon 可以提高稳定性,以防 x 可能接近零,
# but for x starting at 2.0 and moving towards 1.0, it's generally not an issue.
# 但对于从 2.0 开始并移向 1.0 的 x,通常不是问题。
return tf.square(tf.math.log(x))
# Define the optimizer. We'll use Stochastic Gradient Descent (SGD).
# 定义优化器。我们将使用随机梯度下降(SGD)。
# The learning rate controls the step size.
# 学习率控制步长。
learning_rate = 0.5
optimizer = tf.keras.optimizers.SGD(learning_rate=learning_rate)

接下来,我们编写一个循环来执行优化。在每个步骤中,我们计算损失,使用 tf.GradientTape 计算损失相对于 x 的梯度,然后使用优化器应用这些梯度来更新 x。

def optimize(num_steps=10):
print(f"Starting at: x = {x.numpy():.4f}, loss = (log(x))^2 = {loss_function().numpy():.4f}")
for step in range(num_steps):
# Open a GradientTape context to record operations for automatic differentiation.
# 打开一个 GradientTape 上下文,以记录操作进行自动微分。
with tf.GradientTape() as tape:
# Calculate the loss for the current value of x.
# 计算当前 x 值的损失。
current_loss = loss_function()
# Calculate the gradients of the loss with respect to our variable [x].
# 计算损失相对于我们的变量 [x] 的梯度。
# tape.gradient will return a list of gradients, one for each trainable variable.
# tape.gradient 将返回一个梯度列表,每个可训练变量对应一个。
gradients = tape.gradient(current_loss, [x])
# Apply the gradients to the variable(s) using the optimizer.
# 使用优化器将梯度应用于变量。
# optimizer.apply_gradients expects a list of (gradient, variable) pairs.
# optimizer.apply_gradients 需要一个 (梯度, 变量) 对的列表。
optimizer.apply_gradients(zip(gradients, [x]))
# Print the progress
# 打印进度
print(f"Step {step+1:2d}: x = {x.numpy():.4f}, loss = {loss_function().numpy():.4f}")
# Run the optimization
# 运行优化
optimize(num_steps=10)
# For further learning:
# 进一步学习:
# - tf.GradientTape: https://www.tensorflow.org/guide/autodiff
# - Optimizers in Keras: https://www.tensorflow.org/api_docs/python/tf/keras/optimizers

输出将显示每一步中 x 的值和损失 (log(x))^2。你应该会观察到 x 迭代地趋近于 1.0,因此损失值会趋近于 0。这演示了梯度下降算法如何最小化函数。具体的数值和收敛速度取决于 x 的初始值、学习率以及函数本身。