梯度下降优化
TensorFlow - 梯度下降优化
Section titled “TensorFlow - 梯度下降优化”梯度下降(Gradient Descent)是一种基础的优化算法,用于寻找函数的最小值。在机器学习中,它被广泛用于通过迭代调整模型的参数(权重 weights 和偏置 biases)来最小化模型的损失函数(loss function)。
其核心思想是朝着与当前点梯度(或近似梯度)相反的方向迈步,因为这是最陡峭的下降方向。这些步骤的大小由学习率(learning rate)决定。
让我们在 TensorFlow 2.x 中实现一个简单的梯度下降优化,以最小化函数 f(x) = (log(x))^2。这个函数的最小值出现在 x = 1 处,此时 f(1) = (log(1))^2 = 0。
步骤 1:设置变量和优化器
Section titled “步骤 1:设置变量和优化器”首先,我们导入 TensorFlow 并定义我们想要优化的变量 x。我们还定义了损失函数和优化器(optimizer)。
import tensorflow as tf
# Define the variable to optimize. Initialize it to a value, e.g., 2.0.# 定义要优化的变量。将其初始化为一个值,例如 2.0。# We use tf.Variable because its value needs to change during optimization.# 我们使用 tf.Variable 是因为其值需要在优化过程中改变。x = tf.Variable(2.0, name='x', dtype=tf.float32)
# Define the loss function: f(x) = (log(x))^2# 定义损失函数: f(x) = (log(x))^2# Note: log(x) is typically natural logarithm (ln(x)) in TensorFlow.# 注意:在 TensorFlow 中,log(x) 通常指自然对数 (ln(x))。# Ensure x remains positive for log(x) to be defined.# 确保 x 保持正值以使 log(x) 有定义。def loss_function(): # tf.math.log computes the natural logarithm. # tf.math.log 计算自然对数。 # Adding a small epsilon can improve stability if x might approach zero, # 添加一个小的 epsilon 可以提高稳定性,以防 x 可能接近零, # but for x starting at 2.0 and moving towards 1.0, it's generally not an issue. # 但对于从 2.0 开始并移向 1.0 的 x,通常不是问题。 return tf.square(tf.math.log(x))
# Define the optimizer. We'll use Stochastic Gradient Descent (SGD).# 定义优化器。我们将使用随机梯度下降(SGD)。# The learning rate controls the step size.# 学习率控制步长。learning_rate = 0.5optimizer = tf.keras.optimizers.SGD(learning_rate=learning_rate)步骤 2:执行优化步骤
Section titled “步骤 2:执行优化步骤”接下来,我们编写一个循环来执行优化。在每个步骤中,我们计算损失,使用 tf.GradientTape 计算损失相对于 x 的梯度,然后使用优化器应用这些梯度来更新 x。
def optimize(num_steps=10): print(f"Starting at: x = {x.numpy():.4f}, loss = (log(x))^2 = {loss_function().numpy():.4f}")
for step in range(num_steps): # Open a GradientTape context to record operations for automatic differentiation. # 打开一个 GradientTape 上下文,以记录操作进行自动微分。 with tf.GradientTape() as tape: # Calculate the loss for the current value of x. # 计算当前 x 值的损失。 current_loss = loss_function()
# Calculate the gradients of the loss with respect to our variable [x]. # 计算损失相对于我们的变量 [x] 的梯度。 # tape.gradient will return a list of gradients, one for each trainable variable. # tape.gradient 将返回一个梯度列表,每个可训练变量对应一个。 gradients = tape.gradient(current_loss, [x])
# Apply the gradients to the variable(s) using the optimizer. # 使用优化器将梯度应用于变量。 # optimizer.apply_gradients expects a list of (gradient, variable) pairs. # optimizer.apply_gradients 需要一个 (梯度, 变量) 对的列表。 optimizer.apply_gradients(zip(gradients, [x]))
# Print the progress # 打印进度 print(f"Step {step+1:2d}: x = {x.numpy():.4f}, loss = {loss_function().numpy():.4f}")
# Run the optimization# 运行优化optimize(num_steps=10)
# For further learning:# 进一步学习:# - tf.GradientTape: https://www.tensorflow.org/guide/autodiff# - Optimizers in Keras: https://www.tensorflow.org/api_docs/python/tf/keras/optimizers输出将显示每一步中 x 的值和损失 (log(x))^2。你应该会观察到 x 迭代地趋近于 1.0,因此损失值会趋近于 0。这演示了梯度下降算法如何最小化函数。具体的数值和收敛速度取决于 x 的初始值、学习率以及函数本身。