TensorFlow - 优化器
TensorFlow - 优化器
Section titled “TensorFlow - 优化器”优化器(Optimizers)是用于改变神经网络属性(例如权重 weights 和学习率 learning rate)以减少损失(losses)的算法或方法。它们对于有效地训练深度学习模型至关重要。在 TensorFlow 中,优化器通常位于 tf.keras.optimizers 模块中。
Keras 优化器的基类是 tf.keras.optimizers.Optimizer。编译 Keras 模型时,需要指定一个优化器,该优化器将在训练过程 (model.fit()) 中根据计算出的损失函数梯度来更新模型的权重。
优化器有助于最小化损失函数。损失函数就像一个地形图的向导,告诉优化器它是否正朝着山谷底部(即最小损失)的正确方向移动。
在 tf.keras.optimizers 中常用的优化器包括:
- SGD (Stochastic Gradient Descent):一个基础的优化器。常与动量(momentum)一起使用。
- Adam (Adaptive Moment Estimation):一种自适应学习率的优化算法,专门为训练深度神经网络而设计。它通常是一个很好的默认选择。
- RMSprop:为每个参数维护一个学习率,根据该权重梯度近期幅度的平均值进行调整。
- Adagrad:根据参数调整学习率,对不常出现的参数进行较大的更新,对常出现的参数进行较小的更新。
- Adadelta:Adagrad 的扩展,旨在减少其激进的、单调递减的学习率。
- Adamax:基于无穷范数(infinity norm)的 Adam 变体。
- Nadam:带有 Nesterov 动量的 Adam 优化器。
让我们以 SGD 为例,演示如何在 Keras 模型中使用优化器。
要使用优化器,通常需要实例化它,并将其传递给 model.compile() 方法的 optimizer 参数。创建优化器实例时,可以指定学习率等参数。
import tensorflow as tffrom tensorflow import kerasfrom tensorflow.keras import layers
# Example: Using the SGD optimizer# Define a simple modelmodel = keras.Sequential([ layers.Dense(64, activation='relu', input_shape=(784,)), layers.Dense(10, activation='softmax')])
# Instantiate the SGD optimizer with a specific learning rate and momentumsgd_optimizer = tf.keras.optimizers.SGD(learning_rate=0.01, momentum=0.9)
# Compile the model with the chosen optimizermodel.compile(optimizer=sgd_optimizer, loss='categorical_crossentropy', metrics=['accuracy'])
print("Model compiled with SGD optimizer.")
# You can also use optimizers by their string identifiers (with default parameters)# model.compile(optimizer='adam',# loss='categorical_crossentropy',# metrics=['accuracy'])# print("Model compiled with Adam optimizer (using string identifier).")
# --- For demonstration: A conceptual look at what an optimizer does ---# This is NOT how you'd typically implement SGD, but shows the idea:## def conceptual_sgd_update(params, gradients, learning_rate):# updated_params = []# for param, grad in zip(params, gradients):# updated_param = param - learning_rate * grad# updated_params.append(updated_param)# return updated_params## # In a real scenario, 'params' would be model weights,# # 'gradients' would be calculated by backpropagation.# # Keras handles all this automatically when you call model.fit().优化器的选择及其配置(如学习率)可以显著影响训练速度和模型的最终性能。通常需要进行实验来找到特定问题的最佳设置。像 Adam 这样的现代优化器通常在各种问题上使用默认设置都能很好地工作,因此是很好的起点。
有关每个优化器及其参数的更多详细信息,请参阅官方 TensorFlow Keras 文档:https://www.tensorflow.org/api_docs/python/tf/keras/optimizers