单层感知器
TensorFlow - 单层感知机
Section titled “TensorFlow - 单层感知机”要理解单层感知机,首先了解人工神经网络(ANN)的基础知识会很有帮助。人工神经网络是一种信息处理系统,其灵感来源于生物神经网络的功能。ANN 由许多相互连接的处理单元组成,通常称为神经元。这些网络通常有一个输入层、一个或多个隐藏层和一个输出层。
在基本的前馈网络中,信息从输入层流向输出层,经过任何隐藏层。隐藏单元处理来自输入层的信息,并将其传递给输出层。输入和输出单元通常与外部环境或更大系统的其他部分交互,而隐藏层执行中间计算。
神经网络的架构由节点之间的连接模式、总层数、层内节点的排列方式以及每层的神经元数量定义。
神经网络架构可以根据其结构和数据流大致分为两类。两种基本类型包括:
- 单层感知机 (Single Layer Perceptron)
- 多层感知机 (Multi-Layer Perceptron)
单层感知机 (Single Layer Perceptron)
Section titled “单层感知机 (Single Layer Perceptron)”单层感知机(SLP)是最早、最简单的神经网络模型类型之一。它只包含一个输出节点层;输入通过一系列权重直接馈送到输出。计算过程涉及计算输入的加权和。然后将这个和通过激活函数传递,从而产生输出。SLP 没有隐藏层。
我们来使用 TensorFlow 为图像分类问题实现一个单层感知机。逻辑回归(Logistic Regression)是一个经典的例子,它可以被建模为一个 SLP,特别是使用 Softmax 激活函数进行多类别分类时。
训练像逻辑回归(在此作为单层感知机)这样的模型的基本步骤如下:
- 初始化权重:模型的权重(和偏置)通常在训练开始时用小的随机值或零进行初始化。
- 迭代和更新:对于每一批训练数据,模型进行预测。使用损失函数(例如,交叉熵)计算误差(期望输出与实际输出之间的差)。然后使用优化算法(如梯度下降)根据该误差调整权重。
- 重复:重复此过程,直到达到设定的周期数(遍历整个训练数据集的次数),或直到验证集上的误差停止改善或达到期望的阈值。
以下代码演示了如何使用 TensorFlow 2.x 和 Keras 对 MNIST 数字进行逻辑回归分类,这可以作为一个单层感知机的例子:
import tensorflow as tfimport matplotlib.pyplot as pltimport numpy as np
# Load MNIST data using tf.keras.datasets(x_train, y_train), (x_test, y_test) = tf.keras.datasets.mnist.load_data()
# Preprocess the data# Normalize pixel values to be between 0 and 1x_train = x_train.astype('float32') / 255.0x_test = x_test.astype('float32') / 255.0
# Flatten the images (28x28) into 784-dimensional vectorsx_train = x_train.reshape((-1, 784))x_test = x_test.reshape((-1, 784))
# Convert labels to one-hot encoded vectors (e.g., 2 -> [0,0,1,0,0,0,0,0,0,0])y_train = tf.keras.utils.to_categorical(y_train, num_classes=10)y_test = tf.keras.utils.to_categorical(y_test, num_classes=10)
# Parameterslearning_rate = 0.01training_epochs = 25batch_size = 100display_step = 1
# Define the model using Keras Sequential API (a single dense layer is an SLP)model = tf.keras.Sequential([ tf.keras.layers.Dense(10, activation='softmax', input_shape=(784,)) # This Dense layer implements: output = softmax(dot(input, kernel) + bias) # 'kernel' are the weights (W), 'bias' is the bias (b)])
# Compile the model# Optimizer: Stochastic Gradient Descent (SGD)# Loss function: Categorical Crossentropy for multi-class classificationoptimizer = tf.keras.optimizers.SGD(learning_rate=learning_rate)model.compile(optimizer=optimizer, loss='categorical_crossentropy', metrics=['accuracy'])
# Train the modelhistory = model.fit(x_train, y_train, batch_size=batch_size, epochs=training_epochs, verbose=0, # Set to 1 or 2 for more verbose output during training validation_data=(x_test, y_test)) # Optionally track validation loss/accuracy
# Display training progress (loss)print("Training phase finished.")for epoch in range(training_epochs): if epoch % display_step == 0: # Access loss from history object (Keras model.fit returns a History object) # Note: history.history['loss'] is a list of losses per epoch print(f"Epoch: {epoch+1:04d} cost= {history.history['loss'][epoch]:.9f}")
# Plot training costplt.plot(np.arange(1, training_epochs + 1), history.history['loss'], 'o-', label='Logistic Regression Training cost')plt.ylabel('Cost (Categorical Crossentropy)')plt.xlabel('Epoch')plt.title('Training Cost Over Epochs')plt.legend()plt.show()
# Evaluate the model on the test setloss, accuracy = model.evaluate(x_test, y_test, verbose=0)print(f"Model accuracy on test set: {accuracy:.4f}")
# For further exploration:# - `tf.data` API for efficient input pipelines: https://www.tensorflow.org/guide/data# - Keras overview: https://www.tensorflow.org/guide/keras/overview运行代码会首先下载 MNIST 数据集(如果本地不存在)。然后,它会训练逻辑回归模型。在训练期间,它会按指定的周期间隔打印损失(cost)。训练完成后,会显示一个图表,展示训练损失随周期减少的情况。最后,它会打印模型在测试数据集上的准确率(accuracy)。对于这个简单的模型在 MNIST 数据集上,准确率应该相当高(通常 >90%),这表明单层感知机已经学会了对手写数字进行分类。
此处实现的逻辑回归是一种用于分类的线性模型。它描述了输入特征与属于特定类别的概率之间的关系。尽管它很简单,但在机器学习中它是一个基础算法。