Skip to content

TensorFlow - 基础

TensorFlow 的核心是“张量”(tensor)。张量是具有统一数据类型的多维数组。它们是 TensorFlow 中使用的基本数据结构。你可以将张量视为向量和矩阵向更高维度的推广。

张量具有三个关键属性:

张量的秩(rank)描述了它的维度数量。有时也称为阶(order)或度(degree)。

  • 标量(一个单独的数字)是秩为 0 的张量。
  • 向量(一维数组)是秩为 1 的张量。
  • 矩阵(二维数组)是秩为 2 的张量。
  • 依此类推,对于秩 3(例如,数字立方体)、秩 4 等。

张量的形状(shape)是一个整数元组,描述了每个维度的大小。例如:

  • 标量的形状是空的 ()。
  • 长度为 5 的向量的形状是 (5,)。
  • 具有 3 行 4 列的矩阵的形状是 (3, 4)。
  • 秩为 3 的张量可能有形状 (2, 3, 4)。

数据类型(dtype)定义了张量中存储的数据类型(例如,tf.float32、tf.int32、tf.string)。单个张量中的所有元素必须具有相同的数据类型。

创建张量通常涉及从 Python 列表或 NumPy 数组开始,然后使用 tf.constant() 或 tf.Variable() 将它们转换为 tf.Tensor 对象。

让我们看看如何使用 TensorFlow 2.x 创建不同秩的张量,通常为了熟悉起见会从 NumPy 数组开始。

import tensorflow as tf
import numpy as np
# 创建一个标量(秩为 0 的张量)
scalar_np = np.array(10)
scalar_tf = tf.constant(scalar_np)
print("NumPy scalar:", scalar_np)
print("TensorFlow scalar (tensor):")
print(scalar_tf)
print(f"Rank: {tf.rank(scalar_tf).numpy()}, Shape: {scalar_tf.shape}, Dtype: {scalar_tf.dtype}")

输出将显示标量值及其属性(秩 0,空形状)。

# 创建一个向量(秩为 1 的张量)
vector_np = np.array([1.3, 1.0, 4.0, 23.99])
vector_tf = tf.constant(vector_np, dtype=tf.float32)
print("\nNumPy vector:", vector_np)
print("TensorFlow vector (tensor):")
print(vector_tf)
print(f"Rank: {tf.rank(vector_tf).numpy()}, Shape: {vector_tf.shape}, Dtype: {vector_tf.dtype}")
# 访问元素(类似于 Python 列表/NumPy 数组)
print("Element at index 0:", vector_tf[0].numpy())
print("Element at index 2:", vector_tf[2].numpy())

输出将显示向量、其属性(秩 1,形状例如 (4,)),以及元素访问示例。

# 创建一个矩阵(秩为 2 的张量)
matrix_np = np.array([(1, 2, 3, 4), (5, 6, 7, 8), (9, 10, 11, 12)])
matrix_tf = tf.constant(matrix_np, dtype=tf.int32)
print("\nNumPy matrix:\n", matrix_np)
print("TensorFlow matrix (tensor):")
print(matrix_tf)
print(f"Rank: {tf.rank(matrix_tf).numpy()}, Shape: {matrix_tf.shape}, Dtype: {matrix_tf.dtype}")
# 访问元素(例如,第 2 行第 3 列的元素 - 0 索引)
print("Element at [1, 2]:", matrix_tf[1, 2].numpy()) # 访问值 7

输出将显示矩阵及其属性(例如,秩 2,形状 (3, 4))。

TensorFlow 提供了一系列丰富的操作来处理张量。在 TensorFlow 2.x 中,这些操作以 Eager Execution (即时执行) 模式立即执行(与 TF 1.x 的图执行模型和会话不同)。

# 定义两个 TensorFlow 常量矩阵
matrix1_tf = tf.constant([[2, 2, 2], [2, 2, 2], [2, 2, 2]], dtype=tf.int32)
matrix2_tf = tf.constant([[1, 1, 1], [1, 1, 1], [1, 1, 1]], dtype=tf.int32)
print("\nMatrix 1:")
print(matrix1_tf)
print("Matrix 2:")
print(matrix2_tf)
# 逐元素相加
matrix_sum = tf.add(matrix1_tf, matrix2_tf)
print("\nMatrix Sum (matrix1 + matrix2):")
print(matrix_sum)
# 矩阵乘法
matrix_product = tf.matmul(matrix1_tf, matrix2_tf)
print("\nMatrix Product (matrix1 * matrix2):")
print(matrix_product)
# 矩阵的行列式(必须是浮点或复数类型)
matrix_float_tf = tf.constant([[2.0, 7.0, 2.0],
[1.0, 4.0, 2.0],
[9.0, 0.0, 2.0]], dtype=tf.float32)
print("\nFloat Matrix for Determinant:")
print(matrix_float_tf)
matrix_determinant = tf.linalg.det(matrix_float_tf)
print("\nDeterminant of Float Matrix:")
print(matrix_determinant.numpy()) # .numpy() 获取标量值

输出:

代码将打印定义的矩阵、它们的和、它们的乘积以及浮点矩阵的行列式。在 TensorFlow 2.x 中,这些操作直接执行,无需像 TensorFlow 1.x 那样调用 tf.Session().run(...)。

与 TensorFlow 1.x 相比:

  • Eager Execution (即时执行): 操作立即计算。无需 tf.Session 来评估张量。
  • .numpy(): 要从 TensorFlow 张量获取 Python/NumPy 值,请使用 .numpy() 方法。
  • API Updates: 一些函数可能有新的名称或位于不同的模块中(例如,tf.matrix_determinant 变为 tf.linalg.det)。
  • Python 3 Syntax: 使用现代 Python 语法,例如 print() 函数。

理解张量和基本操作是使用任何深度学习框架(包括 TensorFlow)的基础。有关更高级的张量操作和完整的操作列表,请参阅官方 TensorFlow 关于张量的文档:https://www.tensorflow.org/guide/tensor