NumPy - 数据类型
NumPy - 数据类型 (dtype)
Section titled “NumPy - 数据类型 (dtype)”每个 NumPy 数组(ndarray)都是由具有相同数据类型的元素组成的集合。与标准 Python 相比,NumPy 提供了更丰富的数值数据类型,可以对内存使用和精度进行精细控制。这些数据类型信息被封装在每个数组关联的**数据类型对象(dtype)**中。
dtype 对象提供了关于数组内存块中的字节应如何解释的元数据。这包括:
- **类型:**数据的种类(例如,整数、浮点数、布尔值、复数、字符串)。
- **大小:**每个元素的字节数(例如,
int64使用 8 字节)。 - 字节序:数据的字节顺序(小端序或大端序)。
- **结构(针对结构化数组):**字段名、每个字段的数据类型以及内存布局。
- **子数组形状(如适用):**对于表示数组中包含数组的数据类型。
常见的 NumPy 数据类型
Section titled “常见的 NumPy 数据类型”下表列出了一些 NumPy 中最常用的内置标量数据类型。通常可以通过 np.<type_name> 访问它们(例如 np.int64、np.float32)。
| 类型 | 描述 |
|---|---|
np.bool_ | 布尔类型(True 或 False),存储为单个字节。 |
np.int8、np.int16、np.int32、np.int64 | 有符号整数,大小各异(8、16、32、64 位)。np.int_ 是默认整数类型(通常是 np.int64)。 |
np.uint8、np.uint16、np.uint32、np.uint64 | 无符号整数,大小各异。 |
np.float16、np.float32、np.float64 | 浮点数,精度不同(半精度、单精度、双精度)。np.float_ 是默认浮点类型(通常是 np.float64)。 |
np.complex64、np.complex128 | 复数,由一对浮点数表示(两个 32 位或两个 64 位浮点数)。np.complex_ 默认为 np.complex128。 |
np.str_ | Unicode 字符串类型。兼容 Python str。使用变长字符串。 |
np.bytes_ | 字节串类型。兼容 Python bytes。使用变长字符串。 |
np.datetime64 | 日期和时间值,具有各种单位(例如 [Y]、[M]、[D]、[h]、[m]、[s]、[ms]、[us]、[ns])。 |
np.timedelta64 | 表示两个 datetime64 值之间的差值。 |
np.object_ | 表示 Python 对象。允许数组包含不同类型的元素(如果对性能要求高,则不推荐使用)。 |
np.void | 表示原始、非结构化的数据块。常用于作为结构化数组的容器。 |
创建 dtype 对象
Section titled “创建 dtype 对象”可以使用 numpy.dtype() 显式创建 dtype 对象。这通常在定义结构化数组或指定精确类型细节(如字节序)时需要。
numpy.dtype(object, align=False, copy=False)object 参数可以是:
- NumPy 类型对象(例如
np.int32)。 - 类型字符代码字符串(例如,
'i4'表示 4 字节整数,'f8'表示 8 字节浮点数,'U'表示 Unicode 字符串)。 - 指定类型及字节序的字符串(例如,
'>i4'表示大端序 4 字节整数,'<f8'表示小端序 8 字节浮点数)。=表示本机字节序。 - 用于结构化数组的元组列表(见下文)。
示例 1:使用类型对象和字符代码
Section titled “示例 1:使用类型对象和字符代码”import numpy as np
# Using a NumPy type objectdt1 = np.dtype(np.int32)print(f"dtype from np.int32: {dt1}")
# Using a character code string ('i4' = 4-byte integer)dt2 = np.dtype('i4')print(f"dtype from 'i4': {dt2}")
# Specifying big-endian byte orderdt3 = np.dtype('>i4') # '>' means big-endianprint(f"dtype with big-endian: {dt3}")输出:
dtype from np.int32: int32dtype from 'i4': int32dtype with big-endian: >i4结构化数据类型
Section titled “结构化数据类型”结构化数组允许您定义这样的数组:其中每个元素都像 C 语言中的结构体(struct)一样,包含多个命名字段,这些字段可能具有不同的数据类型。这对于表示异构数据记录非常有用。
使用元组列表来定义结构化 dtype,其中每个元组的格式是 ('field_name', field_dtype)。
示例 2:定义和使用结构化数组
Section titled “示例 2:定义和使用结构化数组”import numpy as np
# 定义学生记录的结构化 dtype# 'U20' = 最大长度为 20 的 Unicode 字符串# 'i1' = 1 字节整数 (int8)# 'f4' = 4 字节浮点数 (float32)student_dtype = np.dtype([ ('name', 'U20'), ('age', np.int8), ('marks', np.float32)])print(f"Structured dtype: {student_dtype}\n")
# 使用此结构化 dtype 创建一个数组students = np.array([ ('Alice', 21, 85.5), ('Bob', 19, 92.0), ('Charlie', 22, 78.5)], dtype=student_dtype)
print(f"Structured array:\n{students}\n")
# 按字段名访问数据(返回的是视图,不是副本)ages = students['age']print(f"Ages: {ages}")print(f"Type of 'ages': {type(ages)}")print(f"dtype of 'ages': {ages.dtype}")
# 访问特定记录的数据print(f"\nFirst student's name: {students[0]['name']}")输出:
Structured dtype: [('name', '<U20'), ('age', 'i1'), ('marks', '<f4')]
Structured array:[('Alice', 21, 85.5) ('Bob', 19, 92. ) ('Charlie', 22, 78.5)]
Ages: [21 19 22]Type of 'ages': <class 'numpy.ndarray'>dtype of 'ages': int8
First student's name: Alice注意:虽然结构化数组功能强大,但对于处理具有异构类型的通用表格数据,Pandas 库(基于 NumPy 构建)及其 DataFrame 对象通常更方便、功能更丰富。
更多详细信息,请参阅 NumPy 数据类型文档。