Skip to content

Python Pandas - Series

Pandas 的 Series 是一种基础的一维带标签(labeled)数组结构。它可以容纳各种类型的数据(整数、字符串、浮点数、Python 对象等)。可以将其视为电子表格中的单列,或者带有相关联**索引(index)**的更强大的 Python 列表或 NumPy 数组。

轴标签(axis labels)共同构成索引,为每个数据点提供了有意义的名称或标识符。

您可以使用 pd.Series() 构造函数创建 Series:

pandas.Series(data=None, index=None, dtype=None, name=None, copy=False)

关键参数:

参数描述
data要存储在 Series 中的数据。可以是 Python 列表、NumPy 数组、字典、标量值或另一个 Series。
index数据的索引标签。必须是可哈希的(hashable),并且与 data 长度相同。如果为 None(默认),且 data 是类似数组(array-like)的,则会创建一个 RangeIndex(0, 1, 2, …)。如果 data 是字典,除非指定了 index,否则字典的键将用作索引。
dtypeSeries 的期望数据类型(例如 'int64'、'float64'、'object'、'category'、'datetime64[ns]')。如果为 None,则从 data 推断类型。
nameSeries 的可选名称属性。
copy如果为 True,总是复制输入的 data。默认为 False(可能复制或使用引用,取决于输入)。

您可以从各种输入创建 Series:

  • Python 列表或 NumPy 数组
  • Python 字典
  • 标量值(需要指定索引)

您可以创建一个空 Series,但通常需要指定 dtype 以避免歧义。

# 导入 pandas
import pandas as pd
# 创建一个空 Series(指定 dtype)
s_empty = pd.Series(dtype='object')
print(s_empty)

输出:

Series([], dtype: object)

(注意:对于未指定 dtype 的完全空 Series,其默认 dtype 可能因 Pandas 版本或上下文而异,历史上常默认为 float64,但在明确创建用于稍后存储非数值数据时,使用 object 更安全。)

如果 data 是列表或 NumPy 数组,并且未提供 index,则会创建一个默认的整数索引 [0, 1, ..., n-1]。

import pandas as pd
import numpy as np
data_list = ['a', 'b', 'c', 'd']
data_array = np.array(['a', 'b', 'c', 'd'])
# 从列表创建 Series(默认索引)
s_from_list = pd.Series(data_list)
print("Series from list (default index):")
print(s_from_list)

输出:

Series from list (default index):
0 a
1 b
2 c
3 d
dtype: object

您可以提供一个自定义索引:

# 从数组创建 Series(自定义索引)
custom_index = [100, 101, 102, 103]
s_custom_index = pd.Series(data_array, index=custom_index)
print("\nSeries from array (custom index):")
print(s_custom_index)

输出:

Series from array (custom index):
100 a
101 b
102 c
103 d
dtype: object

如果 data 是字典,其键(keys)默认用作索引标签(对于 Python 3.7+,按插入顺序;对于旧版 Python/Pandas,按排序顺序)。字典的值(values)成为 Series 的数据。

import pandas as pd
data_dict = {'x': 10.0, 'y': 20.0, 'z': 30.0}
# 从字典创建 Series(键成为索引)
s_from_dict = pd.Series(data_dict)
print("Series from dict (default index):")
print(s_from_dict)

输出:

Series from dict (default index):
x 10.0
y 20.0
z 30.0
dtype: float64

如果您从字典创建 Series 时提供了 index 参数,Pandas 将根据提供的索引标签对字典值进行对齐。缺失的标签将导致 NaN 值。

# 从字典创建 Series(指定索引)
explicit_index = ['y', 'z', 'w', 'x'] # 'w' 不在字典键中
s_dict_explicit_index = pd.Series(data_dict, index=explicit_index)
print("\nSeries from dict (explicit index):")
print(s_dict_explicit_index)

输出:

Series from dict (explicit index):
y 20.0
z 30.0
w NaN
x 10.0
dtype: float64

注意:顺序遵循 explicit_index。‘w’ 的值为 NaN,因为 data_dict 中没有 ‘w’ 这个键。

如果 data 是单个标量值(例如数字或字符串),您必须提供一个 index。标量值将对索引中的每个标签重复。

import pandas as pd
# 从标量值创建 Series
s_from_scalar = pd.Series(5, index=[0, 1, 2, 3])
print("Series from scalar:")
print(s_from_scalar)

输出:

Series from scalar:
0 5
1 5
2 5
3 5
dtype: int64

您可以使用类似于 NumPy 数组和 Python 字典的方法来访问 Series 中的数据,主要通过索引标签或整数位置。

**最佳实践:**为了清晰和一致,特别是在处理可能与位置混淆的整数索引时,建议使用 .loc[] 进行基于标签的访问,使用 .iloc[] 进行基于整数位置的访问,即使对于 Series 也是如此。

import pandas as pd
s = pd.Series([10, 20, 30, 40, 50], index=['a', 'b', 'c', 'd', 'e'])
print("Sample Series:")
print(s)

使用 .iloc[] 配合整数索引(基于 0)进行访问。

# 获取第一个元素(位置 0)
print("\nFirst element (position 0):")
print(s.iloc[0])
# 获取前三个元素(切片 0:3)
print("\nFirst three elements (slice 0:3):")
print(s.iloc[0:3]) # 或 s.iloc[:3]
# 获取后三个元素(从位置 -3 开始的切片)
print("\nLast three elements (slice -3:):")
print(s.iloc[-3:])

输出:

Sample Series:
a 10
b 20
c 30
d 40
e 50
dtype: int64
First element (position 0):
10
First three elements (slice 0:3):
a 10
b 20
c 30
dtype: int64
Last three elements (slice -3:):
c 30
d 40
e 50
dtype: int64

使用 .loc[] 配合索引标签进行访问。

# 按标签 'a' 获取单个元素
print("\nElement with label 'a':")
print(s.loc['a'])
# 使用标签列表获取多个元素
print("\nElements with labels 'a', 'c', 'd':")
print(s.loc[['a', 'c', 'd']])
# 使用标签切片获取范围内的元素(包含)
print("\nElements from label 'b' to 'd' (inclusive):")
print(s.loc['b':'d'])

输出:

Element with label 'a':
10
Elements with labels 'a', 'c', 'd':
a 10
c 30
d 40
dtype: int64
Elements from label 'b' to 'd' (inclusive):
b 20
c 30
d 40
dtype: int64

在 Series 上直接使用 [] 索引时,它主要尝试先进行标签查找。如果索引包含整数,这有时可能产生歧义。使用整数进行切片通常指位置。

# 直接索引(主要尝试标签查找)
print("\nDirect access label 'a':", s['a'])
# 直接切片(通常基于位置)
print("\nDirect slice [:3]:")
print(s[:3])
# 如果找不到标签,会引发 KeyError
try:
print(s['f'])
except KeyError as e:
print(f"\n访问标签 'f' 出错:{e}")

输出:

Direct access label 'a': 10
Direct slice [:3]:
a 10
b 20
c 30
dtype: int64
Error accessing label 'f': 'f'

虽然直接使用 [] 访问 Series 很常见,但明确使用 .loc 和 .iloc 可以使您的代码更具可读性,并减少出错的可能性,尤其是在处理整数标签或在 Series 和 DataFrame 之间转换时。