Python Pandas - Series
Pandas - Series
Section titled “Pandas - Series”Pandas 的 Series 是一种基础的一维带标签(labeled)数组结构。它可以容纳各种类型的数据(整数、字符串、浮点数、Python 对象等)。可以将其视为电子表格中的单列,或者带有相关联**索引(index)**的更强大的 Python 列表或 NumPy 数组。
轴标签(axis labels)共同构成索引,为每个数据点提供了有意义的名称或标识符。
pandas.Series 构造函数
Section titled “pandas.Series 构造函数”您可以使用 pd.Series() 构造函数创建 Series:
pandas.Series(data=None, index=None, dtype=None, name=None, copy=False)关键参数:
| 参数 | 描述 |
|---|---|
data | 要存储在 Series 中的数据。可以是 Python 列表、NumPy 数组、字典、标量值或另一个 Series。 |
index | 数据的索引标签。必须是可哈希的(hashable),并且与 data 长度相同。如果为 None(默认),且 data 是类似数组(array-like)的,则会创建一个 RangeIndex(0, 1, 2, …)。如果 data 是字典,除非指定了 index,否则字典的键将用作索引。 |
dtype | Series 的期望数据类型(例如 'int64'、'float64'、'object'、'category'、'datetime64[ns]')。如果为 None,则从 data 推断类型。 |
name | Series 的可选名称属性。 |
copy | 如果为 True,总是复制输入的 data。默认为 False(可能复制或使用引用,取决于输入)。 |
您可以从各种输入创建 Series:
- Python 列表或 NumPy 数组
- Python 字典
- 标量值(需要指定索引)
创建 Series
Section titled “创建 Series”创建一个空 Series
Section titled “创建一个空 Series”您可以创建一个空 Series,但通常需要指定 dtype 以避免歧义。
# 导入 pandasimport pandas as pd
# 创建一个空 Series(指定 dtype)s_empty = pd.Series(dtype='object')print(s_empty)输出:
Series([], dtype: object)(注意:对于未指定 dtype 的完全空 Series,其默认 dtype 可能因 Pandas 版本或上下文而异,历史上常默认为 float64,但在明确创建用于稍后存储非数值数据时,使用 object 更安全。)
从列表或 NumPy 数组创建
Section titled “从列表或 NumPy 数组创建”如果 data 是列表或 NumPy 数组,并且未提供 index,则会创建一个默认的整数索引 [0, 1, ..., n-1]。
import pandas as pdimport numpy as np
data_list = ['a', 'b', 'c', 'd']data_array = np.array(['a', 'b', 'c', 'd'])
# 从列表创建 Series(默认索引)s_from_list = pd.Series(data_list)print("Series from list (default index):")print(s_from_list)输出:
Series from list (default index):0 a1 b2 c3 ddtype: object您可以提供一个自定义索引:
# 从数组创建 Series(自定义索引)custom_index = [100, 101, 102, 103]s_custom_index = pd.Series(data_array, index=custom_index)print("\nSeries from array (custom index):")print(s_custom_index)输出:
Series from array (custom index):100 a101 b102 c103 ddtype: object如果 data 是字典,其键(keys)默认用作索引标签(对于 Python 3.7+,按插入顺序;对于旧版 Python/Pandas,按排序顺序)。字典的值(values)成为 Series 的数据。
import pandas as pd
data_dict = {'x': 10.0, 'y': 20.0, 'z': 30.0}
# 从字典创建 Series(键成为索引)s_from_dict = pd.Series(data_dict)print("Series from dict (default index):")print(s_from_dict)输出:
Series from dict (default index):x 10.0y 20.0z 30.0dtype: float64如果您从字典创建 Series 时提供了 index 参数,Pandas 将根据提供的索引标签对字典值进行对齐。缺失的标签将导致 NaN 值。
# 从字典创建 Series(指定索引)explicit_index = ['y', 'z', 'w', 'x'] # 'w' 不在字典键中s_dict_explicit_index = pd.Series(data_dict, index=explicit_index)print("\nSeries from dict (explicit index):")print(s_dict_explicit_index)输出:
Series from dict (explicit index):y 20.0z 30.0w NaNx 10.0dtype: float64注意:顺序遵循 explicit_index。‘w’ 的值为 NaN,因为 data_dict 中没有 ‘w’ 这个键。
从标量值创建
Section titled “从标量值创建”如果 data 是单个标量值(例如数字或字符串),您必须提供一个 index。标量值将对索引中的每个标签重复。
import pandas as pd
# 从标量值创建 Seriess_from_scalar = pd.Series(5, index=[0, 1, 2, 3])print("Series from scalar:")print(s_from_scalar)输出:
Series from scalar:0 51 52 53 5dtype: int64访问 Series 中的数据
Section titled “访问 Series 中的数据”您可以使用类似于 NumPy 数组和 Python 字典的方法来访问 Series 中的数据,主要通过索引标签或整数位置。
**最佳实践:**为了清晰和一致,特别是在处理可能与位置混淆的整数索引时,建议使用 .loc[] 进行基于标签的访问,使用 .iloc[] 进行基于整数位置的访问,即使对于 Series 也是如此。
import pandas as pd
s = pd.Series([10, 20, 30, 40, 50], index=['a', 'b', 'c', 'd', 'e'])print("Sample Series:")print(s)按整数位置访问(.iloc[])
Section titled “按整数位置访问(.iloc[])”使用 .iloc[] 配合整数索引(基于 0)进行访问。
# 获取第一个元素(位置 0)print("\nFirst element (position 0):")print(s.iloc[0])
# 获取前三个元素(切片 0:3)print("\nFirst three elements (slice 0:3):")print(s.iloc[0:3]) # 或 s.iloc[:3]
# 获取后三个元素(从位置 -3 开始的切片)print("\nLast three elements (slice -3:):")print(s.iloc[-3:])输出:
Sample Series:a 10b 20c 30d 40e 50dtype: int64
First element (position 0):10
First three elements (slice 0:3):a 10b 20c 30dtype: int64
Last three elements (slice -3:):c 30d 40e 50dtype: int64按标签访问(.loc[])
Section titled “按标签访问(.loc[])”使用 .loc[] 配合索引标签进行访问。
# 按标签 'a' 获取单个元素print("\nElement with label 'a':")print(s.loc['a'])
# 使用标签列表获取多个元素print("\nElements with labels 'a', 'c', 'd':")print(s.loc[['a', 'c', 'd']])
# 使用标签切片获取范围内的元素(包含)print("\nElements from label 'b' to 'd' (inclusive):")print(s.loc['b':'d'])输出:
Element with label 'a':10
Elements with labels 'a', 'c', 'd':a 10c 30d 40dtype: int64
Elements from label 'b' to 'd' (inclusive):b 20c 30d 40dtype: int64直接索引([])
Section titled “直接索引([])”在 Series 上直接使用 [] 索引时,它主要尝试先进行标签查找。如果索引包含整数,这有时可能产生歧义。使用整数进行切片通常指位置。
# 直接索引(主要尝试标签查找)print("\nDirect access label 'a':", s['a'])
# 直接切片(通常基于位置)print("\nDirect slice [:3]:")print(s[:3])
# 如果找不到标签,会引发 KeyErrortry: print(s['f'])except KeyError as e: print(f"\n访问标签 'f' 出错:{e}")输出:
Direct access label 'a': 10
Direct slice [:3]:a 10b 20c 30dtype: int64
Error accessing label 'f': 'f'虽然直接使用 [] 访问 Series 很常见,但明确使用 .loc 和 .iloc 可以使您的代码更具可读性,并减少出错的可能性,尤其是在处理整数标签或在 Series 和 DataFrame 之间转换时。