Skip to content

Python Pandas - DataFrame

DataFrame 是 Pandas 中的主力数据结构。它是一个二维、大小可变且潜在异构的表格型数据结构,具有带标签的轴(行和列)。你可以将其视为电子表格、SQL 表或 Series 对象的字典。

  • 列可以持有不同的数据类型(整数、字符串、浮点数、布尔值等)。
  • 行和列都有标签,分别称为 Index(索引)和 Columns(列)。
  • 大小可变:你可以添加或删除列和行。
  • 支持对行和列进行算术运算。
  • 高效处理缺失数据(表示为 NaN)。
  • 提供强大的工具用于切片(slicing)、切块(dicing)、分组(grouping)、合并(merging)和重塑数据(reshaping data)。

想象一个包含学生数据的表格:

行可能由学生 ID(索引)标识,列可能代表 ‘Name’(姓名)、‘Age’(年龄)、‘Major’(专业)、‘GPA’(平均绩点)(列)。每列持有的数据类型可能不同。

Name Age Major GPA
Index
101 Alice 20 Economics 3.8
102 Bob 21 Physics 3.5
103 Charlie 19 History 3.9
...

创建 DataFrame 的主要方式是使用 pandas.DataFrame 构造函数:

pandas.DataFrame(data=None, index=None, columns=None, dtype=None, copy=None)

构造函数参数:

参数描述
data用于填充 DataFrame 的数据。可以是多种类型:NumPy ndarray、字典、列表、Series、另一个 DataFrame 等。
index行的标签。可选;如果未提供,默认为 RangeIndex(0、1、2…)。
columns列的标签。可选;从输入数据推断(例如,字典键),或者如果未提供且无法推断,默认为 RangeIndex。
dtype强制所有列使用特定的数据类型。如果为 None,则按列推断类型。
copy在最新 Pandas 版本中已弃用。之前控制是否复制输入数据(默认为 False)。现在数据复制行为更微妙;请参阅 Pandas 文档了解详细信息。

DataFrame 可以从各种常用的数据结构创建:

  • 列表或 NumPy 数组的字典
  • 字典列表
  • Series 字典
  • 列表的列表或 NumPy 数组
  • 另一个 DataFrame

你可以创建一个空的 DataFrame,稍后可以填充数据。

# 导入 pandas 并将其别名为 pd(标准约定)
import pandas as pd
empty_df = pd.DataFrame()
print(empty_df)

输出:

Empty DataFrame
Columns: []
Index: []

你可以从简单的列表(产生单列)或列表的列表(其中每个内部列表成为一行)创建 DataFrame。

import pandas as pd
data_list_simple = [10, 20, 30, 40, 50]
df_from_list = pd.DataFrame(data_list_simple, columns=['Numbers'])
print(df_from_list)

输出:

Numbers
0 10
1 20
2 30
3 40
4 50
import pandas as pd
data_list_of_lists = [
['Alex', 25, 'Software Engineer'],
['Bob', 30, 'Data Scientist'],
['Clarke', 22, 'Web Developer']
]
# 提供列名
df_from_lol = pd.DataFrame(data_list_of_lists,
columns=['Name', 'Age', 'Job Title'])
print(df_from_lol)

输出:

Name Age Job Title
0 Alex 25 Software Engineer
1 Bob 30 Data Scientist
2 Clarke 22 Web Developer
import pandas as pd
import numpy as np # 用于 np.float64
data_list_of_lists = [
['Alex', 10],
['Bob', 12],
['Clarke', 13]
]
# 在创建时指定 'Age' 列的 dtype(通常在创建后更常用)
# 或者为所有适用的列指定单一 dtype
df_typed = pd.DataFrame(data_list_of_lists,
columns=['Name', 'Age'],
dtype=np.float64) # 在可能的情况下应用 float64
print(df_typed)
print('\n数据类型:')
print(df_typed.dtypes)

输出:

Name Age
0 Alex 10.0
1 Bob 12.0
2 Clarke 13.0
Data Types:
Name object # Name 无法转换为 float,保持为 object
Age float64
dtype: object

注意: dtype 参数会尝试强制转换列。如果列无法转换(例如 ‘Name’ 无法转换为 float),它会保留其原始推断类型(字符串的 ‘object’)。通常更实用的是在创建后使用 .astype() 调整 dtype。

这是创建 DataFrame 的一种非常常见且通常高效的方式。字典键成为列名,列表或数组成为列数据。所有列表/数组必须长度相同。

如果未提供 index,则会创建默认的 RangeIndex(0、1、2…)。

import pandas as pd
data_dict = {
'Student Name': ['Tom', 'Jack', 'Steve', 'Ricky'],
'Score': [85, 92, 78, 88]
}
df_from_dict = pd.DataFrame(data_dict)
print(df_from_dict)

输出:

Student Name Score
0 Tom 85
1 Jack 92
2 Steve 78
3 Ricky 88

注意: 自动分配了默认索引 0、1、2、3。

import pandas as pd
data_dict = {
'Student Name': ['Tom', 'Jack', 'Steve', 'Ricky'],
'Score': [85, 92, 78, 88]
}
custom_index = ['ID_101', 'ID_102', 'ID_103', 'ID_104']
df_indexed = pd.DataFrame(data_dict, index=custom_index)
print(df_indexed)

输出:

Student Name Score
ID_101 Tom 85
ID_102 Jack 92
ID_103 Steve 78
ID_104 Ricky 88

注意: index 参数为行分配了有意义的标签。

列表中的每个字典代表一行。字典键会自动推断为列名。如果某个字典缺少其他字典中存在的键,则会在该单元格中放置 NaN(非数字)。

import pandas as pd
data_list_of_dicts = [
{'col_a': 1, 'col_b': 2},
{'col_a': 5, 'col_b': 10, 'col_c': 20}, # 这个字典包含 'col_c'
{'col_a': 9, 'col_b': 15} # 这个字典不包含
]
df_from_lod = pd.DataFrame(data_list_of_dicts)
print(df_from_lod)

输出:

col_a col_b col_c
0 1 2 NaN
1 5 10 20.0
2 9 15 NaN

注意: 在数据缺失的地方(例如第一行和第三行的 ‘col_c’),会自动插入 NaN。Pandas 会推断出 ‘col_c’ 最适合的数据类型(这里是 float64,因为 NaN 是一个浮点数)。

import pandas as pd
data_list_of_dicts = [
{'col_a': 1, 'col_b': 2},
{'col_a': 5, 'col_b': 10, 'col_c': 20}
]
row_labels = ['Row_1', 'Row_2']
df_lod_indexed = pd.DataFrame(data_list_of_dicts, index=row_labels)
print(df_lod_indexed)

输出:

col_a col_b col_c
Row_1 1 2 NaN
Row_2 5 10 20.0

你可以明确指定列及其顺序。如果指定列在字典中不存在对应的键,则会用 NaN 填充。

import pandas as pd
data_list_of_dicts = [
{'col_a': 1, 'col_b': 2},
{'col_a': 5, 'col_b': 10, 'col_c': 20}
]
row_labels = ['Row_1', 'Row_2']
# 选择现有列并重新排序
df1 = pd.DataFrame(data_list_of_dicts, index=row_labels, columns=['col_b', 'col_a'])
# 选择现有列和一个不存在的列 ('col_d')
df2 = pd.DataFrame(data_list_of_dicts, index=row_labels, columns=['col_a', 'col_c', 'col_d'])
print("DataFrame df1 (选择了现有列):")
print(df1)
print("\nDataFrame df2 (包含不存在的列 'col_d'):")
print(df2)

输出:

DataFrame df1 (Selected existing columns):
col_b col_a
Row_1 2 1
Row_2 10 5
DataFrame df2 (Includes non-existing column 'col_d'):
col_a col_c col_d
Row_1 1 NaN NaN
Row_2 5 20.0 NaN

注意: df1 按照指定顺序显示列 ‘col_b’ 和 ‘col_a’。df2 包含 ‘col_a’ 和 ‘col_c’(缺失处填充 NaN),并添加了一个新列 ‘col_d’,因为它不在输入字典中,所以完全填充了 NaN。

字典键成为列名,Series 值成为列数据。生成的 DataFrame 的索引将是输入 Series 中所有索引的并集。索引不一致的地方,缺失值将填充 NaN。

import pandas as pd
data_dict_of_series = {
'ColOne': pd.Series([1.0, 2.0, 3.0], index=['a', 'b', 'c']),
'ColTwo': pd.Series([1.0, 2.0, 3.0, 4.0], index=['a', 'b', 'c', 'd'])
}
df_from_dos = pd.DataFrame(data_dict_of_series)
print(df_from_dos)

输出:

ColOne ColTwo
a 1.0 1.0
b 2.0 2.0
c 3.0 3.0
d NaN 4.0

注意: 最终索引是 (‘a’, ‘b’, ‘c’, ‘d’)。由于 ‘ColOne’ 没有索引 ‘d’,该单元格填充了 NaN。

选择、添加和删除列是基本的 DataFrame 操作。

使用类似字典的方括号表示法 [] 或属性访问 . 访问列(如果列名是有效的 Python 标识符且与 DataFrame 方法不冲突)。

# 使用上一个示例中的 df_from_dos
print("使用 [] 选择 'ColOne':")
print(df_from_dos['ColOne'])
# print("\n使用属性访问选择 'ColTwo':") # 谨慎使用
# print(df_from_dos.ColTwo)
print("\n选择多列 [['ColTwo', 'ColOne']]:")
print(df_from_dos[['ColTwo', 'ColOne']]) # 传入一个列名列表

输出:

Selecting 'ColOne' using []:
a 1.0
b 2.0
c 3.0
d NaN
Name: ColOne, dtype: float64
Selecting multiple columns [['ColTwo', 'ColOne']]:
ColTwo ColOne
a 1.0 1.0
b 2.0 2.0
c 3.0 3.0
d 4.0 NaN

通过将数据赋给新的列标签来添加新列。

# 使用 df_from_dos
print("原始 DataFrame:")
print(df_from_dos)
# 通过赋给一个 Series 来添加新列(索引对齐)
print("\n通过赋给一个 Series 添加 'ColThree':")
df_from_dos['ColThree'] = pd.Series([10, 20, 30], index=['a', 'b', 'c'])
print(df_from_dos)
# 通过现有列计算添加新列
print("\n添加 'ColFour' = ColOne + ColThree:")
df_from_dos['ColFour'] = df_from_dos['ColOne'] + df_from_dos['ColThree']
print(df_from_dos)
# 添加一个带有标量值的新列(广播到所有行)
print("\n添加带有标量值的 'ColFive':")
df_from_dos['ColFive'] = 100
print(df_from_dos)

输出:

Original DataFrame:
ColOne ColTwo
a 1.0 1.0
b 2.0 2.0
c 3.0 3.0
d NaN 4.0
Adding 'ColThree' by assigning a Series:
ColOne ColTwo ColThree
a 1.0 1.0 10.0
b 2.0 2.0 20.0
c 3.0 3.0 30.0
d NaN 4.0 NaN
Adding 'ColFour' = ColOne + ColThree:
ColOne ColTwo ColThree ColFour
a 1.0 1.0 10.0 11.0
b 2.0 2.0 20.0 22.0
c 3.0 3.0 30.0 33.0
d NaN 4.0 NaN NaN
Adding 'ColFive' with a scalar value:
ColOne ColTwo ColThree ColFour ColFive
a 1.0 1.0 10.0 11.0 100
b 2.0 2.0 20.0 22.0 100
c 3.0 3.0 30.0 33.0 100
d NaN 4.0 NaN NaN 100

使用 del 关键字或 .pop() 或 .drop() 方法。

# 使用上一步骤中的 df_from_dos
print("删除前的 DataFrame:")
print(df_from_dos)
# 使用 del(原地修改)
print("\n使用 del 删除 'ColOne':")
del df_from_dos['ColOne']
print(df_from_dos)
# 使用 pop(移除并返回该列)
print("\n使用 pop 删除 'ColTwo':")
popped_col = df_from_dos.pop('ColTwo')
print("pop 后的 DataFrame:")
print(df_from_dos)
print("\n弹出的列 'ColTwo':")
print(popped_col)
# 使用 drop(返回一个新的 DataFrame,使用 inplace=True 修改原始 DataFrame)
print("\n使用 drop 删除 'ColThree'(返回新 DataFrame):")
df_after_drop = df_from_dos.drop(columns=['ColThree']) # 使用 'columns' 参数
# 或者: df_after_drop = df_from_dos.drop(['ColThree'], axis=1)
print(df_after_drop)
print("\n执行非原地 drop 后的原始 DataFrame(未改变):")
print(df_from_dos)

输出:

DataFrame before deletion:
ColOne ColTwo ColThree ColFour ColFive
a 1.0 1.0 10.0 11.0 100
b 2.0 2.0 20.0 22.0 100
c 3.0 3.0 30.0 33.0 100
d NaN 4.0 NaN NaN 100
Deleting 'ColOne' using del:
ColTwo ColThree ColFour ColFive
a 1.0 10.0 11.0 100
b 2.0 20.0 22.0 100
c 3.0 30.0 33.0 100
d 4.0 NaN NaN 100
Deleting 'ColTwo' using pop:
DataFrame after pop:
ColThree ColFour ColFive
a 10.0 11.0 100
b 20.0 22.0 100
c 30.0 33.0 100
d NaN NaN 100
Popped Column 'ColTwo':
a 1.0
b 2.0
c 3.0
d 4.0
Name: ColTwo, dtype: float64
Deleting 'ColThree' using drop (returns new DataFrame):
ColFour ColFive
a 11.0 100
b 22.0 100
c 33.0 100
d NaN 100
Original DataFrame after non-inplace drop (unchanged):
ColThree ColFour ColFive
a 10.0 11.0 100
b 20.0 22.0 100
c 30.0 33.0 100
d NaN NaN 100

del 和 .pop() 会原地修改原始 DataFrame,而 .drop() 默认返回一个新 DataFrame。你可以使用 inplace=True 使 .drop() 也原地修改。

选择、添加和删除行涉及使用索引。

使用 .loc[] 进行基于标签的选择,使用 .iloc[] 进行基于整数位置的选择。

import pandas as pd
data_dict = {
'ColOne': pd.Series([1.0, 2.0, 3.0], index=['a', 'b', 'c']),
'ColTwo': pd.Series([1.0, 2.0, 3.0, 4.0], index=['a', 'b', 'c', 'd'])
}
df = pd.DataFrame(data_dict)
print("DataFrame:")
print(df)
# 使用 .loc['b'] 选择标签为 'b' 的行
print("\n使用 .loc['b'] 选择标签为 'b' 的行:")
print(df.loc['b'])
# 使用 .iloc[2] 选择位置 2 的行(标签 'c')
print("\n使用 .iloc[2] 选择位置 2 的行(标签 'c'):")
print(df.iloc[2])
# 按标签选择多行
print("\n使用 .loc[['a', 'd']] 选择标签为 'a' 和 'd' 的行:")
print(df.loc[['a', 'd']])
# 按位置选择多行
print("\n使用 .iloc[[0, 3]] 选择位置 0 和 3 的行:")
print(df.iloc[[0, 3]])

输出:

DataFrame:
ColOne ColTwo
a 1.0 1.0
b 2.0 2.0
c 3.0 3.0
d NaN 4.0
Selecting row with label 'b' using .loc['b']:
ColOne 2.0
ColTwo 2.0
Name: b, dtype: float64
Selecting row at position 2 (label 'c') using .iloc[2]:
ColOne 3.0
ColTwo 3.0
Name: c, dtype: float64
Selecting rows 'a' and 'd' using .loc[['a', 'd']]:
ColOne ColTwo
a 1.0 1.0
d NaN 4.0
Selecting rows at positions 0 and 3 using .iloc[[0, 3]]:
ColOne ColTwo
a 1.0 1.0
d NaN 4.0

选择单行返回一个 Series,而选择多行返回一个 DataFrame。

使用 .loc(基于标签,终点包含在内)或 .iloc(基于整数位置,终点不包含在内)进行切片。

# 使用上一个示例中的 df
# 按标签切片行(包含终点)
print("\n使用 .loc['b':'d'] 切片标签为 'b' 到 'd' 的行:")
print(df.loc['b':'d'])
# 按位置切片行(不包含终点)
print("\n使用 .iloc[1:3] 切片位置 1 到(不包括)3 的行:")
print(df.iloc[1:3])
# 直接对 DataFrame 进行整数切片 (df[1:3]) 也是可能的,但通常
# 不推荐;推荐使用显式的 .iloc 以保持清晰性和鲁棒性。
# print("\n使用 df[1:3] 进行切片(不推荐):")
# print(df[1:3])

输出:

Slicing rows 'b' through 'd' using .loc['b':'d']:
ColOne ColTwo
b 2.0 2.0
c 3.0 3.0
d NaN 4.0
Slicing rows from position 1 up to (not including) 3 using .iloc[1:3]:
ColOne ColTwo
b 2.0 2.0
c 3.0 3.0

添加行的推荐方法是使用 pd.concat()。较旧的 .append() 方法已在 Pandas 2.0 中弃用并移除,不应使用。

import pandas as pd
df1 = pd.DataFrame([[1, 2], [3, 4]], columns=['A', 'B'], index=[0, 1])
df2 = pd.DataFrame([[5, 6], [7, 8]], columns=['A', 'B'], index=[2, 3]) # 使用不同的索引
df3 = pd.DataFrame([[9, 10], [11, 12]], columns=['A', 'C'], index=[4, 5]) # 不同的列
print("DataFrame df1:")
print(df1)
print("\nDataFrame df2:")
print(df2)
print("\nDataFrame df3:")
print(df3)
# 连接 df1 和 df2(相同列)
combined_df = pd.concat([df1, df2])
print("\n使用 pd.concat() 连接 df1 和 df2:")
print(combined_df)
# 连接 df1 和 df3(不同列,按索引对齐)
combined_df_cols = pd.concat([df1, df3]) # 默认使用外连接
print("\n使用 pd.concat() 连接 df1 和 df3(外连接):")
print(combined_df_cols)
# 如果使用较旧的示例(之前可能由于 append 产生重复索引):
# df_old_append = pd.DataFrame([[1, 2], [3, 4]], columns=['A', 'B'], index=[0, 1])
# df_old_append2 = pd.DataFrame([[5, 6], [7, 8]], columns=['A', 'B'], index=[0, 1])
# combined_dup_index = pd.concat([df_old_append, df_old_append2], ignore_index=True)
# print("\n使用 ignore_index=True 连接以重置索引:")
# print(combined_dup_index)

输出:

DataFrame df1:
A B
0 1 2
1 3 4
DataFrame df2:
A B
2 5 6
3 7 8
DataFrame df3:
A C
4 9 10
5 11 12
Combined df1 and df2 using pd.concat():
A B
0 1 2
1 3 4
2 5 6
3 7 8
Combined df1 and df3 using pd.concat() (outer join):
A B C
0 1.0 2.0 NaN
1 3.0 4.0 NaN
4 9.0 NaN 10.0
5 11.0 NaN 12.0

pd.concat 智能地组合 DataFrame。如果你想在结果上获得一个全新的默认整数索引,请使用 ignore_index=True。

使用 .drop() 方法,指定要移除的行标签(索引值)并设置 axis=0(或省略 axis,因为它默认为 0)。

# 使用上一个示例中的 combined_df
print("删除行前的 DataFrame:")
print(combined_df)
# 删除标签为 0 和 3 的行
# 默认情况下,drop 返回一个新的 DataFrame。使用 inplace=True 进行修改。
df_dropped = combined_df.drop([0, 3], axis=0)
print("\n删除行 0 和 3 后的 DataFrame:")
print(df_dropped)
# 具有重复标签的示例(如果之前这样创建)
# df_dup = pd.DataFrame([[1,2],[3,4],[5,6]], columns=['A','B'], index=[0,1,0])
# print("\n具有重复索引 0 的 DataFrame:")
# print(df_dup)
# print("\n删除标签 0(移除所有具有该标签的行):")
# print(df_dup.drop(0))

输出:

DataFrame before row deletion:
A B
0 1 2
1 3 4
2 5 6
3 7 8
DataFrame after dropping rows 0 and 3:
A B
1 3 4
2 5 6

如果索引标签重复,drop() 将移除所有匹配该标签的行。