Python Pandas - DataFrame
Pandas - DataFrame
Section titled “Pandas - DataFrame”DataFrame 是 Pandas 中的主力数据结构。它是一个二维、大小可变且潜在异构的表格型数据结构,具有带标签的轴(行和列)。你可以将其视为电子表格、SQL 表或 Series 对象的字典。
DataFrame 的主要特性
Section titled “DataFrame 的主要特性”- 列可以持有不同的数据类型(整数、字符串、浮点数、布尔值等)。
- 行和列都有标签,分别称为 Index(索引)和 Columns(列)。
- 大小可变:你可以添加或删除列和行。
- 支持对行和列进行算术运算。
- 高效处理缺失数据(表示为
NaN)。 - 提供强大的工具用于切片(slicing)、切块(dicing)、分组(grouping)、合并(merging)和重塑数据(reshaping data)。
想象一个包含学生数据的表格:
行可能由学生 ID(索引)标识,列可能代表 ‘Name’(姓名)、‘Age’(年龄)、‘Major’(专业)、‘GPA’(平均绩点)(列)。每列持有的数据类型可能不同。
Name Age Major GPAIndex101 Alice 20 Economics 3.8102 Bob 21 Physics 3.5103 Charlie 19 History 3.9...创建 DataFrame: pd.DataFrame()
Section titled “创建 DataFrame: pd.DataFrame()”创建 DataFrame 的主要方式是使用 pandas.DataFrame 构造函数:
pandas.DataFrame(data=None, index=None, columns=None, dtype=None, copy=None)构造函数参数:
| 参数 | 描述 |
|---|---|
| data | 用于填充 DataFrame 的数据。可以是多种类型:NumPy ndarray、字典、列表、Series、另一个 DataFrame 等。 |
| index | 行的标签。可选;如果未提供,默认为 RangeIndex(0、1、2…)。 |
| columns | 列的标签。可选;从输入数据推断(例如,字典键),或者如果未提供且无法推断,默认为 RangeIndex。 |
| dtype | 强制所有列使用特定的数据类型。如果为 None,则按列推断类型。 |
| copy | 在最新 Pandas 版本中已弃用。之前控制是否复制输入数据(默认为 False)。现在数据复制行为更微妙;请参阅 Pandas 文档了解详细信息。 |
创建 DataFrame 的方法
Section titled “创建 DataFrame 的方法”DataFrame 可以从各种常用的数据结构创建:
- 列表或 NumPy 数组的字典
- 字典列表
- Series 字典
- 列表的列表或 NumPy 数组
- 另一个 DataFrame
创建空 DataFrame
Section titled “创建空 DataFrame”你可以创建一个空的 DataFrame,稍后可以填充数据。
示例: 空 DataFrame
Section titled “示例: 空 DataFrame”# 导入 pandas 并将其别名为 pd(标准约定)import pandas as pd
empty_df = pd.DataFrame()print(empty_df)输出:
Empty DataFrameColumns: []Index: []你可以从简单的列表(产生单列)或列表的列表(其中每个内部列表成为一行)创建 DataFrame。
示例 1: 从单个列表创建
Section titled “示例 1: 从单个列表创建”import pandas as pd
data_list_simple = [10, 20, 30, 40, 50]df_from_list = pd.DataFrame(data_list_simple, columns=['Numbers'])print(df_from_list)输出:
Numbers0 101 202 303 404 50示例 2: 从列表的列表创建
Section titled “示例 2: 从列表的列表创建”import pandas as pd
data_list_of_lists = [ ['Alex', 25, 'Software Engineer'], ['Bob', 30, 'Data Scientist'], ['Clarke', 22, 'Web Developer']]
# 提供列名df_from_lol = pd.DataFrame(data_list_of_lists, columns=['Name', 'Age', 'Job Title'])print(df_from_lol)输出:
Name Age Job Title0 Alex 25 Software Engineer1 Bob 30 Data Scientist2 Clarke 22 Web Developer示例 3: 指定数据类型
Section titled “示例 3: 指定数据类型”import pandas as pdimport numpy as np # 用于 np.float64
data_list_of_lists = [ ['Alex', 10], ['Bob', 12], ['Clarke', 13]]
# 在创建时指定 'Age' 列的 dtype(通常在创建后更常用)# 或者为所有适用的列指定单一 dtypedf_typed = pd.DataFrame(data_list_of_lists, columns=['Name', 'Age'], dtype=np.float64) # 在可能的情况下应用 float64
print(df_typed)print('\n数据类型:')print(df_typed.dtypes)输出:
Name Age0 Alex 10.01 Bob 12.02 Clarke 13.0
Data Types:Name object # Name 无法转换为 float,保持为 objectAge float64dtype: object注意: dtype 参数会尝试强制转换列。如果列无法转换(例如 ‘Name’ 无法转换为 float),它会保留其原始推断类型(字符串的 ‘object’)。通常更实用的是在创建后使用 .astype() 调整 dtype。
从列表或数组字典创建
Section titled “从列表或数组字典创建”这是创建 DataFrame 的一种非常常见且通常高效的方式。字典键成为列名,列表或数组成为列数据。所有列表/数组必须长度相同。
如果未提供 index,则会创建默认的 RangeIndex(0、1、2…)。
示例 1: 基本创建
Section titled “示例 1: 基本创建”import pandas as pd
data_dict = { 'Student Name': ['Tom', 'Jack', 'Steve', 'Ricky'], 'Score': [85, 92, 78, 88]}
df_from_dict = pd.DataFrame(data_dict)print(df_from_dict)输出:
Student Name Score0 Tom 851 Jack 922 Steve 783 Ricky 88注意: 自动分配了默认索引 0、1、2、3。
示例 2: 提供自定义索引
Section titled “示例 2: 提供自定义索引”import pandas as pd
data_dict = { 'Student Name': ['Tom', 'Jack', 'Steve', 'Ricky'], 'Score': [85, 92, 78, 88]}custom_index = ['ID_101', 'ID_102', 'ID_103', 'ID_104']
df_indexed = pd.DataFrame(data_dict, index=custom_index)print(df_indexed)输出:
Student Name ScoreID_101 Tom 85ID_102 Jack 92ID_103 Steve 78ID_104 Ricky 88注意: index 参数为行分配了有意义的标签。
从字典列表创建
Section titled “从字典列表创建”列表中的每个字典代表一行。字典键会自动推断为列名。如果某个字典缺少其他字典中存在的键,则会在该单元格中放置 NaN(非数字)。
示例 1: 缺失数据
Section titled “示例 1: 缺失数据”import pandas as pd
data_list_of_dicts = [ {'col_a': 1, 'col_b': 2}, {'col_a': 5, 'col_b': 10, 'col_c': 20}, # 这个字典包含 'col_c' {'col_a': 9, 'col_b': 15} # 这个字典不包含]
df_from_lod = pd.DataFrame(data_list_of_dicts)print(df_from_lod)输出:
col_a col_b col_c0 1 2 NaN1 5 10 20.02 9 15 NaN注意: 在数据缺失的地方(例如第一行和第三行的 ‘col_c’),会自动插入 NaN。Pandas 会推断出 ‘col_c’ 最适合的数据类型(这里是 float64,因为 NaN 是一个浮点数)。
示例 2: 指定索引
Section titled “示例 2: 指定索引”import pandas as pd
data_list_of_dicts = [ {'col_a': 1, 'col_b': 2}, {'col_a': 5, 'col_b': 10, 'col_c': 20}]row_labels = ['Row_1', 'Row_2']
df_lod_indexed = pd.DataFrame(data_list_of_dicts, index=row_labels)print(df_lod_indexed)输出:
col_a col_b col_cRow_1 1 2 NaNRow_2 5 10 20.0示例 3: 指定列
Section titled “示例 3: 指定列”你可以明确指定列及其顺序。如果指定列在字典中不存在对应的键,则会用 NaN 填充。
import pandas as pd
data_list_of_dicts = [ {'col_a': 1, 'col_b': 2}, {'col_a': 5, 'col_b': 10, 'col_c': 20}]row_labels = ['Row_1', 'Row_2']
# 选择现有列并重新排序df1 = pd.DataFrame(data_list_of_dicts, index=row_labels, columns=['col_b', 'col_a'])
# 选择现有列和一个不存在的列 ('col_d')df2 = pd.DataFrame(data_list_of_dicts, index=row_labels, columns=['col_a', 'col_c', 'col_d'])
print("DataFrame df1 (选择了现有列):")print(df1)print("\nDataFrame df2 (包含不存在的列 'col_d'):")print(df2)输出:
DataFrame df1 (Selected existing columns): col_b col_aRow_1 2 1Row_2 10 5
DataFrame df2 (Includes non-existing column 'col_d'): col_a col_c col_dRow_1 1 NaN NaNRow_2 5 20.0 NaN注意: df1 按照指定顺序显示列 ‘col_b’ 和 ‘col_a’。df2 包含 ‘col_a’ 和 ‘col_c’(缺失处填充 NaN),并添加了一个新列 ‘col_d’,因为它不在输入字典中,所以完全填充了 NaN。
从 Series 字典创建
Section titled “从 Series 字典创建”字典键成为列名,Series 值成为列数据。生成的 DataFrame 的索引将是输入 Series 中所有索引的并集。索引不一致的地方,缺失值将填充 NaN。
import pandas as pd
data_dict_of_series = { 'ColOne': pd.Series([1.0, 2.0, 3.0], index=['a', 'b', 'c']), 'ColTwo': pd.Series([1.0, 2.0, 3.0, 4.0], index=['a', 'b', 'c', 'd'])}
df_from_dos = pd.DataFrame(data_dict_of_series)print(df_from_dos)输出:
ColOne ColTwoa 1.0 1.0b 2.0 2.0c 3.0 3.0d NaN 4.0注意: 最终索引是 (‘a’, ‘b’, ‘c’, ‘d’)。由于 ‘ColOne’ 没有索引 ‘d’,该单元格填充了 NaN。
选择、添加和删除列是基本的 DataFrame 操作。
使用类似字典的方括号表示法 [] 或属性访问 . 访问列(如果列名是有效的 Python 标识符且与 DataFrame 方法不冲突)。
# 使用上一个示例中的 df_from_dosprint("使用 [] 选择 'ColOne':")print(df_from_dos['ColOne'])
# print("\n使用属性访问选择 'ColTwo':") # 谨慎使用# print(df_from_dos.ColTwo)
print("\n选择多列 [['ColTwo', 'ColOne']]:")print(df_from_dos[['ColTwo', 'ColOne']]) # 传入一个列名列表输出:
Selecting 'ColOne' using []:a 1.0b 2.0c 3.0d NaNName: ColOne, dtype: float64
Selecting multiple columns [['ColTwo', 'ColOne']]: ColTwo ColOnea 1.0 1.0b 2.0 2.0c 3.0 3.0d 4.0 NaN通过将数据赋给新的列标签来添加新列。
# 使用 df_from_dosprint("原始 DataFrame:")print(df_from_dos)
# 通过赋给一个 Series 来添加新列(索引对齐)print("\n通过赋给一个 Series 添加 'ColThree':")df_from_dos['ColThree'] = pd.Series([10, 20, 30], index=['a', 'b', 'c'])print(df_from_dos)
# 通过现有列计算添加新列print("\n添加 'ColFour' = ColOne + ColThree:")df_from_dos['ColFour'] = df_from_dos['ColOne'] + df_from_dos['ColThree']print(df_from_dos)
# 添加一个带有标量值的新列(广播到所有行)print("\n添加带有标量值的 'ColFive':")df_from_dos['ColFive'] = 100print(df_from_dos)输出:
Original DataFrame: ColOne ColTwoa 1.0 1.0b 2.0 2.0c 3.0 3.0d NaN 4.0
Adding 'ColThree' by assigning a Series: ColOne ColTwo ColThreea 1.0 1.0 10.0b 2.0 2.0 20.0c 3.0 3.0 30.0d NaN 4.0 NaN
Adding 'ColFour' = ColOne + ColThree: ColOne ColTwo ColThree ColFoura 1.0 1.0 10.0 11.0b 2.0 2.0 20.0 22.0c 3.0 3.0 30.0 33.0d NaN 4.0 NaN NaN
Adding 'ColFive' with a scalar value: ColOne ColTwo ColThree ColFour ColFivea 1.0 1.0 10.0 11.0 100b 2.0 2.0 20.0 22.0 100c 3.0 3.0 30.0 33.0 100d NaN 4.0 NaN NaN 100使用 del 关键字或 .pop() 或 .drop() 方法。
# 使用上一步骤中的 df_from_dosprint("删除前的 DataFrame:")print(df_from_dos)
# 使用 del(原地修改)print("\n使用 del 删除 'ColOne':")del df_from_dos['ColOne']print(df_from_dos)
# 使用 pop(移除并返回该列)print("\n使用 pop 删除 'ColTwo':")popped_col = df_from_dos.pop('ColTwo')print("pop 后的 DataFrame:")print(df_from_dos)print("\n弹出的列 'ColTwo':")print(popped_col)
# 使用 drop(返回一个新的 DataFrame,使用 inplace=True 修改原始 DataFrame)print("\n使用 drop 删除 'ColThree'(返回新 DataFrame):")df_after_drop = df_from_dos.drop(columns=['ColThree']) # 使用 'columns' 参数# 或者: df_after_drop = df_from_dos.drop(['ColThree'], axis=1)print(df_after_drop)print("\n执行非原地 drop 后的原始 DataFrame(未改变):")print(df_from_dos)输出:
DataFrame before deletion: ColOne ColTwo ColThree ColFour ColFivea 1.0 1.0 10.0 11.0 100b 2.0 2.0 20.0 22.0 100c 3.0 3.0 30.0 33.0 100d NaN 4.0 NaN NaN 100
Deleting 'ColOne' using del: ColTwo ColThree ColFour ColFivea 1.0 10.0 11.0 100b 2.0 20.0 22.0 100c 3.0 30.0 33.0 100d 4.0 NaN NaN 100
Deleting 'ColTwo' using pop:DataFrame after pop: ColThree ColFour ColFivea 10.0 11.0 100b 20.0 22.0 100c 30.0 33.0 100d NaN NaN 100
Popped Column 'ColTwo':a 1.0b 2.0c 3.0d 4.0Name: ColTwo, dtype: float64
Deleting 'ColThree' using drop (returns new DataFrame): ColFour ColFivea 11.0 100b 22.0 100c 33.0 100d NaN 100
Original DataFrame after non-inplace drop (unchanged): ColThree ColFour ColFivea 10.0 11.0 100b 20.0 22.0 100c 30.0 33.0 100d NaN NaN 100del 和 .pop() 会原地修改原始 DataFrame,而 .drop() 默认返回一个新 DataFrame。你可以使用 inplace=True 使 .drop() 也原地修改。
选择、添加和删除行涉及使用索引。
行选择: .loc 和 .iloc
Section titled “行选择: .loc 和 .iloc”使用 .loc[] 进行基于标签的选择,使用 .iloc[] 进行基于整数位置的选择。
import pandas as pd
data_dict = { 'ColOne': pd.Series([1.0, 2.0, 3.0], index=['a', 'b', 'c']), 'ColTwo': pd.Series([1.0, 2.0, 3.0, 4.0], index=['a', 'b', 'c', 'd'])}df = pd.DataFrame(data_dict)print("DataFrame:")print(df)
# 使用 .loc['b'] 选择标签为 'b' 的行print("\n使用 .loc['b'] 选择标签为 'b' 的行:")print(df.loc['b'])
# 使用 .iloc[2] 选择位置 2 的行(标签 'c')print("\n使用 .iloc[2] 选择位置 2 的行(标签 'c'):")print(df.iloc[2])
# 按标签选择多行print("\n使用 .loc[['a', 'd']] 选择标签为 'a' 和 'd' 的行:")print(df.loc[['a', 'd']])
# 按位置选择多行print("\n使用 .iloc[[0, 3]] 选择位置 0 和 3 的行:")print(df.iloc[[0, 3]])输出:
DataFrame: ColOne ColTwoa 1.0 1.0b 2.0 2.0c 3.0 3.0d NaN 4.0
Selecting row with label 'b' using .loc['b']:ColOne 2.0ColTwo 2.0Name: b, dtype: float64
Selecting row at position 2 (label 'c') using .iloc[2]:ColOne 3.0ColTwo 3.0Name: c, dtype: float64
Selecting rows 'a' and 'd' using .loc[['a', 'd']]: ColOne ColTwoa 1.0 1.0d NaN 4.0
Selecting rows at positions 0 and 3 using .iloc[[0, 3]]: ColOne ColTwoa 1.0 1.0d NaN 4.0选择单行返回一个 Series,而选择多行返回一个 DataFrame。
使用 .loc(基于标签,终点包含在内)或 .iloc(基于整数位置,终点不包含在内)进行切片。
# 使用上一个示例中的 df
# 按标签切片行(包含终点)print("\n使用 .loc['b':'d'] 切片标签为 'b' 到 'd' 的行:")print(df.loc['b':'d'])
# 按位置切片行(不包含终点)print("\n使用 .iloc[1:3] 切片位置 1 到(不包括)3 的行:")print(df.iloc[1:3])
# 直接对 DataFrame 进行整数切片 (df[1:3]) 也是可能的,但通常# 不推荐;推荐使用显式的 .iloc 以保持清晰性和鲁棒性。# print("\n使用 df[1:3] 进行切片(不推荐):")# print(df[1:3])输出:
Slicing rows 'b' through 'd' using .loc['b':'d']: ColOne ColTwob 2.0 2.0c 3.0 3.0d NaN 4.0
Slicing rows from position 1 up to (not including) 3 using .iloc[1:3]: ColOne ColTwob 2.0 2.0c 3.0 3.0添加行的推荐方法是使用 pd.concat()。较旧的 .append() 方法已在 Pandas 2.0 中弃用并移除,不应使用。
import pandas as pd
df1 = pd.DataFrame([[1, 2], [3, 4]], columns=['A', 'B'], index=[0, 1])df2 = pd.DataFrame([[5, 6], [7, 8]], columns=['A', 'B'], index=[2, 3]) # 使用不同的索引df3 = pd.DataFrame([[9, 10], [11, 12]], columns=['A', 'C'], index=[4, 5]) # 不同的列
print("DataFrame df1:")print(df1)print("\nDataFrame df2:")print(df2)print("\nDataFrame df3:")print(df3)
# 连接 df1 和 df2(相同列)combined_df = pd.concat([df1, df2])print("\n使用 pd.concat() 连接 df1 和 df2:")print(combined_df)
# 连接 df1 和 df3(不同列,按索引对齐)combined_df_cols = pd.concat([df1, df3]) # 默认使用外连接print("\n使用 pd.concat() 连接 df1 和 df3(外连接):")print(combined_df_cols)
# 如果使用较旧的示例(之前可能由于 append 产生重复索引):# df_old_append = pd.DataFrame([[1, 2], [3, 4]], columns=['A', 'B'], index=[0, 1])# df_old_append2 = pd.DataFrame([[5, 6], [7, 8]], columns=['A', 'B'], index=[0, 1])# combined_dup_index = pd.concat([df_old_append, df_old_append2], ignore_index=True)# print("\n使用 ignore_index=True 连接以重置索引:")# print(combined_dup_index)输出:
DataFrame df1: A B0 1 21 3 4
DataFrame df2: A B2 5 63 7 8
DataFrame df3: A C4 9 105 11 12
Combined df1 and df2 using pd.concat(): A B0 1 21 3 42 5 63 7 8
Combined df1 and df3 using pd.concat() (outer join): A B C0 1.0 2.0 NaN1 3.0 4.0 NaN4 9.0 NaN 10.05 11.0 NaN 12.0pd.concat 智能地组合 DataFrame。如果你想在结果上获得一个全新的默认整数索引,请使用 ignore_index=True。
使用 .drop() 方法,指定要移除的行标签(索引值)并设置 axis=0(或省略 axis,因为它默认为 0)。
# 使用上一个示例中的 combined_dfprint("删除行前的 DataFrame:")print(combined_df)
# 删除标签为 0 和 3 的行# 默认情况下,drop 返回一个新的 DataFrame。使用 inplace=True 进行修改。df_dropped = combined_df.drop([0, 3], axis=0)print("\n删除行 0 和 3 后的 DataFrame:")print(df_dropped)
# 具有重复标签的示例(如果之前这样创建)# df_dup = pd.DataFrame([[1,2],[3,4],[5,6]], columns=['A','B'], index=[0,1,0])# print("\n具有重复索引 0 的 DataFrame:")# print(df_dup)# print("\n删除标签 0(移除所有具有该标签的行):")# print(df_dup.drop(0))输出:
DataFrame before row deletion: A B0 1 21 3 42 5 63 7 8
DataFrame after dropping rows 0 and 3: A B1 3 42 5 6如果索引标签重复,drop() 将移除所有匹配该标签的行。