索引与选择数据
Pandas - 索引与数据选择
Section titled “Pandas - 索引与数据选择”选择特定的数据子集(行、列或单个值)是数据分析中的基本操作。Pandas 提供了强大且优化的方法来执行此操作,主要通过 .loc 和 .iloc 索引器实现。
虽然基本的 Python/NumPy 索引 ([]) 和属性访问 (.) 在简单情况下可以工作,但有时可能会产生歧义,尤其是在处理整数标签或对 DataFrame 进行切片时。为了提高清晰度、性能并在生产代码中避免潜在陷阱,强烈建议使用专门的 .loc 和 .iloc 访问器。
Pandas 提供了两个主要索引器用于选择数据:
| 索引器 | 选择方法 | 描述 |
|---|---|---|
| 1 | .loc[] | **基于标签(Label-based)**的选择。根据索引标签(行名)和列名选择数据。 |
| 2 | .iloc[] | **基于整数位置(Integer position-based)**的选择。根据整数位置(从 0 开始索引,类似于 Python 列表)选择数据。 |
重要提示: .ix[]索引器,它试图同时处理基于标签和整数位置的索引,由于其潜在的歧义性和混乱性而**已被弃用**。你应该**始终**使用.loc进行基于标签的索引,并使用.iloc` 进行基于整数位置的索引。
.loc[](基于标签的选择)
Section titled “.loc[](基于标签的选择)”.loc[] 用于通过标签(索引名称、列名称)选择数据。它接受各种输入:
- 单个标签(例如,
'a'或5,如果索引包含整数 5 作为标签)。 - 标签列表或数组(例如,
['a', 'b', 'c'])。 - 带有标签的切片对象(例如,
'a':'f')。关键是,使用.loc通过标签进行切片时,起始和结束边界都包含在内。 - 布尔数组(与要索引的轴具有相同的长度)。
- 一个可调用函数(Callable function),带有一个参数(调用它的 Series 或 DataFrame),该函数返回有效的索引输出。
.loc 通常接受两个用逗号分隔的参数:.loc[行索引器, 列索引器]。你可以使用 : 来选择所有行或所有列。
让我们创建一个示例 DataFrame:
import pandas as pdimport numpy as np
# Sample DataFrame with string index labelsindex_labels = list('abcdefgh')df = pd.DataFrame(np.random.randn(8, 4), index=index_labels, columns=['A', 'B', 'C', 'D'])print("Sample DataFrame:")print(df)(输出会因随机数而异)
示例 1:选择特定列(返回一个 Series)
Section titled “示例 1:选择特定列(返回一个 Series)”# Select all rows for column 'A'print("\nColumn 'A':")print(df.loc[:, 'A'])输出(示例):
Column 'A':a -0.58b 1.23c -0.15d 0.48e -1.09f 0.21g 0.88h -1.52Name: A, dtype: float64示例 2:选择多列(返回一个 DataFrame)
Section titled “示例 2:选择多列(返回一个 DataFrame)”# Select all rows for columns 'A' and 'C'print("\nColumns 'A' and 'C':")print(df.loc[:, ['A', 'C']])输出(示例):
Columns 'A' and 'C': A Ca -0.58 0.77b 1.23 -0.31c -0.15 1.50d 0.48 -0.99e -1.09 0.25f 0.21 1.11g 0.88 -0.40h -1.52 0.65示例 3:选择特定行和列
Section titled “示例 3:选择特定行和列”# Select rows 'a', 'b', 'f', 'h' and columns 'A', 'C'print("\nSpecific rows and columns:")print(df.loc[['a', 'b', 'f', 'h'], ['A', 'C']])输出(示例):
Specific rows and columns: A Ca -0.58 0.77b 1.23 -0.31f 0.21 1.11h -1.52 0.65示例 4:选择行范围(标签切片,包含结束标签)
Section titled “示例 4:选择行范围(标签切片,包含结束标签)”# Select rows from 'c' to 'f' (inclusive) for all columnsprint("\nRow slice 'c' to 'f':")print(df.loc['c':'f'])输出(示例):
Row slice 'c' to 'f': A B C Dc -0.15 0.91 1.50 -0.22d 0.48 -0.55 -0.99 1.05e -1.09 0.33 0.25 -0.78f 0.21 -1.20 1.11 0.14示例 5:基于布尔条件选择
Section titled “示例 5:基于布尔条件选择”# Select rows where column 'A' is greater than 0print("\nRows where A > 0:")print(df.loc[df['A'] > 0])输出(示例):
Rows where A > 0: A B C Db 1.23 -0.67 -0.31 0.50d 0.48 -0.55 -0.99 1.05f 0.21 -1.20 1.11 0.14g 0.88 0.10 -0.40 -0.95示例 6:获取单个标量值
Section titled “示例 6:获取单个标量值”# Get the value at row 'b', column 'B'value = df.loc['b', 'B']print(f"\nValue at ('b', 'B'): {value:.2f}") # Formatted output输出(示例):
Value at ('b', 'B'): -0.67.iloc[](基于整数位置的选择)
Section titled “.iloc[](基于整数位置的选择)”.iloc[] 用于通过整数位置(从 0 到 length-1)选择数据。它的工作方式类似于索引列表或 NumPy 数组。
- 单个整数(例如,
0代表第一行/列)。 - 整数列表或数组(例如,
[0, 2, 3])。 - 带有整数的切片对象(例如,
1:5)。**与.loc的标签切片不同,使用.iloc进行整数切片时,不包含结束边界。**因此1:5选择位置 1、2、3、4。 - 布尔数组(与要索引的轴具有相同的长度)。
- 一个可调用函数(类似于
.loc)。
.iloc 也接受参数,格式为 .iloc[行索引器, 列索引器]。
为了清晰起见,我们使用带有默认整数索引的 DataFrame:
import pandas as pdimport numpy as np
df_int = pd.DataFrame(np.random.randn(8, 4), columns=['A', 'B', 'C', 'D'])print("Sample DataFrame (integer index):")print(df_int)(输出会因随机数而异)
示例 1:选择前 4 行
Section titled “示例 1:选择前 4 行”# Select rows from index 0 up to (but not including) 4print("\nFirst 4 rows:")print(df_int.iloc[0:4]) # Equivalent to df_int.iloc[:4]输出(示例):
First 4 rows: A B C D0 0.69 0.25 -1.27 -0.641 -0.68 0.89 -0.81 0.632 -0.78 -0.53 0.02 0.233 0.53 -1.28 0.82 -0.02示例 2:通过整数位置切片选择行和列
Section titled “示例 2:通过整数位置切片选择行和列”# Select rows 1, 2, 3, 4 and columns 2, 3print("\nSlice [1:5, 2:4]:")print(df_int.iloc[1:5, 2:4])输出(示例):
Slice [1:5, 2:4]: C D1 -0.81 0.632 0.02 0.233 0.82 -0.024 1.42 1.13示例 3:通过整数列表选择特定行和列
Section titled “示例 3:通过整数列表选择特定行和列”# Select rows 1, 3, 5 and columns 1, 3print("\nRows [1, 3, 5], Columns [1, 3]:")print(df_int.iloc[[1, 3, 5], [1, 3]])输出(示例):
Rows [1, 3, 5], Columns [1, 3]: B D1 0.89 0.633 -1.28 -0.025 -0.51 -0.51示例 4:使用切片选择特定行或列
Section titled “示例 4:使用切片选择特定行或列”# Select rows 1 and 2 for all columnsprint("\nRows 1 and 2 (all columns):")print(df_int.iloc[1:3, :])
# Select all rows for columns 1 and 2print("\nColumns 1 and 2 (all rows):")print(df_int.iloc[:, 1:3])输出(示例):
Rows 1 and 2 (all columns): A B C D1 -0.68 0.89 -0.81 0.632 -0.78 -0.53 0.02 0.23
Columns 1 and 2 (all rows): B C0 0.25 -1.271 0.89 -0.812 -0.53 0.023 -1.28 0.824 -0.46 1.425 -0.51 0.586 -1.20 0.097 -0.94 0.64示例 5:通过位置获取单个标量值
Section titled “示例 5:通过位置获取单个标量值”# Get the value at row index 0, column index 1value_iloc = df_int.iloc[0, 1]print(f"\nValue at [0, 1]: {value_iloc:.2f}")输出(示例):
Value at [0, 1]: 0.25基本索引([])和属性访问(.)
Section titled “基本索引([])和属性访问(.)”Pandas 也支持更简单的索引方法,但请谨慎使用:
[] 运算符对于 Series 和 DataFrame 的工作方式不同:
- 对于 Series:
[]主要像.loc[]一样工作(基于标签),但如果在标签未找到且输入是整数时,会回退到.iloc[]。切片通常按位置工作。 - 对于 DataFrames:
-
- 传入单个标签或标签列表会选择列:
df['A']或df[['A', 'B']]。
- 传入单个标签或标签列表会选择列:
-
- 传入切片(
:)会选择行:df[0:3](基于整数位置的切片,不包含结束边界)。这种使用[]进行行切片的行为可能会令人困惑,并且通常不推荐,建议使用.iloc。
- 传入切片(
import pandas as pdimport numpy as npdf = pd.DataFrame(np.random.randn(8, 4), columns = ['A', 'B', 'C', 'D'])
# Select column 'A' using []print("Column 'A' using []:")print(df['A'])
# Select columns 'A', 'B' using []print("\nColumns 'A', 'B' using []:")print(df[['A', 'B']])
# Select first 3 rows using [] slice (integer position)print("\nRows 0-2 using []:")print(df[0:3])输出(示例):
Column 'A' using []:0 -0.471 0.392 0.333 -1.054 -0.165 -0.326 0.567 -0.75Name: A, dtype: float64
Columns 'A', 'B' using []: A B0 -0.47 -0.601 0.39 -0.942 0.33 0.093 -1.05 -0.014 -0.16 1.555 -0.32 -0.226 0.56 -0.317 -0.75 -0.37
Rows 0-2 using []: A B C D0 -0.47 -0.60 -0.11 0.921 0.39 -0.94 1.10 -0.712 0.33 0.09 -0.05 0.58属性访问(.)
Section titled “属性访问(.)”如果列名是有效的 Python 标识符(没有空格,不与 DataFrame 方法冲突),你可以将其作为属性访问:
import pandas as pdimport numpy as npdf = pd.DataFrame(np.random.randn(8, 4), columns = ['A', 'B', 'C', 'D'])
# Select column 'A' using attribute accessprint("Column 'A' using attribute access:")print(df.A)输出(示例):
Column 'A' using attribute access:0 -0.471 0.392 0.333 -1.054 -0.165 -0.326 0.567 -0.75Name: A, dtype: float64警告: 虽然方便,但属性访问(.)有局限性。它不适用于不是有效标识符的列名(例如,‘column name with spaces’)或与现有 DataFrame 方法冲突的列名(如 df.count)。在列选择中,使用 [](例如,df['count'])或 .loc 更安全且更通用。
总结与最佳实践
Section titled “总结与最佳实践”为了在 Pandas 中进行稳健清晰的数据选择:
- 使用
.loc[]基于索引和列的标签选择数据。 - 使用
.iloc[]基于整数位置选择数据。 - 避免使用已弃用的
.ix[]索引器。 - 主要使用
[]通过名称选择列。对使用[]进行行切片保持谨慎,因为它可能产生歧义。 - 仅在交互式访问具有有效标识符名称的列时,为了方便可以使用属性访问(
.),但在可重用代码中更倾向于使用[]或.loc。