Skip to content

索引与选择数据

选择特定的数据子集(行、列或单个值)是数据分析中的基本操作。Pandas 提供了强大且优化的方法来执行此操作,主要通过 .loc 和 .iloc 索引器实现。

虽然基本的 Python/NumPy 索引 ([]) 和属性访问 (.) 在简单情况下可以工作,但有时可能会产生歧义,尤其是在处理整数标签或对 DataFrame 进行切片时。为了提高清晰度、性能并在生产代码中避免潜在陷阱,强烈建议使用专门的 .loc 和 .iloc 访问器。

Pandas 提供了两个主要索引器用于选择数据:

索引器选择方法描述
1.loc[]**基于标签(Label-based)**的选择。根据索引标签(行名)和列名选择数据。
2.iloc[]**基于整数位置(Integer position-based)**的选择。根据整数位置(从 0 开始索引,类似于 Python 列表)选择数据。

重要提示: .ix[]索引器,它试图同时处理基于标签和整数位置的索引,由于其潜在的歧义性和混乱性而**已被弃用**。你应该**始终**使用.loc进行基于标签的索引,并使用.iloc` 进行基于整数位置的索引。

.loc[] 用于通过标签(索引名称、列名称)选择数据。它接受各种输入:

  • 单个标签(例如,'a' 或 5,如果索引包含整数 5 作为标签)。
  • 标签列表或数组(例如,['a', 'b', 'c'])。
  • 带有标签的切片对象(例如,'a':'f')。关键是,使用 .loc 通过标签进行切片时,起始和结束边界都包含在内。
  • 布尔数组(与要索引的轴具有相同的长度)。
  • 一个可调用函数(Callable function),带有一个参数(调用它的 Series 或 DataFrame),该函数返回有效的索引输出。

.loc 通常接受两个用逗号分隔的参数:.loc[行索引器, 列索引器]。你可以使用 : 来选择所有行或所有列。

让我们创建一个示例 DataFrame:

import pandas as pd
import numpy as np
# Sample DataFrame with string index labels
index_labels = list('abcdefgh')
df = pd.DataFrame(np.random.randn(8, 4),
index=index_labels,
columns=['A', 'B', 'C', 'D'])
print("Sample DataFrame:")
print(df)

(输出会因随机数而异)

示例 1:选择特定列(返回一个 Series)

Section titled “示例 1:选择特定列(返回一个 Series)”
# Select all rows for column 'A'
print("\nColumn 'A':")
print(df.loc[:, 'A'])

输出(示例):

Column 'A':
a -0.58
b 1.23
c -0.15
d 0.48
e -1.09
f 0.21
g 0.88
h -1.52
Name: A, dtype: float64

示例 2:选择多列(返回一个 DataFrame)

Section titled “示例 2:选择多列(返回一个 DataFrame)”
# Select all rows for columns 'A' and 'C'
print("\nColumns 'A' and 'C':")
print(df.loc[:, ['A', 'C']])

输出(示例):

Columns 'A' and 'C':
A C
a -0.58 0.77
b 1.23 -0.31
c -0.15 1.50
d 0.48 -0.99
e -1.09 0.25
f 0.21 1.11
g 0.88 -0.40
h -1.52 0.65
# Select rows 'a', 'b', 'f', 'h' and columns 'A', 'C'
print("\nSpecific rows and columns:")
print(df.loc[['a', 'b', 'f', 'h'], ['A', 'C']])

输出(示例):

Specific rows and columns:
A C
a -0.58 0.77
b 1.23 -0.31
f 0.21 1.11
h -1.52 0.65

示例 4:选择行范围(标签切片,包含结束标签)

Section titled “示例 4:选择行范围(标签切片,包含结束标签)”
# Select rows from 'c' to 'f' (inclusive) for all columns
print("\nRow slice 'c' to 'f':")
print(df.loc['c':'f'])

输出(示例):

Row slice 'c' to 'f':
A B C D
c -0.15 0.91 1.50 -0.22
d 0.48 -0.55 -0.99 1.05
e -1.09 0.33 0.25 -0.78
f 0.21 -1.20 1.11 0.14
# Select rows where column 'A' is greater than 0
print("\nRows where A > 0:")
print(df.loc[df['A'] > 0])

输出(示例):

Rows where A > 0:
A B C D
b 1.23 -0.67 -0.31 0.50
d 0.48 -0.55 -0.99 1.05
f 0.21 -1.20 1.11 0.14
g 0.88 0.10 -0.40 -0.95
# Get the value at row 'b', column 'B'
value = df.loc['b', 'B']
print(f"\nValue at ('b', 'B'): {value:.2f}") # Formatted output

输出(示例):

Value at ('b', 'B'): -0.67

.iloc[] 用于通过整数位置(从 0 到 length-1)选择数据。它的工作方式类似于索引列表或 NumPy 数组。

  • 单个整数(例如,0 代表第一行/列)。
  • 整数列表或数组(例如,[0, 2, 3])。
  • 带有整数的切片对象(例如,1:5)。**与 .loc 的标签切片不同,使用 .iloc 进行整数切片时,不包含结束边界。**因此 1:5 选择位置 1、2、3、4。
  • 布尔数组(与要索引的轴具有相同的长度)。
  • 一个可调用函数(类似于 .loc)。

.iloc 也接受参数,格式为 .iloc[行索引器, 列索引器]。

为了清晰起见,我们使用带有默认整数索引的 DataFrame:

import pandas as pd
import numpy as np
df_int = pd.DataFrame(np.random.randn(8, 4), columns=['A', 'B', 'C', 'D'])
print("Sample DataFrame (integer index):")
print(df_int)

(输出会因随机数而异)

# Select rows from index 0 up to (but not including) 4
print("\nFirst 4 rows:")
print(df_int.iloc[0:4]) # Equivalent to df_int.iloc[:4]

输出(示例):

First 4 rows:
A B C D
0 0.69 0.25 -1.27 -0.64
1 -0.68 0.89 -0.81 0.63
2 -0.78 -0.53 0.02 0.23
3 0.53 -1.28 0.82 -0.02

示例 2:通过整数位置切片选择行和列

Section titled “示例 2:通过整数位置切片选择行和列”
# Select rows 1, 2, 3, 4 and columns 2, 3
print("\nSlice [1:5, 2:4]:")
print(df_int.iloc[1:5, 2:4])

输出(示例):

Slice [1:5, 2:4]:
C D
1 -0.81 0.63
2 0.02 0.23
3 0.82 -0.02
4 1.42 1.13

示例 3:通过整数列表选择特定行和列

Section titled “示例 3:通过整数列表选择特定行和列”
# Select rows 1, 3, 5 and columns 1, 3
print("\nRows [1, 3, 5], Columns [1, 3]:")
print(df_int.iloc[[1, 3, 5], [1, 3]])

输出(示例):

Rows [1, 3, 5], Columns [1, 3]:
B D
1 0.89 0.63
3 -1.28 -0.02
5 -0.51 -0.51

示例 4:使用切片选择特定行或列

Section titled “示例 4:使用切片选择特定行或列”
# Select rows 1 and 2 for all columns
print("\nRows 1 and 2 (all columns):")
print(df_int.iloc[1:3, :])
# Select all rows for columns 1 and 2
print("\nColumns 1 and 2 (all rows):")
print(df_int.iloc[:, 1:3])

输出(示例):

Rows 1 and 2 (all columns):
A B C D
1 -0.68 0.89 -0.81 0.63
2 -0.78 -0.53 0.02 0.23
Columns 1 and 2 (all rows):
B C
0 0.25 -1.27
1 0.89 -0.81
2 -0.53 0.02
3 -1.28 0.82
4 -0.46 1.42
5 -0.51 0.58
6 -1.20 0.09
7 -0.94 0.64

示例 5:通过位置获取单个标量值

Section titled “示例 5:通过位置获取单个标量值”
# Get the value at row index 0, column index 1
value_iloc = df_int.iloc[0, 1]
print(f"\nValue at [0, 1]: {value_iloc:.2f}")

输出(示例):

Value at [0, 1]: 0.25

基本索引([])和属性访问(.)

Section titled “基本索引([])和属性访问(.)”

Pandas 也支持更简单的索引方法,但请谨慎使用:

[] 运算符对于 Series 和 DataFrame 的工作方式不同:

  • 对于 Series: [] 主要像 .loc[] 一样工作(基于标签),但如果在标签未找到且输入是整数时,会回退到 .iloc[]。切片通常按位置工作。
  • 对于 DataFrames:
    • 传入单个标签或标签列表会选择列:df['A'] 或 df[['A', 'B']]。
    • 传入切片(:)会选择行:df[0:3](基于整数位置的切片,不包含结束边界)。这种使用 [] 进行行切片的行为可能会令人困惑,并且通常不推荐,建议使用 .iloc。
import pandas as pd
import numpy as np
df = pd.DataFrame(np.random.randn(8, 4), columns = ['A', 'B', 'C', 'D'])
# Select column 'A' using []
print("Column 'A' using []:")
print(df['A'])
# Select columns 'A', 'B' using []
print("\nColumns 'A', 'B' using []:")
print(df[['A', 'B']])
# Select first 3 rows using [] slice (integer position)
print("\nRows 0-2 using []:")
print(df[0:3])

输出(示例):

Column 'A' using []:
0 -0.47
1 0.39
2 0.33
3 -1.05
4 -0.16
5 -0.32
6 0.56
7 -0.75
Name: A, dtype: float64
Columns 'A', 'B' using []:
A B
0 -0.47 -0.60
1 0.39 -0.94
2 0.33 0.09
3 -1.05 -0.01
4 -0.16 1.55
5 -0.32 -0.22
6 0.56 -0.31
7 -0.75 -0.37
Rows 0-2 using []:
A B C D
0 -0.47 -0.60 -0.11 0.92
1 0.39 -0.94 1.10 -0.71
2 0.33 0.09 -0.05 0.58

如果列名是有效的 Python 标识符(没有空格,不与 DataFrame 方法冲突),你可以将其作为属性访问:

import pandas as pd
import numpy as np
df = pd.DataFrame(np.random.randn(8, 4), columns = ['A', 'B', 'C', 'D'])
# Select column 'A' using attribute access
print("Column 'A' using attribute access:")
print(df.A)

输出(示例):

Column 'A' using attribute access:
0 -0.47
1 0.39
2 0.33
3 -1.05
4 -0.16
5 -0.32
6 0.56
7 -0.75
Name: A, dtype: float64

警告: 虽然方便,但属性访问(.)有局限性。它不适用于不是有效标识符的列名(例如,‘column name with spaces’)或与现有 DataFrame 方法冲突的列名(如 df.count)。在列选择中,使用 [](例如,df['count'])或 .loc 更安全且更通用。

为了在 Pandas 中进行稳健清晰的数据选择:

  • 使用 .loc[] 基于索引和列的标签选择数据。
  • 使用 .iloc[] 基于整数位置选择数据。
  • 避免使用已弃用的 .ix[] 索引器。
  • 主要使用 [] 通过名称选择列。对使用 [] 进行行切片保持谨慎,因为它可能产生歧义。
  • 仅在交互式访问具有有效标识符名称的列时,为了方便可以使用属性访问(.),但在可重用代码中更倾向于使用 [] 或 .loc。