Skip to content

Python Pandas - 排序

Pandas 提供了高效的方法来对 Series 和 DataFrame 中的数据进行排序。主要有两种排序类型:

  • 按索引标签排序(Sorting by Index Labels):根据其索引或列标签对行或列进行排列。
  • 按值排序(Sorting by Values):根据一个或多个列中的值对行进行排列。

让我们创建一个未排序的 DataFrame 来演示排序:

import pandas as pd
import numpy as np
# Create a DataFrame with a jumbled index
unsorted_df = pd.DataFrame(np.random.randn(10, 2),
index=[1, 4, 6, 2, 3, 5, 9, 8, 0, 7],
columns=['col2', 'col1'])
print("Unsorted DataFrame:")
print(unsorted_df)

输出 (随机值可能不同):

Unsorted DataFrame:
col2 col1
1 -2.063177 0.537527
4 0.142932 -0.684884
6 0.012667 -0.389340
2 -0.548797 1.848743
3 -1.044160 0.837381
5 0.385605 1.300185
9 1.031425 -1.002967
8 -0.407374 -0.435142
0 2.237453 -1.067139
7 -1.445831 -1.701035

这个 DataFrame 的索引标签(0-9 乱序)和值都未排序。

sort_index() 方法根据索引标签对 DataFrame 或 Series 进行排序。默认情况下,它以升序对行索引 (axis=0) 进行排序。

import pandas as pd
import numpy as np
unsorted_df = pd.DataFrame(np.random.randn(10, 2),
index=[1, 4, 6, 2, 3, 5, 9, 8, 0, 7],
columns=['col2', 'col1'])
# Sort by row index (ascending by default)
sorted_by_index_df = unsorted_df.sort_index()
print("DataFrame Sorted by Row Index (Ascending):")
print(sorted_by_index_df)

输出 (随机值可能不同,但索引按 0-9 排序):

DataFrame Sorted by Row Index (Ascending):
col2 col1
0 0.208464 0.627037
1 0.641004 0.331352
2 -0.038067 -0.464730
3 -0.638456 -0.021466
4 0.014646 -0.737438
5 -0.290761 -1.669827
6 -0.797303 -0.018737
7 0.525753 1.628921
8 -0.567031 0.775951
9 0.060724 -0.322425

使用 ascending 参数(布尔值)来控制排序顺序。ascending=False 以降序排序。

import pandas as pd
import numpy as np
unsorted_df = pd.DataFrame(np.random.randn(10, 2),
index=[1, 4, 6, 2, 3, 5, 9, 8, 0, 7],
columns=['col2', 'col1'])
# Sort by row index in descending order
sorted_desc_df = unsorted_df.sort_index(ascending=False)
print("DataFrame Sorted by Row Index (Descending):")
print(sorted_desc_df)

输出 (随机值可能不同,但索引按 9-0 排序):

DataFrame Sorted by Row Index (Descending):
col2 col1
9 0.825697 0.374463
8 -1.699509 0.510373
7 -0.581378 0.622958
6 -0.202951 0.954300
5 -1.289321 -1.551250
4 1.302561 0.851385
3 -0.157915 -0.388659
2 -1.222295 0.166609
1 0.584890 -0.291048
0 0.668444 -0.061294

要按列标签而不是行索引标签排序,请设置 axis=1。

import pandas as pd
import numpy as np
unsorted_df = pd.DataFrame(np.random.randn(5, 3), # 5 rows, 3 columns
index=[1, 4, 0, 3, 2],
columns=['col_c', 'col_a', 'col_b'])
print("Original DataFrame (Unsorted Columns):")
print(unsorted_df)
# Sort by column labels (ascending: col_a, col_b, col_c)
sorted_by_cols_df = unsorted_df.sort_index(axis=1)
print("\nDataFrame Sorted by Column Labels (Ascending):")
print(sorted_by_cols_df)

输出 (随机值可能不同,但列按字母顺序排序):

Original DataFrame (Unsorted Columns):
col_c col_a col_b
1 -0.291048 0.584890 1.302561
4 0.851385 -0.202951 0.166609
0 0.954300 -1.222295 -1.289321
3 -0.388659 -0.157915 0.825697
2 -1.551250 -1.699509 -0.061294
DataFrame Sorted by Column Labels (Ascending):
col_a col_b col_c
1 0.584890 1.302561 -0.291048
4 -0.202951 0.166609 0.851385
0 -1.222295 -1.289321 0.954300
3 -0.157915 0.825697 -0.388659
2 -1.699509 -0.061294 -1.551250

sort_values() 方法根据一个或多个指定列中的值对 DataFrame 进行排序。by 参数用于指示要排序的列。

import pandas as pd
import numpy as np
# Create a DataFrame with specific values for sorting demonstration
df = pd.DataFrame({
'col1': [2, 1, 1, 1, 3],
'col2': [1, 3, 2, 4, 0]
})
print("Original DataFrame:")
print(df)
# Sort by values in 'col1' (ascending by default)
sorted_by_col1 = df.sort_values(by='col1')
print("\nDataFrame Sorted by 'col1':")
print(sorted_by_col1)

输出:

Original DataFrame:
col1 col2
0 2 1
1 1 3
2 1 2
3 1 4
4 3 0
DataFrame Sorted by 'col1':
col1 col2
1 1 3 # Original index 1
2 1 2 # Original index 2
3 1 4 # Original index 3
0 2 1 # Original index 0
4 3 0 # Original index 4

注意,行现在根据 col1 中的值进行排序。整行(包括 col2 和原始索引)在排序过程中会一起移动。

要按多个列排序,请将列名列表传递给 by 参数。排序将根据列表中列的顺序分层进行。

import pandas as pd
import numpy as np
df = pd.DataFrame({
'col1': [2, 1, 1, 1, 3],
'col2': [1, 3, 2, 4, 0]
})
# Sort first by 'col1' (ascending), then by 'col2' (ascending) for ties in 'col1'
sorted_multi = df.sort_values(by=['col1', 'col2'])
print("DataFrame Sorted by 'col1' then 'col2':")
print(sorted_multi)

输出:

DataFrame Sorted by 'col1' then 'col2':
col1 col2
2 1 2 # Smallest col2 value where col1 is 1
1 1 3
3 1 4 # Largest col2 value where col1 is 1
0 2 1
4 3 0

您可以通过向 ascending 参数传递布尔值列表来为每列指定不同的排序顺序(升序/降序),该列表与 by 中的列相对应。

# Sort by 'col1' ascending, 'col2' descending
sorted_multi_mixed = df.sort_values(by=['col1', 'col2'], ascending=[True, False])
print("\nDataFrame Sorted by 'col1' (Asc), 'col2' (Desc):")
print(sorted_multi_mixed)

输出:

DataFrame Sorted by 'col1' (Asc), 'col2' (Desc):
col1 col2
3 1 4 # Largest col2 value where col1 is 1
1 1 3
2 1 2 # Smallest col2 value where col1 is 1
0 2 1
4 3 0

sort_values() 和 sort_index() 允许使用 kind 参数指定排序算法。可用选项通常包括:

  • 'quicksort' (默认):通常速度很快,但不稳定(not stable)。
  • 'mergesort':稳定(stable)(保留相等元素的原始顺序),在某些情况下可能较慢或使用更多内存。
  • 'heapsort':提供 O(n log n) 性能,与 quicksort 和 mergesort 类似。
  • 'stable':mergesort 的别名,强调稳定性。

稳定性意味着如果两行在排序列中具有相同的值,则它们在输出中的相对顺序与在输入中的相对顺序相同。

import pandas as pd
import numpy as np
df = pd.DataFrame({
'col1': [2, 1, 1, 1, 3],
'col2': [1, 3, 2, 4, 0]
})
# Sort using mergesort (guaranteed stable)
sorted_stable = df.sort_values(by='col1', kind='mergesort')
print("DataFrame Sorted by 'col1' using Mergesort (Stable):")
print(sorted_stable)

输出 (注意 col1=1 的行的原始相对顺序被保留:索引 1,然后 2,然后 3):

DataFrame Sorted by 'col1' using Mergesort (Stable):
col1 col2
1 1 3
2 1 2
3 1 4
0 2 1
4 3 0

选择算法可能会影响性能,但对于大多数用例来说,除非明确需要稳定性,否则默认值 (quicksort) 就足够了。