Python Pandas - 排序
Pandas DataFrame - 排序
Section titled “Pandas DataFrame - 排序”Pandas 提供了高效的方法来对 Series 和 DataFrame 中的数据进行排序。主要有两种排序类型:
- 按索引标签排序(Sorting by Index Labels):根据其索引或列标签对行或列进行排列。
- 按值排序(Sorting by Values):根据一个或多个列中的值对行进行排列。
让我们创建一个未排序的 DataFrame 来演示排序:
import pandas as pdimport numpy as np
# Create a DataFrame with a jumbled indexunsorted_df = pd.DataFrame(np.random.randn(10, 2), index=[1, 4, 6, 2, 3, 5, 9, 8, 0, 7], columns=['col2', 'col1'])
print("Unsorted DataFrame:")print(unsorted_df)输出 (随机值可能不同):
Unsorted DataFrame: col2 col11 -2.063177 0.5375274 0.142932 -0.6848846 0.012667 -0.3893402 -0.548797 1.8487433 -1.044160 0.8373815 0.385605 1.3001859 1.031425 -1.0029678 -0.407374 -0.4351420 2.237453 -1.0671397 -1.445831 -1.701035这个 DataFrame 的索引标签(0-9 乱序)和值都未排序。
按索引标签排序 (sort_index)
Section titled “按索引标签排序 (sort_index)”sort_index() 方法根据索引标签对 DataFrame 或 Series 进行排序。默认情况下,它以升序对行索引 (axis=0) 进行排序。
import pandas as pdimport numpy as np
unsorted_df = pd.DataFrame(np.random.randn(10, 2), index=[1, 4, 6, 2, 3, 5, 9, 8, 0, 7], columns=['col2', 'col1'])
# Sort by row index (ascending by default)sorted_by_index_df = unsorted_df.sort_index()
print("DataFrame Sorted by Row Index (Ascending):")print(sorted_by_index_df)输出 (随机值可能不同,但索引按 0-9 排序):
DataFrame Sorted by Row Index (Ascending): col2 col10 0.208464 0.6270371 0.641004 0.3313522 -0.038067 -0.4647303 -0.638456 -0.0214664 0.014646 -0.7374385 -0.290761 -1.6698276 -0.797303 -0.0187377 0.525753 1.6289218 -0.567031 0.7759519 0.060724 -0.322425控制排序顺序
Section titled “控制排序顺序”使用 ascending 参数(布尔值)来控制排序顺序。ascending=False 以降序排序。
import pandas as pdimport numpy as np
unsorted_df = pd.DataFrame(np.random.randn(10, 2), index=[1, 4, 6, 2, 3, 5, 9, 8, 0, 7], columns=['col2', 'col1'])
# Sort by row index in descending ordersorted_desc_df = unsorted_df.sort_index(ascending=False)
print("DataFrame Sorted by Row Index (Descending):")print(sorted_desc_df)输出 (随机值可能不同,但索引按 9-0 排序):
DataFrame Sorted by Row Index (Descending): col2 col19 0.825697 0.3744638 -1.699509 0.5103737 -0.581378 0.6229586 -0.202951 0.9543005 -1.289321 -1.5512504 1.302561 0.8513853 -0.157915 -0.3886592 -1.222295 0.1666091 0.584890 -0.2910480 0.668444 -0.061294要按列标签而不是行索引标签排序,请设置 axis=1。
import pandas as pdimport numpy as np
unsorted_df = pd.DataFrame(np.random.randn(5, 3), # 5 rows, 3 columns index=[1, 4, 0, 3, 2], columns=['col_c', 'col_a', 'col_b'])
print("Original DataFrame (Unsorted Columns):")print(unsorted_df)
# Sort by column labels (ascending: col_a, col_b, col_c)sorted_by_cols_df = unsorted_df.sort_index(axis=1)
print("\nDataFrame Sorted by Column Labels (Ascending):")print(sorted_by_cols_df)输出 (随机值可能不同,但列按字母顺序排序):
Original DataFrame (Unsorted Columns): col_c col_a col_b1 -0.291048 0.584890 1.3025614 0.851385 -0.202951 0.1666090 0.954300 -1.222295 -1.2893213 -0.388659 -0.157915 0.8256972 -1.551250 -1.699509 -0.061294
DataFrame Sorted by Column Labels (Ascending): col_a col_b col_c1 0.584890 1.302561 -0.2910484 -0.202951 0.166609 0.8513850 -1.222295 -1.289321 0.9543003 -0.157915 0.825697 -0.3886592 -1.699509 -0.061294 -1.551250按值排序 (sort_values)
Section titled “按值排序 (sort_values)”sort_values() 方法根据一个或多个指定列中的值对 DataFrame 进行排序。by 参数用于指示要排序的列。
import pandas as pdimport numpy as np
# Create a DataFrame with specific values for sorting demonstrationdf = pd.DataFrame({ 'col1': [2, 1, 1, 1, 3], 'col2': [1, 3, 2, 4, 0]})print("Original DataFrame:")print(df)
# Sort by values in 'col1' (ascending by default)sorted_by_col1 = df.sort_values(by='col1')
print("\nDataFrame Sorted by 'col1':")print(sorted_by_col1)输出:
Original DataFrame: col1 col20 2 11 1 32 1 23 1 44 3 0
DataFrame Sorted by 'col1': col1 col21 1 3 # Original index 12 1 2 # Original index 23 1 4 # Original index 30 2 1 # Original index 04 3 0 # Original index 4注意,行现在根据 col1 中的值进行排序。整行(包括 col2 和原始索引)在排序过程中会一起移动。
要按多个列排序,请将列名列表传递给 by 参数。排序将根据列表中列的顺序分层进行。
import pandas as pdimport numpy as np
df = pd.DataFrame({ 'col1': [2, 1, 1, 1, 3], 'col2': [1, 3, 2, 4, 0]})
# Sort first by 'col1' (ascending), then by 'col2' (ascending) for ties in 'col1'sorted_multi = df.sort_values(by=['col1', 'col2'])
print("DataFrame Sorted by 'col1' then 'col2':")print(sorted_multi)输出:
DataFrame Sorted by 'col1' then 'col2': col1 col22 1 2 # Smallest col2 value where col1 is 11 1 33 1 4 # Largest col2 value where col1 is 10 2 14 3 0您可以通过向 ascending 参数传递布尔值列表来为每列指定不同的排序顺序(升序/降序),该列表与 by 中的列相对应。
# Sort by 'col1' ascending, 'col2' descendingsorted_multi_mixed = df.sort_values(by=['col1', 'col2'], ascending=[True, False])print("\nDataFrame Sorted by 'col1' (Asc), 'col2' (Desc):")print(sorted_multi_mixed)输出:
DataFrame Sorted by 'col1' (Asc), 'col2' (Desc): col1 col23 1 4 # Largest col2 value where col1 is 11 1 32 1 2 # Smallest col2 value where col1 is 10 2 14 3 0排序算法的选择
Section titled “排序算法的选择”sort_values() 和 sort_index() 允许使用 kind 参数指定排序算法。可用选项通常包括:
'quicksort'(默认):通常速度很快,但不稳定(not stable)。'mergesort':稳定(stable)(保留相等元素的原始顺序),在某些情况下可能较慢或使用更多内存。'heapsort':提供 O(n log n) 性能,与 quicksort 和 mergesort 类似。'stable':mergesort 的别名,强调稳定性。
稳定性意味着如果两行在排序列中具有相同的值,则它们在输出中的相对顺序与在输入中的相对顺序相同。
import pandas as pdimport numpy as np
df = pd.DataFrame({ 'col1': [2, 1, 1, 1, 3], 'col2': [1, 3, 2, 4, 0]})
# Sort using mergesort (guaranteed stable)sorted_stable = df.sort_values(by='col1', kind='mergesort')
print("DataFrame Sorted by 'col1' using Mergesort (Stable):")print(sorted_stable)输出 (注意 col1=1 的行的原始相对顺序被保留:索引 1,然后 2,然后 3):
DataFrame Sorted by 'col1' using Mergesort (Stable): col1 col21 1 32 1 23 1 40 2 14 3 0选择算法可能会影响性能,但对于大多数用例来说,除非明确需要稳定性,否则默认值 (quicksort) 就足够了。