Skip to content

Python Pandas - 连接 (Concatenation)

Pandas 提供了强大的 pd.concat() 函数,用于沿着某个轴向连接 Series 或 DataFrame 对象。它在对象的连接方式和索引处理方面提供了灵活性。

重要提示: Series 和 DataFrame 上较旧的 .append() 方法在 Pandas 2.0 中已弃用并移除。您应仅使用 pd.concat() 来组合对象。

pd.concat(objs, axis=0, join='outer', ignore_index=False, keys=None,
levels=None, names=None, verify_integrity=False,
sort=False, copy=None) # copy is deprecated

主要参数:

  • objs: 要连接的 Series 或 DataFrame 对象序列(通常是列表)或映射。
  • axis ({0 or ‘index’, 1 or ‘columns’}, default 0): 连接的轴。0 意味着按行垂直堆叠,1 意味着按列水平堆叠。
  • join ({‘inner’, ‘outer’}, default ‘outer’): 如何处理另一个轴(未连接的轴)上的索引。‘outer’ 执行并集(保留所有标签),‘inner’ 执行交集(仅保留共享标签)。
  • ignore_index (boolean, default False): 如果为 True,则不保留沿着连接轴的索引值。结果轴将使用新的 RangeIndex (0, …, n-1) 进行标记。
  • keys: 序列(例如,列表或元组)。在连接轴上构建一个层级索引 (MultiIndex),使用提供的键作为最外层。
  • verify_integrity (boolean, default False): 如果为 True,则检查新的连接轴是否包含重复项。如果发现重复项,则引发异常。当索引应唯一时很有用。
  • sort (boolean, default False): 在执行 ‘outer’ join 时,如果另一个轴(axis=0 时是列,axis=1 时是索引)未对齐,则对其进行排序。

最常见的用例是垂直堆叠 DataFrames(沿着 axis=0)。

import pandas as pd
df1 = pd.DataFrame({
'Name': ['Alex', 'Amy', 'Allen', 'Alice'],
'SubjectID': ['sub1', 'sub2', 'sub4', 'sub6'],
'Score': [98, 90, 87, 69]
},
index=[101, 102, 103, 104]) # Meaningful index
df2 = pd.DataFrame({
'Name': ['Billy', 'Brian', 'Bran', 'Bryce'],
'SubjectID': ['sub2', 'sub4', 'sub3', 'sub6'],
'Score': [89, 80, 79, 97]
},
index=[201, 202, 203, 204]) # Different meaningful index
print("DataFrame df1:")
print(df1)
print("\nDataFrame df2:")
print(df2)
# Concatenate along axis=0 (default)
concatenated_df = pd.concat([df1, df2])
print("\nConcatenated DataFrame (axis=0, default):")
print(concatenated_df)

输出:

DataFrame df1:
Name SubjectID Score
101 Alex sub1 98
102 Amy sub2 90
103 Allen sub4 87
104 Alice sub6 69
DataFrame df2:
Name SubjectID Score
201 Billy sub2 89
202 Brian sub4 80
203 Bran sub3 79
204 Bryce sub6 97
Concatenated DataFrame (axis=0, default):
Name SubjectID Score
101 Alex sub1 98
102 Amy sub2 90
103 Allen sub4 87
104 Alice sub6 69
201 Billy sub2 89
202 Brian sub4 80
203 Bran sub3 79
204 Bryce sub6 97

请注意,原始索引被保留了。

如果您想跟踪连接部分的来源,请使用 keys 参数创建一个 MultiIndex。

# Using df1 and df2 from the previous example
concatenated_keys = pd.concat([df1, df2], keys=['Set1', 'Set2'])
print("\nConcatenated DataFrame with keys:")
print(concatenated_keys)
# Accessing data using the MultiIndex
print("\nAccessing Set1:")
print(concatenated_keys.loc['Set1'])
print("\nAccessing specific row in Set2:")
print(concatenated_keys.loc[('Set2', 201)])

输出:

Concatenated DataFrame with keys:
Name SubjectID Score
Set1 101 Alex sub1 98
102 Amy sub2 90
103 Allen sub4 87
104 Alice sub6 69
Set2 201 Billy sub2 89
202 Brian sub4 80
203 Bran sub3 79
204 Bryce sub6 97
Accessing Set1:
Name SubjectID Score
101 Alex sub1 98
102 Amy sub2 90
103 Allen sub4 87
104 Alice sub6 69
Accessing specific row in Set2:
Name Billy
SubjectID sub2
Score 89
Name: (Set2, 201), dtype: object

如果原始索引没有意义,或者您想要一个简单的从 0 开始的索引,请使用 ignore_index=True。

# Using df1 and df2 from the previous example
concatenated_ignore = pd.concat([df1, df2], ignore_index=True)
print("\nConcatenated DataFrame with ignore_index=True:")
print(concatenated_ignore)
# Note: Using ignore_index=True overrides the keys argument if both are provided.

输出:

Concatenated DataFrame with ignore_index=True:
Name SubjectID Score
0 Alex sub1 98
1 Amy sub2 90
2 Allen sub4 87
3 Alice sub6 69
4 Billy sub2 89
5 Brian sub4 80
6 Bran sub3 79
7 Bryce sub6 97

注意:如果同时提供了 ignore_index=True 和 keys 参数,ignore_index=True 会覆盖 keys 参数。

要根据索引将 DataFrames 并排放置,请使用 axis=1。

import pandas as pd
df_a = pd.DataFrame({'A': [1, 2, 3], 'B': [4, 5, 6]}, index=['r1', 'r2', 'r3'])
df_b = pd.DataFrame({'C': [7, 8, 9], 'D': [10, 11, 12]}, index=['r1', 'r2', 'r3'])
df_c = pd.DataFrame({'E': [13, 14]}, index=['r1', 'r2']) # Mismatched index
print("DataFrame A:")
print(df_a)
print("\nDataFrame B:")
print(df_b)
print("\nDataFrame C (mismatched index):")
print(df_c)
# Concatenate A and B horizontally (aligned index)
concat_axis1 = pd.concat([df_a, df_b], axis=1)
print("\nConcatenated A and B (axis=1):")
print(concat_axis1)
# Concatenate A and C horizontally (mismatched index -> NaN)
concat_axis1_nan = pd.concat([df_a, df_c], axis=1) # join='outer' is default
print("\nConcatenated A and C (axis=1, outer join):")
print(concat_axis1_nan)
# Concatenate A and C using inner join
concat_axis1_inner = pd.concat([df_a, df_c], axis=1, join='inner')
print("\nConcatenated A and C (axis=1, inner join):")
print(concat_axis1_inner)

输出:

DataFrame A:
A B
r1 1 4
r2 2 5
r3 3 6
DataFrame B:
C D
r1 7 10
r2 8 11
r3 9 12
DataFrame C (mismatched index):
E
r1 13
r2 14
Concatenated A and B (axis=1):
A B C D
r1 1 4 7 10
r2 2 5 8 11
r3 3 6 9 12
Concatenated A and C (axis=1, outer join):
A B E
r1 1 4 13.0
r2 2 5 14.0
r3 3 6 NaN
Concatenated A and C (axis=1, inner join):
A B E
r1 1 4 13
r2 2 5 14

当 axis=1 时,join 参数控制是基于输入 DataFrames 索引的并集 (outer) 还是交集 (inner) 来保留行。

虽然连接处理的是组合现有对象,但 Pandas 还擅长创建日期和时间序列,这些序列通常在连接或合并之前用作 DataFrames 的索引。

import pandas as pd
current_time = pd.Timestamp.now()
print(f"Current Timestamp: {current_time}")
# Get only the date part
print(f"Current Date: {current_time.date()}")

输出 (会因运行时间而异):

Current Timestamp: 2023-10-27 10:30:00.123456
Current Date: 2023-10-27
import pandas as pd
# From a string
ts_from_string = pd.Timestamp('2024-01-15 09:00:00')
print(f"Timestamp from string: {ts_from_string}")
# From epoch seconds (specify unit)
ts_from_epoch = pd.Timestamp(1705309200, unit='s') # 15 Jan 2024 09:00:00 GMT
print(f"Timestamp from epoch: {ts_from_epoch}")
# With timezone
ts_with_tz = pd.Timestamp('2024-01-15 09:00:00', tz='America/New_York')
print(f"Timestamp with timezone: {ts_with_tz}")

输出:

Timestamp from string: 2024-01-15 09:00:00
Timestamp from epoch: 2024-01-15 09:00:00
Timestamp with timezone: 2024-01-15 09:00:00-05:00

pd.date_range 功能非常强大,用于创建日期或时间序列。

import pandas as pd
# Daily frequency for 5 periods
daily_index = pd.date_range(start='2024-03-01', periods=5, freq='D')
print("Daily Date Range:")
print(daily_index)
# Hourly frequency between two times
hourly_index = pd.date_range(start='2024-03-01 10:00', end='2024-03-01 14:00', freq='H')
print("\nHourly Date Range:")
print(hourly_index)
# Business days
business_days = pd.date_range(start='2024-03-01', periods=7, freq='B')
print("\nBusiness Days:")
print(business_days)

输出:

Daily Date Range:
DatetimeIndex(['2024-03-01', '2024-03-02', '2024-03-03', '2024-03-04',
'2024-03-05'],
dtype='datetime64[ns]', freq='D')
Hourly Date Range:
DatetimeIndex(['2024-03-01 10:00:00', '2024-03-01 11:00:00',
'2024-03-01 12:00:00', '2024-03-01 13:00:00',
'2024-03-01 14:00:00'],
dtype='datetime64[ns]', freq='H')
Business Days:
DatetimeIndex(['2024-03-01', '2024-03-04', '2024-03-05', '2024-03-06',
'2024-03-07', '2024-03-08', '2024-03-11'],
dtype='datetime64[ns]', freq='B')

pd.to_datetime 可以智能地将各种表示形式(字符串、epoch 秒、列表)转换为日期时间对象。如果传入类似列表的对象,它会返回一个 DatetimeIndex;如果传入 Series,则返回一个包含日期时间的 Series。

import pandas as pd
# From a Series containing various formats and None
date_series = pd.Series(['Jul 31, 2024', '2024-01-10 15:30', None, '2024/05/20'])
converted_series = pd.to_datetime(date_series)
print("Converted Series:")
print(converted_series)
# From a list
date_list = ['2024/11/23', '2025.12.31', None]
converted_index = pd.to_datetime(date_list)
print("\nConverted List to DatetimeIndex:")
print(converted_index)

输出:

Converted Series:
0 2024-07-31 00:00:00
1 2024-01-10 15:30:00
2 NaT
3 2024-05-20 00:00:00
dtype: datetime64[ns]
Converted List to DatetimeIndex:
DatetimeIndex(['2024-11-23', '2025-12-31', 'NaT'], dtype='datetime64[ns]', freq=None)

NaT (Not a Time) 表示缺失的时间戳值,类似于数值数据中的 NaN。