Python Pandas - IO 工具
Pandas - 输入/输出工具 (I/O 工具)(读写数据)
Section titled “Pandas - 输入/输出工具 (I/O 工具)(读写数据)”Pandas 提供了丰富的函数集,用于从各种文件格式中读取数据到 Pandas 对象(主要是 DataFrame),并将数据写回文件。这些 I/O (Input/Output,输入/输出) 工具对于加载真实世界数据和保存结果至关重要。
用于读取基于文本的表格数据(通常称为平面文件)的最常用函数是 pd.read_csv() 和 pd.read_table()。
pd.read_csv():专为读取逗号分隔值 (CSV) 文件设计。这可以说是最常用的 I/O 函数。
pd.read_table():一个更通用的函数,用于读取分隔符文件,默认为使用制表符 (\t) 作为分隔符。通常,会使用指定了 sep 参数的 pd.read_csv() 来代替。
read_csv 的基本签名包含许多有用的参数:
pd.read_csv(filepath_or_buffer, sep=',', # Delimiter to use (default is comma) delimiter=None, # Alias for sep header='infer', # Row number(s) to use as column names (default: 0 if names not provided) names=None, # List of column names to use index_col=None, # Column(s) to use as row index labels usecols=None, # Subset of columns to read dtype=None, # Specify data type for columns (e.g., {'col': str}) engine=None, # Parser engine ('c' or 'python') skiprows=None, # Number of lines to skip at the start nrows=None, # Number of rows to read from the file parse_dates=False, # Specify columns to parse as dates ... # Many other options available )假设我们有一个名为 data.csv 的 CSV 文件,内容如下:
S.No,Name,Age,City,Salary1,Tom,28,Toronto,200002,Lee,32,HongKong,30003,Steven,43,Bay Area,83004,Ram,38,Hyderabad,3900(您需要将这些文本保存到一个名为 data.csv 的文件中,该文件应与您的脚本/notebook 位于同一目录,以便运行示例。)
使用 read_csv 进行基本读取
Section titled “使用 read_csv 进行基本读取”读取标准的 CSV 文件很简单:
import pandas as pd
try: # Read the CSV file into a DataFrame df = pd.read_csv("data.csv") print("DataFrame read from data.csv:") print(df)except FileNotFoundError: print("Error: data.csv not found. Please create the file with the sample data.")输出(如果 data.csv 存在):
DataFrame read from data.csv: S.No Name Age City Salary0 1 Tom 28 Toronto 200001 2 Lee 32 HongKong 30002 3 Steven 43 Bay Area 83003 4 Ram 38 Hyderabad 3900Pandas 自动从第一行推断出了头部,并使用了默认的逗号分隔符。
指定索引列 (index_col)
Section titled “指定索引列 (index_col)”如果 CSV 文件中的某一列应该用作 DataFrame 的索引,请使用 index_col 参数。
import pandas as pd
try: # Use the 'S.No' column as the index df_indexed = pd.read_csv("data.csv", index_col='S.No') print("DataFrame with 'S.No' as index:") print(df_indexed)except FileNotFoundError: print("Error: data.csv not found.")输出:
DataFrame with 'S.No' as index: Name Age City SalaryS.No1 Tom 28 Toronto 200002 Lee 32 HongKong 30003 Steven 43 Bay Area 83004 Ram 38 Hyderabad 3900指定数据类型 (dtype)
Section titled “指定数据类型 (dtype)”Pandas 会尝试推断数据类型,但有时您需要显式指定它们,特别是对于可能被错误解释的列(例如,ID 被读取为数字而非字符串)。使用 dtype 参数,并提供一个将列名映射到类型的字典。
import pandas as pdimport numpy as np # Needed for np.float64
try: # Explicitly set 'Salary' to float and 'Age' to string df_typed = pd.read_csv("data.csv", dtype={'Salary': np.float64, 'Age': str}) print("DataFrame with specified dtypes:") print(df_typed) print("\nData types:") print(df_typed.dtypes)except FileNotFoundError: print("Error: data.csv not found.")输出:
DataFrame with specified dtypes: S.No Name Age City Salary0 1 Tom 28 Toronto 20000.01 2 Lee 32 HongKong 3000.02 3 Steven 43 Bay Area 8300.03 4 Ram 38 Hyderabad 3900.0
Data types:S.No int64Name objectAge object # Note: 'object' often indicates string typeCity objectSalary float64dtype: object请注意,正如指定的那样,Salary 现在是 float64 类型,而 Age 是 object(字符串)类型。
自定义列名和处理头部 (names, header)
Section titled “自定义列名和处理头部 (names, header)”如果文件没有头部行,或者您想提供自己的列名,请使用 names 参数。如果文件确实有头部但您想覆盖它,您还需要指定 header=0 来指示第一行应该被视为数据(如果使用了 skiprows 则跳过)。
import pandas as pd
try: # Provide custom names AND specify that row 0 is the original header to replace custom_names = ['ID', 'EmployeeName', 'AgeYrs', 'Location', 'Pay'] df_named = pd.read_csv("data.csv", names=custom_names, header=0) print("DataFrame with custom names (original header replaced):") print(df_named)except FileNotFoundError: print("Error: data.csv not found.")输出:
DataFrame with custom names (original header replaced): ID EmployeeName AgeYrs Location Pay0 1 Tom 28 Toronto 200001 2 Lee 32 HongKong 30002 3 Steven 43 Bay Area 83003 4 Ram 38 Hyderabad 3900如果文件没有头部,您只需使用 names=custom_names 并省略 header=0(或设置为 header=None)。如果头部在不同的行(例如第 3 行),您应该使用 header=2(基于 0 的索引)。
跳过行 (skiprows)
Section titled “跳过行 (skiprows)”使用 skiprows 跳过文件开头的特定行数,或者提供一个要跳过的行索引列表(基于 0 的索引)。
import pandas as pd
try: # Skip the first two rows (header and first data row) # Note: This will likely cause the next row (Lee's data) to become the header df_skipped = pd.read_csv("data.csv", skiprows=2) print("DataFrame after skipping first 2 rows:") print(df_skipped)except FileNotFoundError: print("Error: data.csv not found.")输出:
DataFrame after skipping first 2 rows: 2 Lee 32 HongKong 30000 3 Steven 43 Bay Area 83001 4 Ram 38 Hyderabad 3900要跳过行并仍然使用原始头部或提供名称,请适当地将 skiprows 与 header 或 names 结合使用。例如,要跳过前 2 行并使用自定义名称,而不将第 2 行视为头部:pd.read_csv("data.csv", skiprows=2, header=None, names=custom_names)
将 DataFrame 写入文件 (.to_csv() 等)
Section titled “将 DataFrame 写入文件 (.to_csv() 等)”就像您读取数据一样,您也可以将 DataFrame 写入各种格式的文件。.to_csv() 方法是 read_csv 的对应操作。
import pandas as pd
# Assuming df_indexed exists from a previous stepdf_indexed = pd.DataFrame({'Name': ['Tom', 'Lee'], 'Age': [28, 32]}, index=[1, 2])df_indexed.index.name = 'S.No'
try: # Write df_indexed to a new CSV file # index=True (default) writes the index # index=False omits the index df_indexed.to_csv("output.csv", index=True) print("\nDataFrame successfully written to output.csv")
# Read it back to verify df_read_back = pd.read_csv("output.csv", index_col='S.No') print("\nReading back output.csv:") print(df_read_back)
except Exception as e: print(f"An error occurred during writing/reading: {e}")to_csv() 的关键参数包括 sep、header(布尔值或字符串列表)、index(布尔值)、mode('w' 用于写入,'a' 用于追加)、encoding(编码)等。
其他常用格式
Section titled “其他常用格式”Pandas 支持许多其他格式:
- Excel:
pd.read_excel()、df.to_excel()(需要openpyxl或xlrd库)。 - JSON:
pd.read_json()、df.to_json()。 - HTML:
pd.read_html()(从网页读取表格)、df.to_html()。 - SQL:
pd.read_sql()、df.to_sql()(需要SQLAlchemy和数据库驱动程序)。 - Parquet:
pd.read_parquet()、df.to_parquet()(需要pyarrow或fastparquet)。高效的列式格式,适用于大型数据集。 - Feather:
pd.read_feather()、df.to_feather()(需要pyarrow)。快速、轻量级的二进制格式。
格式的选择取决于数据源、存储需求、性能需求以及与其他系统的互操作性。CSV 仍然是基于文本交换的常用格式,而像 Parquet 这样的格式在处理大型数据集时性能更佳。