Skip to content

机器学习项目的数据加载

任何机器学习项目的第一步都是加载数据。数据格式多种多样,但对于表格数据(类似于电子表格,按行和列组织的数据)来说,最常见和通用的格式之一是 CSV (Comma-Separated Values,逗号分隔值)。

CSV 文件是一种纯文本文件,其中列中的值通常用逗号分隔,每行代表一个数据记录。虽然简单,但有效加载 CSV 文件需要考虑文件结构的几个方面。

在加载 CSV 文件之前,检查它(例如,在文本编辑器或电子表格软件中打开)并考虑以下几点会很有帮助:

文件是否包含表头行?第一行通常包含每列的名称。了解是否存在表头至关重要:

  • 存在表头: 大多数加载函数可以自动使用表头行来指定列名。
  • 无表头: 如果没有表头,您可能需要手动提供列名,或让加载函数指定默认名称(例如,数字索引)。

通常需要指定是否期望存在表头。

虽然“逗号”是 CSV 的标准分隔符,但有时文件会使用其他字符来分隔值,例如制表符(\t)、分号(;)或空格。加载数据时必须指定正确的分隔符;否则,文件将无法正确解析。

一些数据文件包含注释,通常以特定字符(如 #)开头。这些行在加载过程中通常应该被忽略。加载函数通常有参数来指定注释字符。

包含分隔符本身的字段(例如,包含逗号的文本字段)通常用引号括起来,通常是双引号(")。加载函数需要正确处理这些引号,以准确解析字段。如果引号字符不是标准的,您可能需要指定它。

CSV 文件将所有内容存储为文本。加载时,您通常希望将数值列解释为数字(整数或浮点数),而不是字符串。优秀的加载工具通常会尝试自动推断数据类型,但有时需要手动指定,特别是对于日期或特定的数值格式。

文件中如何表示缺失值(例如,空字符串、‘NA’、‘NaN’、’?’)?加载函数通常允许指定哪些字符串应被视为缺失值(在 Pandas 和 NumPy 等库中表示为 NaN - Not a Number)。

Python 提供了几种加载 CSV 数据的方法。虽然标准库提供了基本工具,但像 NumPy,尤其是 Pandas 这样的库提供了更强大和方便的方法来进行数据分析和机器学习准备。

1. 使用 Pandas(推荐用于机器学习)

Section titled “1. 使用 Pandas(推荐用于机器学习)”

Pandas 是 Python 中用于机器学习数据处理的基石。它的 read_csv() 函数非常灵活和高效,可将数据直接加载到 DataFrame 对象中,这非常适合后续的分析、预处理以及馈送给机器学习模型。

示例 1:加载 Iris 数据集(带表头)

Iris 数据集是经典的机器学习数据集,通常存储时包含表头行。

import pandas as pd
# URL for the Iris dataset
url_iris = 'https://raw.githubusercontent.com/mwaskom/seaborn-data/master/iris.csv'
try:
# Load the CSV file directly from the URL
# Pandas automatically detects the header and delimiter (comma)
iris_df = pd.read_csv(url_iris)
# --- Inspect the loaded data ---
print("-- Iris Dataset (using Pandas) ---")
print(f"Shape: {iris_df.shape}") # (rows, columns)
print("\nFirst 5 rows:\n", iris_df.head())
print("\nData Types:\n", iris_df.dtypes)
print("\nBasic Statistics:\n", iris_df.describe())
except Exception as e:
print(f"Error loading Iris dataset: {e}")

示例 2:加载 Pima Indians Diabetes 数据集(无表头)

此数据集通常不包含表头行,因此我们需要提供列名。

import pandas as pd
# URL for the Pima Indians Diabetes dataset
url_pima = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/pima-indians-diabetes.data.csv'
# Define column names as the file doesn't have a header
names = ['preg', 'plas', 'pres', 'skin', 'test', 'mass', 'pedi', 'age', 'class']
try:
# Load the CSV, specifying no header and providing names
pima_df = pd.read_csv(url_pima, header=None, names=names)
# --- Inspect the loaded data ---
print("\n--- Pima Indians Diabetes Dataset (using Pandas) ---")
print(f"Shape: {pima_df.shape}")
print("\nFirst 5 rows:\n", pima_df.head())
print("\nData Types:\n", pima_df.dtypes)
except Exception as e:
print(f"Error loading Pima dataset: {e}")

Pandas 的优点: 返回 DataFrame,处理混合数据类型,强大的数据处理功能,与 Scikit-learn 和其他机器学习库无缝集成。对于机器学习工作流程而言,通常是最佳选择。

Pandas read_csv 文档:https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_csv.html

NumPy 提供了 loadtxt() 函数,适用于将纯数值数据加载到 NumPy 数组中。与 Pandas 的 read_csv 相比,它的灵活性较低,难以处理混合数据类型或表头。

示例:加载 Pima Indians Diabetes 数据集(仅限数值数据)

import numpy as np
from urllib.request import urlopen # Needed to open URL for loadtxt
# URL for the Pima Indians Diabetes dataset
url_pima = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/pima-indians-diabetes.data.csv'
try:
# Open the URL and load data using loadtxt
# Requires specifying the delimiter
raw_data = urlopen(url_pima)
# Note: loadtxt expects bytes, so decode might be needed depending on source/python version
# We'll load directly assuming numerical data
data = np.loadtxt(raw_data, delimiter=",")
# --- Inspect the loaded data ---
print("\n--- Pima Indians Diabetes Dataset (using NumPy) ---")
print(f"Shape: {data.shape}")
print(f"Data Type: {data.dtype}")
print("\nFirst 3 rows:\n", data[:3])
except Exception as e:
print(f"Error loading Pima dataset with NumPy: {e}")

NumPy 的优点: 对于纯数值数据来说很简单,直接加载到 NumPy 数组中,对于数值计算很高效。

缺点: 对于带有表头、注释或混合数据类型的文件不太方便。错误处理可能不如 Pandas 信息丰富。

3. 使用 Python 标准库(csv 模块)

Section titled “3. 使用 Python 标准库(csv 模块)”

Python 内置的 csv 模块提供了按行或按字符串列表读取和写入 CSV 文件的基本工具。它需要更多手动工作来处理表头、数据类型转换以及加载到 NumPy 数组或 Pandas DataFrame 等结构化格式中。

示例:加载 Iris 数据集(手动转换)

import csv
import numpy as np
from urllib.request import urlopen
import io # To handle decoding from URL stream
# URL for the Iris dataset
url_iris = 'https://raw.githubusercontent.com/mwaskom/seaborn-data/master/iris.csv'
data_list = []
header = []
try:
response = urlopen(url_iris)
# Decode bytes to text and use StringIO to treat it like a file
csvfile = io.StringIO(response.read().decode('utf-8'))
# Use csv.reader
reader = csv.reader(csvfile, delimiter=',')
# Read header
header = next(reader)
# Read data rows (needs manual conversion for numeric types)
for row in reader:
# Attempt to convert numeric columns to float, handle species separately
if len(row) == 5: # Basic check for expected number of columns
try:
# Convert first 4 columns to float
numeric_part = [float(item) for item in row[:4]]
# Keep species name as string
species = row[4]
# Combine (example: store numeric part only for simplicity here)
data_list.append(numeric_part)
except ValueError:
print(f"Skipping row due to conversion error: {row}")
# Convert list of lists to NumPy array if needed
data_np = np.array(data_list)
# --- Inspect the loaded data ---
print("\n--- Iris Dataset (using csv module) ---")
print(f"Header: {header}")
print(f"Shape of numeric data: {data_np.shape}")
print("\nFirst 3 rows (numeric):\n", data_np[:3])
except Exception as e:
print(f"Error loading Iris dataset with csv module: {e}")

优点: 内置,无需外部依赖。对读取过程有精细控制。

缺点: 需要大量手动编码进行解析、类型转换和数据结构化以供分析。对于典型的机器学习任务来说,远不如 Pandas 方便。

总而言之,虽然 Python 提供了多种加载 CSV 数据的方法,但对于大多数机器学习项目,强烈推荐使用 Pandas 的 read_csv,因为它具有灵活性、高效性,并且为后续的数据处理和分析提供了强大的 DataFrame 结构。