Python Pandas - 基本功能
Pandas - 基本功能
Section titled “Pandas - 基本功能”现在我们了解了如何创建 Pandas Series 和 DataFrame,接下来让我们探讨一些用于检查和理解这些数据结构的基本属性(attribute)和方法(method)。由于 DataFrame 在数据分析中更为普遍,我们将主要关注它。
Series 的基本属性和方法
Section titled “Series 的基本属性和方法”以下是 Pandas Series 的一些重要属性和方法:
| 属性/方法 | 描述 |
|---|---|
.index / .axes | 返回 Series 的索引 (轴标签)。.axes 返回一个包含索引的列表。 |
.dtype | 返回 Series 中元素的 数据类型 (dtype)。 |
.empty | 如果 Series 为空 (元素数量为零),返回 True,否则返回 False。 |
.ndim | 返回维度数量 (Series 始终为 1)。 |
.size | 返回 Series 中元素的总数量。 |
.shape | 返回表示维度的元组 (例如,长度为 n 的 Series 为 (n,))。 |
.values / .to_numpy() | 将 Series 数据作为 NumPy 数组返回。.to_numpy() 是推荐的现代方法。 |
.head(n=5) | 返回前 n 个元素 (默认为 5)。 |
.tail(n=5) | 返回后 n 个元素 (默认为 5)。 |
我们来创建一个示例 Series 并演示这些属性和方法:
import pandas as pdimport numpy as np
# Create a sample Seriess = pd.Series(np.random.randn(5), index=['a', 'b', 'c', 'd', 'e'], name='SampleData')print("Sample Series:")print(s)输出(值会变化):
Sample Series:a 0.967853b -0.148368c -1.395906d -1.758394e 0.562421Name: SampleData, dtype: float64.index / .axes
Section titled “.index / .axes”print("\nIndex:", s.index)print("Axes:", s.axes)输出:
Index: Index(['a', 'b', 'c', 'd', 'e'], dtype='object')Axes: [Index(['a', 'b', 'c', 'd', 'e'], dtype='object')].dtype
Section titled “.dtype”print("\nDtype:", s.dtype)输出:
Dtype: float64.empty
Section titled “.empty”print("\nIs empty?", s.empty)print("Is pd.Series(dtype='float64').empty empty?", pd.Series(dtype='float64').empty)输出:
Is empty? FalseIs pd.Series(dtype='float64').empty empty? Trueprint("\nNumber of dimensions:", s.ndim)输出:
Number of dimensions: 1.size / .shape
Section titled “.size / .shape”print("\nSize (number of elements):", s.size)print("Shape (dimensionality tuple):", s.shape)输出:
Size (number of elements): 5Shape (dimensionality tuple): (5,).values / .to_numpy()
Section titled “.values / .to_numpy()”print("\nValues (as NumPy array using .values):") # 值 (.values 返回的 NumPy 数组):print(s.values)print("\nValues (as NumPy array using .to_numpy()):") # 值 (.to_numpy() 返回的 NumPy 数组):print(s.to_numpy())输出(值将与 Series 匹配):
Values (as NumPy array using .values):[ 0.967853 -0.148368 -1.395906 -1.758394 0.562421 ]
Values (as NumPy array using .to_numpy()):[ 0.967853 -0.148368 -1.395906 -1.758394 0.562421 ]注意:在现代 Pandas 中,通常推荐使用 .to_numpy() 而非 .values,以获得更好的清晰度和一致性。
.head() & .tail()
Section titled “.head() & .tail()”print("\nFirst 2 elements (.head(2)):") # 前 2 个元素 (.head(2)):print(s.head(2))print("\nLast 2 elements (.tail(2)):") # 后 2 个元素 (.tail(2)):print(s.tail(2))输出:
First 2 elements (.head(2)):a 0.967853b -0.148368Name: SampleData, dtype: float64
Last 2 elements (.tail(2)):d -1.758394e 0.562421Name: SampleData, dtype: float64DataFrame 的基本属性和方法
Section titled “DataFrame 的基本属性和方法”DataFrame 与 Series 共享许多属性,但也有与其二维特性相关的特定属性:
| 属性/方法 | 描述 |
|---|---|
.T | 转置 DataFrame (交换行和列)。 |
.index | 返回 DataFrame 的索引 (行标签)。 |
.columns | 返回 DataFrame 的列标签。 |
.axes | 返回一个包含索引 (轴 0) 和列 (轴 1) 的列表。 |
.dtypes | 返回一个 Series,包含每列的数据类型。 |
.empty | 如果 DataFrame 为空 (0 行或 0 列),返回 True。 |
.ndim | 返回维度数量 (DataFrame 始终为 2)。 |
.size | 返回元素的总数量 (行 * 列)。 |
.shape | 返回表示维度的元组 (行数, 列数)。 |
.values / .to_numpy() | 将 DataFrame 数据作为 2D NumPy 数组返回。推荐使用 .to_numpy()。 |
.head(n=5) | 返回前 n 行。 |
.tail(n=5) | 返回后 n 行。 |
.info() | 打印一个简洁的摘要,包括索引 dtype、列 dtypes、非空值数量和内存使用情况。 |
我们来创建一个示例 DataFrame:
import pandas as pdimport numpy as np
# Create a dictionary for the DataFramedata = { 'Name': pd.Series(['Tom', 'James', 'Ricky', 'Vin', 'Steve', 'Smith', 'Jack']), 'Age': pd.Series([25, 26, 25, 23, 30, 29, 23]), 'Rating': pd.Series([4.23, 3.24, 3.98, 2.56, 3.20, 4.6, 3.8])}
# Create the DataFramedf = pd.DataFrame(data)print("Sample DataFrame:")print(df)输出:
Sample DataFrame: Name Age Rating0 Tom 25 4.231 James 26 3.242 Ricky 25 3.983 Vin 23 2.564 Steve 30 3.205 Smith 29 4.606 Jack 23 3.80.T (转置)
Section titled “.T (转置)”print("\nTransposed DataFrame (.T):") # 转置后的 DataFrame (.T):print(df.T)输出:
Transposed DataFrame (.T): 0 1 2 3 4 5 6Name Tom James Ricky Vin Steve Smith JackAge 25 26 25 23 30 29 23Rating 4.23 3.24 3.98 2.56 3.2 4.6 3.8.index、.columns、.axes
Section titled “.index、.columns、.axes”print("\nIndex (Row Labels):", df.index) # 索引 (行标签):print("Column Labels:", df.columns) # 列标签:print("Axes (List of Index, Columns):", df.axes) # 轴 (索引和列的列表):输出:
Index (Row Labels): RangeIndex(start=0, stop=7, step=1)Column Labels: Index(['Name', 'Age', 'Rating'], dtype='object')Axes (List of Index, Columns): [RangeIndex(start=0, stop=7, step=1), Index(['Name', 'Age', 'Rating'], dtype='object')].dtypes
Section titled “.dtypes”返回一个 Series,显示每列的数据类型。
print("\nData types of each column (.dtypes):") # 每列的数据类型 (.dtypes):print(df.dtypes)输出:
Data types of each column (.dtypes):Name objectAge int64Rating float64dtype: object.empty、.ndim、.size、.shape
Section titled “.empty、.ndim、.size、.shape”print("\nIs empty?", df.empty) # 是否为空?:print("Number of dimensions:", df.ndim) # 维度数量:print("Size (total elements):", df.size) # 大小 (总元素数):print("Shape (rows, columns):", df.shape) # 形状 (行数, 列数):输出:
Is empty? FalseNumber of dimensions: 2Size (total elements): 21Shape (rows, columns): (7, 3).values / .to_numpy()
Section titled “.values / .to_numpy()”print("\nValues (as 2D NumPy array using .to_numpy()):") # 值 (使用 .to_numpy() 返回的二维 NumPy 数组):print(df.to_numpy())输出:
Values (as 2D NumPy array using .to_numpy()):[['Tom' 25 4.23] ['James' 26 3.24] ['Ricky' 25 3.98] ['Vin' 23 2.56] ['Steve' 30 3.2] ['Smith' 29 4.6] ['Jack' 23 3.8]].head() & .tail()
Section titled “.head() & .tail()”print("\nFirst 2 rows (.head(2)):") # 前 2 行 (.head(2)):print(df.head(2))print("\nLast 2 rows (.tail(2)):") # 后 2 行 (.tail(2)):print(df.tail(2))输出:
First 2 rows (.head(2)): Name Age Rating0 Tom 25 4.231 James 26 3.24
Last 2 rows (.tail(2)): Name Age Rating5 Smith 29 4.66 Jack 23 3.8.info()
Section titled “.info()”提供 DataFrame 结构的快速概览。
print("\nDataFrame Info (.info()):") # DataFrame 信息 (.info()):df.info()输出:
DataFrame Info (.info()):<class 'pandas.core.frame.DataFrame'>RangeIndex: 7 entries, 0 to 6Data columns (total 3 columns): # Column Non-Null Count Dtype--- ------ -------------- ----- 0 Name 7 non-null object 1 Age 7 non-null int64 2 Rating 7 non-null float64dtypes: float64(1), int64(1), object(1)memory usage: 296.0+ bytes