R - 数据类型
R 语言:数据类型和结构
Section titled “R 语言:数据类型和结构”在 R 语言中,一切皆对象。变量不以特定数据类型声明;相反,它们的数据类型由赋值给它们的对象决定。理解基本数据类型和容纳它们的数据结构是有效使用 R 的关键。
R 有两种主要的数据对象:原子向量(atomic vectors)和列表(lists)。所有其他数据结构,如矩阵和数据框,都建立在它们之上。
原子向量类型
Section titled “原子向量类型”原子向量是 R 中最简单的数据结构,包含相同类型的数据元素。有六种主要的原子类型:
| 数据类型 | 描述与示例 | 使用 class() 检查 |
|---|---|---|
| logical | 布尔值。TRUE、FALSE。 | v <- TRUE print(class(v)) # [1] “logical” |
| numeric | 实数(浮点数)。12.3、5、-99。 | v <- 23.5 print(class(v)) # [1] “numeric” |
| integer | 整数。使用 L 后缀。2L、34L。 | v <- 2L print(class(v)) # [1] “integer” |
| character | 文本字符串。"hello"、'R'。 | v <- “R is fun” print(class(v)) # [1] “character” |
| complex | 带有虚部的复数。3 + 2i。 | v <- 2+5i print(class(v)) # [1] “complex” |
| raw | 存储原始字节。在数据分析中很少使用。 | v <- charToRaw(“Hi”) print(class(v)) # [1] “raw” |
核心数据结构
Section titled “核心数据结构”这些结构用于组织和存储原子类型的数据。
向量(Vectors)
Section titled “向量(Vectors)”向量是一维的数据元素序列,且其所有元素必须为相同类型。使用 c() 函数(用于组合)创建包含多个元素的向量。
# 创建一个字符向量fruit_colors <- c('red', 'green', 'yellow')print(fruit_colors)
# 获取向量的类print(class(fruit_colors))[1] "red" "green" "yellow"[1] "character"列表(Lists)
Section titled “列表(Lists)”列表是一种通用向量,可以包含不同类型的数据元素,包括其他列表、函数或数据框。
# 创建一个包含数值向量、字符串和函数的列表my_list <- list(numbers = c(2, 5, 3), name = "My List", summary_func = summary)
# 打印列表print(my_list)$numbers[1] 2 5 3
$name[1] "My List"
$summary_funcfunction (object, ..., maxsum = 7, digits = max(3L, getOption("digits") - 3L), ...) .Primitive("summary")矩阵(Matrices)
Section titled “矩阵(Matrices)”矩阵是二维矩形数据集。与向量类似,其所有元素必须是相同类型。
# 从向量创建一个 2x3 矩阵M <- matrix(c(1:6), nrow = 2, ncol = 3)print(M) [,1] [,2] [,3][1,] 1 3 5[2,] 2 4 6因子(Factors)
Section titled “因子(Factors)”因子是一种特殊类型的向量,用于存储分类数据。它将数据存储为整数编码的向量,以及一组相应的字符标签(水平/层级)。因子对于统计建模和一些绘图函数非常有用。
# 创建一个教育水平向量education_levels <- c('High School', 'Bachelor', 'Bachelor', 'Master', 'PhD', 'Master')
# 创建一个因子对象factor_education <- factor(education_levels)
# 打印因子及其水平print(factor_education)print(levels(factor_education))[1] High School Bachelor Bachelor Master PhD MasterLevels: Bachelor High School Master PhD[1] "Bachelor" "High School" "Master" "PhD"数据框和 Tibble
Section titled “数据框和 Tibble”数据框是 R 中存储表格数据最常见的方式。它们本质上是等长向量的列表,这意味着每列可以有不同的数据类型。tibble 是数据框的一种现代重构,由 tidyverse 包提供。Tibble 更适合现代数据科学,因为它们更严格并具有更方便的打印方法。
# 加载 tidyverse 包(或仅 tibble 包)# install.packages("tidyverse")library(tidyverse)
# 创建一个 tibble(现代数据框)patient_data <- tibble( gender = c("Male", "Male", "Female"), height_cm = c(152, 171.5, 165), weight_kg = c(81, 93, 78), age = c(42L, 38L, 26L) # 对整数使用 L 后缀)
print(patient_data)请注意 tibble 如何整洁地打印每列的数据类型:
# A tibble: 3 × 4 gender height_cm weight_kg age <chr> <dbl> <dbl> <int>1 Male 152 81 422 Male 171. 93 383 Female 165 78 26