R - 数据框
R - 使用 Tidyverse 的现代数据帧
Section titled “R - 使用 Tidyverse 的现代数据帧”数据帧是 R 存储表格数据的主要结构。可以将其视为电子表格或数据库表,其中列代表变量,行代表观测值。列可以包含不同的数据类型(例如,数字、文本、日期)。
现代 R 生态系统,特别是 tidyverse,引入了 tibble,这是一种下一代数据帧。Tibbles 更用户友好:它们打印效果更好,避免更改变量名,并且默认情况下永远不会将字符串转换为因子。对于所有新的数据分析项目,强烈建议使用 tibbles。
创建现代数据帧 (Tibble)
Section titled “创建现代数据帧 (Tibble)”我们将使用 tibble 包(tidyverse 的一部分)中的 tibble() 函数。
# 首先,确保已安装并加载 tidyverse# install.packages("tidyverse")library(tibble)
# 创建一个 tibbleemp_data <- tibble( emp_id = 1:5, emp_name = c("Rick", "Dan", "Michelle", "Ryan", "Gary"), salary = c(623.3, 515.2, 611.0, 729.0, 843.25), start_date = as.Date(c("2012-01-01", "2013-09-23", "2014-11-15", "2014-05-11", "2015-03-27")))
# 打印 tibble。注意输出整洁!print(emp_data)执行此代码会生成整洁的格式化输出:
# A tibble: 5 × 4 emp_id emp_name salary start_date <int> <chr> <dbl> <date>1 1 Rick 623. 2012-01-012 2 Dan 515. 2013-09-233 3 Michelle 611 2014-11-154 4 Ryan 729 2014-05-115 5 Gary 843. 2015-03-27检查数据帧的结构
Section titled “检查数据帧的结构”基础 R 的 str()(结构)和 summary() 函数仍然完美运行。然而,tidyverse 提供了 glimpse(),它提供了一个紧凑且高度可读的摘要。
library(dplyr) # glimpse 在 dplyr 包中
# 使用 glimpse() 获取结构glimpse(emp_data)glimpse() 的输出旨在方便阅读:
Rows: 5Columns: 4$ emp_id <int> 1, 2, 3, 4, 5$ emp_name <chr> "Rick", "Dan", "Michelle", "Ryan", "Gary"$ salary <dbl> 623.30, 515.20, 611.00, 729.00, 843.25$ start_date <date> 2012-01-01, 2013-09-23, 2014-11-15, 2014-05-11, 2015-03-27使用 dplyr 操纵数据
Section titled “使用 dplyr 操纵数据”dplyr 包提供了一套强大且直观的动词,用于数据操作。我们经常使用管道运算符 |>(或 %>%)将这些动词链接在一起。管道将其左侧表达式的结果作为第一个参数传递给其右侧的函数。
使用 select() 选择列
Section titled “使用 select() 选择列”# 加载 dplyrlibrary(dplyr)
# 选择特定列name_and_salary <- emp_data |> select(emp_name, salary)
print(name_and_salary)使用 filter() 筛选行
Section titled “使用 filter() 筛选行”# 查找薪水超过 700 的员工high_earners <- emp_data |> filter(salary > 700)
print(high_earners)使用 mutate() 添加/修改列
Section titled “使用 mutate() 添加/修改列”# 添加新的部门列和计算出的奖金列emp_data_expanded <- emp_data |> mutate( dept = c("IT", "Operations", "IT", "HR", "Finance"), bonus = salary * 0.10 )
print(emp_data_expanded)使用 bind_rows() 添加行
Section titled “使用 bind_rows() 添加行”# 创建一个新的 tibble,包含新员工数据new_hires <- tibble( emp_id = 6:7, emp_name = c("Rasmi", "Pranab"), salary = c(578.0, 722.5), start_date = as.Date(c("2023-05-21", "2023-07-30")), dept = c("IT", "Operations"))
# 组合原始数据和新数据# bind_rows 比 rbind() 更安全all_employees <- bind_rows(emp_data_expanded, new_hires)
print(all_employees)综合运用:一个 dplyr 工作流程
Section titled “综合运用:一个 dplyr 工作流程”dplyr 的真正强大之处在于通过链式操作以可读的方式执行复杂查询。让我们找出 IT 部门中在 2013 年之后入职的员工姓名。
it_employees_post_2013 <- all_employees |> filter(dept == "IT", start_date > as.Date("2013-12-31")) |> select(emp_name, salary)
print(it_employees_post_2013)这个可读的命令链产生了最终的、经过筛选的结果:
# A tibble: 2 × 2 emp_name salary <chr> <dbl>1 Michelle 6112 Rasmi 578这仅仅是个开始!掌握 dplyr 的下一步包括学习如何对数据进行分组和汇总:
group_by():按一个或多个变量对数据进行分组,以便后续操作针对每个组单独执行。summarise():将每个组折叠成一行,计算诸如均值、总和或计数等汇总统计量。arrange():按一个或多个列对数据行进行排序。