R - 速查表
现代 R 速查表
Section titled “现代 R 速查表”此速查表提供了现代 R 编程的简明概述,重点关注 tidyverse 生态系统,以实现高效且易读的数据分析。它涵盖了从基本语法到数据可视化和建模的关键概念。
1. R 环境
Section titled “1. R 环境”与您的 R 会话交互的核心命令。
# 获取函数帮助?meanhelp(mean)
# 从 CRAN 安装包install.packages("dplyr")
# 将包加载到会话中library(dplyr)
# 列出当前工作区中的对象ls()
# 移除对象rm(my_variable)2. 基本语法与运算符
Section titled “2. 基本语法与运算符”R 代码的构建块。
# 这是一个注释
# 变量赋值(<- 是首选约定)my_variable <- "Hello R!"my_number <- 100
# 管道运算符(来自 magrittr 包,随 tidyverse 加载)# 将一个函数的结果传递给下一个# x %>% f(y) 等同于 f(x, y)# 自 R 4.1 起,也提供了原生管道运算符 |>
# --- 算术运算符 ---# +、-、*、/、^(指数)、%%(取模)、%/%(整除)result <- (5 + 3) * 2 # 16
# --- 关系与逻辑运算符 ---# >、<、>=、<=、==(等于)、!=(不等于)# &(逐元素与)、|(逐元素或)、!(非)# &&(短路与)、||(短路或)TRUE & FALSE # FALSEx > 5 & x < 103. 核心数据结构
Section titled “3. 核心数据结构”R 存储数据的主要对象。
| 数据结构 | 描述 | 示例 |
|---|---|---|
| 向量 | 相同类型元素的 1D 序列。 | c(1, 2, 3) (numeric), c("a", "b", "c") (character) |
| 列表 | 任何类型元素的 1D 集合。 | list(name = "John", age = 30, scores = c(88, 95)) |
| 矩阵 | 相同类型元素的 2D 数组。 | matrix(1:6, nrow = 2, ncol = 3) |
| 数据框 | 一种 2D、类似表格的结构,其中列可以具有不同的类型。 | data.frame(id = 1:3, name = c("A", "B", "C")) |
| Tibble | 一种现代的、增强型数据框(来自 tidyverse)。提供更好的打印和子集操作。 | tibble(id = 1:3, name = c("A", "B", "C")) |
| 因子 | 带有已定义水平的分类数据向量。 | factor(c("low", "high", "low"), levels = c("low", "medium", "high")) |
4. 数据导入与导出
Section titled “4. 数据导入与导出”使用 readr 和 readxl 包进行快速可靠的数据导入。
library(readr)library(readxl)library(writexl)
# 读取 CSV 文件(返回一个 tibble)my_data_csv <- read_csv("path/to/my_file.csv")
# 读取 Excel 文件(按名称或编号指定工作表)my_data_excel <- read_excel("path/to/my_file.xlsx", sheet = "Sheet1")
# 将数据写入 CSV 文件write_csv(my_data_csv, "output/data.csv")
# 将数据写入 Excel 文件write_xlsx(my_data_excel, "output/data.xlsx")5. 使用 dplyr 进行数据转换
Section titled “5. 使用 dplyr 进行数据转换”dplyr 用于数据操作的核心动词,常与管道 %>% 一起使用。
library(dplyr)
# 示例数据集data <- starwars
# 过滤行droids <- data %>% filter(species == "Droid")
# 选择列name_and_height <- data %>% select(name, height, mass)
# 创建或修改列heavy_hitters <- data %>% mutate(bmi = mass / ((height/100)^2))
# 排序行sorted_by_height <- data %>% arrange(desc(height))
# 汇总数据,常与 group_by 结合使用species_summary <- data %>% filter(!is.na(mass)) %>% group_by(species) %>% summarize( count = n(), avg_mass = mean(mass) ) %>% arrange(desc(count))6. 使用 ggplot2 进行数据可视化
Section titled “6. 使用 ggplot2 进行数据可视化”ggplot2 使用分层图形语法。您通过添加图层来构建绘图。
library(ggplot2)
# ggplot 的结构:# ggplot(data = <数据>, mapping = aes(<映射>)) + <几何函数>()
# 散点图ggplot(data = starwars, aes(x = mass, y = height)) + geom_point(aes(color = gender)) + # 添加颜色映射 labs( title = "Height vs. Mass of Star Wars Characters", x = "Mass (kg)", y = "Height (cm)", caption = "Data from the dplyr package" ) + theme_minimal()
# 物种计数柱状图starwars %>% count(species, sort = TRUE) %>% filter(n > 1) %>% ggplot(aes(x = reorder(species, n), y = n)) + geom_col() + coord_flip() + # 翻转坐标以提高可读性 labs(title = "Most Common Species", x = "Species", y = "Count")7. 控制流与函数
Section titled “7. 控制流与函数”# If-else 语句if (x > 10) { print("x is large")} else { print("x is small")}
# For 循环for (i in 1:5) { print(i)}
# 编写自定义函数calculate_bmi <- function(mass_kg, height_m) { if (!is.numeric(mass_kg) || !is.numeric(height_m)) { stop("Inputs must be numeric.") } bmi <- mass_kg / (height_m ^ 2) return(bmi)}
# 使用函数calculate_bmi(70, 1.75)8. 基本建模
Section titled “8. 基本建模”拟合和检查简单的线性模型。
# 拟合线性模型:根据重量预测 mpgmodel <- lm(mpg ~ wt, data = mtcars)
# 查看模型摘要summary(model)
# 使用 'broom' 获取整洁的模型输出library(broom)tidy(model) # 系数以 tibble 形式呈现glance(model) # 模型摘要统计数据以 tibble 形式呈现9. 可复现分析
Section titled “9. 可复现分析”| 工具 | 用途 |
|---|---|
| R Markdown (.Rmd) | 一种用于创建 R 动态文档的文件格式。将代码、其输出(图表、表格)和叙述性文本结合成高质量的报告(HTML、PDF、Word)。 |
renv 包 | 一个依赖管理工具。它创建项目特定的包库,确保您的代码现在和将来都能为您和他人以相同的方式运行。 |