R - 字符串
R - 现代字符串操作
Section titled “R - 现代字符串操作”处理文本数据是任何编程语言中的常见任务。在 R 中,字符串是使用单引号 (') 或双引号 (") 创建的一段文本。虽然基础 R 提供了字符串操作函数,但现代且推荐的方法是使用 stringr 包,它是 tidyverse 的一部分。stringr 提供了一套更一致、直观和强大的工具。
入门:基础 R 字符串
Section titled “入门:基础 R 字符串”首先,让我们了解基础知识。R 在内部将所有字符串存储为双引号。
# 这些都是有效的字符串string1 <- 'A string with single quotes.'string2 <- "A string with double quotes."string3 <- "You can include 'single' quotes inside double quotes."string4 <- 'You can include "double" quotes inside single quotes.'
print(string1)print(string4)输出:
[1] "A string with single quotes."[1] "You can include \"double\" quotes inside single quotes."使用 stringr 进行现代字符串操作
Section titled “使用 stringr 进行现代字符串操作”让我们探索 stringr 包中与基础 R 函数对应的常用功能。stringr 函数很容易识别,因为它们都以 str_ 开头。首先,确保 stringr 已安装并加载。
# install.packages("stringr") # 如果您没有安装,请运行一次library(stringr)字符串拼接: str_c()
Section titled “字符串拼接: str_c()”str_c() 是 paste() 的现代替代方案。它在将向量合并为单个字符串时更直观。
first_name <- "Ada"last_name <- "Lovelace"
# 使用分隔符拼接full_name <- str_c(first_name, last_name, sep = " ")print(full_name)
# 将字符串向量合并为一个字符串colors <- c("red", "green", "blue")color_list <- str_c(colors, collapse = ", ")print(color_list)输出:
[1] "Ada Lovelace"[1] "red, green, blue"字符串长度和子字符串提取:str_length() 和 str_sub()
Section titled “字符串长度和子字符串提取:str_length() 和 str_sub()”my_string <- "R for Data Science"
# 获取字符数量print(str_length(my_string))
# 提取从位置 7 到 10 的子字符串# 注意: str_sub() 支持负数索引,表示从字符串末尾开始计数print(str_sub(my_string, start = 7, end = 10))
# 提取最后 7 个字符print(str_sub(my_string, start = -7))输出:
[1] 18[1] "Data"[1] "Science"更改大小写和去除空白字符
Section titled “更改大小写和去除空白字符”这些对于清理混乱的文本数据至关重要。
messy_string <- " Some Text with Mixed Case and whitespace "
# 转换为小写print(str_to_lower(messy_string))
# 转换为大写print(str_to_upper(messy_string))
# 移除开头和结尾的空白字符clean_string <- str_trim(messy_string)print(clean_string)输出:
[1] " some text with mixed case and whitespace "[1] " SOME TEXT WITH MIXED CASE AND WHITESPACE "[1] "Some Text with Mixed Case and whitespace"使用正则表达式进行模式匹配
Section titled “使用正则表达式进行模式匹配”stringr 擅长使用正则表达式 (regex) 进行模式匹配。这是一种查找、提取和替换文本的强大方法。
sentences <- c("The price is $19.99.", "The time is 10:30 PM.", "No price here.")
# 检测字符串是否包含某个模式(例如,美元符号)# '\$' 需要用另一个 '\' 转义,因为 '$' 是一个特殊的正则表达式字符print(str_detect(sentences, "\\$"))
# 提取模式的第一个匹配项(例如,一个数字)# '\\d+' 表示“一个或多个数字”print(str_extract(sentences, "\\d+"))
# 替换模式print(str_replace(sentences, "price", "cost"))输出:
[1] TRUE FALSE TRUE[1] "19" "10" NA[1] "The cost is $19.99." "The time is 10:30 PM." "No cost here."最佳实践:用于格式化的 glue 包
Section titled “最佳实践:用于格式化的 glue 包”虽然 format() 可用,但 glue 包提供了一种更现代、更易读的方式来将变量直接嵌入字符串中。这通常被称为字符串插值 (string interpolation)。
# install.packages("glue") # 运行一次library(glue)
user_name <- "Alex"item_count <- 5item_cost <- 15.75total_cost <- item_count * item_cost
# 使用 glue 创建格式化字符串# {} 内部的变量会被自动评估和嵌入message <- glue( "Hello, {user_name}!\n" ,"You have ordered {item_count} items.\n" ,"The total cost will be ${total_cost}.")
cat(message) # 使用 cat() 进行清晰打印输出:
Hello, Alex!You have ordered 5 items.The total cost will be $78.75.