Skip to content

R - 字符串

处理文本数据是任何编程语言中的常见任务。在 R 中,字符串是使用单引号 (') 或双引号 (") 创建的一段文本。虽然基础 R 提供了字符串操作函数,但现代且推荐的方法是使用 stringr 包,它是 tidyverse 的一部分。stringr 提供了一套更一致、直观和强大的工具。

首先,让我们了解基础知识。R 在内部将所有字符串存储为双引号。

# 这些都是有效的字符串
string1 <- 'A string with single quotes.'
string2 <- "A string with double quotes."
string3 <- "You can include 'single' quotes inside double quotes."
string4 <- 'You can include "double" quotes inside single quotes.'
print(string1)
print(string4)

输出:

[1] "A string with single quotes."
[1] "You can include \"double\" quotes inside single quotes."

使用 stringr 进行现代字符串操作

Section titled “使用 stringr 进行现代字符串操作”

让我们探索 stringr 包中与基础 R 函数对应的常用功能。stringr 函数很容易识别,因为它们都以 str_ 开头。首先,确保 stringr 已安装并加载。

# install.packages("stringr") # 如果您没有安装,请运行一次
library(stringr)

str_c() 是 paste() 的现代替代方案。它在将向量合并为单个字符串时更直观。

first_name <- "Ada"
last_name <- "Lovelace"
# 使用分隔符拼接
full_name <- str_c(first_name, last_name, sep = " ")
print(full_name)
# 将字符串向量合并为一个字符串
colors <- c("red", "green", "blue")
color_list <- str_c(colors, collapse = ", ")
print(color_list)

输出:

[1] "Ada Lovelace"
[1] "red, green, blue"

字符串长度和子字符串提取:str_length() 和 str_sub()

Section titled “字符串长度和子字符串提取:str_length() 和 str_sub()”
my_string <- "R for Data Science"
# 获取字符数量
print(str_length(my_string))
# 提取从位置 7 到 10 的子字符串
# 注意: str_sub() 支持负数索引,表示从字符串末尾开始计数
print(str_sub(my_string, start = 7, end = 10))
# 提取最后 7 个字符
print(str_sub(my_string, start = -7))

输出:

[1] 18
[1] "Data"
[1] "Science"

这些对于清理混乱的文本数据至关重要。

messy_string <- " Some Text with Mixed Case and whitespace "
# 转换为小写
print(str_to_lower(messy_string))
# 转换为大写
print(str_to_upper(messy_string))
# 移除开头和结尾的空白字符
clean_string <- str_trim(messy_string)
print(clean_string)

输出:

[1] " some text with mixed case and whitespace "
[1] " SOME TEXT WITH MIXED CASE AND WHITESPACE "
[1] "Some Text with Mixed Case and whitespace"

stringr 擅长使用正则表达式 (regex) 进行模式匹配。这是一种查找、提取和替换文本的强大方法。

sentences <- c("The price is $19.99.", "The time is 10:30 PM.", "No price here.")
# 检测字符串是否包含某个模式(例如,美元符号)
# '\$' 需要用另一个 '\' 转义,因为 '$' 是一个特殊的正则表达式字符
print(str_detect(sentences, "\\$"))
# 提取模式的第一个匹配项(例如,一个数字)
# '\\d+' 表示“一个或多个数字”
print(str_extract(sentences, "\\d+"))
# 替换模式
print(str_replace(sentences, "price", "cost"))

输出:

[1] TRUE FALSE TRUE
[1] "19" "10" NA
[1] "The cost is $19.99." "The time is 10:30 PM." "No cost here."

最佳实践:用于格式化的 glue 包

Section titled “最佳实践:用于格式化的 glue 包”

虽然 format() 可用,但 glue 包提供了一种更现代、更易读的方式来将变量直接嵌入字符串中。这通常被称为字符串插值 (string interpolation)。

# install.packages("glue") # 运行一次
library(glue)
user_name <- "Alex"
item_count <- 5
item_cost <- 15.75
total_cost <- item_count * item_cost
# 使用 glue 创建格式化字符串
# {} 内部的变量会被自动评估和嵌入
message <- glue(
"Hello, {user_name}!\n"
,"You have ordered {item_count} items.\n"
,"The total cost will be ${total_cost}."
)
cat(message) # 使用 cat() 进行清晰打印

输出:

Hello, Alex!
You have ordered 5 items.
The total cost will be $78.75.