Python 3 - 字符串
Python 3 - 字符串
Section titled “Python 3 - 字符串”字符串是 Python 中最基本和最常用的数据类型之一。它们表示字符序列。你可以通过将字符括在单引号 (’) 或双引号 (”) 中来创建字符串。Python 将这两种引号视为相同。将字符串赋给变量非常直接:
message = 'Hello World!'subject = "Python Programming"你也可以使用三引号 (''' 或 """) 来创建多行字符串或包含单引号和双引号而无需转义的字符串。
multi_line_string = '''This is a stringthat spans multiplelines.'''
quoted_string = """He said, 'Python is fun!'"""访问字符串中的值
Section titled “访问字符串中的值”字符串是序列,这意味着它们的字符是有序的。你可以使用索引和切片(使用方括号 [])来访问单个字符或字符串的一部分(子字符串)。
索引从 0 开始,代表第一个字符。负数索引从末尾计数(-1 是最后一个字符)。
切片使用 [start:end] 提取字符串的一部分。start 索引是包含的,而 end 索引是排除的。省略 start 默认为开头,省略 end 默认为末尾。
#!/usr/bin/env python3
message = 'Hello World!'subject = "Python Programming"
# 访问单个字符print(f"message[0]: {message[0]}") # Output: Hprint(f"message[-1]: {message[-1]}") # Output: !
# 切片以获取子字符串print(f"subject[0:6]: {subject[0:6]}") # Output: Pythonprint(f"subject[7:]: {subject[7:]}") # Output: Programmingprint(f"subject[:6]: {subject[:6]}") # Output: Pythonprint(f"message[::2]: {message[::2]}") # Output: HloWrd (every second char)执行上面的代码会产生:
message[0]: Hmessage[-1]: !subject[0:6]: Pythonsubject[7:]: Programmingsubject[:6]: Pythonmessage[::2]: HloWrdPython 中的字符串是不可变的,这意味着你不能直接更改现有字符串(例如,你不能将新字符赋给特定索引)。要“更新”字符串,你需要创建一个基于原始字符串的新字符串,可能将其部分内容与新内容组合。
#!/usr/bin/env python3
message = 'Hello World!'print(f"Original message: {message}")
# 通过切片和连接创建新字符串updated_message = message[:6] + 'Python!'print(f"Updated message: {updated_message}")
# 尝试就地修改(会引发错误)# message[0] = 'J' # 这行代码会引发 TypeError执行有效代码会产生:
Original message: Hello World!Updated message: Hello Python!某些字符在前面加上反斜杠 (\) 时具有特殊含义。这些被称为转义序列,允许你在字符串字面量中包含换行符、制表符或字面反斜杠和引号等字符。
常见的转义序列:
| 序列 | 含义 |
|---|---|
| \ | 反斜杠 () |
| ‘ | 单引号 (’) |
| “ | 双引号 (”) |
| \n | ASCII 换行 (Newline) |
| \t | ASCII 横向制表符 (Horizontal Tab) |
| \r | ASCII 回车 (Carriage Return) |
| \b | ASCII 退格 (Backspace) |
| \f | ASCII 换页 (Formfeed) |
| \ooo | 八进制值为 ooo 的字符 |
| \xhh | 十六进制值为 hh 的字符 |
示例:
print("This string contains a newline\nand a tab\t character.")print('You can include \'single quotes\' or \"double quotes\".')print("To print a literal backslash, use \\.")字符串特殊运算符
Section titled “字符串特殊运算符”字符串支持几个运算符:
| 运算符 | 描述 | 示例 (a = 'Hello', b = 'Python') |
|---|---|---|
| + | 连接 - 将字符串连接在一起。 | a + ' ' + b 结果为 'Hello Python' |
| * | 重复 - 重复字符串。 | a * 2 结果为 'HelloHello' |
| [] | 索引 - 访问单个字符。 | a[1] 结果为 'e' |
| [:] | 切片 - 提取子字符串。 | a[1:4] 结果为 'ell' |
| in | 成员测试 - 检查子字符串是否存在于字符串中。 | 'H' in a 结果为 True |
| not in | 成员测试 - 检查子字符串是否不存在。 | 'M' not in a 结果为 True |
如果你需要阻止转义序列的解释,可以在开头的引号前加上 r 或 R 来创建一个原始字符串。这对于正则表达式或 Windows 文件路径特别有用。
# 普通字符串,其中 \n 是换行符print('C:\some\name')
# 原始字符串,其中 \n 被视为字面字符print(r'C:\some\name')输出:
C:\someameC:\some\name注意:原始字符串不能以奇数个反斜杠结尾。
字符串格式化
Section titled “字符串格式化”Python 提供了几种格式化字符串的方法,可以将值嵌入其中。虽然旧的 % 运算符仍然存在(类似于 C 语言的 printf),但现代且首选的方法是 f-字符串(格式化字符串字面量)和 str.format() 方法。
f-字符串 (Formatted String Literals)
Section titled “f-字符串 (Formatted String Literals)”f-字符串于 Python 3.6 引入,以 f 或 F 为前缀,允许你直接在字符串字面量中使用 {} 嵌入表达式。它们简洁且通常是速度最快的方法。
#!/usr/bin/env python3
name = 'Alice'age = 30height = 1.68
print(f"My name is {name} and I am {age} years old.")print(f"{name}'s height is {height:.2f} meters.") # 将 height 格式化为 2 位小数print(f"In 5 years, {name} will be {age + 5}.")输出:
My name is Alice and I am 30 years old.Alice's height is 1.68 meters.In 5 years, Alice will be 35.str.format() 方法
Section titled “str.format() 方法”format() 方法可以在字符串字面量或变量上调用。在字符串中使用占位符 {},这些占位符将由传递给 format() 的参数填充。
#!/usr/bin/env python3
name = 'Bob'age = 25
# 使用位置参数print("My name is {} and I am {} years old.".format(name, age))
# 使用关键字参数print("My name is {n} and I am {a} years old.".format(n=name, a=age))
# 使用索引参数print("My name is {0} and I am {1} years old. {0} likes Python.".format(name, age))输出:
My name is Bob and I am 25 years old.My name is Bob and I am 25 years old.My name is Bob and I am 25 years old. Bob likes Python.f-字符串和 str.format() 都支持强大的迷你语言,用于控制对齐、填充、数字格式化等。你可以在官方 Python 文档中找到详细信息。
旧式格式化 (%)
Section titled “旧式格式化 (%)”% 运算符也可以用于格式化字符串。它使用格式说明符,如 %s(字符串)、%d(整数)和 %f(浮点数)。虽然你可能会在旧代码中遇到这种方式,但对于新代码,推荐使用 f-字符串或 str.format(),因为它们具有更好的可读性和灵活性。
#!/usr/bin/env python3
name = 'Charlie'age = 42
print("My name is %s and weight is %d kg!" % (name, age))输出:
My name is Charlie and weight is 42 kg!Unicode 字符串
Section titled “Unicode 字符串”在 Python 3 中,所有字符串默认都是 Unicode 字符串。这意味着 Python 可以无缝处理来自不同语言和符号的各种字符。你无需执行任何特殊操作(如 Python 2 中使用的 u'' 前缀)来创建 Unicode 字符串。
#!/usr/bin/env python3
unicode_string = "你好世界"emoji_string = "Python is fun 👍"
print(unicode_string)print(emoji_string)虽然字符串是 Unicode 的,但你可能需要将字符串 encode(编码)为特定的字节表示形式(如 UTF-8),以便存储到文件或通过网络发送,或者将字节串 decode(解码)回字符串。
#!/usr/bin/env python3
text = "résumé"encoded_utf8 = text.encode('utf-8') # 编码为字节串decoded_text = encoded_utf8.decode('utf-8') # 解码回字符串
print(f"Original: {text}")print(f"Encoded (bytes): {encoded_utf8}")print(f"Decoded: {decoded_text}")常用内置字符串方法
Section titled “常用内置字符串方法”Python 字符串自带了许多用于常见操作的内置方法。请记住,这些方法会返回新字符串,因为字符串是不可变的。
以下是一些常用的方法:
| 方法签名 | 描述 |
|---|---|
capitalize() | 返回一个副本,其中第一个字符大写,其余小写。 |
lower() | 返回一个副本,其中所有字符转换为小写。 |
upper() | 返回一个副本,其中所有字符转换为大写。 |
title() | 返回一个标题格式的副本(每个单词的首字母大写,其余小写)。 |
swapcase() | 返回一个副本,其中大写字符转换为小写,小写字符转换为大写。 |
strip([chars]) | 返回一个副本,移除开头和结尾的空白字符(或指定的 chars)。 |
lstrip([chars]) | 返回一个副本,移除开头的空白字符(或指定的 chars)。 |
rstrip([chars]) | 返回一个副本,移除结尾的空白字符(或指定的 chars)。 |
replace(old, new[, count]) | 返回一个副本,将 old 子字符串的出现替换为 new。可选的 count 限制替换次数。 |
split([sep[, maxsplit]]) | 返回一个由单词/子字符串组成的列表,按指定的 sep 分隔符(默认为空白字符)进行分割。可选的 maxsplit 限制分割次数。 |
join(iterable) | 使用字符串本身作为分隔符,将 iterable(如列表)中的元素连接成一个字符串。 |
find(sub[, start[, end]]) | 返回子字符串 sub 在切片 [start:end] 中找到的最低索引。如果未找到,返回 -1。 |
index(sub[, start[, end]]) | 类似于 find(),但如果找不到 sub,则引发 ValueError。 |
count(sub[, start[, end]]) | 返回子字符串 sub 在切片 [start:end] 中非重叠出现的次数。 |
startswith(prefix[, start[, end]]) | 如果字符串以指定的 prefix 开头,返回 True,否则返回 False。 |
endswith(suffix[, start[, end]]) | 如果字符串以指定的 suffix 结尾,返回 True,否则返回 False。 |
isdigit() | 如果所有字符都是数字且至少有一个字符,返回 True。 |
isalpha() | 如果所有字符都是字母且至少有一个字符,返回 True。 |
isalnum() | 如果所有字符都是字母或数字且至少有一个字符,返回 True。 |
islower() | 如果所有大小写字母都是小写且至少有一个大小写字母,返回 True。 |
isupper() | 如果所有大小写字母都是大写且至少有一个大小写字母,返回 True。 |
isspace() | 如果所有字符都是空白字符且至少有一个字符,返回 True。 |
istitle() | 如果字符串是标题格式且至少有一个字符,返回 True。 |
使用示例:
#!/usr/bin/env python3
text = " Python programming is FUN! "
print(f"Original: '{text}'")print(f"Stripped: '{text.strip()}'")print(f"Lower: '{text.lower()}'")print(f"Upper: '{text.upper()}'")print(f"Replaced: '{text.replace('FUN', 'Awesome')}'")print(f"Split: {text.split()}")
words = ['Python', 'is', 'cool']separator = ' 'print(f"Joined: '{separator.join(words)}'")
print(f"Starts with ' P': {text.startswith(' P')}")print(f"Ends with '! ': {text.endswith('! ')}")print(f"Count of 'm': {text.count('m')}")print(f"Find 'program': {text.find('program')}")