Ruby 多线程
Ruby - 多线程与并发
Section titled “Ruby - 多线程与并发”并发(Concurrency)允许程序看似同时处理多个任务,从而提高响应性和吞吐量。多线程(Multithreading)是实现并发的一种方式,它将程序分成多个可以独立运行的执行线程(threads)。
Ruby 提供了 Thread 类用于创建和管理线程。虽然 Ruby 中的线程可以并发执行代码的不同部分,但了解在标准 Ruby 实现 MRI (Matz’s Ruby Interpreter) 中的全局解释器锁(Global Interpreter Lock, GIL)的作用非常重要。
MRI 中的全局解释器锁(GIL):
Section titled “MRI 中的全局解释器锁(GIL):”在 MRI 中,GIL 确保在单个进程中,任何给定时刻只有一个线程可以执行 Ruby 代码。这简化了 C 扩展 API 和一些 Ruby 内部工作,但也意味着 CPU 密集型任务(大部分时间用于计算的任务)在使用单个 MRI 进程内的线程时,无法在多核处理器上实现真正的并行性。然而,I/O 密集型任务(大部分时间等待外部操作,如网络请求或文件访问的任务)可以显著受益于线程,因为一个线程在等待时可以释放 GIL,允许其他线程运行。
对于 MRI 中 CPU 密集型任务的真正并行性,您可以考虑使用多个进程(例如,通过 fork 或 parallel 等库)或探索 Ractors(Ruby 3.0 中引入),这是一种为并行性设计的新并发模型。其他 Ruby 实现,如 JRuby 或 TruffleRuby,没有 GIL,可以并行运行多个线程。
创建 Ruby 线程:
Section titled “创建 Ruby 线程:”要启动一个新线程,将一个代码块关联到 Thread.new。原始线程会在创建新线程后立即继续执行。
# Main thread is running hereThread.new do # New thread (Thread #2) runs this code # ...end# Main thread (Thread #1) continues to run this code#!/usr/bin/env ruby
def task_one 3.times do |i| puts "Task One at iteration #{i}: #{Time.now}" sleep(0.5) # 模拟 I/O 密集型工作 endend
def task_two 3.times do |j| puts "Task Two at iteration #{j}: #{Time.now}" sleep(0.3) # 模拟 I/O 密集型工作 endend
puts "Program Started At #{Time.now}"
t1 = Thread.new { task_one() }t2 = Thread.new { task_two() }
# 等待线程完成,然后主程序退出t1.joint2.join
puts "Program Ended At #{Time.now}"这将产生显示交错执行的输出(确切时间会有所不同):
Program Started At 2023-10-27 10:00:00 +0000Task One at iteration 0: 2023-10-27 10:00:00 +0000Task Two at iteration 0: 2023-10-27 10:00:00 +0000Task Two at iteration 1: 2023-10-27 10:00:00 +0000Task One at iteration 1: 2023-10-27 10:00:01 +0000Task Two at iteration 2: 2023-10-27 10:00:01 +0000Task One at iteration 2: 2023-10-27 10:00:01 +0000Program Ended At 2023-10-27 10:00:02 +0000线程生命周期:
Section titled “线程生命周期:”线程通过 Thread.new(或其别名 Thread.start、Thread.fork)创建。它们会自动开始执行。Thread#value 方法会等待线程完成并返回其代码块中最后一个表达式的值。Thread#join 也等待完成,但返回线程对象本身。
Thread.current 返回当前正在执行线程的 Thread 对象。Thread.main 返回主线程对象。
线程与异常:
Section titled “线程与异常:”如果一个未处理的异常在线程(主线程除外)中发生,该线程默认会静默终止。如果另一个线程在该终止的线程上调用 join 或 value,原始异常将在调用线程中重新抛出。
要使任何线程中的任何未处理异常都终止整个 Ruby 程序,请设置 Thread.abort_on_exception = true。或者,为了更精细的控制,您可以按每个线程设置:a_thread_instance.abort_on_exception = true。
更好的做法通常是在每个线程内处理异常或使用健壮的错误报告机制。
线程局部变量:
Section titled “线程局部变量:”线程共享全局变量和实例变量。线程代码块内的局部变量是该线程独有的。对于需要通过名称访问的线程特定数据,Ruby 线程提供了类似哈希的接口,使用 [] 和 []=:Thread.current[:my_key] = value。
#!/usr/bin/env ruby
shared_counter = 0threads = []
5.times do |i| threads << Thread.new do # 使用线程局部变量存储初始值 Thread.current[:initial_shared_value] = shared_counter # 如果不对 shared_counter 本身进行同步,这里会发生竞态条件 local_copy = shared_counter sleep(rand * 0.01) # 短暂延迟,增加发生竞态条件的机会 shared_counter = local_copy + 1 puts "Thread #{i} set shared_counter. Initial: #{Thread.current[:initial_shared_value]}, Final attempt: #{shared_counter}" endend
threads.each(&:join)puts "Final shared_counter = #{shared_counter}" # Likely not 5 due to race condition上面的示例说明了线程局部存储,但也强调了当多个线程访问共享可变数据而不同步时可能出现的竞态条件。最终的 shared_counter 可能不是 5。
线程优先级:
Section titled “线程优先级:”线程优先级(thread.priority = value)可以影响调度,但其行为可能依赖于操作系统,通常不是控制执行顺序的可靠方法。更推荐依赖显式同步机制。
线程排他:互斥锁(Mutex)与同步
Section titled “线程排他:互斥锁(Mutex)与同步”当多个线程共享和修改数据时,您需要同步以防止竞态条件并确保数据一致性。Mutex(互斥,Mutual Exclusion)是实现此目的的常用工具。同一时间只有一个线程可以持有 Mutex 的锁。
#!/usr/bin/env rubyrequire 'thread' # 在现代 Ruby 中,Mutex/ConditionVariable 不一定需要显式 require 'thread'
mutex = Mutex.newsafe_counter = 0threads = []
5.times do |i| threads << Thread.new do 1000.times do # 临界区:同一时间只有一个线程可以执行此处的代码 mutex.synchronize do safe_counter += 1 end end endend
threads.each(&:join)puts "Final safe_counter = #{safe_counter}" # Should be 5000mutex.synchronize 块确保其中的代码相对于使用同一个 mutex 的其他线程是原子执行的。
处理死锁与条件变量:
Section titled “处理死锁与条件变量:”死锁发生在两个或多个线程被永久阻塞,每个线程都在等待另一个线程持有的资源时。仔细设计锁的顺序可以防止某些死锁。ConditionVariable 与 Mutex 一起使用,用于管理复杂的同步场景,其中线程需要等待某个条件变为真。
一个线程可以在条件变量上等待(cv.wait(mutex)),这会原子地释放 mutex 并使线程进入睡眠。另一个线程在改变条件后,可以通知一个等待的线程(cv.signal)或所有等待的线程(cv.broadcast)唤醒。被唤醒的线程随后会重新获取 mutex。
#!/usr/bin/env rubyrequire 'thread' # Mutex/ConditionVariable 不一定需要显式 require 'thread'
mutex = Mutex.newresource_available = ConditionVariable.newitem = nil
producer = Thread.new do puts "Producer: Waiting to produce..." sleep(1) # 模拟工作 mutex.synchronize do item = "Hello from producer!" puts "Producer: Produced item. Signaling consumer." resource_available.signal # 发出信号表示物品已准备好 endend
consumer = Thread.new do mutex.synchronize do puts "Consumer: Waiting for item..." while item.nil? resource_available.wait(mutex) # 等待生产者发出的信号 end puts "Consumer: Received item: '#{item}'" item = nil # 消费物品 endend
producer.joinconsumer.joinputs "Finished."线程可以处于各种状态。thread.status 返回一个表示其状态的字符串("run"、"sleep"、"aborting"),如果是正常终止则返回 false,如果因未处理的异常终止则返回 nil。
| Status Value | 描述 |
|---|---|
| run | 线程当前正在运行或处于可运行状态。 |
| sleep | 线程处于睡眠状态或等待 I/O 操作或同步原语。 |
| aborting | 线程正在中止过程中。 |
| false (Boolean) | 线程已正常终止。 |
| nil (NilClass) | 线程因未处理的异常而终止。 |
Ractors:真正的并行性(Ruby 3.0+)
Section titled “Ractors:真正的并行性(Ruby 3.0+)”Ruby 3.0 引入了 Ractors (Ruby Actors),这是一种基于 Actor 模型的并发机制,旨在通过允许多个 Ractors 在不同的 CPU 核心上同时运行 Ruby 代码来实现真正的并行性,它们之间没有 GIL。Ractors 通过消息传递进行通信,并且对共享对象有严格限制以确保线程安全。虽然全面讨论 Ractors 超出了本概览的范围,但它们代表了 Ruby 中 CPU 密集型并行处理的一个重大进步。
Thread 类和实例方法:
Section titled “Thread 类和实例方法:”Thread 类提供了各种类方法(例如,Thread.list、Thread.stop)和实例方法(例如,thr.join、thr.value、thr.kill)。请查阅官方 Ruby 文档以获取完整列表:https://ruby-doc.org/core/Thread.html
- 由于 GIL 的存在,MRI 中的线程适用于 I/O 密集型并发。
- 使用
Mutex来同步对共享数据的访问。 - 使用
ConditionVariable实现线程间复杂的协作。 - 在线程内部处理异常,或谨慎配置
abort_on_exception。 - 对于 MRI 中的 CPU 密集型并行处理,考虑使用多个进程或 Ractors(Ruby 3.0+)。
- 其他 Ruby 实现(JRuby, TruffleRuby)提供真正的线程并行性。