Skip to content

使用 Python 进行 AI – 计算机视觉

使用 Python 实现 AI – 计算机视觉

Section titled “使用 Python 实现 AI – 计算机视觉”

计算机视觉 (Computer Vision) 是人工智能 (Artificial Intelligence, AI) 的一个领域,它使计算机和系统能够从数字图像、视频和其他视觉输入中获取有意义的信息,并基于这些信息采取行动或提出建议。它旨在利用计算机软件和硬件来模拟和复制人类视觉。

计算机视觉旨在从二维图像或图像序列中重建、解释和理解三维场景或其组成部分。它涉及获取、处理、分析和理解数字图像,以从现实世界中提取高维数据,从而产生数值或符号信息,例如以决策的形式。

计算机视觉层次结构(处理级别)

Section titled “计算机视觉层次结构(处理级别)”

计算机视觉任务可以大致分为不同的复杂性级别:

  • 低级视觉 (Low-level Vision):涉及基本的图像处理操作,如降噪、对比度增强和特征提取(例如,边缘、角点、纹理)。
  • 中级视觉 (Mid-level Vision):侧重于图像分割(将图像分割成多个区域)、物体识别(识别物体)和运动跟踪等任务。
  • 高级视觉 (High-level Vision):处理场景理解、活动识别以及解释视觉场景中描绘的意图或行为。

图像处理 (Image Processing) 通常涉及图像到图像的转换。输入是图像,输出也是图像(例如,增强图像、去除噪声)。

计算机视觉 (Computer Vision) 通常将图像处理技术作为第一步,但更进一步。其目标是从图像或视频中生成物理对象和场景的明确、有意义的描述或解释。输出通常是符号信息、决策或场景模型。

计算机视觉在各个行业都有广泛的应用:

  • 定位与导航 (Localization & Navigation):帮助机器人确定位置并自主导航(例如,SLAM - 即时定位与地图构建)。
  • 避障 (Obstacle Avoidance):检测并避开环境中的障碍物。
  • 装配与操作 (Assembly & Manipulation):引导机械臂完成焊接、喷漆或抓取放置物体等任务。
  • 人机交互 (Human-Robot Interaction, HRI):使机器人能够更自然地感知和与人互动。

医疗保健与医学 (Healthcare & Medicine)

Section titled “医疗保健与医学 (Healthcare & Medicine)”
  • 医学图像分析 (Medical Image Analysis):分类和检测异常(例如,X 光、CT 扫描、MRI 中的肿瘤)。
  • 2D/3D 分割 (Segmentation):勾勒医学图像中的器官或组织。
  • 3D 器官重建 (Reconstruction):从 MRI 或超声波等扫描创建 3D 模型。
  • 手术辅助 (Surgical Assistance):视觉引导的机器人手术。
  • 生物识别技术 (Biometrics):人脸识别、虹膜扫描、指纹识别。
  • 监控 (Surveillance):检测可疑活动、未经授权的访问或人群监控。
  • 自动驾驶汽车 (Autonomous Vehicles):自动驾驶汽车使用计算机视觉进行感知、车道检测、交通标志识别和行人检测。
  • 驾驶辅助系统 (ADAS):诸如车道偏离警告、自适应巡航控制等功能。
  • 交通流量监控 (Traffic Monitoring):分析交通流量并检测事件。
  • 质量控制 (Quality Control):制造业中用于缺陷检测的自动化检查。
  • 流水线自动化 (Assembly Line Automation):引导机器人并验证装配步骤。
  • 条形码与包装标签读取 (Barcode & Package Label Reading):自动化分拣和跟踪。
  • 文档分析 (OCR):用于数字化文本的光学字符识别 (Optical Character Recognition)。

安装计算机视觉常用的 Python 软件包

Section titled “安装计算机视觉常用的 Python 软件包”

Python 中用于计算机视觉任务的主要库是 OpenCV (Open Source Computer Vision Library)。它提供了大量的算法和功能,用于实时计算机视觉。

OpenCV 使用 C++ 编写,并提供了 Python、Java 和 MATLAB 的接口。要安装 OpenCV 的 Python 绑定,请使用 pip:

pip install opencv-python

此包包含 OpenCV 的主要模块。对于附加模块(例如,一些专利算法),您可能需要 opencv-contrib-python:

pip install opencv-contrib-python

如果您正在使用 Anaconda,可以从 conda-forge 通道安装 OpenCV:

conda install -c conda-forge opencv

我们还将使用 NumPy 进行数组操作,NumPy 通常作为 OpenCV 的依赖项安装;以及 Matplotlib 用于在某些情况下显示图像(尽管 OpenCV 有自己的显示功能)。

大多数计算机视觉应用程序都涉及读取图像作为输入,处理它们,有时还会生成图像作为输出。让我们介绍一下基本操作。

OpenCV 提供了用于这些任务的函数:

  • cv2.imread(filepath, flags):从文件读取图像。flags 参数指定颜色模式(例如,cv2.IMREAD_COLOR、cv2.IMREAD_GRAYSCALE)。
  • cv2.imshow(window_name, image):在窗口中显示图像。需要配合 cv2.waitKey() 和 cv2.destroyAllWindows() 进行适当处理。
  • cv2.imwrite(filepath, image):将图像保存到文件。

首先,确保您的工作目录中有一个图像文件(例如,my_image.jpg),或者提供完整路径。

import cv2
import numpy as np
# Define image path (replace with your image file)
image_path = 'my_image.jpg' # Make sure this image exists
# Try to load the image
# As a fallback, create a dummy image if the specified one is not found
img = cv2.imread(image_path, cv2.IMREAD_COLOR)
if img is None:
print(f"Error: Could not read image at {image_path}. Using a dummy image instead.")
# Create a dummy 300x400 color image (BGR format)
img = np.zeros((300, 400, 3), dtype=np.uint8)
cv2.putText(img, 'Dummy Image', (50, 150), cv2.FONT_HERSHEY_SIMPLEX,
1.5, (0, 255, 0), 2, cv2.LINE_AA)
# Display the image
cv2.imshow('Original Image', img)
print("Press any key to close the image window...")
# waitKey(0) waits indefinitely for a key press
cv2.waitKey(0)
# Save the image in a different format (e.g., PNG)
output_path = 'output_image.png'
success = cv2.imwrite(output_path, img)
if success:
print(f"Image successfully saved to {output_path}")
else:
print(f"Error: Failed to save image to {output_path}")
# Destroy all OpenCV windows
cv2.destroyAllWindows()

运行这段代码时,一个图像(要么是 my_image.jpg,要么是一个带有文本的绿色虚拟图像)将显示在一个标题为 ‘Original Image’ 的窗口中。按下任意键将关闭窗口,脚本将尝试将图像保存为 output_image.png。

色彩空间转换 (Color Space Conversion)

Section titled “色彩空间转换 (Color Space Conversion)”

OpenCV 默认以 BGR (Blue, Green, Red) 顺序读取图像,这与 Matplotlib 等库使用的更常见的 RGB 顺序不同。通常需要将图像在不同色彩空间之间进行转换(例如,BGR 转灰度,BGR 转 HSV)。cv2.cvtColor(input_image, flag) 函数用于此目的。

示例:将 BGR 转换为灰度 (Grayscale)

Section titled “示例:将 BGR 转换为灰度 (Grayscale)”
import cv2
import numpy as np
image_path = 'my_image.jpg' # Use the same image or a new one
img_bgr = cv2.imread(image_path, cv2.IMREAD_COLOR)
if img_bgr is None:
print(f"Error: Could not read image at {image_path}. Using a dummy BGR image.")
img_bgr = np.random.randint(0, 256, (300, 400, 3), dtype=np.uint8)
# Convert BGR image to Grayscale
img_gray = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2GRAY)
# Display both images
cv2.imshow('BGR Image', img_bgr)
cv2.imshow('Grayscale Image', img_gray)
print("Press any key to close image windows...")
cv2.waitKey(0)
cv2.destroyAllWindows()

这个脚本将显示两个窗口:一个显示原始彩色图像(OpenCV 加载的 BGR 格式),另一个显示其灰度版本。灰度图像是单通道的,通常用于简化处理或用于不需要颜色信息的算法。

边缘是图像强度中显著的局部变化,代表重要的结构信息。边缘检测是许多计算机视觉应用中的基本步骤。Canny 边缘检测器是一种流行且有效的算法。

import cv2
import numpy as np
image_path = 'my_image.jpg'
img_color = cv2.imread(image_path, cv2.IMREAD_COLOR)
if img_color is None:
print(f"Error: Could not read image at {image_path}. Using a dummy image.")
# Create a dummy image with some shapes for edge detection
img_color = np.zeros((300, 400, 3), dtype=np.uint8)
cv2.rectangle(img_color, (50, 50), (150, 150), (0, 255, 0), -1) # Green square
cv2.circle(img_color, (250, 100), 50, (0, 0, 255), -1) # Red circle
# Convert to grayscale first, as Canny works on single-channel images
img_gray_for_canny = cv2.cvtColor(img_color, cv2.COLOR_BGR2GRAY)
# Apply Canny edge detection
# The two thresholds (100 and 200) are for hysteresis thresholding.
# Adjust these thresholds based on your image content.
edges = cv2.Canny(img_gray_for_canny, threshold1=100, threshold2=200)
# Display the original and edge-detected images
cv2.imshow('Original Image for Edges', img_color)
cv2.imshow('Canny Edges', edges)
print("Press any key to close image windows...")
cv2.waitKey(0)
cv2.destroyAllWindows()
# Optionally save the edges image
cv2.imwrite('edges_output.jpg', edges)

这个程序将显示原始图像和另一个图像,后者在黑色背景上以白色突出显示检测到的边缘。Canny 算法包括降噪、梯度计算、非极大值抑制和滞后阈值处理,以产生清晰的边缘。

使用 Haar Cascades 进行人脸检测 (Face Detection)

Section titled “使用 Haar Cascades 进行人脸检测 (Face Detection)”

人脸检测是一个经典的计算机视觉任务。OpenCV 提供了预训练的 Haar cascade 分类器,用于检测人脸、眼睛和微笑等对象。这些分类器基于机器学习,并使用 Haar-like 特征。

OpenCV 包含各种预训练 Haar cascades 的 XML 文件。您需要提供相应 XML 文件的路径(例如,haarcascade_frontalface_default.xml)。

import cv2
import numpy as np
# Load the pre-trained Haar Cascade for frontal face detection
# Use cv2.data.haarcascades to get the path to OpenCV's built-in cascades
cascade_path = cv2.data.haarcascades + 'haarcascade_frontalface_default.xml'
face_cascade = cv2.CascadeClassifier(cascade_path)
if face_cascade.empty():
raise IOError('Unable to load the face cascade classifier xml file')
# Load an image with faces (replace 'image_with_faces.jpg')
image_path_faces = 'image_with_faces.jpg'
img_faces = cv2.imread(image_path_faces)
if img_faces is None:
print(f"Error: Could not read image at {image_path_faces}. Creating a dummy image.")
# Create a dummy image (you'd ideally use an image with actual faces for testing)
img_faces = np.zeros((480, 640, 3), dtype=np.uint8)
# Draw a few dummy "faces" (simple rectangles)
cv2.rectangle(img_faces, (100, 100), (200, 220), (200,200,200), -1) # A gray box
cv2.rectangle(img_faces, (300, 150), (450, 300), (200,200,200), -1) # Another gray box
# Convert to grayscale (face detection is typically done on grayscale images)
img_gray_faces = cv2.cvtColor(img_faces, cv2.COLOR_BGR2GRAY)
# Detect faces in the image
# detectMultiScale(image, scaleFactor, minNeighbors)
# scaleFactor: How much the image size is reduced at each image scale.
# minNeighbors: How many neighbors each candidate rectangle should have to retain it.
faces_detected = face_cascade.detectMultiScale(img_gray_faces, scaleFactor=1.1, minNeighbors=5)
print(f"Found {len(faces_detected)} faces!")
# Draw rectangles around the detected faces
for (x, y, w, h) in faces_detected:
cv2.rectangle(img_faces, (x, y), (x + w, y + h), (0, 255, 0), 3) # Green rectangle
# Display the image with detected faces
cv2.imshow('Faces Detected', img_faces)
print("Press any key to close the image window...")
cv2.waitKey(0)
cv2.destroyAllWindows()
cv2.imwrite('faces_detected_output.jpg', img_faces)

当您运行此代码时,如果在 image_with_faces.jpg(或虚拟图像)中存在人脸,它们将用绿色矩形突出显示。scaleFactor 和 minNeighbors 参数可能需要根据图像内容进行调整以获得最佳性能。

使用 Haar Cascades 进行眼睛检测 (Eye Detection)

Section titled “使用 Haar Cascades 进行眼睛检测 (Eye Detection)”

与人脸检测类似,也可以使用 Haar cascades 进行眼睛检测。这通常在人脸检测之后作为辅助步骤进行,以在检测到的人脸区域内定位眼睛,从而提高准确性。

import cv2
import numpy as np
face_cascade_path = cv2.data.haarcascades + 'haarcascade_frontalface_default.xml'
eye_cascade_path = cv2.data.haarcascades + 'haarcascade_eye.xml'
# For eyes with glasses, 'haarcascade_eye_tree_eyeglasses.xml' might be better
face_cascade_eyes = cv2.CascadeClassifier(face_cascade_path)
eye_cascade = cv2.CascadeClassifier(eye_cascade_path)
if face_cascade_eyes.empty() or eye_cascade.empty():
raise IOError('Unable to load Haar cascade xml files')
image_path_eyes = 'image_with_faces.jpg' # Use an image with clear faces and eyes
img_for_eyes = cv2.imread(image_path_eyes)
if img_for_eyes is None:
print(f"Error: Could not read {image_path_eyes}. Using dummy.")
img_for_eyes = np.zeros((480, 640, 3), dtype=np.uint8)
# Dummy face region
cv2.rectangle(img_for_eyes, (100,100), (300,350), (200,200,200), -1)
# Dummy eye regions within the face
cv2.circle(img_for_eyes, (160,180), 15, (50,50,50), -1)
cv2.circle(img_for_eyes, (240,180), 15, (50,50,50), -1)
img_gray_for_eyes = cv2.cvtColor(img_for_eyes, cv2.COLOR_BGR2GRAY)
# First, detect faces
faces = face_cascade_eyes.detectMultiScale(img_gray_for_eyes, 1.1, 5)
# Then, for each detected face, detect eyes
for (fx, fy, fw, fh) in faces:
cv2.rectangle(img_for_eyes, (fx, fy), (fx + fw, fy + fh), (255, 0, 0), 2) # Blue for face
roi_gray = img_gray_for_eyes[fy:fy+fh, fx:fx+fw] # Region of Interest (face) in grayscale
roi_color = img_for_eyes[fy:fy+fh, fx:fx+fw] # Region of Interest (face) in color
eyes = eye_cascade.detectMultiScale(roi_gray, scaleFactor=1.1, minNeighbors=5)
for (ex, ey, ew, eh) in eyes:
cv2.rectangle(roi_color, (ex, ey), (ex + ew, ey + eh), (0, 255, 0), 2) # Green for eyes
cv2.imshow('Eyes Detected', img_for_eyes)
print("Press any key to close...")
cv2.waitKey(0)
cv2.destroyAllWindows()
cv2.imwrite('eyes_detected_output.jpg', img_for_eyes)

这个脚本首先检测人脸(蓝色矩形),然后仅在这些人脸区域内搜索眼睛(绿色矩形)。这种分层方法通常更鲁棒。

现代计算机视觉:深度学习方法

Section titled “现代计算机视觉:深度学习方法”

虽然像 Haar Cascades 这样的经典方法很有用,但现代计算机视觉严重依赖于深度学习 (Deep Learning),特别是卷积神经网络 (Convolutional Neural Networks, CNNs)。CNNs 在以下任务中取得了最先进的性能:

  • 图像分类 (Image Classification):为整个图像分配一个标签(例如,‘猫’、‘狗’、‘汽车’)。
  • 目标检测 (Object Detection):识别并定位图像中的多个对象(例如,YOLO、SSD、Faster R-CNN)。
  • 图像分割 (Image Segmentation):对图像中的每个像素进行分类(语义分割)或区分对象实例(实例分割)。
  • 人脸识别 (Facial Recognition):先进的人脸识别和验证系统。
  • 图像生成 (Image Generation):创建新图像(例如,生成对抗网络 - GANs)。

诸如 TensorFlow、Keras 和 PyTorch 等库用于构建和训练这些深度学习模型。OpenCV 还包含一个 DNN (Deep Neural Network) 模块,可以运行预训练的深度学习模型进行推理 (inference),让您无需从头开始训练即可利用强大的模型。

本介绍提供了一个起点。计算机视觉领域广阔且发展迅速,深度学习正在开辟新的前沿。