Skip to content

实现

Python 深度学习 - 实现示例:客户流失预测

Section titled “Python 深度学习 - 实现示例:客户流失预测”

在此实现示例中,我们将使用 Python 构建一个简单的深度神经网络 (DNN) 来预测银行客户的流失。我们希望根据客户过去的银行行为和特征来识别可能停止使用银行服务的客户。

我们将使用 TensorFlow 2.x 中的 Keras API,这是一个流行且用户友好的深度学习框架。数据集(Churn_Modelling.csv)相对较小(10,000 行,14 列),适合用于基础分类任务。

假设: 您已经设置了 Python 环境(例如 Anaconda),并安装了必要的库。如果尚未安装,通常使用 pip 安装 TensorFlow(其中包括 Keras):

# Ensure TensorFlow is installed (includes Keras)
# pip install tensorflow pandas numpy scikit-learn matplotlib

首先,我们导入必要的库,并使用 pandas 加载数据集。

# Importing the libraries
import numpy as np
import pandas as pd
import tensorflow as tf
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.compose import ColumnTransformer
from sklearn.metrics import confusion_matrix, accuracy_score
# Load the dataset
dataset = pd.read_csv('Churn_Modelling.csv')
# Display first few rows and info
print(dataset.head())
print(dataset.info())
# Define features (X) and target (y)
# We drop RowNumber, CustomerId, Surname as they are irrelevant identifiers
X = dataset.iloc[:, 3:13] # Columns from CreditScore to EstimatedSalary
y = dataset.iloc[:, 13] # 'Exited' column

神经网络需要数值输入。我们需要处理类别特征(‘Geography’、‘Gender’)并缩放数值特征。

# Identify categorical and numerical features
categorical_features = ['Geography', 'Gender']
numerical_features = ['CreditScore', 'Age', 'Tenure', 'Balance', 'NumOfProducts', 'HasCrCard', 'IsActiveMember', 'EstimatedSalary']
# Create preprocessing pipelines for numerical and categorical features
preprocessor = ColumnTransformer(
transformers=[
('num', StandardScaler(), numerical_features),
('cat', OneHotEncoder(handle_unknown='ignore'), categorical_features)
],
remainder='passthrough' # Keep other columns (if any), though we selected all relevant ones
)
# Apply preprocessing
X = preprocessor.fit_transform(X)
# Check the shape of processed features
print(f"Shape of processed X: {X.shape}") # Shape will change due to OneHotEncoding
# Display some processed data (optional)
# Note: Output is a NumPy array, column names are lost in transformation
print("\nSample processed features:")
print(X[:2])

解释:StandardScaler 将数值特征缩放到零均值和单位方差。OneHotEncoder 通过为每个类别创建二进制列,将类别特征转换为数值格式。ColumnTransformer 将这些转换应用于正确的列。

步骤 3:将数据分割为训练集和测试集

Section titled “步骤 3:将数据分割为训练集和测试集”

我们分割数据以便在其中一部分上训练模型,并在未见过的数据上评估其性能。

# Splitting the dataset into the Training set and Test set (80% train, 20% test)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.2, random_state = 42) # 用于结果重现性
print(f"\nTraining set shape: X={X_train.shape}, y={y_train.shape}")
print(f"Test set shape: X={X_test.shape}, y={y_test.shape}")

我们使用 TensorFlow/Keras 定义一个序贯 DNN。

# Initialize the ANN
model = tf.keras.models.Sequential()
# Get the number of input features after preprocessing
input_dim = X_train.shape[1]
# Add the input layer and the first hidden layer
# - units: 层中的神经元数量(超参数)
# - activation: 激活函数(ReLU 常用于隐藏层)
# - kernel_initializer: 初始化权重的方法('he_uniform' 常用于 ReLU)
# - input_shape: 第一层必需,用于指定输入维度
model.add(tf.keras.layers.Dense(units=6, activation='relu', kernel_initializer='he_uniform', input_shape=(input_dim,)))
# Add the second hidden layer
model.add(tf.keras.layers.Dense(units=6, activation='relu', kernel_initializer='he_uniform'))
# Add the output layer
# - units=1: 用于二分类
# - activation='sigmoid': 输出介于 0 和 1 之间的概率
# - kernel_initializer='glorot_uniform': 常适合用于 sigmoid/输出层
model.add(tf.keras.layers.Dense(units=1, activation='sigmoid', kernel_initializer='glorot_uniform'))
# 显示模型摘要
model.summary()

选择层数和每层的单元数通常是实验性的。我们从一个简单的 2 个隐藏层网络开始,每层有 6 个单元。

在训练之前,我们配置学习过程。

# 编译 ANN
# - optimizer: 用于更新权重的算法(Adam 是流行选择)
# - loss: 用于衡量误差的函数(二分类使用 'binary_crossentropy')
# - metrics: 在训练/测试期间评估的性能指标
model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])

我们将模型拟合到训练数据上。

# 在训练集上训练 ANN
# - batch_size: 每次梯度更新的样本数量
# - epochs: 迭代整个训练数据集的次数
history = model.fit(X_train, y_train,
batch_size = 32, # 常用批量大小
epochs = 50, # 训练周期数
validation_split=0.1, # 使用 10% 的训练数据作为训练期间的验证集
verbose=1) # 显示进度

batch_size 和 epochs 是超参数。32 是一个常用的批量大小,对于这个小型数据集,50-100 个 epoch 可能就足够了,但关键是监控验证集损失(可以添加 EarlyStopping 回调)。

我们使用训练好的模型在测试集上进行预测并评估其性能。

# 预测测试集结果
y_pred_proba = model.predict(X_test)
y_pred = (y_pred_proba > 0.5) # 将概率转换为二分类预测(0 或 1)
# 打印一些预测值与实际值对比
print("\nSample Predictions vs Actual:")
print(np.concatenate((y_pred[:10].reshape(-1,1), y_test[:10].values.reshape(-1,1)), axis=1))
# 创建混淆矩阵
cm = confusion_matrix(y_test, y_pred)
acc = accuracy_score(y_test, y_pred)
print("\nConfusion Matrix:")
print(cm)
print(f"\nAccuracy on Test Set: {acc:.4f}")

混淆矩阵显示了真阴性 (TN)、假阳性 (FP)、假阴性 (FN) 和真阳性 (TP)。准确率 (Accuracy) = (TN + TP) / 总数。

步骤 8:预测单个新观测值(示例)

Section titled “步骤 8:预测单个新观测值(示例)”

让我们预测一个假设的新客户是否会流失。

# 示例:预测单个客户
# 数据:地理位置:西班牙,信用分数:500,性别:女性,年龄:40,在职年限:3,余额:50000,
# 产品数量:2,有信用卡:是 (1),活跃会员:是 (1),预估薪资:40000
# 为单个观测值创建 DataFrame 以使用预处理器
new_customer_data = pd.DataFrame({
'CreditScore': [500],
'Geography': ['Spain'],
'Gender': ['Female'],
'Age': [40],
'Tenure': [3],
'Balance': [50000],
'NumOfProducts': [2],
'HasCrCard': [1],
'IsActiveMember': [1],
'EstimatedSalary': [40000]
})
# 重要:使用与训练数据拟合的相同预处理器
new_customer_processed = preprocessor.transform(new_customer_data)
# 进行预测
new_prediction_proba = model.predict(new_customer_processed)
new_prediction = (new_prediction_proba > 0.5)
print(f"\nPrediction for new customer (Probability): {new_prediction_proba[0][0]:.4f}")
print(f"Prediction for new customer (Churn?): {'Yes' if new_prediction[0][0] else 'No'}")

至此,一个基本的深度学习分类任务实现流程就完成了。进一步的改进可以包括超参数调优、添加正则化(如 Dropout)、使用回调(如 EarlyStopping)或探索更复杂的架构。

为了更好地理解网络在预测期间内部发生了什么(前向传播),让我们使用 NumPy 手动计算一个非常简单的网络的输出。这纯粹是为了说明;TensorFlow 等框架会自动高效地处理这些过程。

考虑一个具有 2 个输入神经元、1 个包含 2 个神经元(使用 ReLU 激活)的隐藏层和 1 个输出神经元(此处为简单起见不使用激活,或者使用线性激活)的网络。

输入 -> 隐藏层 (ReLU) -> 输出层 (线性)

import numpy as np
def relu(input_val):
# 修正线性单元激活函数
return np.maximum(0, input_val)
# 样本输入数据(例如,一个客户的 2 个特征)
input_data = np.array([2, 3])
# 样本权重(例如随机初始化或预定义)
weights = {
'node_0': np.array([1.0, -0.5]), # 第一个隐藏神经元的权重
'node_1': np.array([-1.5, 2.0]), # 第二个隐藏神经元的权重
'output': np.array([0.5, -1.0]) # 连接隐藏层到输出的权重
}
# 偏置(常使用,此处为简化省略或假设它们是权重的一部分)
# 计算隐藏层输入
node_0_input = (input_data * weights['node_0']).sum()
node_1_input = (input_data * weights['node_1']).sum()
# 计算隐藏层输出(应用 ReLU)
node_0_output = relu(node_0_input)
node_1_output = relu(node_1_input)
hidden_layer_outputs = np.array([node_0_output, node_1_output])
print(f"Hidden Layer Outputs: {hidden_layer_outputs}")
# 计算最终输出层输入
output_layer_input = (hidden_layer_outputs * weights['output']).sum()
# 最终输出(本示例假设为线性激活)
model_output = output_layer_input
print(f"Manual Forward Pass Output: {model_output}")

这个手动计算展示了数据流:加权求和后应用激活函数,逐层进行。真实的网络使用矩阵来处理批量数据和更复杂的操作,所有这些都由底层框架进行了优化。