Skip to content

MATLAB - 分类数组

MATLAB - 分类数组 (Categorical Arrays)

Section titled “MATLAB - 分类数组 (Categorical Arrays)”

一个分类数组(Categorical Array)是一种专门的 MATLAB 数据类型,旨在高效存储和管理来自有限的离散类别集合的数据。例如,“性别”(男、女、非二元性别)、“大小”(小、中、大)或调查回答(同意、中立、不同意)等数据。

虽然您可以使用 string 数组或字符元胞数组来实现这一点,但分类数组在内存使用、性能和分析能力方面具有显著优势。

  • 内存效率:分类数组不是将“United States of America”这样长而重复的字符串存储数千次,而是将每个唯一字符串存储一次,并使用轻量级的整数代码来表示数据。这大大减少了大型数据集的内存使用。
  • 性能提升:比较整数比比较字符串快得多。在分类数组上进行排序、分组和比较等操作要快得多。
  • 受控词汇表:分类数组强制执行一组固定的类别。这可以防止拼写错误(例如 ‘red’、‘Red’、‘redd’)被视为不同的类别,从而确保数据完整性。
  • 增强的分析和绘图功能:许多用于统计和绘图的 MATLAB 函数专门设计用于处理分类数据,自动创建分组汇总和带标签的图表。

您可以使用 categorical() 函数从现有数据创建分类数组,或者使用 discretize() 函数将连续数据分箱。

categorical() 函数将字符串数组、元胞数组或数值数组转换为分类数组。

%% 示例 1:从字符串数组创建
colorData = ["Red", "Blue", "Green", "Red", "Green"];
% 转换为分类数组
primaryColors = categorical(colorData);
disp(primaryColors);
% 查看唯一且已排序的类别
disp(categories(primaryColors));
% --- Output ---
% 1x5 categorical array
%
% Red Blue Green Red Green
%
% 3x1 cell array
% {'Blue' }
% {'Green'}
% {'Red' }
%% 示例 2:从带有指定标签的数值数据创建
% 想象一个调查,其中 1=低,2=中,3=高
responseData = [1 3 2; 2 1 3; 3 1 2];
valueSet = [1, 2, 3];
categoryNames = {"Low", "Medium", "High"};
% 映射基于顺序:valueSet(i) 映射到 categoryNames{i}
% 因此,1 -> "Low",2 -> "Medium",3 -> "High"
C = categorical(responseData, valueSet, categoryNames);
disp(C);
% --- Output ---
% 3x3 categorical array
%
% Low High Medium
% Medium Low High
% High Low Medium
%% 示例 3:创建有序分类数组
% 有序意味着类别具有自然顺序(例如,小 < 中 < 大)
sizeData = {"Medium"; "Small"; "Large"; "Medium"};
% 定义类别的正确顺序
sizeOrder = {"Small", "Medium", "Large"};
% 创建分类数组并指定它是有序的
B = categorical(sizeData, sizeOrder, 'Ordinal', true);
disp(B);
% 现在您可以根据顺序执行逻辑比较
% 查找所有大于“Medium”的尺寸
isLargerThanMedium = B > "Medium";
disp(isLargerThanMedium);
% --- Output ---
% 4x1 categorical array
%
% Medium
% Small
% Large
% Medium
%
% 4x1 logical array
%
% 0
% 0
% 1
% 0

discretize 函数非常适合将连续数值数据转换为离散的“分箱”(bins)。

%% 示例:将学生分数分箱到成绩类别
scores = [68, 75, 82, 90, 55, 78, 92, 60, 88, 72];
% 定义分箱的边界:(0, 60]、(60, 80]、(80, 100]
binEdges = [0, 60, 80, 100];
binNames = {"Fail", "Pass", "Merit"};
% 将分数离散化为命名类别
gradeCategories = discretize(scores, binEdges, 'Categorical', binNames);
disp(gradeCategories);
% --- Output ---
% 1x10 categorical array
%
% Pass Pass Merit Merit Fail Pass Merit Fail Merit Pass

分类数组最强大的用途之一是使用 groupcounts 和 groupsummary 等函数执行分组计算。

%% 示例:按产品类别分析销售数据
% 创建销售数据表
salesTbl = table(
categorical(["Laptop";"Monitor";"Laptop";"Keyboard";"Monitor";"Laptop"]),
[1200; 350; 1350; 75; 400; 1150],
'VariableNames', {'ProductType', 'SalePrice'}
);
disp('原始销售数据:');
disp(salesTbl);
% 使用 groupcounts 查找每种产品类型的销售数量
productCounts = groupcounts(salesTbl, 'ProductType');
disp('每种产品的销售数量:');
disp(productCounts);
% 使用 groupsummary 查找每种产品类型的平均销售价格
productMeanPrice = groupsummary(salesTbl, 'ProductType', 'mean', 'SalePrice');
disp('每种产品的平均价格:');
disp(productMeanPrice);
% --- Output ---
% 原始销售数据:
% ProductType SalePrice
% ___________ _________
%
% Laptop 1200
% Monitor 350
% Laptop 1350
% Keyboard 75
% Monitor 400
% Laptop 1150
%
% 每种产品的销售数量:
% ProductType GroupCount
% ___________ __________
%
% Keyboard 1
% Laptop 3
% Monitor 2
%
% 每种产品的平均价格:
% ProductType GroupCount mean_SalePrice
% ___________ __________ ______________
%
% Keyboard 1 75
% Laptop 3 1233.3
% Monitor 2 375}