1.National Innovation Center for Digital Fishery,China Agricultural University,Beijing 100083,China
2.College of Information and Electrical Engineering/Key Laboratory of Smart Farming Technologies for Aquatic Animal and Livestock,Ministry of Agriculture and Rural Affairs,China Agricultural University,Beijing 100083,China
3.School of Information Science and Technology,Beijing Forestry University,Beijing 100083,China
To solve the problem of difficulty in fine classification of ornamental fish, this study conducted a PETFish-MViT model based on MobileViT. By introducing a mutual labeling enhancement module to optimize attention allocation, reconstruct the feature fusion structure, and adopt a label smoothing loss function, the model’s classification performance for ornamental fish is improved, alleviating the problem of uneven data distribution. The results showed that:1) The classification of the model on the self-built ornament fish data reached 91.4%, with an F1 value of 90.7%. Compared to the original model, the accuracy increased by 2.8%, and the F1-score increased by 2.8%. 2) On the public dataset WildFish test, the model achieves accuracy rate of 77.0%, indicating that the PETFish-MViT model exhibits certain generalization performance. In summary, this model can provide a classification method and tool for ornamental fish enthusiasts, playing a significant role in the ornamental fish industry.
对同时融合3种改进的完整模型进行分析可以发现,PETFish-MViT模型在PETFish数据集上准确率、精确率、召回率和 F1 分数分别较基线模型提升了1.5%、1.2%、1.7% 和1.4%,性能提升较为明显,主要得益于引入MTE模块扩大关键token对结果的影响,采用MFF模块强化局部信息融合,使用标签平滑损失函数有效缓解类别不平衡问题。这3部分改进互相配合,使模型在观赏鱼图像分类任务中表现出较为优异的性能。
2.3 在不同层中应用 MTE 的影响
为了探究在模型不同层中应用MTE模块的影响,在PETFish-MViT的3种不同规模模型上开展了对比试验(S:small,小型;XS:extra small,更小型;XXS:extra extra small,最小型)。如图2所示,本模型共包含5个层级,其中第3层、第4层和第5层包含MobileViT模块。分别在这3层中引入MTE模块,实验结果如表3所示。
梯度加权类激活映射(Gradient-weighted class activation mapping, Grad-CAM)是深度学习结果评估中一种有效的可视化技术,用于区分网络对不同区域的关注程度。本研究对PETFish-MViT中第4层自注意力权重分布图进行可视化,并与原始MobileViT模型直接对比,结果如图7所示。
PountneyS M.Survey indicates large proportion of fishkeeping hobbyists engaged in producing ornamental fish[J].Aquaculture Reports,2023,29:101503
[2]
Food and Agriculture Organization of the United Nations. Trade and Market News: 3rd International Ornamental Fish Trade and Technical Conference[EB/OL]. [--].
[3]
SalmanA, JalalA, ShafaitF, MianA, ShortisM, SeagerJ, HarveyE.Fish species classification in unconstrained underwater environments based on deep learning[J].Limnology and Oceanography:Methods,2016,14(9):570-585
[4]
YehC H, LinM H, ChangP C, KangL W.Enhanced visual attention-guided deep neural networks for image classification[J].IEEE Access,2020,8:163447-163457
[5]
XuX L, LiW S, DuanQ L. Transfer learning and SE-ResNet152 networks-based for small-scale unbalanced fish species identification [J].Computers and Electronics in Agriculture,2021,180:105878
[6]
HuJ, ShenL, SunG. Squeeze-and-excitation networks [C].In:2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Salt Lake City,UT,USA:IEEE,2018:7132-7141
[7]
HeK M, ZhangX Y, RenS Q, SunJ.Deep residual learning for image recognition[C].In:2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR).Las Vegas,NV,USA:IEEE,2016:770-778
[8]
LiL P, ShiF P, WangC X.Fish image recognition method based on multi-layer feature fusion convolutional network[J].Ecological Informatics,2022,72:101873
[9]
BhanumathiM, ArthiB. FishRNFuseNET:Development of heuristic-derived recurrent neural network with feature fusion strategy for fish species classification [J].Knowledge and Information Systems,2024,66:1997-2038
[10]
LiZ P, HuJ, WuK Y, MiaoJ W, ZhaoZ X, WuJ S.Local feature acquisition and global context understanding network for very high-resolution land cover classification[J].Scientific Reports,2024,14:12597
[11]
LiT P, ZhangZ Y, ZhuM D, CuiZ T, WeiD M.Combining transformer global and local feature extraction for object detection[J].Complex & Intelligent Systems,2024,10:4897-4920
[12]
QinR R, WangC Z, WuY M, DuH F, LvM Y. A U-shaped convolution-aided transformer with double attention for hyperspectral image classification [J].Remote Sensing,2024,16(2):288
[13]
DosovitskiyA, BeyerL, KolesnikovA, WeissenbornD, ZhaiX H, UnterthinerT, DehghaniM, MindererM, HeigoldG, GellyS, UszkoreitJ, HoulsbyN. An image is worth 16x16 words:Transformers for image recognition at scale[EB/OL].[--].
[14]
WangW, YangX, TangJ H.Vision transformer with hybrid shifted windows for gastrointestinal endoscopy image classification[J].IEEE Transactions on Circuits and Systems for Video Technology,2023,33(9):4452-4461
[15]
KhalilM, KhalilA, NgomA.A comprehensive study of vision transformers in image classification tasks[EB/OL].[--].
GongB, DaiK Y, ShaoJ, JingL, ChenY Y. Fish-TViT:A novel fish species classification method in multi water areas based on transfer learning and vision transformer [J].Heliyon,2023,9(6):e16761
[18]
SiG Z, XiaoY, WeiB, BullockL B, WangY Y, WangX D. Token-Selective Vision Transformer for fine-grained image recognition of marine organisms [J].Frontiers in Marine Science,2023,10:1174347
[19]
ManikandanD L, SanthanamS M.Parallel desires:Unifying local and semantic feature representations in marine species images for classification[J].Marine Geophysical Research,2024,45:16
[20]
HuangG, LiuZ, Van Der MaatenL, WeinbergerK Q.Densely connected convolutional networks[C].In:2(CVPR)017 IEEE Conference on Computer Vision and Pattern Recognition,Honolulu,HI,USA:IEEE,2017:2261-2269
DurdenJ M, HoskingB, BettB J, ClineD, RuhlH A. Automated classification of fauna in seabed photographs: The impact of training and validation dataset size, with considerations for the class imbalance [J].Progress in Oceanography,2021,196:102612
[29]
OpitzJ.A closer look at classification evaluation metrics and a critical reflection of common evaluation practice[C].In:Transactions of the Association for Computational Linguistics,Cambrideg:MAAssociatioon for Computational Linguistics,2024,12:820-836
[30]
Abou BakerN, HandmannU.One size does not fit all in evaluating model selection scores for image classification[J].Scientific Reports,2024,14:30239
[31]
KrizhevskyA, SutskeverI, HintonG E.ImageNet classification with deep convolutional neural networks[J].Communications of the ACM,2017,60(6):84-90
[32]
SandlerM, HowardA, ZhuM L, ZhmoginovA, ChenL C.MobileNetV2: Inverted residuals and linear bottlenecks[C].In:2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition,Salt Lake City,UT,USA:IEEE,2018:4510-4520
[33]
MaN N, ZhangX Y, ZhengH T, SunJ.ShuffleNet V2:Practical guidelines for efficient CNN architecture design[C].In:Computer Vision-ECCV 2018.Cham:Springer,2018:122-138
[34]
IandolaF N, HanS, MoskewiczM W, AshrafK, DallyW J, KeutzerK. SqueezeNet:AlexNet-Level Accuracy with 50x Fewer Parameters and <0.5MB Model Size [EB/OL].[--].
[35]
ZhangZ C, ChenZ D, WangY X, LuoX, XuX S.A vision transformer for fine-grained classification by reducing noise and enhancing discriminative information[J].Pattern Recognition,2024,145:109979
[36]
LiB C, WuY, LiuQ Y, ChenY, HeK, YuY. MCM-ViT:Mask-guided context-enhanced multi-scale transformer for fine-grained visual classification [J].Computers and Electrical Engineering,2025,122:109888
[37]
HuX B, ZhuS N, PengT L.Hierarchical attention vision transformer for fine-grained visual classification[J].Journal of Visual Communication and Image Representation,2023,91:103755
[38]
Shwartz-ZivR, GoldblumM, LiY L, BrussC B, WilsonA G.Simplifying neural network training under class imbalance[EB/OL].[--].