基于深度学习的可食用野菜种类识别

吴玉强 ,  孙荀 ,  季呈明 ,  胡乃娟

中国瓜菜 ›› 2024, Vol. 37 ›› Issue (11) : 57 -66.

PDF (2593KB)
中国瓜菜 ›› 2024, Vol. 37 ›› Issue (11) : 57 -66. DOI: 10.16861/j.cnki.zggc.2024.0325
试验研究

基于深度学习的可食用野菜种类识别

作者信息 +

Identification of edible wild vegetable species based on deep learning

Author information +
文章历史 +
PDF (2654K)

摘要

可食用野菜兼具营养价值和药用价值,然而传统采摘可食用野菜的分辨主要依赖人为主观经验,效率低且错误风险高,因此对可食用野菜快速准确的识别对实现野菜产业开发和保障食用安全具有重要意义。以南京地区“七头一脑”共8种可食用野菜为研究对象,构建了8种野菜的2400张图像数据集,采用3种具有代表性的卷积神经网络(convolutional neural network, CNN)模型(AlexNet、VGG16和ResNet50)和3种视觉自注意力(vision transformer, ViT)模型(ViT、CaiT和DeiT)共6种不同的深度学习模型进行训练和验证,并通过梯度加权类激活映射(gradient-weighted class activation mapping, Grad-CAM)来分析深度学习模型的决策机制。结果表明,ResNet50在验证集上的准确率达到94.68%,精确率、召回值和F1分数分别为97.66%、97.74%和97.70%,在6个模型中表现最佳。随后,在最优模型ResNet50基础上添加卷积模块的注意力机制(convolutional block attention module, CBAM)和坐标注意力机制(coordinate attention, CA)模块进行模型优化,结果显示,CBAM-ResNet50准确率达到了97.67%,CA-ResNet50准确率达到了98.34%,分别提高了2.99个百分点和3.66个百分点。以上研究结果证实了CNN模型在数据集上能取得比ViT更好的结果,利用深度学习识别可食用野菜种类是可行的,且添加注意力模块能够实现更高的识别准确率。

Abstract

Edible wild vegetables possess both nutritional and medicinal values. However, the traditional identification of wild edible vegetables mainly relies on subjective human experience, which is inefficient and carries a high risk of error. Therefore, rapid and accurate identification of edible wild vegetables is of great significance for the development of the wild vegetable industry and the assurance of food safety. Eight types of edible wild vegetables known as the "Seven Heads and One Brain" in the Nanjing region were selected as the research subjects and a database of 2400 images were constructed. Training and validation were conducted using 6 different deep learning models, including 3 representative convolutional neural network (CNN) models (AlexNet, VGG16 and ResNet50) and 3 vision transformers (ViT) models (ViT, CaiT and DeiT). Furthermore, the decision-making mechanisms of the deep learning models were analyzed using Gradient-Weighted Class Activation Mapping. The results showed that ResNet50 achieved an accuracy rate of 94.68% on the validation set, with precision, recall value, and F1-score of 97.66%, 97.74%, and 97.70%, respectively, and performed the best among the 6 models. Subsequently, the attention mechanism modules, convolutional block attention module and coordinate attention module were added to the optimal ResNet50 model for further optimization. The results showed that the accuracy of CBAM-ResNet50 and CA-ResNet50 models achieved 97.67% and 98.34%, respectively, representing enhancements of 2.99 and 3.66 percent point. The above research results confirmed that the CNN model can achieve better results than ViT on the dataset in this paper. It is feasible to use deep learning to identify edible wild vegetable species, and adding attention modules can lead to higher recognition accuracy.

关键词

可食用野菜 / 种类识别 / 卷积神经网络 / 视觉自注意力 / 注意力机制模块

Key words

Edible wild vegetables / Species identification / Convolutional neural networks / Vision transformer / Attention mechanism modules

引用本文

引用格式 ▾
吴玉强,孙荀,季呈明,胡乃娟. 基于深度学习的可食用野菜种类识别[J]. 中国瓜菜, 2024, 37(11): 57-66 DOI:10.16861/j.cnki.zggc.2024.0325

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

卢超. 长沙地区野菜资源开发利用研究[D]. 长沙: 湖南农业大学, 2017.

[2]

查金平. 利用野菜资源开展校本研究提升学生核心素养[J]. 科技风, 2019(8): 37.

[3]

何进, 刘琳, 朱姝, 等. 贵州省2016-2021年有毒植物及其毒素中毒暴发事件监测情况分析[J]. 现代预防医学, 2022, 49(21): 4009-4013.

[4]

刘文斌, 庹先国, 张贵宇, 等. 基于卷积神经网络的白酒上甑探汽方法[J]. 食品研究与开发, 2024, 45(5): 139-144.

[5]

WANG M S, MA H B, WANG Y L, et al. Design of smart home system speech emotion recognition model based on ensemble deep learning and feature fusion[J]. Applied Acoustics, 2024, 218: 109886.

[6]

XU G, YUE Q R, LIU X G. Real-time multi-object detection model for cracks and deformations based on deep learning[J]. Advanced Engineering Informatics, 2024, 61: 102578.

[7]

LECUN Y, BOTTOU L, BENGIO Y, et al. Gradient-based learning applied to document recognition[J]. Proceedings of the Ieee, 1998, 86(11): 2278-2324.

[8]

DOSOVITSKIY A, BEYER L, KOLESNIKOV A, et al. An image is worth 16×16 Words:Transformers for image recognition at scale,May 04, 2021[C].Vienna: International Computer on Learning, 2021.

[9]

KRIZHEVSKY A, SUTSKEVER I, HINTON G E. ImageNet classification with deep convolutional neural networks[J]. Communications of the Acm, 2017, 60(6): 84-90.

[10]

林伟, 仲伟波, 袁毓, 等. 基于改进AlexNet与CUDA的大豆快速三分类方法[J]. 计算机与数字工程, 2023, 51(12): 2997-3003.

[11]

王圆, 祝俊辉, 周贤勇, 等. 基于改进ResNet模型的番茄叶片病虫害识别[J]. 激光杂志, 2024, 45(5): 209-214.

[12]

ZHOU B, YU X, LIU J, AN D, et al. Effective vision transformer training:A data-centric perspective[J]. Computer Vision and Pattern Recognition, 2022, 2209: 15006.

[13]

王杨, 李迎春, 许佳炜, 等. 基于改进Vision Transformer网络的农作物病害识别方法[J]. 小型微型计算机系统, 2024, 45(4): 887-893.

[14]

CASTELLANO G, MARINIS P D, VESSIO G. Weed mapping in multispectral drone imagery using lightweight vision transformers[J]. Neurocomputing, 2023, 562: 126914.

[15]

SELVARAJU R R, COGSWELL M, DAS A, et al. Grad-CAM:Visual explanations from deep networks via gradient-based localization[C]. Ieee International Conference on Computer Vision (ICCV), 2017: 618-626.

[16]

SIMONYAN K, ZISSERMAN A. Very deep convolutional networks for large-scale image recognition[J]. Computer Science, 2014.

[17]

HE K M, ZHANG X Y, REN S Q, et al. Deep residual learning for image recognition[C]. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016: 770-778.

[18]

VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[J]. Computation and Language, 2017, 30: 6000-6010.

[19]

TOUVRON H, CORD M, SABLAYROLLES A, et al. Going deeper with image transformers[C]. IEEE/CVF International Conference on Computer Vision (ICCV), 2021: 32-42.

[20]

LIU Y, ZHANG Y, WANG Y, et al. A survey of visual transformers[J]. IEEE Transactions on Neural Networks and Learning Systems, 2024, 35: 7478-7498.

[21]

TOUVRON H, CORD M, DOUZE M, et al. Training data-efficient image transformers & distillation through attention[C]. International Conference on Machine Learning, 2021, 139: 7358-7367.

[22]

赵婷婷, 高欢, 常玉广, 等. 基于知识蒸馏与目标区域选取的细粒度图像分类方法[J]. 计算机应用研究, 2023, 40(9): 2863-2868.

[23]

曹明亮, 尹蜜, 王庆彬, 等. 基于深度学习算法联合Grad-CAM的宫腔镜子宫内膜病变诊断模型研究[J]. 实用妇产科杂志, 2024, 40(5): 409-413.

[24]

谢瑞麟, 崔展齐, 陈翔, 等. IATG:基于解释分析的自动驾驶软件测试方法[J]. 软件学报, 2024, 35(6): 2753-2774.

[25]

WOO S H, PARK J, LEE J Y, et al. CBAM:Convolutional block attention module[J]. Computer Vision, 2018, 11211: 3-19.

[26]

HOU Q B, ZHOU D Q, FENG J S, et al. Coordinate attention for efficient mobile network design[C]. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021: 13708-13717.

[27]

PRAKASH J A, ASSWIN C R, KUMAR K S D. A, et al. Transfer learning approach for pediatric pneumonia diagnosis using channel attention deep CNN architectures[J]. Engineering Applications of Artificial Intelligence, 2023, 123: 106416.

[28]

XIONG B P, CHEN W S, NIU Y X, et al. A Global and Local Feature fused CNN architecture for the sEMG-based hand gesture recognition[J]. Computers in Biology and Medicine, 2023, 166: 107497.

[29]

ZHOU D, KANG B, JIN X, et al. DeepViT:Towards deeper vision transformer[J]. Computer Vision and Pattern Recognition, 2021.

[30]

LI X C, LI X H, ZHANG M Q, et al. SugarcaneGAN:A novel dataset generating approach for sugarcane leaf diseases based on lightweight hybrid CNN-Transformer network[J]. Computers and Electronics in Agriculture, 2024, 219: 108762.

[31]

LI X P, XIANG Y Y, LI S Q. Combining convolutional and vision transformer structures for sheep face recognition[J]. Computers and Electronics in Agriculture, 2023, 205: 107651.

[32]

LI Y X, HUANG Y W, HE N J, et al. Improving vision transformer for medical image classification via token-wise perturbation[J]. Journal of Visual Communication and Image Representation, 2023, 98: 104022.

[33]

LIN T Y, DOLLAR P, GIRSHICK R, et al. Feature pyramid networks for object detection[C]. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, 106: 936-944.

[34]

KIM W, JUNG W S, CHOI H K. Lightweight driver monitoring system based on multi-task mobilenets[J]. Sensors, 2019, 19(14): 3200.

基金资助

江苏省重点研发计划项目(BE2019762)

中央高校基本科研业务费专项资金项目(LGZD202408)

国家自然科学基金(32201923)

“十四五”江苏省重点学科“公安技术”(苏教研函﹝2022﹞2号)

AI Summary AI Mindmap
PDF (2593KB)

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/

〈 〉