With the widespread application of artificial intelligence in fields such as security, industry, and agriculture,the demand for edge devices on vision reasoning tasks continues to grow.However, due to hardware constraints,deployment schemes of vision-language models designed for STM32 microcontrollers remain relatively scarce.To address this problem,this paper proposes an STM32-oriented vision-language model, MCUVLM-RWKV.The model integrates three core modules: a lightweight vision encoder, a lightweight vision feature mapper, and an RWKV decoder with a dual-mode operation mechanism, enabling image captioning tasks.Experimental results show that under the memory and storage limitations of STM32, MCUVLM-RWKV outperforms several mainstream models in evaluation metrics such as BLEU-4, ROUGE-L, and METEOR.Specifically, the ROUGE-L score reaches 55.7, which is significantly higher than that of other comparative models,indicating stronger modeling capability in long-sequence reasoning tasks.In addition, MCUVLM-RWKV demonstrates excellent performance in terms of parameter scale and inference memory consumption, further verifying its reasoning efficiency and deployment feasibility in MCU scenarios.
SHARSHARA, KHANL U, ULLAHW, et al.Vision-language models for edge networks: A comprehensive survey[J].IEEE Internet of Things Journal, 2025, 12(16): 32701-32724.
[2]
RUANH, LIUY, WANGB,et al.Learning semantic-aware representation in visual-language models for multi-label recognition with partial labels[J]. ACM Transactions on Multimedia Computing, Communications, and Applications,2025,21(3): 1-19.
[3]
RADFORDA, KIMJ W, HALLACYC,et al. Learning transferable visual models from natural language supervision[C]//Proceedings of the 38th International Conference on Machine Learning, 2021: 8748-8763.
[4]
WANGJ F, HUX W, ZHANGP C,et al.MiniVLM: A smaller and faster vision-language model[DB/OL]. (2021-08-09)[2025-09-23].
[5]
VASUP K A, POURANSARIH, FAGHRIF, et al. MobileCLIP: Fast image-text models through multi-modal reinforced training[C]//IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR), 2024: 15963-15974.
[6]
CHUX X, QIAOL M, LINX Y, et al. MobileVLM: A fast,strong and open vision language assistant for mobile devices[DB/OL]. (2023-12-28)[2025-09-23].
[7]
ARIFK H I, YOONJ, NIKOLOPOULOSD S,et al. HiRED: Attention-guided token dropping for efficient inference of high-resolution vision-language models[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2025, 39(2): 1773-1781.
[8]
FARINAM, MANCINIM, CUNEGATTIE, et al. MULTIFLOW: Shifting towards task-agnostic vision-language pruning[C]//IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR), 2024: 16185-16195.
[9]
RANGM, BIZ, LIUC,et al.Eve: Efficient multimodal vision language models with elastic visual experts[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2025, 39(7): 6694-6702.
[10]
LIJ, LID, SAVARESES, et al. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models[C]//Proceedings of the 40th International Conference on Machine Learning, 2023: 19730-19742.
[11]
YAOY, YUT Y, ZHANGA, et al.Efficient GPT-4V level multimodal large language model for deployment on edge devices[J].Nature Communications, 2025, 16: 5509.
[12]
PENGB, ALCAIDEE, ANTHONYQ, et al.RWKV: Reinventing RNNs for the transformer era[C]//Findings of the Association for Computational Linguistics: EMNLP 2023. Singapore: Association for Computational Linguistics, 2023: 14048-14077.
[13]
HOUH W, ZENGP G, MAF, et al. VisualRWKV: Exploring recurrent neural networks for visual language models[C]//Proceedings of the 31st International Conference on Computational Linguistics, 2025: 10423-10434.
[14]
PHANA H, SOBOLEVK, SOZYKINK, et al. Stable low-rank tensor decomposition for compression of convolutional neural network[C]//Computer Vision-ECCV 2020. Cham: Springer, 2020: 522-539.
[15]
CAOY Z, CAOW W, WANGZ Y, et al.A light-weight rectangular decomposition large kernel convolution network for deformable medical image registration[J].Biomedical Signal Processing and Control, 2024, 95: 106476.
[16]
SILVAJ D, MAGALHÃESJ, TUIAD, et al. Multilingual vision-language pre-training for the remote sensing domain[C]//Proceedings of the 32nd ACM International Conference on Advances in Geographic Information Systems, 2024: 220-232.
[17]
LIJ, LID, XIONGC, et al.BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation[C]//International Conference on Machine Learning, 2022: 12888-12900.
[18]
LIUH T, LIC Y, WUQ Y, et al.Visual instruction tuning[DB/OL].(2023-12-11)[2025-09-23].
SUNGonglingyun, ZHANGJingyu, LIANJunbo, et al.Research on identification of succulent based on lightweight convolutional neural network[J].Chinese Journal of Sensors and Actuators, 2023, 36(12): 1916-1927.(in Chinese)
[21]
WUH X, XUJ H, WANGJ M, et al.Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting[C]//35th Conference on Neural Information Processing Systems, 2021: 22419-22430.
[22]
DAVIDR, DUKEJ, JAINA, et al.TensorFlow lite micro: Embedded machine learning for TinyML systems[C]//Proceedings of Machine Learning and Systems, 2021: 800-811.
[23]
LUOG, CHENGL, JINGC, et al.A thorough review of models, evaluation metrics, and datasets on image captioning[J].IET Image Processing, 2022, 16(2): 311-332.
[24]
LINJ, CHENW M, LINY J, et al.MCUNet: Tiny deep learning on IoT devices[C]//34th International Conference on Neural Information Processing Systems, 2020: 11711-11722.