This paper proposes a video semantic feature extraction method from the perspective of time-space with conditional random domain model and ontology, which includes video low-level feature model parameter estimation (model parameter estimation,MPE) algorithm and high-level object semantic model update (model update,MU) algorithm. This method can realize the automatic extraction and annotation of video semantic concept ontology. It can provide support for semantic feature extraction and analysis. The experiment’s results show that the method improves the precision and recall rate of video semantic feature extraction.
结合国内外针对视频语义特征提取方法的各类研究[18,19,20],本文针对视频语义检索的实际需要与目标,基于条件随机场[21,22](conditional random field, CRF)方法及相关的概率统计学知识,探索使用机器自动标注视频语义特征概念本体,提出了一个融合本体与条件随机场的视频语义特征提取方法。
条件随机场是给定一组输入随机变量条件下另一组输出随机变量的条件概率分布模型,其特点是假设输出随机变量构成马尔科夫随机场[25,26,27](Markov random field, MRF)。条件随机场适用于时间序列随机变量的标注问题[28],而视频数据的语义提取正好满足这种数据特性和应用目标[29]。因而,本文的研究工作主要是基于条件随机场来构建视频语义特征提取的方法,得到视频语义信息的描述。
在对象识别的领域,级联分类器是常用的方法,例如基于梯度方向直方图(HOG)[31,32]特征的级联分类器。这类分类器的功能已由简单的人脸识别扩展到车牌、指纹识别和车辆、行人检测等方面。近几年,比较流行的对象识别算法是基于可变形部件模型(deformable part models,DPM)的Latent SVM [33,34,35]算法。
WEIW, YOUJ, LIUF Y, et al. A Survey on semantic-based video retrieval techniques[J]. Computer Science, 2006, 33(2):1-7.DOI: 10.3969/j.issn.1002-137X.2006.02.001(Ch).
KIMA, OH K, JUNGJ Y,et al. Imbalanced classification of manufacturing quality conditions using cost-sensitive decision tree ensembles[J]. International Journal of Computer Integrated Manufacturing, 2018, 31(8): 701-717. DOI: 10.1080/0951192X.2017.1407447 .
[7]
BISHOPC. Pattern Recognition and Machine Learning[M]. New York:Morgan Kaufmann, 1971.
[8]
MICHAILT. Bayesian network learning with the PC algorithm: An improved and correct variation[J]. Applied Artificial Intelligence, 2019, 33(2): 101-123. DOI: 10.1080/08839514.2018.1526760 .
[9]
AMJADM, AKBARK, ABDULH H M, et al. ANTSC: An intelligent naïve Bayesian probabilistic estimation practice for traffic flow to form stable clustering in vanet[J]. IEEE Access, 2018, 6: 4452-4461. DOI: 10.1109/ACCESS.2017.2732727 .
ZHIQ H, YUZ, LEIS, et al. A new bearing fault diagnosis method based on fine-to-coarse multiscale permutation entropy, Laplacian score and SVM [J]. IEEE Access, 2019,7: 17050-17066. DOI: 10.1109/ACCESS.2019.2893497 .
[12]
CHENY X, TRYPHONT G, ALLENT. Optimal transport for Gaussian mixture models[J]. IEEE Access, 2019, 7:6269-6278. DOI: 10.1109/ACCESS.2018.2889 838.
[13]
STAUFFERC, ERICW, GRIMSONW. Learning patterns of activity using real-time tracking [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2000, 22(8):747-757. DOI: 10.1109/34.868677 .
[14]
RABINERL, JUANGB H. An introduction to hidden Markov models [J]. IEEE ASSP Magazine, 1986, 3(1):4-16. DOI: 10.1109/MASSP.1986.1165342 .
[15]
KUZNETSOVM A, OSELEDETSI V. Tensor train spectral method for learning of hidden Markov models (HMM)[J]. Computational Methods in Applied Mathematics, 2019, 19(1): 93-99. DOI: 10.1515/cmam-2018 -0027.
[16]
METZEF, DINGD, YOUNESSIANE, et al. Beyond audio and video retrieval: Topic-oriented multimedia summarization [J]. International Journal of Multimedia Information Retrieval, 2013, 2(2):131-144. DOI: 10.1007/s13735-012-0028-y .
[17]
LIUX, SONGM L, ZHANGL M, et al. Joint shot boundary detection and key frame extraction[C]// Proceedings of the International Conference on Pattern Recognition. Piscataway: IEEE, 2012:2565-2568.
[18]
HO S Y, CHENH M, HO S J, et al. Design of accurate classifiers with a compact fuzzy-rule base using an evolutionary scatter partition of feature space[J]. IEEE Transactions on Systems Man & Cybernetics Part B, 2004, 34(2):1031-1044. DOI:10.1109/TSMCB.2003.819160 .
[19]
WORRINGM, SANTINIS, GUPTAA, et al. Content-based image retrieval at the end of the early years [J]. IEEE Transactions on Pattern Analysis & Machine Intelligence, 2000, 22(12):1349-1380. DOI:10.1109/34.895972 .
SHIY X, WENH, GONGP, et al. Projection of semantics and retrieval in natural scenery images based on fuzzy nerve network[J]. Computer Science, 2013, 40(12):122-126.DOI: 10.3969/j.issn.1002-137X.2013.12.026(Ch).
[22]
高文.网络视频服务的关键问题[J]. 中国计算机学会通讯, 2013, 9(1):25-29.
[23]
GAOW. The key issue in the network video service[J]. Communications of the CCF, 2013, 9(1):25-29 (Ch).
[24]
ZHONGX, LIL. Video semantic feature extraction model[C]//Computer Intelligent Computing and Education Technology. Boca Raton:CRC Press, 2014:137-141.
[25]
YANGZ G, YUH S, SUNW, et al. Locally shared features: An efficient alternative to conditional random field for semantic segmentation[J]. IEEE Access, 2019, 7: 2263-2272. DOI: 10.1109/ACCESS.2018.2886524 .
[26]
KOLLERD, FRIEDMANN. Probabilistic Graphical Models: Principles and Techniques [M]. Cambridge: MIT Press, 2019: 149-151.
[27]
LINC Y, TSENGB L . Segmentation, classification and watermarking for image/video semantic authentication[C]// 2002 IEEE Workshop on Multimedia Signal Processing. Piscataway:IEEE,2002:359-362. DOI: 10.1109/MMSP.2002.1203320 .
[28]
钟忺.视频图像语义分析及检索方法[M].北京:科学出版社,2017.
[29]
ZHONGX. Semantic Analysis and Search Method in Video Image [M].Beijing:Sciences Press,2017(Ch).
[30]
LIS Z. Markov Random Field Modeling in Computer Vision [M]. Berlin:Springer-Verlag, 1995.
[31]
WANGS F, SARATHK, SHIL,et al. Two-stage road terrain identification approach for land vehicles using feature-based and Markov random field algorithm[J]. IEEE Intelligent Systems, 2018, 33(1): 29-39. DOI:10.1109/MIS.2017.2581327 .
[32]
POPOOLAO, WANGK. Video-based abnormal human behavior recognition—A review[J].IEEE Transactions on Systems Man & Cybernetics Part C, 2012, 42(6):865-878. DOI:10.1109/TSMCC.2011.2178594 .
[33]
BERTINIM, BIMBOA, FERRACANIA, et al. Interactive multi-user video retrieval systems[J]. Multimedia Tools & Applications, 2013, 62(1):111-137. DOI: 10.1007/s11042-011-0888-9 .
[34]
柯佳. 基于语义的视频事件检测分析方法研究[D].镇江:江苏大学, 2013.
[35]
KEJ. Research on Detection and Analysis Method for Video Semantic Events [D]. Zhenjiang:Jiangsu University, 2013(Ch).
[36]
ZHONGX, LUY S, LIL. Video key frame extraction for semantic retrieval[C]// Proceedings of the International Conference on Information Computing and Applications(ICICA 2013). Berlin:Springer-Verlag, 2013:531-540.
[37]
SAMSULS, AZMINS S. Difference of gaussian oriented gradient histogram for face sketch to photo matching[J]. IEEE Access, 2018, 6: 39344-39352. DOI: 10.1109/ACCESS.2018.2855208 .
LIUZ, CHENK, ZHENGZ W. Object tracking algorithm based on HOG and multiple-instance online learning[J].Computer Engineering,2015,41(1):158-163.DOI: 10.3969/j.issn.1000-3428.2015.01.029(Ch).
[40]
FELZENSZWALBP. Object detection grammars[C]// 2011 IEEE International Conference on Computer Vision Workshops(ICCV Workshops). Piscataway: IEEE, 2011: 691~705.
[41]
钟忺.视频语义特征提取方法研究[D]. 武汉:华中科技大学,2013.
[42]
ZHONGX. Research on Video Semantic Feature Extraction Method [D]. Wuhan:Huazhong University of Science and Technology, 2013(Ch).
[43]
THIBAUTD, NICOLAST, MATTHIEUC. SyMIL: Minmax Latent SVM for weakly labeled data[J]. IEEE Transactions on Neural Networks and Learning Systems, 2018, 29(12): 6099-6112. DOI: 10.1109/TNNLS.2018.2820055 .