基于爬虫技术的电商网页关联数据自适应挖掘算法

王琪

吉林大学学报(信息科学版) ›› 2026, Vol. 44 ›› Issue (4) : 991 -997.

PDF (2384KB)
吉林大学学报(信息科学版) ›› 2026, Vol. 44 ›› Issue (4) : 991 -997.

基于爬虫技术的电商网页关联数据自适应挖掘算法

作者信息 +

Adaptive Mining Algorithm for Association Data of E-commerce Web Page Based on Crawler Technology

Author information +
文章历史 +
PDF (2440K)

摘要

为有效挖掘和利用网页数据, 以爬虫技术为支持, 提出一种关联数据自适应挖掘算法。利用爬虫技术爬取电商网页, 突破数据获取限制, 将所需数据以结构化格式存储至数据库, 为后续处理提供有序数据基础。基于存储在数据库中的稀疏数据, 依据最小支持度并运用 Apriori 算法, 并通过剪枝策略, 在减少数据处理量的同时, 能在稀疏数据中精准找出频繁项集, 成功解决传统方法因数据稀疏难以获取频繁项集的问题。以树形结构为框架, 将 Apriori 算法得到的频繁项集作为结点, 通过自适应结合旧结点与新结点构建挖掘树, 直至无结点可结合, 使其生成的挖掘树全面呈现数据关联关系, 进而得到充分且精准的电商网页关联数据挖掘结果。经在大型电商网站的网页上验证后表明, 所提算法能确保网页爬取的完整性, 准确获得以购买行为为目标的频繁项集, 精准挖掘电商网页中的关联数据, 在为电商平台运营提供有益参考的同时, 给予商家和消费者更好的决策支持。

Abstract

To effectively mine and utilize web data, a correlation data adaptive mining algorithm is proposed with the support of web crawling technology. Web crawling technology is used to crawl e-commerce web pages, breaking through data acquisition limitations, storing the required data in a structured format in a database, and providing an orderly data foundation for subsequent processing. Based on sparse data stored in the database, the Apriori algorithm is applied according to the minimum support degree. Through pruning strategy, this algorithm can accurately find frequent itemsets in sparse data while reducing data processing, successfully solving the problem of traditional methods being difficult to obtain frequent itemsets due to data sparsity. Using a tree structure as a framework, the frequent itemsets obtained by the Apriori algorithm are used as nodes. By adaptively combining old and new nodes, a mining tree is constructed until there are no nodes to combine. The generated mining tree comprehensively presents data association relationships, thereby obtaining sufficient and accurate e-commerce webpage association data mining results. After verification on the web pages of large e-commerce websites, it is found that the proposed algorithm can ensure the integrity of web crawling, accurately obtain frequent itemsets targeting purchasing behavior, and precisely mine associated data in e-commerce web pages. While providing useful references for e-commerce platform operation, it also provides better decision support for merchants and consumers.

关键词

电商网页数据 / 爬虫技术 / 网页数据爬取 / 频繁项集 / 树形结构 / 自适应挖掘

Key words

E-commerce webpage data / crawling technology / Web data crawling / frequent itemsets / tree structure / adaptive mining

引用本文

引用格式 ▾
王琪. 基于爬虫技术的电商网页关联数据自适应挖掘算法[J]. 吉林大学学报(信息科学版), 2026, 44(4): 991-997 DOI:

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

TANG W, LI G. Enhancing Competitiveness in Cross-Border E-Commerce through Knowledge-Based Consumer Perception Theory: An Exploration of Translation Ability[J]. Journal of the Knowledge Economy, 2023, 15(3): 14935-14968.

[2]

GUANGKUAN D, JIANYU Z, YING X, et al. Revisiting E-Commerce Platforms’ Strategies of Exercising Channel Power: A Contingency Perspective[J]. Journal of Business & Industrial Marketing, 2024, 39(10): 2239-2256.

[3]

UMAMAHESWARI N, KUMUTHA K, KALAIVANI S M, et al. Backpropagation Shuffled Leaping Neural Network Implementation in Classification of DDoS Packet Flow Traffic in Data Mining[J]. SN Computer Science, 2024, 5(8): 1131-1131.

[4]

韩永印, 王侠, 王志晓. 基于决策树的社交网络隐式用户行为数据挖掘方法[J]. 沈阳工业大学学报, 2024, 46(3): 312-317.

[5]

HAN Y Y, WANG X, WANG Z X. Data Mining Method Based on Decision Tree for Implicit User Behavior in Social Network[J]. Journal of Shenyang University of Technology, 2024, 46(3): 312-317.

[6]

陈万志, 赵帅, 方圆, . 改进 PrefixSpan 的行为轨迹数据挖掘算法[J]. 辽宁工程技术大学学报(自然科学版), 2023, 42(4): 506-512.

[7]

CHEN W Z, ZHAO S, FANG Y, et al. Behavioral Trajectory Data Mining Algorithm with Improved PrefixSpan[J]. Journal of Liaoning Technical University (Natural Science Edition), 2023, 42(4): 506-512.

[8]

夏小雅, 赵生宇, 韩凡宇, . 面向开源协作数字生态的信息服务与数据挖掘[J]. 计算机科学, 2024, 51(10): 187-195.

[9]

XIA X Y, ZHAO S Y, HAN F Y, et al. Data Mining and Information Service for Open Collaboration Digital Ecosystem[J]. Computer Science, 2024, 51(10): 187-195.

[10]

曹俊彬, 邵航, 姜坤, . 军用航空器事故关键质量特性的数据挖掘模型[J]. 计算机工程与设计, 2024, 45(2): 562-570.

[11]

CAO J B, SHAO H, JIANG K, et al. Data Mining Model for Critical-to-Quality Characteristics of Military Aviation Accidents[J]. Computer Engineering and Design, 2024, 45(2): 562-570.

[12]

刘多林, 吕苗. Scrapy 框架下分布式网络爬虫数据采集算法仿真[J]. 计算机仿真, 2023, 40(6): 504-508.

[13]

LIU D L, M. Simulation of Distributed Web Crawler Data Collection Algorithm under Scrapy Framework[J]. Computer Simulation, 2023, 40(6): 504-508.

[14]

GAVIN R, THORSTEN W, MARKUS S, et al. TomoTwin: Generalized 3D Localization of Macromolecules in Cryo-Electron Tomograms with Structural Data Mining[J]. Nature Methods, 2023, 20(6): 871-880.

[15]

李远航, 王劲林, 韩锐. 信息中心网络中一种基于内容热度的分区缓存替换方法[J]. 电子设计工程, 2023, 31(6): 133-138, 143.

[16]

LI Y H, WANG J L, HAN R. A Partition Cache Replacement Method Based on Content Popularity in Information Centric Networking[J]. Electronic Design Engineering, 2023, 31(6): 133-138, 143.

[17]

周艺腾, 唐鑫, 金路超. 基于自适应 MSB 可逆信息隐藏的图像云数据密文安全去重机制[J]. 计算机科学, 2024, 51(12): 352-360.

[18]

ZHOU Y T, TANG X, JIN L C. Adaptive MSB Reversible Data Hiding Based Security Deduplication for Encrypted Images in Cloud Storage[J]. Computer Science, 2024, 51(12): 352-360.

[19]

杨阳蕊, 朱亚萍, 刘雪梅, . 水利工程文本中抢险实体和关系的智能分析与提取[J]. 水利学报, 2023, 54(7): 818-828.

[20]

YANG Y R, ZHU Y P, LIU X M, et al. Intelligent Analysis and Joint Extraction of Rescue Entities and Relationships in Water Project Texts[J]. Journal of Hydraulic Engineering, 2023, 54(7): 818-828.

[21]

张磊, 焦晶, 李勃昕, . 融合机器学习和深度学习的大容量半结构化数据抽取算法[J]. 吉林大学学报(工学版), 2024, 54(9): 2631-2637.

[22]

ZHANG L, JIAO J, LI B X, et al. Large Capacity Semi Structured Data Extraction Algorithm Combining Machine Learning and Deep Learning[J]. Journal of Jilin University (Engineering and Technology Edition), 2024, 54(9): 2631-2637.

[23]

SOMU K, VELU C. A Novel Prediction of Sales and Purchase Forecasting for Festival Season of Hypermarkets with Customer Dataset Using Apriori Algorithm Instead of FP-Growth Algorithm to Improve the Accuracy[J]. Electrochemical Society Transactions, 2022, 107(1): 12647-12659.

[24]

韩虎, 孔博, 何勇禧, . 基于剪枝策略的知识增强方面级情感分析[J]. 华中科技大学学报(自然科学版), 2024, 52(11): 140-146.

[25]

HAN H, KONG B, HE Y X, et al. Aspect Based Sentiment Analysis Based on Knowledge Enhancement of Pruning Strategy[J]. Journal of Huazhong University of Science and Technology (Nature Science Edition), 2024, 52(11): 140-146.

[26]

张敏, 沈嘉裕. 突发公共卫生事件中政务短视频主题与用户行为的关联演化研究[J]. 情报杂志, 2023, 42(3): 181-189.

[27]

ZHANG M, SHEN J Y. Research on the Evolution of the Association between Governmental Short Video Themes and User Behavior in Public Health Emergency[J]. Journal of Intelligence, 2023, 42(3): 181-189.

[28]

朱宇婧, 陈芳, 王学昭. 基于商业管制清单-专利网络映射的关键核心技术识别研究 以工业软件为例[J]. 数据分析与知识发现, 2024, 8(10): 1-13.

[29]

ZHU Y J, CHEN F, WANG X Z. Identifying Core Technologies Based on the Commerce Control List-Patent Network Mapping: Case Study of Industrial Software[J]. Data Analysis and Knowledge Discovery, 2024, 8(10): 1-13.

[30]

贺帆, 刘漫丹, 钟超. 基于动态最小支持度的增量频繁序列挖掘[J]. 华东理工大学学报(自然科学版), 2024, 50(2): 257-263.

[31]

HE F, LIU M D, ZHONG C. Dynamic Minimum Support Incremental Frequent Trajectory Mining Algorithm[J]. Journal of East China University of Science and Technology (Natural Science Edition), 2024, 50(2): 257-263.

[32]

杨建平, 向月, 刘俊勇. 面向配电网投资决策的小样本关联规则自适应迁移学习方法[J]. 中国电机工程学报, 2022, 42(16): 5823-5834,6159.

[33]

YANG J P, XIANG Y, LIU J Y. Adaptive Transfer Learning of Small Sample Correlation Rules for Distribution Network Investment Decision[J]. Proceedings of the CSEE, 2022, 42(16): 5823-5834,6159.

[34]

刘昌春, 张凯, 包美凯, . 基于语境与文本结构融合的中文拼写纠错方法[J]. 南京大学学报(自然科学), 2024, 60(3): 451-463.

[35]

LIU C C, ZHANG K, BAO M K, et al. Research on Chinese Spelling Correction Based on the Integration of Context and Text Structure[J]. Journal of Nanjing University (Natural Sciences), 2024, 60(3): 451-463.

基金资助

黑龙江省教育科学规划重点课题基金资助项目(GJB1424318)

AI Summary AI Mindmap
PDF (2384KB)

2

访问

0

被引

详细

导航
相关文章

AI思维导图

/