污水处理厂作为水资源保护的重要设施,其核心任务之一是监控和预测再生水的排放指标。由于污水排放数据具有高复杂度和场景一致性差等特点,传统的时序预测方法效果有限。为此,提出一种新型工业时序预测模型SSP-GLM(Sewage-oriented Serial data Processing Generative Large Model),该模型通过将时间序列划分为子序列,利用深度卷积神经网络(CNN)提取局部特征,并结合生成式大模型(GLM)进行时序推理,提高了预测准确度。基于西安市两处污水处理厂的真实出水数据进行了实验,采用均方误差(MSE)和平均绝对误差(MAE)作为评估指标。结果表明,SSP-GLM在全量样本和小样本学习场景下均优于GRU、DLinear和Autoformer等基线方法,特别是在小样本条件下,其对复杂时序特征的捕捉能力显著增强。SSP-GLM在不同污水处理厂的数据上展现出良好的泛化能力,为工业污水处理的智能化管理提供了有力的技术支持。
Abstract
As a vital facility for water resource protection, monitoring and predicting the discharge indicators of reclaimed water is one of the core tasks of a wastewater treatment plant. Given the high complexity and poor scene consistency of wastewater discharge data, traditional time-series prediction methods are limited in their effectiveness. To address this, a novel industrial time-series prediction model, the Sewage-oriented Serial data Processing Generative Large Model (SSP-GLM), was proposed. This model segments time-series data into subsequences, extracts local features using deep convolutional neural networks (CNNs), and performs time-series reasoning with a generative large model (GLM), thereby enhancing the prediction accuracy. Experiments were conducted using real effluent data from two wastewater treatment plants in Xi’an, with the mean squared error (MSE) and mean absolute error (MAE) as evaluation metrics. The results show that SSP-GLM outperforms baseline methods, such as GRU, DLinear, and Autoformer, in both full-sample and few-shot learning scenarios. In particular, under few-shot conditions, SSP-GLM demonstrates significantly stronger capabilities in capturing complex temporal features. The model also exhibited good generalization across different wastewater treatment plants, providing robust technical support for the intelligent management of industrial wastewater treatment.
近年来,以时序挖掘为代表的深度学习技术已在污水处理指标预测方面展现出显著优势。其中,门控循环单元(Gated Recurrent Unit, GRU)等时序神经网络,凭借强大的信息捕获与存储能力,有效应对长时间序列数据,缓解了梯度消失/爆炸问题,成为该领域内广泛应用的模型[4-5]。例如,Ni等[6]构建了生物过程模型以预测N2O排放,成功处理了反应过程中的非线性复杂特征。Sun等[7]的研究则指出,时序数据中的噪声可能削弱模型预测精度,提出了一种数据驱动的降噪算法,提升了污水处理指标预测的准确性。
生成式人工智能已成为解决小样本学习任务的重要范式之一[13-14]。近年来,基于Transformer模型[15]的文本序列处理技术推动了大语言模型(Large Language Model,LLM)的兴起,以GPT-4[16]为代表的LLM已全面覆盖文本、图像等多模态信息的生成。学术界正积极挖掘LLM在时间序列分析领域的潜力[17],初步尝试将LLM在离散文本挖掘方面的优势应用于时间序列的连续数值处理任务中。
由此,本文采用子序列分块技术,提出了基于时序大语言模型的污水处理排放指标预测模型,称该模型为SSP-GLM(Sewage-oriented Serial data Processing Generative Large Model)。SSP-GLM首先将连续时间序列分割成多个子序列,利用深度卷积神经网络(Convolutional Neural Network,CNN)深入捕捉时序数据的局部特征与复杂动态语义;随后,针对LLM引擎,通过提示词工程将污水时序数据的上下文信息编码为模型可理解的形式,以激发其强大的时序推理能力;最后,利用深度神经网络融合上述特征,并依托LLM引擎生成预测输出。
本文使用了GPT2[25-27]作为SSP-GLM的大模型引擎底座,LLAMA[28]和BERT[29]作为替换的两个大模型底座,在4.3小节会详细探讨其预测性能差异。实验环境为CPU:13th Gen Intel(R) Core(TM) i9-13900K,GPU:GeForce RTX 4090(显存:24 GB),操作系统:Windows Server 2016。所有数据处理脚本均使用Python语言编写,基于PyTorch深度学习框架实现。
HAMOOD ALTOWAYTI WALI, SHAHIRS, OTHMANN, et al. The role of conventional methods and artificial intelligence in the wastewater treatment: A comprehensive review[J]. Processes, 2022, 10(9): 1832. DOI:10.3390/pr10091832 .
HUH X, SUNX Y. Application of artificial intelligence technology in simulation,prediction and optimization of wastewater treatment [J]. Environmental Pollution & Control, 2023, 45(11): 1587-1590. DOI:10.15985/j.cnki.1001-3865.2023.11.018(Ch ).
[4]
WANGY Q, WANGH C, SONGY P, et al. Machine learning framework for intelligent aeration control in wastewater treatment plants: Automatic feature engineering based on variation sliding layer[J]. Water Research, 2023, 246: 120676. DOI:10.1016/j.watres.2023.120676 .
[5]
WANK Y, DUB X, WANGJ H, et al. Deep learning-based intelligent management for sewage treatment plants[J]. Journal of Central South University, 2022, 29(5): 1537-1552. DOI:10.1007/s11771-022-5036-3 .
[6]
YUY, SIX S, HUC H, et al. A review of recurrent neural networks: LSTM cells and network architectures[J]. Neural Computation, 2019, 31(7): 1235-1270. DOI:10.1162/neco_a_01199 .
[7]
NIB J, YEL, LAWY, et al. Mathematical modeling of nitrous oxide (N2O) emissions from full-scale wastewater treatment plants[J]. Environmental Science & Technology, 2013, 47(14): 7795-7803. DOI:10.1021/es4005398 .
[8]
SUNS C, BAOZ Y, LIR Y, et al. Reduction and prediction of N2O emission from an Anoxic/Oxic wastewater treatment plant upon DO control and model simulation[J]. Bioresource Technology, 2017, 244: 800-809. DOI:10.1016/j.biortech.2017.08.054 .
[9]
THEODORISC V, XIAOL, CHOPRAA, et al. Transfer learning enables predictions in network biology[J]. Nature, 2023, 618(7965): 616-624. DOI:10.1038/s41586-023-06139-9 .
[10]
ARDALANZ, SUBBIANV. Transfer learning approaches for neuroimaging analysis: A scoping review[J]. Frontiers in Artificial Intelligence, 2022, 5: 780405. DOI:10.3389/frai.2022.780405 .
[11]
PISAI, MORELLA, VILANOVAR, et al. Transfer learning in wastewater treatment plant control design: From conventional to long short-term memory-based controllers[J]. Sensors, 2021, 21(18): 6315. DOI:10.3390/s21186315 .
[12]
CHENJ G, HANX B, SUNT, et al. Analysis and prediction of battery aging modes based on transfer learning[J]. Applied Energy, 2024, 356: 122330. DOI:10.1016/j.apenergy.2023.122330 .
[13]
ZHANGJ W, LIANGX Y, ZENGL Z, et al. Deep transfer learning for groundwater flow in heterogeneous aquifers using a simple analytical model[J]. Journal of Hydrology, 2023, 626: 130293. DOI:10.1016/j.jhydrol.2023.130293 .
[14]
MASUKOT, TOKUDAK, KOBAYASHIT, et al. Voice characteristics conversion for HMM-based speech synthesis system[DB/OL]. [2024-06-25]. DOI: 10.1109/icassp.1997.598807 .
[15]
SOSIAWANA Y, NOORAENIR, SARIL K. Implementation of using HMM-GA in time series data[J]. Procedia Computer Science, 2021, 179: 713-720. DOI:10.1016/j.procs.2021.01.060 .
[16]
VASWANIA, SHAZEERN, PARMARN, et al. Attention is all you need[DB/OL].[2023-07-18]. DOI: 10.1007/978-3-031-84300-6_13 .
WUJ, LIAOM C. Research on prediction of total nitrogen concentration in wastewater treatment plant effluent based on CA-GRU[J]. Process Automation Instrumentation, 2024, 45(4): 97-100. DOI:10.16086/j.cnki.issn1000-0380.2023040088(Ch ).
[25]
ZENGA L, CHENM X, ZHANGL, et al. Are transformers effective for time series forecasting?[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2023, 37(9): 11121-11128. DOI:10.1609/aaai.v37i9.26317 .
[26]
WUH, XUJ, WANGJ, et al. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting[DB/OL].[2024-07-05].
[27]
GUOY C, DINGG G, HANJ G, et al.Zero-Shot learning with transferred samples[DB/OL]. [2024-07-10]. DOI: 10.1109/tip.2017.2696747 .
[28]
GRAVEE, JOULINA, USUNIERN. Improving neural language models with a continuous cache[EB/OL]. 2016: 1612.04426.
[29]
BUDZIANOWSKIP, VULIĆI. Hello, it’s GPT-2 —How can I help you? Towards the use of pretrained language models for task-oriented dialogue systems[EB/OL]. 2019: 1907.05774. DOI: 10.18653/v1/d19-5602 .
[30]
TOUVRONH, LAVRILT, IZACARDG, et al. LLaMA: Open and efficient foundation language models[EB/OL]. 2023: 2302.13971.
[31]
DEVLINJ, CHANGM W, LEEK, et al. BERT: Pre-training of deep bidirectional transformers for language understanding[EB/OL]. 2018: 1810.04805. DOI: 10.18653/v1/n18-2 .
[32]
LIUX, MCDUFFD, KOVACSG, et al. Large language models are few-shot health learners[EB/OL]. 2023: 2305.15525.
[33]
ZHOUT, NIUP, WANGX, et al. One fits all: Power general time series analysis by pretrained LM[J]. Advances in Neural Information Processing Systems, 2023, 36: 43322-43355.