

FOLLOWUS
School of Computer Science and Artificial Intelligence, Zhengzhou University, Zhengzhou 450001, China
International College of Zhengzhou University, Zhengzhou University, Zhengzhou 450001, China
Shandong Bosuan Zhixin Information Technology Co., Ltd., Jinan 250101, China
✉ Shizhe HU, ieshizhehu@zzu.edu.cn
✉ Mingliang XU, iexumingliang@zzu.edu.cn
Received:23 December 2025,
Revised:2026-06-10,
Online First:22 July 2026,
Published:01 August 2026
Scan QR Code
Yanzheng WANG, Yujun WANG, Fengshuo DAI, et al. Multi-stage contrastive multi-modal clustering with consistency retained[J]. ENGINEERING Information Technology & Electronic Engineering, 2026, 27(8): 1-13.
Yanzheng WANG, Yujun WANG, Fengshuo DAI, et al. Multi-stage contrastive multi-modal clustering with consistency retained[J]. ENGINEERING Information Technology & Electronic Engineering, 2026, 27(8): 1-13. DOI: 10.1631/ENG.ITEE.2025.0185.
Contrastive learning has received extensive attention because it can effectively strengthen inter-modal alignment and information complementarity in multi-modal clustering. However
there are two main problems with existing contrastive multi-modal clustering methods: (1) Current methods usually perform contrastive learning at only one or two stages and have imprecise selection strategies for negative sample pairs. (2) Most methods focus only on contrastive learning and ignore the retention of consistency information. To address the above challenges
we propose a framework of multi-stage contrastive multi-modal clustering with consistency retained. First
we design a systematic contrastive learning framework
which includes early–middle–late contrastive learning between single features
fused features
and clusters
respectively
which improves the semantic integrity and information fidelity of the fused representation. Second
we propose a new negative sample pair selection strategy to discriminately select the correct negative sample pair for each sample. Finally
we add a new loss term to late-stage contrastive learning
which considers the consistency of clustering results of the same sample under different modalities for the first time. Therefore
compared to the baseline model
this module has improved the accuracy by an average of 7.4 percentage points (PPs). In addition
we design a consistency retained module to ensure that the model can learn the information of each modality without being disturbed by high/low quality modalities
which improves the accuracy by an average of 2.2 PPs. Experiments conducted on multiple multi-modal datasets demonstrate that the accuracy of our method is 4.1 PPs higher on average compared to the second-best multi-modal clustering method
proving the excellent potential of our method.
Cai X , Nie F , Huang H , 2013 . Multi-view K -means clustering on big data . Proc 23 rd Int Joint Conf on Artificial Intelligence , p. 2598 - 2604 .
Chua TS , Tang J , Hong R , et al. , 2009 . NUS-WIDE: a real-world web image database from National University of Singapore . Proc ACM Int Conf on Image Video Retrieval , Article 48 . https://doi.org/10.1145/1646396.1646452 https://doi.org/10.1145/1646396.1646452
Cui J , Li Y , Huang H , et al. , 2024 . Dual contrast-driven deep multi-view clustering . IEEE Trans Image Process , 33 : 4753 - 4764 . https://doi.org/10.1109/TIP.2024.3444269 https://doi.org/10.1109/TIP.2024.3444269
Dai Z , Wang G , Yuan W , et al. , 2022 . Cluster contrast for unsupervised person re-identification . 16 th Asian Conf on Computer Vision , p. 319 - 337 . https://doi.org/doi/10.1007/978-3-031-26351-4-20 https://doi.org/doi/10.1007/978-3-031-26351-4-20
Friedman A , 1979 . Framing pictures: the role of knowledge in automatized encoding and memory for gist . J Exp Psychol-Gen , 108 ( 3 ): 316 - 355 . https://doi.org/10.1037/0096-3445.108.3.316 https://doi.org/10.1037/0096-3445.108.3.316
Grubinger M , Clough P , Müller H , et al. , 2006 . The IAPR TC-12 benchmark: a new evaluation resource for visual information systems . https://www-i6.informatik.rwth-aachen.de/publications/download/34/Grubinger-LREC-2006.pdf https://www-i6.informatik.rwth-aachen.de/publications/download/34/Grubinger-LREC-2006.pdf [Accessed on Dec. 22, 2025 ]
Gu WJ , Zhu CM , 2024 . Global and local combined contrastive learning for multi-view clustering . Multim Syst , 30 ( 5 ): 307 . https://doi.org/10.1007/s00530-024-01512-8 https://doi.org/10.1007/s00530-024-01512-8
Hartigan JA , Wong MA , 1979 . Algorithm AS 136: a K -means clustering algorithm . J R Stat Soc , 28 ( 1 ): 100 - 108 . https://doi.org/10.2307/2346830 https://doi.org/10.2307/2346830
Herath A , Kobti Z , 2024 . Subtype-MMCC: multimodal contrastive clustering approach for cancer subtype discovery with multi-omics data . Proc Comput Sci , 246 : 696 - 705 . https://doi.org/10.1016/j.procs.2024.09.488 https://doi.org/10.1016/j.procs.2024.09.488
Hu SZ , Zou GL , Zhang CY , et al. , 2023 . Joint contrastive triple-learning for deep multi-view clustering . Inform Process Manag , 60 ( 3 ): 103284 . https://doi.org/10.1016/j.ipm.2023.103284 https://doi.org/10.1016/j.ipm.2023.103284
Hu SZ , Zhang CK , Zou GL , et al. , 2025a . Deep multiview clustering by pseudo-label guided contrastive learning and dual correlation learning . IEEE Trans Neur Netw Learn Syst , 36 ( 2 ): 3646 - 3658 . https://doi.org/10.1109/TNNLS.2024.3354731 https://doi.org/10.1109/TNNLS.2024.3354731
Hu SZ , Fan JH , Zou GL , et al. , 2025b . Multi-aspect self-guided deep information bottleneck for multi-modal clustering . Proc 39 th AAAI Conf on Artificial Intelligence , p. 17314 - 17322 . https://doi.org/10.1609/aaai.v39i16.33903 https://doi.org/10.1609/aaai.v39i16.33903
Huiskes MJ , Lew MS , 2008 . The MIR Flickr retrieval evaluation . Proc 1 st ACM Int Conf on Multimedia Information Retrieval , p. 39 - 43 . https://doi.org/10.1145/1460096.1460104 https://doi.org/10.1145/1460096.1460104
Joyce JM , 2011 . Kullback–Leibler Divergence . Springer Berlin , Heidelberg , p. 720 - 722 . https://doi.org/10.1007/978-3-642-04898-2_327 https://doi.org/10.1007/978-3-642-04898-2_327
Kampffmeyer M , Løkse S , Bianchi FM , et al. , 2019 . Deep divergence-based approach to clustering . Neur Netw , 113 : 91 - 101 . https://doi.org/https://doi.org/10.1016/j.neunet.2019.01.015 https://doi.org/https://doi.org/10.1016/j.neunet.2019.01.015
Ke GZ , Hong ZY , Zeng ZQ , et al. , 2021 . CONAN: contrastive fusion networks for multi-view clustering . Proc IEEE Int Conf on Big Data , p. 653 - 660 . https://doi.org/10.1109/BigData52589.2021.9671851 https://doi.org/10.1109/BigData52589.2021.9671851
Kingma DP , Ba J , 2017 . Adam: a method for stochastic optimization . https://doi.org/10.48550/arXiv.1412.6980 https://doi.org/10.48550/arXiv.1412.6980
Kumar A , Rai P , III Daumé H , 2011 . Co-regularized multi-view spectral clustering . Proc 25 th Int Conf on Neural Information Processing Systems , p. 1413 - 1421 .
Lao JH , Huang D , Wang CD , et al. , 2024 . Towards scalable multi-view clustering via joint learning of many bipartite graphs . IEEE Trans Big Data , 10 ( 1 ): 77 - 91 . https://doi.org/10.1109/TBDATA.2023.3325045 https://doi.org/10.1109/TBDATA.2023.3325045
Li FF , Fergus R , Perona P , 2004 . Learning generative visual models from few training examples: an incremental Bayesian approach tested on 101 object categories . Proc IEEE Conf on Computer Vision and Pattern Recognition Workshop , p. 178 . https://doi.org/10.1109/CVPR.2004.383 https://doi.org/10.1109/CVPR.2004.383
Lin K , Wang D , Xia FZ , et al. , 2018 . Device clustering algorithm based on multimodal data correlation in cognitive Internet of Things . IEEE Int Things J , 5 ( 4 ): 2263 - 2271 . https://doi.org/10.1109/JIOT.2017.2728705 https://doi.org/10.1109/JIOT.2017.2728705
Lou ZZ , Xue H , Wang YZ , et al. , 2025 . Parameter-free deep multi-modal clustering with reliable contrastive learning . IEEE Trans Image Process , 34 : 2628 - 2640 . https://doi.org/10.1109/TIP.2025.3562083 https://doi.org/10.1109/TIP.2025.3562083
Lu YD , Lin YJ , Yang MX , et al. , 2024 . Decoupled contrastive multi-view clustering with high-order random walks . Proc 38 th AAAI Conf on Artificial Intelligence , p. 14193 - 14201 . https://doi.org/10.1609/aaai.v38i13.29330 https://doi.org/10.1609/aaai.v38i13.29330
Nie FP , Li J , Li XL , 2017 . Self-weighted multiview clustering with multiple graphs . Proc 26 th Int Joint Conf on Artificial Intelligence , p. 2564 - 2570 .
Pan L , Sun GD , Chang BF , et al. , 2023 . Visual interactive image clustering: a target-independent approach for configuration optimization in machine vision measurement . Front Inform Technol Electron Eng , 24 ( 3 ): 355 - 372 . https://doi.org/10.1631/FITEE.2200547 https://doi.org/10.1631/FITEE.2200547
Shen DG , Ip HHS , 1999 . Discriminative wavelet shape descriptors for recognition of 2-D patterns . Patt Recog , 32 ( 2 ): 151 - 165 . https://doi.org/10.1016/S0031-3203(98)00137-X https://doi.org/10.1016/S0031-3203(98)00137-X
Shi JB , Malik J , 2000 . Normalized cuts and image segmentation . IEEE Trans Patt Anal Mach Intell , 22 ( 8 ): 888 - 905 . https://doi.org/10.1109/34.868688 https://doi.org/10.1109/34.868688
Tian YL , Krishnan D , Isola P , 2020 . Contrastive multiview coding . Proc 16 th European Conf on Computer Vision , p. 776 - 794 . https://doi.org/10.1007/978-3-030-58621-8-45 https://doi.org/10.1007/978-3-030-58621-8-45
Tong Y , 2025 . Research on deep clustering and multimodal deep learning service recommendation system for large-scale user data . 3 rd Int Conf on Integrated Circuits and Communication Systems , p. 1 - 5 . https://doi.org/10.1109/ICICACS65178.2025.10967931 https://doi.org/10.1109/ICICACS65178.2025.10967931
Trosten DJ , Løkse S , Jenssen R , et al. , 2021 . Reconsidering representation alignment for multi-view clustering . Proc IEEE/CVF Conf on Computer Vision and Pattern , p. 1255 - 1265 . https://doi.org/10.1109/CVPR46437.2021.00131 https://doi.org/10.1109/CVPR46437.2021.00131
Wolf L , Hassner T , Taigman Y , 2008 . Descriptor based methods in the wild . Workshop on Faces in ‘Real-Life’ Images: Detection, Alignment, and Recognition .
Wu JX , Rehg JM , 2011 . CENTRIST: a visual descriptor for scene categorization . IEEE Trans Patt Anal Mach Intell , 33 ( 8 ): 1489 - 1501 . https://doi.org/10.1109/TPAMI.2010.224 https://doi.org/10.1109/TPAMI.2010.224
Wu S , Zheng Y , Ren YZ , et al. , 2024 . Self-weighted contrastive fusion for deep multi-view clustering . IEEE Trans Multim , 26 : 9150 - 9162 . https://doi.org/10.1109/TMM.2024.3387298 https://doi.org/10.1109/TMM.2024.3387298
Xu J , Ren YZ , Li GF , et al. , 2021 . Deep embedded multi-view clustering with collaborative training . Inform Sci , 573 : 279 - 290 . https://doi.org/10.1016/j.ins.2020.12.073 https://doi.org/10.1016/j.ins.2020.12.073
Xu J , Tang HY , Ren YZ , et al. , 2022 . Multi-level feature learning for contrastive multi-view clustering . Proc IEEE/CVF Conf on Computer Vision and Pattern Recognition , p. 16030 - 16039 . https://doi.org/10.1109/CVPR52688.2022.01558 https://doi.org/10.1109/CVPR52688.2022.01558
Xu J , Chen S , Ren YZ , et al. , 2023 . Self-weighted contrastive learning among multiple views for mitigating representation degeneration . Proc 37 th Int Conf on Neural Information Processing Systems , Article 54 .
Yan WQ , Yang TY , Tang C , 2025 . Self-supervised semantic soft label learning network for deep multi-view clustering . IEEE Trans Multim , 27 : 4971 - 4983 . https://doi.org/10.1109/TMM.2025.3543075 https://doi.org/10.1109/TMM.2025.3543075
Yang KL , Zhang TL , Alhuzali H , et al. , 2023 . Cluster-level contrastive learning for emotion recognition in conversations . IEEE Trans Affect Comput , 14 ( 4 ): 3269 - 3280 . https://doi.org/10.1109/TAFFC.2023.3243463 https://doi.org/10.1109/TAFFC.2023.3243463
Yin YJ , Ruan H , Chen Y , et al. , 2025 . Prototypical clustered federated learning for heart rate prediction . Front Inform Technol Electron Eng , 26 ( 10 ): 1896 - 1912 . https://doi.org/10.1631/FITEE.2500062 https://doi.org/10.1631/FITEE.2500062
Zhang FD , Kuang K , Chen L , et al. , 2023 . Federated unsupervised representation learning . Front Inform Technol Electron Eng , 24 ( 8 ): 1181 - 1193 . https://doi.org/10.1631/FITEE.2200268 https://doi.org/10.1631/FITEE.2200268
Zhou RW , Shen YD , 2020 . End-to-end adversarial-attention network for multi-modal clustering . IEEE/CVF Conf on Computer Vision and Pattern Recognition , p. 14607 - 14616 . https://doi.org/10.1109/CVPR42600.2020.01463 https://doi.org/10.1109/CVPR42600.2020.01463
Zhou SH , Liu XW , Liu JY , et al. , 2020 . Multi-view spectral clustering with optimal neighborhood Laplacian matrix . Proc 34 th AAAI Conf on Artificial Intelligence , p. 6965 - 6972 . https://doi.org/10.1609/aaai.v34i04.6180 https://doi.org/10.1609/aaai.v34i04.6180
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010802024621