

FOLLOWUS
School of Communication and Electronic Engineering, Jishou University, Jishou 416000, China
Tencent Inc., Shenzhen 518057, China
Ant Group Co., Ltd., Changsha 410000, China
School of Electronic and Information Engineering, Huaihua University, Huaihua 418000, China
✉Renmin ZHANG, rzhang1981@163.com
Received:26 May 2026,
Revised:2026-08-04,
Published:01 September 2026
Scan QR Code
Lili WU, Renmin ZHANG, Bin ZHANG, et al. Decoupled prompt-guided mixture-of-experts dynamic distillation for multimodal recommendation[J]. ENGINEERING Information Technology & Electronic Engineering, 2026, 27(9): 260160-15.
Lili WU, Renmin ZHANG, Bin ZHANG, et al. Decoupled prompt-guided mixture-of-experts dynamic distillation for multimodal recommendation[J]. ENGINEERING Information Technology & Electronic Engineering, 2026, 27(9): 260160-15. DOI: 10.1631/ENG.ITEE.2026.0160.
Multimodal recommendation aims to enrich preference modeling by leveraging visual and textual features. However
integrating high-dimensional pretrained features introduces substantial computational overhead. While knowledge distillation provides an effective compression strategy
existing frameworks face three intertwined challenges: rank bottlenecks caused by low-dimensional projections
cross-modal interference induced by shared fusion spaces
and optimization instability under static distillation temperatures. To address these issues
we propose ProMoE-DTS
a decoupled prompt-guided mixture-of-experts framework with dynamic temperature scheduling. Using an asymmetric teacher–student architecture
the teacher model leverages modality-aware soft prompts as semantic anchors to route heterogeneous features into parameter-disjoint expert networks
thereby alleviating cross-modal conflicts and resolving the rank bottlenecks. To ensure stable knowledge transfer
a feedback-driven dynamic temperature scheduler adaptively regulates the distillation intensity based on epoch-wise signals. This asymmetric design confines intensive multimodal operations to the offline teacher
leaving the online student model with a highly efficient
pure identifier-based structure. Extensive experiments on three benchmark datasets demonstrate that ProMoE-DTS improves Recall@20 by 2.24%–3.96% over state-of-the-art baselines
while requiring only 3.28%–3.55% of the teacher’s parameters.
Bao KQ , Zhang JZ , Zhang Y , et al. , 2023 . TALLRec: an effective and efficient tuning framework to align large language model with recommendation . Proc 17 th ACM Conf on Recommender Systems , p. 1007 - 1014 . https://doi.org/10.1145/3604915.3608857 https://doi.org/10.1145/3604915.3608857
Chen JY , Zhang HW , He XN , et al. , 2017 . Attentive collaborative filtering: multimedia recommendation with item- and component-level attention . Proc 40 th Int ACM SIGIR Conf on Research and Development in Information Retrieval , p. 335 - 344 . https://doi.org/10.1145/3077136.3080797 https://doi.org/10.1145/3077136.3080797
Dai BQ , Du ZC , Zhu JM , et al. , 2024 . UniEmbedding: learning universal multi-modal multi-domain item embeddings via user-view contrastive learning . Proc 33 rd ACM Int Conf on Information and Knowledge Management , p. 4446 - 4453 . https://doi.org/10.1145/3627673.3680098 https://doi.org/10.1145/3627673.3680098
Deng JX , Wang SY , Cai K , et al. , 2025 . OneRec: unifying retrieve and rank with generative recommender and iterative preference alignment . https://arxiv.org/abs/2502.18965 https://arxiv.org/abs/2502.18965
Dong XM , Huang O , Thulasiraman P , et al. , 2023 . Improved knowledge distillation via teacher assistants for sentiment analysis . Proc IEEE Symp Series on Computational Intelligence , p. 300 - 305 . https://doi.org/10.1109/SSCI52147.2023.10371965 https://doi.org/10.1109/SSCI52147.2023.10371965
Du YP , Sun Z , Wang ZY , et al. , 2025 . Active large language model-based knowledge distillation for session-based recommendation . Proc 39 th AAAI Conf on Artificial Intelligence , p. 11607 - 11615 . https://doi.org/10.1609/aaai.v39i11.33263 https://doi.org/10.1609/aaai.v39i11.33263
Geng X , Zhang HW , Bian JW , et al. , 2015 . Learning image and user features for recommendation in social networks . Proc IEEE Int Conf on Computer Vision , p. 4274 - 4282 . https://doi.org/10.1109/ICCV.2015.486 https://doi.org/10.1109/ICCV.2015.486
He RN , McAuley J , 2016 . VBPR: visual Bayesian personalized ranking from implicit feedback . Proc 30 th AAAI Conf on Artificial Intelligence , p. 144 - 150 . https://doi.org/10.1609/aaai.v30i1.9973 https://doi.org/10.1609/aaai.v30i1.9973
He XN , Deng K , Wang X , et al. , 2020 . LightGCN: simplifying and powering graph convolution network for recommendation . Proc 43 rd Int ACM SIGIR Conf on Research and Development in Information Retrieval , p. 639 - 648 . https://doi.org/10.1145/3397271.3401063 https://doi.org/10.1145/3397271.3401063
Hinton G , Vinyals O , Dean J , 2015 . Distilling the knowledge in a neural network . https://arxiv.org/abs/1503.02531 https://arxiv.org/abs/1503.02531
Hu XH , Zhang HT , 2025 . Invariant representation learning in multimedia recommendation with modality alignment and model fusion . Entropy , 27 ( 1 ): 56 . https://doi.org/10.3390/e27010056 https://doi.org/10.3390/e27010056
Huo FS , Xu WC , Guo JC , et al. , 2024 . C 2 KD: bridging the modality gap for cross-modal knowledge distillation . Proc IEEE/CVF Conf on Computer Vision and Pattern Recognition , p. 16006 - 16015 . https://doi.org/10.1109/CVPR52733.2024.01515 https://doi.org/10.1109/CVPR52733.2024.01515
Islam SU , Ahad JI , Rahman F , et al. , 2025 . Dynamic temperature scheduler for knowledge distillation . https://arxiv.org/abs/2511.13767 https://arxiv.org/abs/2511.13767
Jian M , Wang T , Yang MJ , et al. , 2026 . Hierarchy-aware multimodal distillation for recommendation . IEEE Trans Multim , 28 : 2279 - 2290 . https://doi.org/10.1109/TMM.2026.3651049 https://doi.org/10.1109/TMM.2026.3651049
Kang S , Kweon W , Lee D , et al. , 2023 . Distillation from heterogeneous models for top- K recommendation . Proc ACM Web Conf , p. 801 - 811 . https://doi.org/10.1145/3543507.3583209 https://doi.org/10.1145/3543507.3583209
Lan PX , Xu HY , Yang EN , et al. , 2025 . Efficient and effective prompt tuning via prompt decomposition and compressed outer product . Proc Conf of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies , p. 4406 - 4421 . https://doi.org/10.18653/v1/2025.naacl-long.225 https://doi.org/10.18653/v1/2025.naacl-long.225
Lee JW , Choi M , Lee J , et al. , 2019 . Collaborative distillation for top- N recommendation . Proc IEEE Int Conf on Data Mining , p. 369 - 378 . https://doi.org/10.1109/ICDM.2019.00047 https://doi.org/10.1109/ICDM.2019.00047
Li L , Zhang YF , Chen L , 2023 . Personalized prompt learning for explainable recommendation . ACM Trans Inform Syst , 41 ( 4 ): 103 . https://doi.org/10.1145/3580488 https://doi.org/10.1145/3580488
Liang JH , Zhao XY , Li MY , et al. , 2023 . MMMLP: multi-modal multilayer perceptron for sequential recommendations . Proc ACM Web Conf , p. 1109 - 1117 . https://doi.org/10.1145/3543507.3583378 https://doi.org/10.1145/3543507.3583378
Lin JH , Chen B , Wang HY , et al. , 2024 . ClickPrompt: CTR models are strong prompt generators for adapting language models to CTR prediction . Proc ACM Web Conf , p. 3319 - 3330 . https://doi.org/10.1145/3589334.3645396 https://doi.org/10.1145/3589334.3645396
Lin YF , He JY , 2026 . A reinforcement learning framework for multimodal recommendation explanation generation . Authorea . https://doi.org/10.22541/au.177368713.30806120/v1 https://doi.org/10.22541/au.177368713.30806120/v1
Liu F , Chen HL , Cheng ZY , et al. , 2023 . Disentangled multimodal representation learning for recommendation . IEEE Trans Multim , 25 : 7149 - 7159 . https://doi.org/10.1109/TMM.2022.3217449 https://doi.org/10.1109/TMM.2022.3217449
Liu MR , Zhang SX , Long C , 2025 . Facet-aware multi-head mixture-of-experts model for sequential recommendation . Proc 18 th ACM Int Conf on Web Search and Data Mining , p. 127 - 135 . https://doi.org/10.1145/3701551.3703552 https://doi.org/10.1145/3701551.3703552
Liu YF , Zhang KN , Ren XY , et al. , 2024 . AlignRec: aligning and training in multimodal recommendations . Proc 33 rd ACM Int Conf on Information and Knowledge Management , p. 1503 - 1512 . https://doi.org/10.1145/3627673.3679626 https://doi.org/10.1145/3627673.3679626
Loshchilov I , Hutter F , 2019 . Decoupled weight decay regularization . https://doi.org/10.48550/arXiv.1711.05101 https://doi.org/10.48550/arXiv.1711.05101
Ma WY , Xia HB , Liu Y , 2025 . DiffKD: collaborative graph diffusion with knowledge distillation for multimodal recommendation . J Intell Inform Syst , 63 ( 5 ): 1487 - 1510 . https://doi.org/10.1007/s10844-025-00946-4 https://doi.org/10.1007/s10844-025-00946-4
McAuley J , Targett C , Shi QF , et al. , 2015 . Image-based recommendations on styles and substitutes . Proc 38 th Int ACM SIGIR Conf on Research and Development in Information Retrieval , p. 43 - 52 . https://doi.org/10.1145/2766462.2767755 https://doi.org/10.1145/2766462.2767755
Mo F , Xiao L , Song QY , et al. , 2025 . FGCM: modality-behavior fusion model integrated with graph contrastive learning for multimodal recommendation . IEEE Multim , 32 ( 3 ): 27 - 37 . https://doi.org/10.1109/MMUL.2025.3542757 https://doi.org/10.1109/MMUL.2025.3542757
Nguyen NH , Nguyen TA , Nguyen T , et al. , 2024 . Towards efficient communication and secure federated recommendation system via low-rank training . Proc ACM Web Conf , p. 3940 - 3951 . https://doi.org/10.1145/3589334.3645702 https://doi.org/10.1145/3589334.3645702
Qiu RH , Wang S , Chen Z , et al. , 2021 . CausalRec: causal inference for visual debiasing in visually-aware recommendation . Proc 29 th ACM Int Conf on Multimedia , p. 3844 - 3852 . https://doi.org/10.1145/3474085.3475266 https://doi.org/10.1145/3474085.3475266
Reimers N , Gurevych I , 2019 . Sentence-BERT: sentence embeddings using Siamese BERT-networks . Proc Conf o n Empirical Methods in Natural Language Processing and the 9 th Int Joint Conf on Natural Language Processing , p. 3982 - 3992 . https://doi.org/10.18653/v1/D19-1410 https://doi.org/10.18653/v1/D19-1410
Rendle S , Freudenthaler C , Gantner Z , et al. , 2009 . BPR: Bayesian personalized ranking from implicit feedback . Proc 25 th Conf on Uncertainty in Artificial Intelligence , p. 452 - 461 .
Tang JX , Wang K , 2018 . Ranking distillation: learning compact ranking models with high performance for recommender system . Proc 24 th ACM SIGKDD Int Conf on Knowledge Discovery and Data Mining , p. 2289 - 2298 . https://doi.org/10.1145/3219819.3220021 https://doi.org/10.1145/3219819.3220021
Tian YJ , Zhang CX , Guo ZC , et al. , 2022 . NOSMOG: learning noise-robust and structure-aware MLPs on graphs . https://doi.org/10.48550/arXiv.2208.10010 https://doi.org/10.48550/arXiv.2208.10010
Wang J , Zhu L , Dai T , et al. , 2021 . Low-rank and sparse matrix factorization with prior relations for recommender systems . Appl Intell , 51 ( 6 ): 3435 - 3449 . https://doi.org/10.1007/s10489-020-02023-5 https://doi.org/10.1007/s10489-020-02023-5
Wang QY , Yin HZ , Wang H , et al. , 2019 . Enhancing collaborative filtering with generative augmentation . Proc 25 th ACM SIGKDD Int Conf on Knowledge Discovery and Data Mining , p. 548 - 556 . https://doi.org/10.1145/3292500.3330873 https://doi.org/10.1145/3292500.3330873
Wang X , He XN , Wang M , et al. , 2019 . Neural graph collaborative filtering . Proc 42 nd Int ACM SIGIR Conf on Research and Development in Information Retrieval , p. 165 - 174 . https://doi.org/10.1145/3331184.3331267 https://doi.org/10.1145/3331184.3331267
Wei RX , Lan JM , Li KK , et al. , 2025 . MFDB: multimodal feature fusion and dynamic behavior modeling for interactive recommendation systems . Knowl-Based Syst , 326 : 114047 . https://doi.org/10.1016/j.knosys.2025.114047 https://doi.org/10.1016/j.knosys.2025.114047
Wei W , Huang C , Xia LH , et al. , 2023 . Multi-modal self-supervised learning for recommendation . Proc ACM Web Conf , p. 790 - 800 . https://doi.org/10.1145/3543507.3583206 https://doi.org/10.1145/3543507.3583206
Wei W , Tang JB , Xia LH , et al. , 2024 . PromptMM: multi-modal knowledge distillation for recommendation with prompt-tuning . Proc ACM Web Conf , p. 3217 - 3228 . https://doi.org/10.1145/3589334.3645359 https://doi.org/10.1145/3589334.3645359
Wei YW , Wang X , Nie LQ , et al. , 2019 . MMGCN: multi-modal graph convolution network for personalized recommendation of micro-video . Proc 27 th ACM Int Conf on Multimedia , p. 1437 - 1445 . https://doi.org/10.1145/3343031.3351034 https://doi.org/10.1145/3343031.3351034
Wei YW , Wang X , Nie LQ , et al. , 2020 . Graph-refined convolutional network for multimedia recommendation with implicit feedback . Proc 28 th ACM Int Conf on Multimedia , p. 3541 - 3549 . https://doi.org/10.1145/3394171.3413556 https://doi.org/10.1145/3394171.3413556
Wei YW , Wang X , He XN , et al. , 2022 . Hierarchical user intent graph network for multimedia recommendation . IEEE Trans Multim , 24 : 2701 - 2712 . https://doi.org/10.1109/TMM.2021.3088307 https://doi.org/10.1109/TMM.2021.3088307
Xun JH , Zhang SY , Zhao Z , et al. , 2021 . Why do we click: visual impression-aware news recommendation . Proc 29 th ACM Int Conf on Multimedia , p. 3881 - 3890 . https://doi.org/10.1145/3474085.3475514 https://doi.org/10.1145/3474085.3475514
Yu HF , Qi ZY , Jang LK , et al. , 2024 . MMoE: enhancing multimodal models with mixtures of multimodal interaction experts . Proc Conf on Empirical Methods in Natural Language Processing , p. 10006 - 10030 . https://doi.org/10.18653/v1/2024.emnlp-main.558 https://doi.org/10.18653/v1/2024.emnlp-main.558
Zhang JH , Zhu YQ , Liu Q , et al. , 2021 . Mining latent structures for multimedia recommendation . Proc 29 th ACM Int Conf on Multimedia , p. 3872 - 3880 . https://doi.org/10.1145/3474085.3475259 https://doi.org/10.1145/3474085.3475259
Zhang ZJ , Liu SC , Yu JA , et al. , 2024 . M 3 oE: multi-domain multi-task mixture-of-experts recommendation framework . Proc 47 th Int ACM SIGIR Conf on Research and Development in Information Retrieval , p. 893 - 902 . https://doi.org/10.1145/3626772.3657686 https://doi.org/10.1145/3626772.3657686
Zhou X , Zhou HY , Liu Y , et al. , 2023 . Bootstrap latent representations for multi-modal recommendation . Proc ACM Web Conf , p. 845 - 854 . https://doi.org/10.1145/3543507.3583251 https://doi.org/10.1145/3543507.3583251
Zhu NJ , Ren YQ , Liu Y , et al. , 2025 . MMGCL: multi-scale and multi-channel graph contrastive learning for flight anomaly detection . Knowl-Based Syst , 329 : 114275 . https://doi.org/10.1016/j.knosys.2025.114275 https://doi.org/10.1016/j.knosys.2025.114275
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010802024621