MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications

About

Large Language Models (LLMs) have emerged as powerful tools for automating complex reasoning and decision-making tasks. In telecommunications, they hold the potential to transform network optimization, automate troubleshooting, enhance customer support, and ensure regulatory compliance. However, their deployment in telecom is hindered by domain-specific challenges that demand specialized adaptation. To overcome these challenges and to accelerate the adaptation of LLMs for telecom, we propose MM-Telco, a comprehensive suite of multimodal benchmarks and models tailored for the telecom domain. The benchmark introduces various tasks (both text based and image based) that address various practical real-life use cases such as network operations, network management, improving documentation quality, and retrieval of relevant text and images. Further, we perform baseline experiments with various LLMs and VLMs. The models fine-tuned on our dataset exhibit a significant boost in performance. Our experiments also help analyze the weak areas in the working of current state-of-art multimodal LLMs, thus guiding towards further development and research.

Anshul Kumar, Gagan Raj Gupta, Manish Rai, Apu Chakraborty, Ashutosh Modi, Abdelaali Chaoub, Soumajit Pramanik, Moyank Giri, Yashwanth Holla, Sunny Kumar, M. V. Kiran Sooraj• 2025

Related benchmarks

Task	Dataset	Result
Multiple-choice Question Answering	MM-Telco	CT WG1 Accuracy82.8	9
Information Retrieval	MM-Telco	--	7
Long Question Answering	MM-Telco Long QA	--	6
Long-form Question Answering	MM-Telco Telecom Blog	--	6
Filter Generation	Scenario-based Filter Generation Benchmark	--	4
Image Based MCQ	Image Based MCQ	--	3
Image Based QA	MM-Telco	--	3
Image caption generation	MM-Telco	--	3
Image Retrieval	MM-Telco	--	2

Showing 9 of 9 rows

Other info

Follow for update

@wizwand_team Discord