Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation

About

Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external APIs, including adaptation to platform-specific data distributions, policy-specific objectives, and product-level safety constraints. As a result, platforms must undertake internal model development, naturally turning to shared public research for guidance. However, existing multimodal foundation-model studies primarily report architectures, training recipes, data scaling strategies, and benchmark results, but provide less systematic guidance on how failures should be localized and translated into targeted model-development interventions. Interventions are essential because deployment failures are rarely self-explanatory. Similar failures can originate from different causes. Without targeted interventions, improvement reduces to heuristic trial-and-error, where benchmark improvements are weakly attributable, and failures are difficult to trace to their underlying causes. To address this gap, we present a diagnostic methodology for industry-scale Audio-Visual-Language Models AVLM development. The methodology maps model failures into a taxonomy of observable failure signatures and links each class of failure to an intervention space. We instantiate this methodology across the development and alignment lifecycle of an AVLM foundation model for a large-scale video and live-streaming platform. The resulting system supports over 100 regions and is designed for noisy, ambiguous, and highly diverse content drawn from global platform traffic.

Shuchang Ye, Jinqiang Yu, Zhujun Xiao, Yajing Kong, Yist Y. Lin, Yang Ma, Jiaxi Liu, Xiaolei Xu, Zheng Yu• 2026

Related benchmarks

TaskDatasetResultRank
Video UnderstandingMVBench
Accuracy59
635
Visual Question AnsweringScienceQA
Accuracy85.03
525
Video UnderstandingVideo-MME without subtitles--
145
Multi-image ReasoningMuirBench
Accuracy46.88
112
Video UnderstandingLVBench
Overall Accuracy38.22
95
OCR & Document UnderstandingOCRBench
Score75.9
77
Audio UnderstandingMMAU
Accuracy68.3
57
Multi-Image Visual ReasoningBLINK
Accuracy46.93
51
General Visual Question AnsweringMMStar
Score57.1
50
General Visual Question AnsweringGQA
Accuracy58.4
35
Showing 10 of 27 rows

Other info

Follow for update