I am a CS PhD student at the School of Computing, National University of Singapore. I am supervised by Prof. Mong-Li Lee (Director of CTIC) and Prof. Wynne Hsu (Director of IDS) at Center for Trusted Internet and Community, and I also work with Dr. Hao Fei, Dr. Shengqiong Wu, Dr. Bobo Li, Dr. Hongzhan Lin and Dr. Tianjie Ju.
Prior to this, I received my master’s degree from NUS and my bachelor’s degree from Wuhan University, where I also completed a minor in Business Administration as part of the Ziqiang Entrepreneurship Program.
My research interest includes Bridging Physical and Mental Worlds toward Human-Like Intelligence through Multimodal (Video) Understanding, Reasoning, and Generation.
I am always exploring new collaboration opportunities. I am always happy to discuss potential collaborations — feel free to drop me an email at mluo@u.nus.edu.
News
- Accepted at NeurIPS 2026
From Finding to Linking: Benchmarking and Advancing Cross-Long-Video Reasoning for Multimodal LLMs
- Will Release a Survey on Video World Model
- Release a Survey on AI for Games
- Accepted at EMNLP (Findings) 2026
RIDGE: Region-Informed Derivative-Guided Evidence Selection for Long Video Understanding
- Accepted at EMNLP (Findings) 2026
OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models
- Accepted at ECCV 2026
From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models
- Accepted at ECCV 2026
No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs
- Accepted at ICML 2026 DL4C Workshop
Nexus: Execution-Grounded Multi-Agent Test Oracle Synthesi
- Accepted at IJCV
Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
- Preprint on ArXiv
Story2Proposal: A Scaffold for Structured Scientific Paper Writing
- Accepted at ICLR 2026
Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
2025 and earlier (17 updates)Hide earlier updates
- Grand Challenge Summary Paper at ACM MM 2025
The ACM Multimedia 2025 Grand Challenge of Multimodal Conversational Aspect-based Sentiment Analysis
- Workshop Summary Paper at ACM MM 2025
CogMAEC’25: The 1st Workshop on Cognition-oriented Multimodal Affective and Empathetic Computing
- Accepted at TIFS
Poisoning Attacks to Knowledge Distillation-based Federated Learning under Robust Aggregation Rules
- Accepted at ACL 2025 FEVER Workshop
EMULATE: A Multi-Agent Framework for Determining the Veracity of Atomic Claims by Emulating Human Actions
- Accepted at ACL 2025 (Oral)
Aristotle: Mastering Logical Reasoning with A Logic-Complete Decompose-Search-Resolve Framework
- Accepted at ICML 2025 (Oral, Spotlight)
On Path to Multimodal Generalist: Levels and Benchmarks
- Accepted at ICML 2025
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
- Accepted at ICML 2025
SWIFTCODE: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning
- Co-organizing a Grand Challenge at ACM MM 2025
Multimodal Conversational Aspect-based Sentiment Analysis (MCABSA 2025)
- Co-organizing a Workshop at ACM MM 2025
The 1st Cognition-oriented Multimodal Affective and Empathetic Computing (CogMAEC 2025) Workshop
- Accepted at ICLR 2025
PAD: Personalized Alignment at Decoding-Time
- Accepted at WWW 2025
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
- New Paper Published on arxiv
A Survey on Benchmarks of Multimodal Large Language Models
- Accepted at ACM MM Workshop (MIS24) (Best Paper Award)
Fine-grained Structural Hallucination Detection for Unified Visual Comprehension and Generation in Multimodal LLM
- Accepted at ACM MM 2024 (Oral)
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis
- 2nd Place at SemEval-2024
NUS-Emo at SemEval-2024 Task 3: Instruction-Tuning LLM for Multimodal Emotion-Cause Analysis in Conversations
- Accepted at TDSC
Towards Class-Balanced Privacy Preserving Heterogeneous Model Aggregation
Selected Works
Google ScholarAI for Games in the Foundation Model Era
Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models
No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs
On Path to Multimodal Generalist: General-Level and General-Bench
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis
Fine-grained Structural Hallucination Detection for Unified Visual Comprehension and Generation in Multimodal LLM
Professional Activity
Academic Service and Honors
- Reviewer for TPAMI, NeurIPS, ICLR, ICML, CVPR, ICCV, ACL, ACM MM, AAAI, WWW, ECCV, EMNLP, Neurocomputing, ACM TOMM, KBS, ACM TALLIP, and various workshops.
Honors and Awards
During my undergraduate studies
Academic Achievements
- Huawei Scholarship, Wuhan University (Top 5%)
- First-class Excellence Scholarship, Wuhan University (Ranked 2nd)
- Merit Student, Wuhan University (Top 10%)
- Outstanding Student, Wuhan University
Competitions and Recognitions
- Silver Award, Hubei Challenge Cup, Wuhan University
- Gold Award, Ziqiang Cup College, Wuhan University
- National First Prize, Citi Cup Financial Innovation Application Contest
- Bole Award, ByteTop Summit Project, ByteDance
- Top 10 Book Ambassador, Wuhan University Library
- The First Prize, HP Dream Factory Innovation Hackathon Wuhan Station, HP
Leadership and Social Activities
- Chairman, Wuhan University Campus Ambassador, ByteDance
- Excellent Campus Ambassador, WePie Team
- Online Course on Interdisciplinary Communication, University of Cambridge




















