
Han Fang (方瀚)
I am currently a Research Scientist at the Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd., where I focus on developing generative AI and advancing its applications in downstream domains. As a core contributor, I led the research and development of the 'Xinghe' Vision Model series, the text-to-image model TeleImage-1.0, and the multimodal large model TeleMM-2.0.
My research aims to bridge the gap between vision and language by enabling models to perform vision-centric reasoning, with a focus on long video understanding, free-form multimodal grounding, and vision-language pretraining.
I am open to academic collaborations and welcome motivated research interns. Please feel free to reach out at fanghan1996@outlook.com.
News
- 2026.07Two papers are accepted by ACM MM 2026.
- 2026.01TeleMM-2.0-Thinking ranked 2nd on the OpenCompass Multi-modal Academic Leaderboard.
- 2025.11One paper is accepted by AAAI 2026.
- 2025.10Ranked 3rd in the ICCV 2025 Multimodal Large Language Model Visual Reasoning Localization Challenge MARS2: VG-RS.
- 2025.04One paper is accepted by TAFFC 2025.
- 2024.12One paper is accepted by AAAI 2025.
- 2024.07One paper is accepted by ACM MM 2024.
- 2024.03Two papers are accepted by ICME 2024 (oral presentation).
- 2023.11Served as a key Contributor to the release of TeleImage 1.0, part of the 'Xingchen' Model Family, at Tianyi Expo 2023.
- 2023.07Two papers are accepted by ACM MM 2023.
- 2023.04Served as a key Contributor to the release of the 'Xinghe' Vision Model 2.0 at the China Telecom Cloud Ecosystem Conference 2023.
- 2022.12One paper is accepted by TMM 2022.
- 2022.07Joined China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd. as a research scientist, focusing on text-to-image models, video-language models and multimodal large language models.
- 2021.03–2021.09Research internship at PCG, Tencent, Beijing.
- 2020.12–2021.02Research internship at MIG, SenseTime, Beijing.
Publications





ACM MM 2026Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding, Tianyi Gao, Han Fang, Tianyi Ding, Hao Li, Xin Wei, Hongbo Sun, xiaodong dong, Ye Yuan, Jinglin Xu, Kongming Liang, Hao Sun* and Jingmin Xin*.ACM MM 2026GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding, Hao Li, Han Fang, Zixin Pan, Xin Wei, Hongbo Sun, Jinglin Xu, Zhiyu Lin, Ye Yuan, Zhongjiang He, Yu Yu* and Hao Sun*.arXiv preprint 2026SSVP: Synergistic Semantic-Visual Prompting for Industrial Zero-Shot Anomaly Detection, Chenhao Fu, Han Fang, Xiuzheng Zheng, Wenbo Wei, Yonghua Li* and Hao Sun*.arXiv preprint 2026Structure-Aware Prototype Guided Trusted Multi-View Classification, Haojian Huang, Jiahao Shi, Zhe Liu, Harold Haodong Chen, Han Fang, Hao Sun* and Zhongjiang He*.AAAI 2026Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval, Haojian Huang, Kaijing Ma, Jin Chen, Haodong Chen, Zhou Wu, Xianghao Zang, Han Fang, Chao Ban, Hao Sun*, Mulin Chen* and Zhongjiang He.ICDAR 2025SelectVision: Adaptive Vision Resolution Selection for Visual Document Understanding, Zhongjiang He, An Zhao, Ye Yuan, Han Fang, Hao Sun, Kongming Liang and Zhanyu Ma*.TAFFC 2025DDL: Dynamic Direction Learning for Semi-Supervised Facial Expression Recognition, Yaqi Li, Jing Jiang, Yuhang Zhang, Han Fang, Jiani Hu and Weihong Deng*.AAAI 2025Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification, Haojian Huang, Chuanyu Qin, Zhe Liu, Kaijing Ma, Jin Chen, Han Fang, Chao Ban, Hao Sun* and Zhongjiang He*.arXiv preprint 2025Bovila: Bootstrapping Video-Language Alignment via LLM-Based Self-Questioning and Answering, Jin Chen, Kaijing Ma, Haojian Huang, Jiayu Shen, Han Fang, Xianghao Zang, Chao Ban, Zhongjiang He, Hao Sun* and Yanmei Kang*.ICASSP 2025FASTER: Face Attribute Sliders with Semantic Rewards, Jingyan Chen, Lanxiang Zhou, Han Fang, Zerun Feng, Chao Ban, Yaqi Li, Hao Sun* and Jiani Hu*.ICASSP 2025ViCo: A Multitask Video-Enhanced and Cognition-Preserving Modality Alignment Training Framework, Zhenda Yu, Jin Chen, Jiayu Shen, Lanxiang Zhou, Han Fang, Xianghao Zang, Chao Ban, Jingfen Chen, Zhongjiang He, Hao Sun, Zerui Li* and Yanmei Kang*.ACM MM 2024GOAL: Grounded Text-to-Image Synthesis with Joint Layout Alignment Tuning, Yaqi Li, Han Fang, Zerun Feng, Kaijing Ma, Chao Ban, Xianghao Zang, Lanxiang Zhou, Zhongjiang He, Jingyan Chen, Jiani Hu*, Hao Sun* and Huayu Zhang.ICME 2024ProTA: Probabilistic Token Aggregation for Text-Video Retrieval (Oral), Han Fang, Xianghao Zang, Chao Ban, Zerun Feng, Lanxiang Zhou, Zhongjiang He, Yongxiang Li and Hao Sun*.ICME 2024Disentangle and Denoise: Tackling Context Misalignment for Video Moment Retrieval (Oral), Kaijing Ma, Han Fang, Xianghao Zang, Chao Ban, Lanxiang Zhou, Zhongjiang He, Yongxiang Li, Hao Sun, Zerun Feng and Xingsong Hou*.ACM MM 2023A Baseline Investigation: Transformer-Based Cross-View Baseline for Text-Based Person Search, Xianghao Zang, Wei Gao*, Ge Li, Han Fang, Chao Ban, Zhongjiang He and Hao Sun*.ACM MM 2023Mask to Reconstruct: Cooperative Semantics Completion for Video-Text Retrieval, Han Fang, Zhifei Yang, Xianghao Zang, Chao Ban, Zhongjiang He, Hao Sun* and Lanxiang Zhou.TMM 2022Transferring Image-CLIP to Video-Text Retrieval via Temporal Relations, Han Fang, Pengfei Xiong*, Luhui Xu and Wenhan Luo.TMM 2021Dynamic Training Data Dropout for Robust Deep Face Recognition, Yaoyao Zhong, Weihong Deng*, Han Fang, Jiani Hu, Dongyue Zhao, Xian Li and Dongchao Wen.ICASSP 2021FASTER: Face Attribute Sliders with Semantic Rewards, Hongyu Chen, Ruifang Liu*, Han Fang and Ximing Zhang.ECCV 2020Generate to Adapt: Resolution Adaptation Network for Surveillance Face Recognition, Han Fang, Weihong Deng*, Yaoyao Zhong and Jiani Hu.
Honors and Awards
- Ranked 3rd in the ICCV 2025 Multimodal Large Language Model Visual Reasoning Localization Challenge, MARS2: VG-RS, 2025.
- Beijing Excellent Graduate Award (Top 1%), 2022.
- Beijing Excellent Graduate Award (Top 1%), 2019.
- Beijing Excellent Bachelor Dissertation Award (Top 3%), 2019.
Education
- 2019.09–2022.06 M.S. in Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing.
- 2015.09–2019.06 B.Eng. in Telecommunications Engineering with Management, Beijing University of Posts and Telecommunications and Queen Mary University of London, Beijing.