A research notebook研究手记

Research scientist · MiniMax研究科学家 · MiniMax

Xiang An安翔

Seeing more.
Computing less.
看见更多。
计算更少。

I build efficient vision encoders and multimodal models, from visual representations to large-scale learning.探索高效的视觉编码器与多模态模型,从视觉表征到大规模学习。

We’re hiring. Come build with us.我们正在招人,欢迎一起探索。

A soft, paper-toned portrait of a rabbit being lifted against a mountain sunset.暖白纸色中的小兔子,在山间夕阳前被轻轻举起。 Use the wheel or swipe to advance the animation while the page stays still. At the end, scrolling continues through the page. Scroll up to rewind, or use Skip to continue immediately.滚轮或轻划先推进动画,页面保持不动;动画结束后继续滚动页面。向上滚动倒放,也可点击跳过直接浏览。
A moment, in motion一瞬,流动之间 Scroll through the film, then the page先滚动动画,再浏览页面

He has 21 main-conference papers across ICCV (4), CVPR (3), AAAI (3), NeurIPS (3), ECCV (3), EMNLP (2), ICLR (1), COLM (1), and ACM MM (1).

目前共有 21 篇顶会主会论文:ICCV 4 篇、CVPR 3 篇、AAAI 3 篇、NeurIPS 3 篇、ECCV 3 篇、EMNLP 2 篇,ICLR 1 篇、COLM 1 篇、ACM MM 1 篇。

More about my research directions进一步了解研究方向

Publications发表论文 §

The following is a selection of notable publications. For a complete list, see All Publications.

以下为代表性论文精选。完整列表请参见所有发表论文。

  1. Technical Report 2026
    Xiang An, Yin Xie, Feilong Tang, Yunyao Yan, Huajie Tan, Didi Zhu, Changrui Chen, Xiuwei Zhao, Bin Qin, Kaicheng Yang, Yifei Shen, Yuanhan Zhang, Kaichen Zhang, Wenkang Zhang, Zheng Cheng, Nansen Zhang, Chunsheng Wu, Chunjiang Ge, Zimin Ran, Dehua Song, Chunyuan Li, Shikun Feng, Ming Hu, Zhangquan Chen, Junbo Niu, Bo Li, Ziyong Feng, Ziwei Liu, Zongyuan Ge, Jiankang Deng
  2. NeurIPS 2026
    Feilong Tang, Xiang An, Yunyao Yan, Yin Xie, Bin Qin, Kaicheng Yang, Yifei Shen, Yuanhan Zhang, Chunyuan Li, Shikun Feng, Changrui Chen, Huajie Tan, Ming Hu, Manyuan Zhang, Bo Li, Ziyong Feng, Ziwei Liu, Zongyuan Ge, Jiankang Deng
  3. Technical Report, 2025
    Xiang An, Yin Xie, Kaicheng Yang, Wenkang Zhang, Xiuwei Zhao, Zheng Cheng, Yirui Wang, Songcen Xu, Changrui Chen, Didi Zhu, Chunsheng Wu, Huajie Tan, Chunyuan Li, Jing Yang, Jie Yu, Xiyao Wang, Bin Qin, Yumeng Wang, Zizhen Yan, Ziyong Feng, Ziwei Liu, Bo Li, Jiankang Deng
  4. Technical Report 2026
    Senqiao Yang, Kaichen Zhang, Zhaoyang Jia, Jinghao Guo, Yifei Shen, Xinjie Zhang, Xiaoyi Zhang, Haoqing Wang, Xiao Li, Peng Zhang, Xiang An, Yin Xie, Zhening Liu, Xun Guo, Jiahao Li, Shicheng Zheng, Jinglu Wang, Zongyu Guo, Wenxuan Xie, Zihan Zheng, Yuxuan Luo, Bin Li, Yan Lu
  5. AAAI, 2026 (Oral)
    Tiancheng Gu, Kaicheng Yang, Kaichen Zhang, Xiang An, Ziyong Feng, Yueyi Zhang, Weidong Cai, Jiankang Deng, Lidong Bing
  6. ICCV, 2025 (Highlight)
    Yin Xie, Kaicheng Yang, Xiang An (Project Leader), Kun Wu, Yongle Zhao, Weimo Deng, Zimin Ran, Yumeng Wang, Ziyong Feng, Jiankang Deng
    Highlight Presentation
  7. ECCV, 2024
    Xiang An, Kaicheng Yang, Xiangzi Dai, Ziyong Feng, Jiankang Deng
  8. ICLR, 2023
    Xiang An, Jiankang Deng, Kaicheng Yang, Jiawei Li, Ziyong Feng, Jia Guo, Jing Yang, Tongliang Liu
  9. CVPR, 2022
    Xiang An, Jiankang Deng, Jia Guo, Ziyong Feng, Xuhan Zhu, Jing Yang, Tongliang Liu
  10. ICCVW, 2021
    Xiang An, Xuhan Zhu, Yuan Gao, Yang Xiao, Yongle Zhao, Ziyong Feng, Lan Wu, Bin Qin, Ming Zhang, Debing Zhang, Ying Fu

Awards & Competitions §

荣誉与竞赛 §

  • ICCV 2025 Outstanding Reviewer
  • CVPR 2024 Outstanding Reviewer
  • Ranked 1st in NIST FRVT Competition, Visa Track 1:1
  • 2024 中国年度力量人物提名
  • Ranked 1st in the graduate entrance examination (major)
  • First Place in Vehicle Re-Identification, PRCV 2019
  • ICCV 2025 杰出审稿人
  • CVPR 2024 杰出审稿人
  • NIST FRVT 竞赛 Visa Track 1:1 第一名
  • 2024 中国年度力量人物提名
  • 研究生入学考试(专业课)第一名
  • PRCV 2019 车辆重识别第一名

Open Source §

开源项目 §

  1. Open Source Library
    #2 contributor to the open-source 2D & 3D deep face analysis library. Author of Glint360K (the largest open-source face recognition training dataset) and Partial FC (enabling training 10 million identities on a single machine). Also organized the ICCV 2021 Workshop on masked face recognition challenge.
    开源2D/3D深度人脸分析库的第二贡献者。Glint360K(最大开源人脸识别训练数据集)和Partial FC(实现单机训练千万级身份)的作者。还组织了ICCV 2021口罩人脸识别挑战赛Workshop。
  2. Multimodal LLM Framework
    Team Leader of this fully open framework designed to democratize multimodal training. Released mid-training and instruct data for community use, and developed offline sampling pack for efficient training. Implemented RiceViT with native resolution support.
    该完全开放框架的团队负责人,旨在推动多模态训练的民主化。向社区发布了中期训练数据和指令数据,并开发了离线采样包以提高训练效率。实现了支持原生分辨率的RiceViT。
  3. Vision Encoder
    Project leader of this next-generation vision encoder that introduces codec-aligned sparsity as a foundational principle for multimodal intelligence. Achieves state-of-the-art performance on 16 image, video, and document understanding benchmarks while using substantially fewer visual tokens. Demonstrates 4.1% average improvement over Qwen3-ViT on video understanding tasks.
    下一代视觉编码器的项目负责人,提出编解码器对齐稀疏性作为多模态智能的基础原则。在 16 个图像、视频和文档理解基准上取得最先进性能,同时显著减少视觉 token 数量。在视频理解任务上比 Qwen3-ViT 平均提升 4.1%。
  4. Image Retrieval Framework
    Lead author and maintainer of Universal and Compact Representation Learning framework for universal image representations. Designed the novel cluster discrimination approach for representation learning. Developed the multi-label and region-based extensions (published at ECCV 2024 and ICCV 2025 (Highlight)).
    通用紧凑表征学习框架的项目负责人和主要作者,用于通用图像表征。设计了新颖的聚类判别方法用于表征学习。开发了多标签和区域级扩展(分别发表于ECCV 2024和ICCV 2025 (Highlight))。
  5. Large Multimodal Model
    Vision module contributor to the next-generation large multimodal model. Enhanced the OCR capability of the vision module for better text recognition in images. Optimized the visual encoder for processing text-rich and document images.
    下一代大型多模态模型的视觉模块贡献者。增强了视觉模块的OCR能力以改善图像中的文字识别。优化了视觉编码器以处理富文本和文档图像。
  6. Educational Project
    Author and maintainer of this educational project for semantic segmentation on remote sensing and satellite imagery. Designed a simple single-file training approach for accessibility and integrated popular pretrained models. Created comprehensive tutorials and documentation for beginners.
    该教育项目的作者和维护者,用于遥感和卫星图像的语义分割。设计了简洁的单文件训练方法以提高可用性,并集成了流行的预训练模型。为初学者编写了全面的教程和文档。

Citation Map §

引用地图 §

Publication affiliations of citing papers, exported from the local citation library.

引用论文发表时的机构分布,来自本地引用库快照。

1×

Preparing your atlas…正在展开世界地图…

About this map地图说明

Click a country or point to explore. Drag to pan; use +/− or Ctrl + scroll to zoom. On touch screens, pinch to zoom.点选国家或地点查看详情;拖动平移,使用 +/− 或 Ctrl+滚轮缩放。触屏可双指缩放。

Counts are distinct citing papers per country, location, or institution, using resolved work IDs. A paper can appear in multiple places, so counts are not additive. Unresolved duplicates may remain. Non-excluded reported citations and publication affiliations are included, including rule-extracted PDF affiliations; they are not all human-verified.各国家、地点及机构按库中论文 ID 去重计数。同一论文可能出现在多个地点,各项不能直接相加;尚未合并的重复论文可能仍存在。快照包含未排除的来源报告引用及发表时机构,也包含 PDF 规则提取的机构关系,并非全部经过人工核验。


This page is styled after Wikipedia.

本页面样式参考自维基百科。

This page was last edited on .  |  Powered by pixels & caffeine  ♥ 本页面最后编辑于。  |  由像素和咏啡因驱动  ♥