My name is Yulin Luo (罗峪霖). I’m a Ph.D. candidate (since 2023) at the School of Computer Science, Peking University, advised by Prof. Shanghang Zhang. I received my Bachelor’s degree in Automation from Shanghai Jiao Tong University in 2023. My current vision is open-world generalization for embodied intelligence. More details are available in my curriculum vitae.

Sections

Embodied Researcher

  • Generalizable models: designing architectures and learning algorithms for embodied foundation models that generalize across objects, skills, embodiments, scenes, and tasks.
  • Goal-aligned benchmarks: evaluating model capabilities under realistic conditions and against the long-term goal of open-world robot intelligence.
  • Evaluation-driven co-evolution: turning diagnosed gaps into model and data redesign, then using improved models and data to define the next evaluation cycle.
Open-world generalization hierarchy for embodied intelligence

🔥 News

  • 2026.07:  🎉 DeepVision-VLA, a vision-representation enhancement method for VLA models, is accepted to ACM MM 2026. (first author)
  • 2026.07:  📊 RoboBench is included in the official evaluation suite of HY-Embodied (RoboBench-MCQ & RoboBench-Planning).
  • 2026.07:  🎉 RoboBench, a robot-scenario benchmark for evaluating MLLMs as embodied brains, is accepted to ECCV 2026. (first author)
  • 2025.11:  🎉 MoASE, a sparse MoE adaptation method for robust continual test-time vision models, is accepted as Oral at AAAI 2026. (co-first author)
  • 2025.09:  🎉 SEEA-R1, a tree-structured RL fine-tuning method for self-evolving embodied agents, is accepted to NeurIPS 2025.
  • 2025.04:  🎉 RoboMIND, a multi-embodiment robot manipulation benchmark with normative data, is accepted to RSS 2025.
  • 2025.01:  🎉 Draw-and-Understand, a visual-prompt method that helps MLLMs understand user intent, is accepted to ICLR 2025.
  • 2024.07:  🎉 SSD-LLM, an LLM-based dataset analyst for discovering subpopulation structure, is accepted to ECCV 2024. (first author)
  • 2023.12:  🎉 MoFME, an uncertainty-aware mixture-of-experts method for efficient image deweathering, is accepted to AAAI 2024.
  • 2023.09:  🎓 Started Ph.D. at Peking University, advised by Prof. Shanghang Zhang.
  • 2023.06:  🎓 Graduated with B.Eng. in Automation from Shanghai Jiao Tong University.
  • 2022.06:  🎉 RandStainNA, a stain augmentation and normalization method for robust histology slide features, is accepted to MICCAI 2022. (co-first author)

🏫 Education

2023.09 - Present

Peking University

Ph.D. Candidate
School of Computer Science
Advisor: Prof. Shanghang Zhang
2019.09 - 2023.06

Shanghai Jiao Tong University

B.Eng. in Automation
School of Electronic Information and Electrical Engineering
GPA: 3.87/4.30; Major rank: 2/89.

💼 Research Experience

2026.07 - Present

EvoPhys (宇万智能)

Chief AI Officer (CAIO)
Unified foundation models integrating world engine and world policy.
2025.12 - 2026.07

Simplexity Robotics (至简动力)

Research Intern
Embodied intelligence and generalizable robot foundation models.
2024.08 - 2025.10

Beijing Academy of Artificial Intelligence (BAAI)

Research Intern
Embodied-brain benchmarks for multimodal large language models and robot intelligence.
2024.03 - 2024.08

ByteDance

Research Intern
Trust & Safety Applied Algorithm Team for Ads and E-commerce.

📃 Publications

Research Trajectory

Selected first / co-first works since 2024. Hollow dots mark arXiv v1 release; solid dots mark acceptance.

arXiv v1 Accepted Oral / highlight
ECCV 2026
RoboBench teaser
TeaserBenchmark overview across embodied abilities
RoboBench construction pipeline
MethodData construction pipeline for robot scenarios

RoboBench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

Yulin Luo⭐️, Chun-Kai Fan⭐️, Menghang Dong⭐️, Jiayu Shi⭐️, Xiangju Mi⭐️, Mengdi Zhao⭐️†, Bo-Wen Zhang⭐️, Cheng Chi⭐️†, et al., Shanghang Zhang📧
European Conference on Computer Vision (ECCV), 2026
arXiv v1: 2025.10 · Accepted: 2026.07
Claim: Can an MLLM serve as a robot's brain? RoboBench turns the question into measurable diagnosis — 6K QA across embodied cognitive dimensions — and shows cognitive scores genuinely predict downstream VLA performance; now the official evaluation suite of Tencent HY-Embodied.
ECCV 2024
SSD-LLM teaser
TeaserWorkflow for discovering subpopulation structure
SSD-LLM method diagram
MethodCaption, criteria generation, refinement, and assignment

LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model

Short name: SSD-LLM
Yulin Luo⭐️, Ruichuan An⭐️, Bocheng Zou, Yiming Tang, Jiaming Liu, Shanghang Zhang📧
European Conference on Computer Vision (ECCV), 2024
arXiv v1: 2024.05 · Accepted: 2024.07
Claim: The first systematic exploration of dataset subpopulation structure: SSD-LLM casts the LLM as a dataset analyst that discovers hidden subpopulations by itself — exposing dataset bias and prescribing which data to add next, a concrete mechanism for model–data co-evolution.
AAAI 2026 · Oral
MoASE teaser
TeaserSparse activations motivate expert decomposition
MoASE method diagram
MethodMoASE routes activation patterns into experts

Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test-Time Adaptation

Short name: MoASE
Rongyu Zhang⭐️, Aosong Cheng⭐️, Yulin Luo⭐️, Gaole Dai, Huanrui Yang, et al., Shanghang Zhang📧
AAAI Conference on Artificial Intelligence (AAAI), 2026 · Oral
arXiv v1: 2024.05 · Accepted: 2025.11
Claim: A new reading of continual test-time adaptation: each domain fires its own sparse neuron pattern, so routing activations into decomposed experts turns catastrophic forgetting into a routing problem — SOTA on four CTTA benchmarks without revisiting old data.
ACM MM 2026
DeepVision-VLA teaser
TeaserVision foundation representations before robot action
DeepVision-VLA method diagram
MethodRepresentation enhancement pipeline for VLA policies

Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models

Short name: DeepVision-VLA
Yulin Luo⭐️, Hao Chen⭐️†, Zhuangzhe Wu⭐️, Bowen Sui⭐️, Jiaming Liu⭐️†, et al., Shanghang Zhang📧
ACM International Conference on Multimedia (ACM MM), 2026
arXiv v1: 2026.03 · Accepted: 2026.07
Claim: A systematic look inside the VLA backbone reveals that visual-token sensitivity decays with depth: injecting multi-level visual features into deeper layers (VL-MoT + action-guided visual pruning) yields +9.0% / +7.5% over prior SOTA on simulated and real-world tasks — look before acting.
Co-authored Papers5
MICCAI 2022
RandStainNA teaser
TeaserStain augmentation and normalization are bridged
RandStainNA method diagram
MethodRandom virtual templates for stain-agnostic training

RandStainNA: Learning Stain-Agnostic Features from Histology Slides by Bridging Stain Augmentation and Normalization

Yiqing Shen⭐️, Yulin Luo⭐️, Dinggang Shen, Jing Ke📧
Medical Image Computing and Computer Assisted Intervention (MICCAI), 2022
arXiv v1: 2022.06 · Accepted: 2022.06
Claim: Histology models break when the stain changes. RandStainNA unifies stain normalization and augmentation into one scheme — random color-space sampling within a practicable range — consistently improving generalization across diagnostic tasks and backbones.
RSS 2025
RoboMIND teaser
TeaserMulti-embodiment robot manipulation benchmark
RoboMIND annotation process
MethodLanguage-guided task process and annotation examples

RoboMIND: Benchmark on Multi-Embodiment Intelligence Normative Data for Robot Manipulation

Kun Wu⭐️, Chengkai Hou⭐️, Jiaming Liu⭐️, Zhengping Che⭐️†, Xiaozhu Ju⭐️†, ..., Yulin Luo, ..., Shanghang Zhang📧, Jian Tang📧
Robotics: Science and Systems (RSS), 2025
arXiv v1: 2024.12 · Accepted: 2025.04
Claim: Skills should transfer across robot bodies, not be re-learned per platform. RoboMIND is the largest multi-embodiment teleoperation dataset on a unified platform (107k trajectories, 479 tasks, 4 embodiments) — a data foundation to test and push general manipulation.
ICLR 2025
Draw-and-Understand teaser
TeaserVisual prompting for fine-grained user intent
Draw-and-Understand training strategy and model architecture
MethodTwo-stage training and the SPHINX-V architecture

Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Weifeng Lin⭐️, Xinyu Wei⭐️, Ruichuan An, Peng Gao, Bocheng Zou, Yulin Luo, Siyuan Huang, Shanghang Zhang, Hongsheng Li📧
International Conference on Learning Representations (ICLR), 2025
arXiv v1: 2024.03 · Accepted: 2025.01
Claim: Language is a lossy way to point. This work lets users draw what they mean — points, boxes, free-form shapes — with 1.2M image–visual-prompt–text triplets (MDVP-Instruct-Data) and MDVP-Bench to make MLLMs follow fine-grained visual intent.
NeurIPS 2025
SEEA-R1 teaser
TeaserSelf-evolving embodied agents through tree search
SEEA-R1 data and model evolution framework
MethodCoupled data and model evolution with relative advantages

SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents

Wanxin Tian⭐️, Shijie Zhang⭐️, Kevin Zhang⭐️, Xiaowei Chi, Chun-Kai Fan, Junyu Lu, Yulin Luo, et al., Shanghang Zhang📧, Jian Tang📧
Advances in Neural Information Processing Systems (NeurIPS), 2025
arXiv v1: 2025.06 · Accepted: 2025.09
Claim: The first RFT framework for self-evolving embodied agents: Tree-GRPO turns sparse delayed rewards into dense intermediate signals, and a multi-modal generative reward model frees self-evolution from hand-crafted rewards — surpassing GPT-4o on ALFWorld even without ground-truth rewards.
AAAI 2024
MoFME teaser
TeaserDeweathering pipeline for downstream perception
MoFME method diagram
MethodFeature modulation experts with uncertainty-aware routing

Efficient Deweather Mixture-of-Experts with Uncertainty-Aware Feature-Wise Linear Modulation

Short name: MoFME
Rongyu Zhang, Yulin Luo, Jiaming Liu, Huanrui Yang, Zhen Dong, et al., Yuan Du📧, Shanghang Zhang📧
AAAI Conference on Artificial Intelligence (AAAI), 2024
arXiv v1: 2023.12 · Accepted: 2023.12
Claim: One model should face all weather. MoFME implicitly instantiates experts via feature modulation on a shared block with an uncertainty-aware router — SOTA-compatible restoration while saving 72% of parameters and 39% of inference time over conventional MoE.

🏆 Awards

Shanghai Jiao Tong University
  • Outstanding Graduate2023.06
  • Weichai Power Scholarship2022.10
  • Huawei Scholarship2021.12
  • B-Class Merit Scholarship2020 · 2021 · 2022

👓 Services

Conference Reviewer
AAAI 2026 ICLR 2026 CVPR 2026 ECCV 2026 NeurIPS 2026 Datasets & Benchmarks

❤️ Love

My lovely girlfriend
My Girl
Lucky beyond words to have her — my lovely girlfriend, who makes the good days sweeter and the hard research days lighter.

⚽ Interests

Manuel Neuer in the FC Bayern goalkeeper kit
Football · Goalkeeper
Huge fan of Manuel Neuer (Germany & FC Bayern München) — the sweeper-keeper who redefined the position. I guard the goal myself: starting goalkeeper for SEIEE Team B at Shanghai Jiao Tong University during my undergrad, and currently first-choice goalkeeper for the Changping United team — a joint squad of the Integrated Circuits, Computer Science, and Electronics schools at Peking University.
SEIEE Team B group photo, Shanghai Jiao Tong University
SEIEE Team B, SJTU · 2019 Freshman Cup
Changping United team photo (IC / CS / EE joint squad), Peking University
Changping United (IC · CS · EE joint team), PKU · 2026
Lab team with silver medals at the school badminton tournament
Lab team at the school badminton tournament, PKU · silver medal
Badminton · Lab Team
Regular badminton player with my lab mates — we entered the school badminton tournament together and took silver. Nothing resets the brain after a long research day like a few hard-fought rallies on court.

✈️ Travel

Academic Conferences
Conferences double as my way to see the world: ICLR 2025 in Singapore, and ECCV 2024 in Milan, Italy — with hopefully many more stamps in the passport to come.
ICLR 2025 · Singapore
NUS campus and city walk, Singapore
Singapore coastline views
Sentosa seaside and water park, Singapore
Beach swings and seaside in Singapore
ECCV 2024 · Milan, Italy
Ponte Vecchio and the Arno river, Italy
Varenna and Lake Como views, Italy
Mandello del Lario railway station, Italy
Lake Como ferry docks and lakeside, Italy

🔗 Academic Friends

My academic friends include Chun-Kai Fan, Gaole Dai, Ruichuan An, Jiaming Liu, Siyuan Qian, and Sixiang Chen from PKU HMI 🇨🇳; Rongyu Zhang from NJU 🇨🇳.

Too many friends to list everyone — let me know if I forgot you 😂