My name is (罗峪霖). I’m a Ph.D. candidate (since 2023) at the School of Computer Science, Peking University, advised by Prof. Shanghang Zhang. I received my Bachelor’s degree in Automation from Shanghai Jiao Tong University in 2023. My current vision is open-world generalization for embodied intelligence. More details are available in my curriculum vitae.
Sections
Embodied Researcher
- Generalizable models: designing architectures and learning algorithms for embodied foundation models that generalize across objects, skills, embodiments, scenes, and tasks.
- Goal-aligned benchmarks: evaluating model capabilities under realistic conditions and against the long-term goal of open-world robot intelligence.
- Evaluation-driven co-evolution: turning diagnosed gaps into model and data redesign, then using improved models and data to define the next evaluation cycle.
🔥 News
- 2026.07: 🎉 DeepVision-VLA, a vision-representation enhancement method for VLA models, is accepted to ACM MM 2026. (first author)
- 2026.07: 📊 RoboBench is included in the official evaluation suite of HY-Embodied (RoboBench-MCQ & RoboBench-Planning).
- 2026.07: 🎉 RoboBench, a robot-scenario benchmark for evaluating MLLMs as embodied brains, is accepted to ECCV 2026. (first author)
- 2025.11: 🎉 MoASE, a sparse MoE adaptation method for robust continual test-time vision models, is accepted as Oral at AAAI 2026. (co-first author)
- 2025.09: 🎉 SEEA-R1, a tree-structured RL fine-tuning method for self-evolving embodied agents, is accepted to NeurIPS 2025.
- 2025.04: 🎉 RoboMIND, a multi-embodiment robot manipulation benchmark with normative data, is accepted to RSS 2025.
- 2025.01: 🎉 Draw-and-Understand, a visual-prompt method that helps MLLMs understand user intent, is accepted to ICLR 2025.
- 2024.07: 🎉 SSD-LLM, an LLM-based dataset analyst for discovering subpopulation structure, is accepted to ECCV 2024. (first author)
- 2023.12: 🎉 MoFME, an uncertainty-aware mixture-of-experts method for efficient image deweathering, is accepted to AAAI 2024.
- 2023.09: 🎓 Started Ph.D. at Peking University, advised by Prof. Shanghang Zhang.
- 2023.06: 🎓 Graduated with B.Eng. in Automation from Shanghai Jiao Tong University.
- 2022.06: 🎉 RandStainNA, a stain augmentation and normalization method for robust histology slide features, is accepted to MICCAI 2022. (co-first author)
🏫 Education
2023.09 - Present
2019.09 - 2023.06
Shanghai Jiao Tong University
B.Eng. in Automation
School of Electronic Information and Electrical Engineering
GPA: 3.87/4.30; Major rank: 2/89.
💼 Research Experience
2026.07 - Present
EvoPhys (宇万智能)
Chief AI Officer (CAIO)
Unified foundation models integrating world engine and world policy.
2025.12 - 2026.07
Simplexity Robotics (至简动力)
Research Intern
Embodied intelligence and generalizable robot foundation models.
2024.08 - 2025.10
Beijing Academy of Artificial Intelligence (BAAI)
Research Intern
Embodied-brain benchmarks for multimodal large language models and robot intelligence.
2024.03 - 2024.08
📃 Publications
Research Trajectory
Selected first / co-first works since 2024. Hollow dots mark arXiv v1 release; solid dots mark acceptance.
arXiv v1
Accepted
Oral / highlight
Evaluation
Model-Data Co-evolution
ECCV 2026
TeaserBenchmark overview across embodied abilities
MethodData construction pipeline for robot scenarios
RoboBench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
European Conference on Computer Vision (ECCV), 2026
arXiv v1: 2025.10 · Accepted: 2026.07
Claim: Can an MLLM serve as a robot's brain? RoboBench turns the question into measurable diagnosis — 6K QA across embodied cognitive dimensions — and shows cognitive scores genuinely predict downstream VLA performance; now the official evaluation suite of Tencent HY-Embodied.
ECCV 2024
TeaserWorkflow for discovering subpopulation structure
MethodCaption, criteria generation, refinement, and assignment
LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model
Short name: SSD-LLM
European Conference on Computer Vision (ECCV), 2024
arXiv v1: 2024.05 · Accepted: 2024.07
Claim: The first systematic exploration of dataset subpopulation structure: SSD-LLM casts the LLM as a dataset analyst that discovers hidden subpopulations by itself — exposing dataset bias and prescribing which data to add next, a concrete mechanism for model–data co-evolution.
AAAI 2026 · Oral
TeaserSparse activations motivate expert decomposition
MethodMoASE routes activation patterns into experts
Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test-Time Adaptation
Short name: MoASE
AAAI Conference on Artificial Intelligence (AAAI), 2026 · Oral
arXiv v1: 2024.05 · Accepted: 2025.11
Claim: A new reading of continual test-time adaptation: each domain fires its own sparse neuron pattern, so routing activations into decomposed experts turns catastrophic forgetting into a routing problem — SOTA on four CTTA benchmarks without revisiting old data.
ACM MM 2026
TeaserVision foundation representations before robot action
MethodRepresentation enhancement pipeline for VLA policies
Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models
Short name: DeepVision-VLA
ACM International Conference on Multimedia (ACM MM), 2026
arXiv v1: 2026.03 · Accepted: 2026.07
Claim: A systematic look inside the VLA backbone reveals that visual-token sensitivity decays with depth: injecting multi-level visual features into deeper layers (VL-MoT + action-guided visual pruning) yields +9.0% / +7.5% over prior SOTA on simulated and real-world tasks — look before acting.
Co-authored Papers5
MICCAI 2022
TeaserStain augmentation and normalization are bridged
MethodRandom virtual templates for stain-agnostic training
RandStainNA: Learning Stain-Agnostic Features from Histology Slides by Bridging Stain Augmentation and Normalization
Medical Image Computing and Computer Assisted Intervention (MICCAI), 2022
arXiv v1: 2022.06 · Accepted: 2022.06
Claim: Histology models break when the stain changes. RandStainNA unifies stain normalization and augmentation into one scheme — random color-space sampling within a practicable range — consistently improving generalization across diagnostic tasks and backbones.
RSS 2025
TeaserMulti-embodiment robot manipulation benchmark
MethodLanguage-guided task process and annotation examples
RoboMIND: Benchmark on Multi-Embodiment Intelligence Normative Data for Robot Manipulation
Robotics: Science and Systems (RSS), 2025
arXiv v1: 2024.12 · Accepted: 2025.04
Claim: Skills should transfer across robot bodies, not be re-learned per platform. RoboMIND is the largest multi-embodiment teleoperation dataset on a unified platform (107k trajectories, 479 tasks, 4 embodiments) — a data foundation to test and push general manipulation.
ICLR 2025
TeaserVisual prompting for fine-grained user intent
MethodTwo-stage training and the SPHINX-V architecture
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
International Conference on Learning Representations (ICLR), 2025
arXiv v1: 2024.03 · Accepted: 2025.01
Claim: Language is a lossy way to point. This work lets users draw what they mean — points, boxes, free-form shapes — with 1.2M image–visual-prompt–text triplets (MDVP-Instruct-Data) and MDVP-Bench to make MLLMs follow fine-grained visual intent.
NeurIPS 2025
TeaserSelf-evolving embodied agents through tree search
MethodCoupled data and model evolution with relative advantages
SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents
Advances in Neural Information Processing Systems (NeurIPS), 2025
arXiv v1: 2025.06 · Accepted: 2025.09
Claim: The first RFT framework for self-evolving embodied agents: Tree-GRPO turns sparse delayed rewards into dense intermediate signals, and a multi-modal generative reward model frees self-evolution from hand-crafted rewards — surpassing GPT-4o on ALFWorld even without ground-truth rewards.
AAAI 2024
TeaserDeweathering pipeline for downstream perception
MethodFeature modulation experts with uncertainty-aware routing
Efficient Deweather Mixture-of-Experts with Uncertainty-Aware Feature-Wise Linear Modulation
Short name: MoFME
AAAI Conference on Artificial Intelligence (AAAI), 2024
arXiv v1: 2023.12 · Accepted: 2023.12
Claim: One model should face all weather. MoFME implicitly instantiates experts via feature modulation on a shared block with an uncertainty-aware router — SOTA-compatible restoration while saving 72% of parameters and 39% of inference time over conventional MoE.
🏆 Awards
Shanghai Jiao Tong University
- Outstanding Graduate2023.06
- Weichai Power Scholarship2022.10
- Huawei Scholarship2021.12
- B-Class Merit Scholarship2020 · 2021 · 2022
👓 Services
Conference Reviewer
❤️ Love
My Girl
Lucky beyond words to have her — my lovely girlfriend,
who makes the good days sweeter and the hard research days lighter.
⚽ Interests
Football · Goalkeeper
Huge fan of Manuel Neuer (Germany & FC Bayern München) — the sweeper-keeper who redefined the position.
I guard the goal myself: starting goalkeeper for SEIEE Team B at Shanghai Jiao Tong University during my undergrad,
and currently first-choice goalkeeper for the Changping United team — a joint squad of the Integrated Circuits,
Computer Science, and Electronics schools at Peking University.
Badminton · Lab Team
Regular badminton player with my lab mates — we entered the school badminton tournament together and took silver.
Nothing resets the brain after a long research day like a few hard-fought rallies on court.
✈️ Travel
Academic Conferences
Conferences double as my way to see the world: ICLR 2025 in Singapore, and ECCV 2024 in Milan, Italy —
with hopefully many more stamps in the passport to come.
ICLR 2025 · Singapore
ECCV 2024 · Milan, Italy
🔗 Academic Friends
My academic friends include Chun-Kai Fan, Gaole Dai, Ruichuan An, Jiaming Liu, Siyuan Qian, and Sixiang Chen from PKU HMI 🇨🇳; Rongyu Zhang from NJU 🇨🇳.
Too many friends to list everyone — let me know if I forgot you 😂