Notes /
PhD student in AI · In progress

Ziheng
He.

PhD student in artificial intelligence. I work on embodied world models, video generation for robotics, and training manipulation policies from generated data.

World Models Embodied AI Robotics
Scroll down
Embodied World Models· Video Generation· Robot Learning· Diffusion· LIBERO
— 01 / ABOUT

About

Ziheng He
Status · PhD student in AI
Affiliation · SAIS, UCAS · Institute of Automation, CAS

I'm Ziheng, a PhD student in artificial intelligence. I work on embodied world models — predicting how a robot's actions change the scene around it — and on making video generation efficient enough to train real manipulation policies.

I like working across the whole stack — the generative models themselves and the tooling to train and evaluate them — and I care about results that hold up on real robots, not just on benchmarks.

04
Papers
2026
First Paper
04
Featured Repos
∞
Curiosity
— 02 / RESEARCH

Research

Centered on embodied world models — from the world model itself, to video generation for robotics, to robot policy learning.

01

Embodied World Models

Modeling how a robot's actions reshape the surrounding scene, and making rollout inference for long-horizon manipulation efficient.

World ModelRoboticsRollout
02

Video Generation & Diffusion

Sparse keyframe synthesis, video diffusion, and action-conditioned interpolation — cutting generation cost while preserving task-critical events.

DiffusionKeyframeVideo
03

Robot Learning & Manipulation

Training manipulation policies (e.g. VLA, π0.5) on generated video data, validated on benchmarks like LIBERO and on real robots.

VLAPolicyLIBERO
— 03 / TECH STACK

Stack

The tools and languages I work with day to day.

AI / Machine Learning
Languages
Tools & Environment
— 04 / ACTIVITY

Activity

Commits and milestones over the past year.

— contributions in the past year
Less More

Loading GitHub contributions…

● Longest streak · —
● Current streak · —
● Busiest day · —
Recent Milestones
Improved agent capability loading and execution, with support for YAML / TOML attachments.
A local-first AI workspace connecting conversations, project files, source-linked knowledge, and tasks.
First-author paper on efficient embodied world models through event-preserving sparse keyframe generation and interpolation.
Open Source Footprint
Local-first AI workspace · Creator & Maintainer
★ 328
Forks · 41
Unified Claude usage dashboard
★ 71
Forks · 5
Link preview & reader-mode Chrome extension
★ 5
Forks · 0
Desktop client for AI agents · Contributor
★ 32.2k
Forks · 3311
— 05 / EDUCATION

Education

A path through textbooks, labs, and open-source communities.

2026 — Present

Ph.D. in Artificial Intelligence

School of Advanced Interdisciplinary Sciences, UCAS

Ph.D. in artificial intelligence at the School of Advanced Interdisciplinary Sciences (SAIS), University of Chinese Academy of Sciences.

AIUCAS
2026 — Present

Ph.D. in Artificial Intelligence

Institute of Automation, CAS · NLPR

Research in artificial intelligence at the Institute of Automation (NLPR), Chinese Academy of Sciences.

AINLPR
2022 — 2026

B.Eng. in Software Engineering

Southeast University

B.Eng. in Software Engineering at Southeast University.

SoftwareEngineering
— 06 / EXPERIENCE

Experience

Embodied-AI research on the industry frontline.

2025.09 — Present

Research Intern

GigaAI (极佳视界)

Research on embodied world models and video generation; member of the GigaBrain team.

World ModelsGigaAI
2025.09 — Present

Research Intern

FiveAges (中科第五纪)

Research internship in embodied intelligence.

Embodied AI
— 07 / RESEARCH WORK

Research Work

First-author and collaborative research. Future work will appear on Google Scholar.

2026 SKIP — paper first page
CoRL 2026 First Author cs.RO

SKIP: Sparse Keyframe Interpolation Paradigm for Efficient Embodied World Models

Ziheng He, Yixiang Chen, Ning Yang, et al. · Conference on Robot Learning (CoRL) 2026 · Accepted Sep 5, 2026

Sparse Keyframe Interpolation (SKIP) is an event-preserving, sparse-to-dense framework for embodied world models: it synthesizes only task-relevant keyframes with a sparse video diffusion model, then interpolates the missing intervals conditioned on robot actions — avoiding dense frame-by-frame rollout. On LIBERO it runs 4.16× faster than a dense baseline while cutting FVD by 89.0%, and its generated videos work directly as policy-training data.

2026 GigaBrain-0.7 — paper first page
arXiv 2026 cs.RO

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

GigaBrain Team (incl. Ziheng He) · arXiv preprint · cs.RO

An embodied foundation model built on a three-system architecture, scaling the VLA paradigm to larger, more heterogeneous data regimes with substantially improved generalization across tasks and robot embodiments.

2026 XEWorld — paper first page
arXiv 2026 cs.RO

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

Yixiang Chen, Jiabing Yang, Yuan Xu, Qisen Ma, Keji He, Peiyan Li, Kai Wang, Ziheng He, et al. · arXiv preprint · cs.RO

A controlled cross-embodiment testbed that evaluates world models on held-out robots within physically identical scenes. Systematic analysis uncovers a shared bottleneck: current models act primarily as 2D visual pattern matchers rather than capturing physical dynamics.

2026 WAM-Nav — paper first page
arXiv 2026 cs.RO

WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation

Ning Yang, Yan Huang, Kaiwen Peng, Ziheng He, et al. · arXiv preprint · cs.RO

A latent world-action model for embodied visual navigation that jointly learns scene prediction and policy, giving navigation anticipatory reasoning while avoiding the error accumulation and slow inference of modular pipelines.

— 08 / NOTES

Notes

Inspired by Super.so — notes open in an in-site reader styled to match this page, not a raw Notion page. Content is written in Notion and synced here.

Auto-synced from Notion