Zijie Yan 颜子杰

I work on training infra at Periodic Labs.

Previously, I was a tech leader at NVIDIA DevTech, where I built and led the MoE training system in Megatron-LM, enabling efficient training of trillion-parameter MoE models.

Before that, I led training infra at SenseTime and researched distributed training during my master's at Sun Yat-sen University.

Zijie Yan hiking in the mountains

Selected work

2025

MoE Parallel Folding

Dennis Liu*, Zijie Yan*, Xin Yao et al.

* Equal contribution. First authors listed alphabetically.

Introduces Parallel Folding, which decouples attention and MoE parallelism so each can be optimized independently. The framework achieves up to 49.3% MFU for Mixtral 8x22B on H100 and scales to 1,024 GPUs.

All publications on Google Scholar