Zhenghai Xue
Tencent Hy. zhenghai001@e.ntu.edu.sg
I am a member of technical staff at Tencent Hy, in charge of production-level post training with OPD and RL. I received my Ph.D. from the College of Computing and Data Science, Nanyang Technological University, Singapore, supervised by Prof. Bo An. Previously, I obtained my B.Sc. of Artificial Intelligence from Nanjing University in 2022. In my undergraduate study, I worked with Prof. Yang Yu at LAMDA and Prof. Bolei Zhou at MMLab.
I have been fortunate to work as a research intern at the following companies:
- Moonshot AI with Dr. Chenjun Xiao. We worked on improving the long-horizon instruction-following abilities of coding agents.
- TikTok with Dr. Qian Liu. We worked on scalable and stable algorithms for agentic RL training.
- Kuaishou Technology with Dr. Qingpeng Cai. We worked on RL for improving user retention.
- Kunlun 2050 Research with Prof. Shuicheng Yan. We worked on robust and human-in-the-loop RL.
News
| Jul 4, 2025 | We release SimpleTIR, an end-to-end solution for stable multi-turn tool use RL training. |
|---|---|
| May 3, 2025 | One paper is accepted at ICML 2025 as Spotlight Poster! |
| Jan 23, 2025 | One paper is accepted at WWW 2025. Two papers are accepted to ICLR 2025. |
| Nov 8, 2023 | I will give a talk on Optimizing Long-term User Engagement in the Applied Artificial Intelligence Workshop of DAI 2023! |
| Oct 19, 2023 | Our paper “State Regularized Policy Optimization on Data with Dynamics Shift” is accepted by NeurIPS 2023! |
Selected publications
-
Policy Optimization under Imperfect Human Interactions with Agent-Gated Shared AutonomyIn The Thirteenth International Conference on Learning Representations, 2025 -
Policy Regularization on Globally Accessible States in Cross-Dynamics Reinforcement LearningIn Forty-Second International Conference on Machine Learning (Spotlight), 2025 -
AURO: Reinforcement Learning for Adaptive User Retention Optimization in Recommender SystemsIn Proceedings of the ACM Web Conference (Oral), 2025




