Biography
I am a M.S. student at Tsinghua University (expecting to graduate in 2027), advised by Prof. Yong Li and Prof. Fengli Xu. I'm interested in LLM post-training, LLM Agent.
Previously, I received my B.E. degree from the School of Electronics and Information Engineering, Harbin Institute of Technology (Shenzhen) in June 2024, ranking 1/227 with a GPA of 4.0/4.0.
My research focuses on LLM post-training and LLM Agent, with particular interest in AI4AI/RSI, Self-evolving Agent, Harness engineering.
Research Publications
-
AgentExpt: Automating AI Experiment Design with Resource Retrieval Agent
Yu Li, Lehui Li, Qingmin Liao, Fengli Xu, Yong Li
International Conference on Machine Learning (ICML), 2026
Cited by 3
-
AgentSwift: Efficient LLM Agent Design via Value-guided Hierarchical Search
Yu Li, Lehui Li, Zhihao Wu, Qingmin Liao, Jianye Hao, Kun Shao, Fengli Xu, Yong Li
Association for the Advancement of Artificial Intelligence (AAAI), 2026
Cited by 24
-
Agentsquare: Automatic LLM Agent Search in Modular Design Space
Yu Shang*, Yu Li*, Keyu Zhao, Likai Ma, Jiahe Liu, Fengli Xu, Yong Li
International Conference on Learning Representations (ICLR), 2025
Cited by 164
-
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Fengli Xu*, Qianyue Hao*, Chenyang Shao*, Zefang Zong*, Yu Li*, Jingwei Wang, Yunke Zhang, Jingyi Wang, Xiaochong Lan, Jiahui Gong, Tianjian Ouyang, Fanjin Meng, Yuwei Yan, Qinglong Yang, Yiwen Song, Sijian Ren, Xinyuan Hu, Jie Feng, Chen Gao, Yong Li
Patterns (Cell), 2025
Cited by 563
-
AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery
Yu Li*, Chenyang Shao*, Xinyang Liu*, Ruotong Zhao, Peijie Liu, Hongyuan Su, Zhibin Chen, Qinglong Yang, Anjie Xu, Yi Fang, Qingbin Zeng, Tianxing Li, Jingbo Xu, Fengli Xu, Yong Li, Tie-Yan Liu
Technical Report
Cited by 21
⭐ 680 stars
-
Synergy-of-Thoughts: Eliciting Efficient Reasoning in Hybrid Language Models
Yu Shang*, Yu Li*, Fengli Xu, Yong Li
preprint
Cited by 25
-
AgentSociety Challenge: Designing LLM Agents for User Modeling and Recommendation on Web Platforms
Yuwei Yan, Yu Shang, Qingbin Zeng, Yu Li, Keyu Zhao, Zhiheng Zheng, Xuefei Ning, Tianji Wu, Shengen Yan, Yu Wang, Fengli Xu, Yong Li
ACM Web Conference (WWW), 2025
Cited by 20
-
OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
Chenyang Shao, Dehao Huang, Yu Li, Keyu Zhao, Weiquan Lin, Yining Zhang, Qingbin Zeng, Zhiyu Chen, Tianxing Li, Yifei Huang, Taozhong Wu, Xinyang Liu, Ruotong Zhao, Mengsheng Zhao, Jiaoyang Li, Xuhua Zhang, Yue Wang, Yuanyi Zhen, Fengli Xu, Yong Li, Tie-Yan Liu
Technical Report
Cited by 20
-
AI agent behavioral science
Lin Chen, Yunke Zhang, Jie Feng, Haoye Chai, Honglin Zhang, Bingbing Fan, Yibo Ma, Shiyuan Zhang, Nian Li, Tianhui Liu, Nicholas Sukiennik, Keyu Zhao, Yu Li, Ziyi Liu, Fengli Xu, Yong Li
Humanities and Social Sciences Communications (Nature), 2026
Cited by 31
* indicates co-first authors.
Internship Experience
-
Research Intern, Hunyuan
Tencent | June 2026 – Present
AI4AI related work.
-
Research Intern, 2012 Lab
Huawei Technologies Co., Ltd. | May 2025 – August 2025
Proposed a value-guided hierarchical agent search framework (AgentSwift) with hierarchical search space, value model training via balanced Bayesian sampling, and uncertainty-based hierarchical Monte Carlo tree search.
Results: Achieved 8.34% improvement over best human-designed agents across 7 benchmarks. Deployed in Huawei smartphone agent scenarios with 21.6% improvement on SPA-bench.
Selected Projects
-
AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery
Project Lead; one of five featured outcomes of Zhongguancun Academy at the Zhongguancun Forum.
⭐ 680 stars
Built an end-to-end automated research system that turns conference papers into executable code and performs open-ended, long-horizon optimization, covering resource preparation, experiment reproduction, and method optimization. Developed long-running execution with autonomous repair (environment initialization, execution tracking, deadlock prevention, and state rollback), along with skills for dependency resolution and path repair as the agent execution environment. Designed an algorithm iteration framework with an idea library, ideation workflow, and a supervisor agent that maintains a checklist to prevent cheating.
Results: Independently discovered 105 new SOTA models that surpassed the original conference paper baselines within one week; over 60% involved novel architectural designs, with an average performance gain of nearly 10%.
Education
-
M.S. in Electronic and Communication Engineering
Tsinghua University | 2024.9 – 2027.6 (Expected)
GPA: 3.95/4.0 (Rank 3/73)
-
B.E. in Information and Communication Engineering
Harbin Institute of Technology (Shenzhen) | 2020.9 – 2024.6
GPA: 4.0/4.0 (Rank 1/227)
Honors and Awards
National Scholarship, 2021, 2023
First-Class Academic Scholarship, 2021, 2022, 2023
Outstanding Student Model (优秀学生标兵)
Outstanding Youth League Member Model (优秀团员标兵)
Outstanding Undergraduate Graduate (本科优秀毕业生)
First Prize, National Undergraduate Mathematical Contest, 2022
Second Prize, National Undergraduate Mathematical Modeling Contest, 2022
Second Prize, National Undergraduate Electronic Design Contest, 2021
Skills
- Frameworks: vllm, Megatron-LM, verl, DeepSpeed
- Algorithms: DPO, PPO, GRPO, LoRA, SFT
Academic Services
Reviewer of NIPS, ICLR, ICML.
© 2026 Yu Li