Jiawen Tao

Jiawen Tao 陶佳文

Master's student at Peking University

Intern at Tencent Hunyuan

Advised by Prof. Tong Yang

About

I am a master's student at Peking University, advised by Prof. Tong Yang. I grew up in Changzhou, Jiangsu, and received my B.S. in Information and Computing Science from Peking University in 2025, with a dual degree in Economics from the National School of Development.

My research focuses on large language models, especially data synthesis for pre-/mid-training and LLM evaluation. I am currently interning at Tencent Hunyuan and previously interned at Moonshot AI (Kimi) and Tongdeng Asset Management (quantitative research).

Outside of research, I play Go and the guitar, and I am usually up for badminton, table tennis, or billiards with friends. I also love traveling — a few days somewhere new is still the surest way for me to reset.

Feel free to reach out if you would like to discuss research or explore potential collaborations.

Education

Publications

  1. Jiawen Tao*, Miao Peng*, Yaoming Li, Xiaokun Yuan, Mengzhou Wu, Wenhan Yu, Guoan Wang, Nuo Chen, Tong Yang, Maxm Pan. Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training. arXiv:2607.28109, 2026. [pdf]
  2. Eileen Ye*, Jiawen Tao*, Yaoming Li, Chenxu Liu, Wenhan Yu, Yaxin Fan, Xiaokun Yuan, Mengzhou Wu, Yanbing Jiang, Maxm Pan. Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding. arXiv:2607.29196, 2026. [pdf]
  3. Zihang Li, Wenjun Liu, Yikun Zong, Jiawen Tao, Siying Dai, Songcheng Ren, Zirui Liu, Yuhang Wang, Yanbing Jiang, Tong Yang. Bridge-RAG: An Abstract Bridge Tree Based Retrieval Augmented Generation Algorithm. arXiv:2603.26668, 2026. [pdf]
  4. Kimi Team. Kimi K2.5: Visual Agentic Intelligence. arXiv:2602.02276, 2026. [pdf]
  5. Zhewen Tan, Wenhan Yu, Jianfeng Si, Tongxin Liu, Kaiqi Guan, Huiyan Jin, Jiawen Tao, Xiaokun Yuan, Duohe Ma, Xiangzheng Zhang, Tong Yang, Lin Sun. TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment. ACL 2026. [arXiv] [ACL]
  6. Kimi Team. Kimi Linear: An Expressive, Efficient Attention Architecture. arXiv:2510.26692, 2025. [pdf]
  7. Kimi Team. Kimi K2: Open Agentic Intelligence. arXiv:2507.20534, 2025. [pdf]

* Equal contribution.

All papers listed above are publicly available on arXiv / Google Scholar.

Experience

Honors