Yite Wang

Research Scientist at Snowflake AI Research

avatar.png

Bellevue, WA

I’m a Research Scientist at Snowflake AI Research, where I train reasoning models and LLM agents, including data science agents and SWE agents. Before that, I was a Research Scientist on ByteDance’s Seed-Foundation-Code team.

I earned my Ph.D. from the University of Illinois at Urbana-Champaign under Prof. Ruoyu Sun, with co-advising from Prof. Naira Hovakimyan. In my first year I collaborated with Prof. Justin Sirignano on deep-learning for computational fluid dynamics (CFD), and during my master’s studies I worked with Prof. Kyle Smith on numerical simulation for energy storage systems.

Research interests: Reasoning Models; SWE Agents; Data Science Agents; Efficient Deep Learning; Natural Language Processing; Computer Vision; Numerical Methods.

news

Jul 09, 2026 We have three papers accepted to COLM 2026. Papers coming soon!
Apr 30, 2026 Our paper Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning is accepted to ICML 2026. [arXiv] [code]
Feb 27, 2026 Our paper DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science is accepted to ICLR 2026. [arXiv] [code]
Apr 01, 2025 I joined Snowflake AI Research, working on reasoning models and LLM agents.
May 01, 2024 I joined ByteDance Seed-Foundation-Code as a Research Scientist working on LLM for code.

blogs

selected publications

* denotes equal contribution. denotes project lead.

  1. Bridging Databases and Documents: Data-Algorithm Co-Design for Hybrid Question Answering
    Ruofan Wu, Boyi Liu, Fan Shu, Yite Wang, Zhewei Yao, Yuxiong He, and Feng Yan
    In COLM, 2026
    RL DataAgent RL TrainingBenchmark
  2. MidTool: Mid-training Data Synthesis for Agentic Tool Use
    Fengqing Jiang, Yite Wang, Boyi Liu, Zhaoyang Wang, Canwen Xu, Zhewei Yao, Radha Poovendran, and Yuxiong He
    In COLM, 2026
    Midtrain DataTool Use AgentAgent Training
  3. The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows
    Junbo Li, Boyi Liu, Canwen Xu, Yite Wang, Zhewei Yao, Zhangyang Wang, Qiang Liu, and Yuxiong He
    In COLM, 2026
    Auto Research
  4. Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
    Zhaoyang Wang, Canwen Xu, Boyi Liu, Yite Wang, Siwei Han, Zhewei Yao, Huaxiu Yao, and Yuxiong He
    In ICML, 2026
    RL DataAgent RL Training
  5. DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
    Fan Shu*, Yite Wang*, Ruofan Wu, Boyi Liu, Zhewei Yao, Yuxiong He, and Feng Yan
    In ICLR, 2026
    RL DataAgent RL TrainingBenchmark
  6. FullStack Bench: Evaluating LLMs as Full-Stack Coders
    ByteDance Seed-Foundation-Code Team
    Preprint, 2025
    Evaluation
  7. LEMON: Lossless Model Expansion
    Yite Wang, Jiahao Su, Hanlin Lu, Cong Xie, Tianyi Liu, Jianbo Yuan, Haibin Lin, Ruoyu Sun, and Hongxia Yang
    In ICLR, 2024
    Model Expansion
  8. Balanced Training for Sparse GANs
    Yite Wang*, Jing Wu*, Naira Hovakimyan, and Ruoyu Sun
    In NeurIPS, 2023
    Pruning
  9. NTK-SAP: Improving Neural Network Pruning by Aligning Training Dynamics
    Yite Wang, Dawei Li, and Ruoyu Sun
    In ICLR, 2023
    Pruning
  10. Global Convergence of MAML and Theory-Inspired Neural Architecture Search for Few-Shot Learning
    Haoxiang Wang*, Yite Wang*, Ruoyu Sun, and Bo Li
    In CVPR, 2022
    Neural Architecture SearchNTK