Yite Wang
Research Scientist at Snowflake AI Research
Bellevue, WA
I’m a Research Scientist at Snowflake AI Research, where I train reasoning models and LLM agents, including data science agents and SWE agents. Before that, I was a Research Scientist on ByteDance’s Seed-Foundation-Code team.
I earned my Ph.D. from the University of Illinois at Urbana-Champaign under Prof. Ruoyu Sun, with co-advising from Prof. Naira Hovakimyan. In my first year I collaborated with Prof. Justin Sirignano on deep-learning for computational fluid dynamics (CFD), and during my master’s studies I worked with Prof. Kyle Smith on numerical simulation for energy storage systems.
Research interests: Reasoning Models; SWE Agents; Data Science Agents; Efficient Deep Learning; Natural Language Processing; Computer Vision; Numerical Methods.
news
| Jul 09, 2026 | We have three papers accepted to COLM 2026. Papers coming soon! |
|---|---|
| Apr 30, 2026 | Our paper Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning is accepted to ICML 2026. [arXiv] [code] |
| Feb 27, 2026 | Our paper DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science is accepted to ICLR 2026. [arXiv] [code] |
| Apr 01, 2025 | I joined Snowflake AI Research, working on reasoning models and LLM agents. |
| May 01, 2024 | I joined ByteDance Seed-Foundation-Code as a Research Scientist working on LLM for code. |
blogs
- Arctic RL: Open-Source Backend Open-source infrastructure for training and evaluating reinforcement-learning agents.
- Hybrid Deep Research Benchmark A benchmark for evaluating long-horizon research agents across web and enterprise data.
- ArcticSwarm: Hybrid Deep Research A hybrid agent system for deep research over private and public information sources.
- ArcticSwarm: Multi-Agent System Architecture Architecture notes for coordinating specialized agents in enterprise research workflows.
- Enterprise Text-to-SQL with Arctic R2 Reasoning models for production-grade text-to-SQL over enterprise data.
- DARE-bench: LLM Data Science Workflows A benchmark for measuring modeling and instruction fidelity in data science workflows.
- Agent World Model for Agentic Reinforcement Learning Synthetic environments for training agentic reinforcement-learning systems.
selected publications
* denotes equal contribution. † denotes project lead.
- Bridging Databases and Documents: Data-Algorithm Co-Design for Hybrid Question AnsweringIn COLM, 2026
- MidTool: Mid-training Data Synthesis for Agentic Tool UseIn COLM, 2026
- The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML WorkflowsIn COLM, 2026
-