Hello👋! I am Yiqiao Jin (靳轶乔, Ahren), a Computer Science Ph.D. candidate at Georgia Tech, advised by Prof. Srijan Kumar.
My research sits at the intersection of multimodal foundation models, intelligent agents, and spatial/physical intelligence: building general-purpose models that can perceive, reason, learn, and act in interactive virtual and physical environments. I work across four threads — vision-language and multimodal models (VLMs/MLLMs); world models, embodied AI, and spatial reasoning; LLM agents and multi-agent systems; and post-training, reinforcement learning, and self-improvement. My work has appeared at ACL, EMNLP, ICML, NeurIPS, ICLR, KDD, AAAI, CIKM, and The Web Conference, including several oral presentations.
Most recently at Waymo (Perception team), I developed a unified VLM for autonomous driving that learns dense metric scene geometry and sensor-grounded spatial reasoning from camera and sparse LiDAR (patent application pending), providing core perceptual capabilities for vision-language-action (VLA) systems. I am currently an Applied Scientist Intern at Amazon, building a unified agentic framework based on diffusion language models (dLMs) for task completion and world modeling. Earlier, I built agentic and multimodal systems at J.P. Morgan AI Research (SlideAgent, ACL'26), Visa Research (SARA, ACL'26), Adobe Research, and Microsoft Research Asia; before Georgia Tech, I worked with Prof. Yizhou Sun and Prof. Wei Wang at the UCLA Scalable Analytics Institute (ScAi).
Selected honors include the MLCommons ML and Systems Rising Stars (2026), Best Paper Award at Good-Data @ AAAI 2025 Workshop, Roblox Graduate Fellowship Finalist (2024), and the Microsoft Research “Star of Tomorrow” Award. My work on cross-lingual LLM evaluation has been featured by Scientific American, The World, and Georgia Tech News.
Ph.D. in Computer Science
Georgia Institute of Technology (Georgia Tech)
B.S. in Computer Science
University of California, Los Angeles (UCLA)
Representative work across vision-language and multimodal models, world models and spatial reasoning, LLM agents, and post-training. The full publication list includes all venues.
Atlanta, GA (during the academic year) — currently in Santa Clara, CA (Amazon internship, Aug–Nov 2026)
yjin328[AT]gatech.edu
756 W Peachtree St NW, Atlanta, GA 30308
CODA 13th Floor
Monday - Sunday 9:00 to 20:00