Yiqiao Jin 🍀
Yiqiao Jin

Ph.D. Candidate in Computer Science

Google Scholar 2084Citations 19h-index 28i10-index

About Me

Hello👋! I am Yiqiao Jin (靳轶乔, Ahren), a Computer Science Ph.D. candidate at Georgia Tech, advised by Prof. Srijan Kumar.

My research sits at the intersection of multimodal foundation models, intelligent agents, and spatial/physical intelligence: building general-purpose models that can perceive, reason, learn, and act in interactive virtual and physical environments. I work across four threads — vision-language and multimodal models (VLMs/MLLMs); world models, embodied AI, and spatial reasoning; LLM agents and multi-agent systems; and post-training, reinforcement learning, and self-improvement. My work has appeared at ACL, EMNLP, ICML, NeurIPS, ICLR, KDD, AAAI, CIKM, and The Web Conference, including several oral presentations.

Most recently at Waymo (Perception team), I developed a unified VLM for autonomous driving that learns dense metric scene geometry and sensor-grounded spatial reasoning from camera and sparse LiDAR (patent application pending), providing core perceptual capabilities for vision-language-action (VLA) systems. I am currently an Applied Scientist Intern at Amazon, building a unified agentic framework based on diffusion language models (dLMs) for task completion and world modeling. Earlier, I built agentic and multimodal systems at J.P. Morgan AI Research (SlideAgent, ACL'26), Visa Research (SARA, ACL'26), Adobe Research, and Microsoft Research Asia; before Georgia Tech, I worked with Prof. Yizhou Sun and Prof. Wei Wang at the UCLA Scalable Analytics Institute (ScAi).

Selected honors include the MLCommons ML and Systems Rising Stars (2026), Best Paper Award at Good-Data @ AAAI 2025 Workshop, Roblox Graduate Fellowship Finalist (2024), and the Microsoft Research “Star of Tomorrow” Award. My work on cross-lingual LLM evaluation has been featured by Scientific American, The World, and Georgia Tech News.

Download CV
Interests
  • Vision-Language & Multimodal Models (VLMs/MLLMs)
  • World Models, Embodied AI & Spatial Reasoning
  • LLM Agents & Multi-Agent Systems
  • Post-training, Reinforcement Learning & Self-Improvement
Education
  • Ph.D. in Computer Science

    Georgia Institute of Technology (Georgia Tech)

  • B.S. in Computer Science

    University of California, Los Angeles (UCLA)

Experience

  1. Amazon logo

    Applied Scientist Intern

    Amazon
    • Develop a unified agentic framework based on diffusion language models (dLMs) for task completion and world modeling.
    • Mentor: Anwesan Pal. Manager: Narayanan Sadagopan.
  2. Waymo logo

    Ph.D. Research Intern

    Waymo
    • Perception Team. Developed a unified VLM for autonomous driving perception (patent application pending) that learns dense metric scene geometry and sensor-grounded spatial reasoning from camera observations and sparse LiDAR.
    • Built core perceptual capabilities for vision-language-action (VLA) systems — dense depth, object-centric distance, driving relevance, depth ordering, and time-to-reach reasoning — through a shared geometry-language scene representation.
    • Mentors: Mayank Singal, Prasanna Krishnasamy. Managers: Ming Zou, Christian Lauterbach.
  3. J.P. Morgan AI Research logo

    Research Scientist Intern

    J.P. Morgan AI Research
    • Led SlideAgent, a hierarchical agentic framework for multi-page visual document understanding (ACL 2026 main).
    • Built MLLM-based agentic pipelines for parsing high-resolution slide decks and financial reports, integrating context compression and retrieval-augmented generation.
    • Mentors: Rachneet Kaur, Zhen Zeng. Managers: Sumitra Ganesh, Manuela Veloso.
  4. Visa Research logo

    Research Collaborator (Visa Research Grant)

    Visa Research
    • Led SARA, a selective and adaptive RAG framework with context compression (ACL 2026 main).
    • Developed a hybrid compression strategy combining fine-grained spans with compact semantic vectors, achieving consistent gains across Mistral, Llama, and Gemma.
    • Mentors: Vineeth Rakesh Mohan, Yingtong Dou, Menghai Pan. Manager: Mahashweta Das.
  5. Adobe Research logo

    Research Scientist Intern

    Adobe Research
    • Built ScreenLLM, a multimodal LLM for GUI understanding and action prediction (WebConf'25 MM4SG Workshop).
    • Designed a stateful screen schema that summarizes dynamic UI sessions as time-aware textual context, plus a key-frame extractor for significant UI transitions.
    • Mentors: Gang Wu, Yu Shen, Stefano Petrangeli. Managers: Saayan Mitra, Vishy Swaminathan.
  6. Microsoft Research Asia (Social Computing Group) logo

    Research Scientist Intern

    Microsoft Research Asia (Social Computing Group)
    • Mentors: Xiting Wang, Jindong Wang, Xing Xie.
    • Published papers across LLMs (ICML'24, ICML'23, AAAI'23), LLM agents (EMNLP'24, ICML'24), misinformation detection (KDD'22, AAAI'22), few-shot learning (ACL'24, AAAI'23), and explainable AI (AAAI'22).
    • Received Microsoft Research “Star of Tomorrow” Award (2021).
  7. Amazon (Fulfillment By Amazon) logo

    Software Engineer Intern

    Amazon (Fulfillment By Amazon)
    • Designed and implemented IAR Manual Analysis, a scalable workflow on AWS Step Functions and Lambda that automates aggregation of S3 and DynamoDB data for SageMaker training, handling >16k requests per summary stage.
    • Deployed the workflow across all AWS regions (EU/FE/NA) via CloudFormation; built DataCraft pipeline for ingesting DynamoDB tables into the Andes catalog.
  8. IBM China Development Laboratories logo

    Software Engineer Intern

    IBM China Development Laboratories
    • Built Compass DataRouter (Go + MongoDB) for the Compass project, reducing memory usage and accelerating data retrieval.
    • Improved the Compass monitoring dashboard with React.js.

Education

  1. Georgia Institute of Technology (Georgia Tech) logo

    Ph.D. in Computer Science

    Georgia Institute of Technology (Georgia Tech)
    • Advisor: Prof. Srijan Kumar.
    • Research on multimodal foundation models, world models and spatial reasoning, LLM agents and multi-agent systems, and post-training and self-improvement.
    • Publications include EMNLP'26, ACL'26, ICLR'26, ICWSM'26, EMNLP'25, ACL'25, WebConf'24, ACL'24, KDD'23, CIKM'24, AAAI'23.
    • GPA: 4.0/4.0.
  2. University of California, Los Angeles (UCLA) logo

    B.S. in Computer Science

    University of California, Los Angeles (UCLA)
    • Advisors: Prof. Yizhou Sun, Prof. Wei Wang.
    • Published 4 papers at top-tier ML and data mining venues (AAAI, KDD, Web Conference). Continuing collaborator on LLMs and graph neural networks (GNNs) (EMNLP'25, WebConf'23).
    • GPA: 3.82/4.0. Dean’s Honor List (5 times).
What’s new?
Featured Publications

Representative work across vision-language and multimodal models, world models and spatial reasoning, LLM agents, and post-training. The full publication list includes all venues.

Recent Publications

Browse all papers on the Publications page, or jump straight to Google Scholar.

(2026). MASCOT: Towards Multi-Agent Socio-Collaborative Companion Systems. ACL'26 TrustNLP Workshop.
(2026). Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations. ACL'26.
(2026). MedHalu: Hallucinations in Responses to Healthcare Queries by Large Language Models. ICWSM'26.
(2026). Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings. Preprint.
(2025). Protein Large Language Models: A Comprehensive Survey. EMNLP'25.
(2025). Empowering Interdisciplinary Insights with Dynamic Graph Embedding Trajectories. KDD'25 TGL Workshop.
(2025). Deconstructing The Ethics of Large Language Models from Long-standing Issues to New-emerging Dilemmas. AI and Ethics.
(2025). ProteinGPT: Multimodal LLM for Protein Property Prediction and Structure Understanding. ICML'25 FM4LS Workshop.
(2025). Topological Structure Learning Should Be A Research Priority for LLM-Based Multi-Agent Systems. Preprint.
(2025). UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models. AAAI'25 DAI Workshop.
(2024). RNA-GPT: Multimodal Generative System for RNA Sequence Understanding. NeurIPS'24 MLSB Workshop.
(2024). PrivacyMind: Large Language Models Can Be Contextual Privacy Protection Learners. EMNLP'24.
(2024). Towards Fair Graph Anomaly Detection: Problem, Benchmark Datasets, and Evaluation. CIKM'24.
(2024). Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare Queries. WWW'24.
(2023). Code Recommendation for Open Source Project Developers. WWW'23.
(2023). Predicting Information Pathways Across Online Communities. KDD'23.
(2022). Reinforcement Subgraph Reasoning for Fake News Detection. KDD'22.
(2022). Towards Fine-Grained Reasoning for Fake News Detection. AAAI'22.
Recent & Upcoming Talks
Skills
Research Areas
Vision-Language & Multimodal Models
World Models, Embodied AI & Spatial Reasoning
LLM Agents & Multi-Agent Systems
Post-training, RL & Self-Improvement
Hobbies
Rabbit
Photography
Hiking
Contact

📍 Location

Atlanta, GA (during the academic year) — currently in Santa Clara, CA (Amazon internship, Aug–Nov 2026)

📧 Email

yjin328[AT]gatech.edu

🏢 Office

756 W Peachtree St NW, Atlanta, GA 30308
CODA 13th Floor

🕒 Office Hours

Monday - Sunday 9:00 to 20:00

🌐 Connect with Me