Xinran Zhao

I am a fourth-year PhD candidate at CMU LTI, advised by Prof. Sherry Wu. My PhD is supported by the Amazon AI PhD Fellowship.

I spent two wonderful summers (2024 and 2026) at Google DeepMind as a student researcher, working on LLM planning, long-horizon code generation, and continual learning with Dr. Azade Nova and Dr. Hanie Sedghi.

I had a great summer 2025 at Ai2 working on deep research agents with Varsha, Aakanksha, Joseph, Jay, and Jena.

I did my MSCS at Stanford University, where I was fortunate to work with Prof. Christopher Manning.

I did my bachelors at HKUST. At HKUST and during my exchange at Cornell, I was fortunate to work with Prof. Yangqiu Song, Prof. Dit-Yan Yeung, and Prof. Claire Cardie.

I always feel lucky to meet great mentors along my journey, including but not limited to, Dr. Shikhar Murty, Dr. Hongming Zhang, and Dr. Esin Durmus.

I love to discuss old and new ideas related to or beyond NLP, ML, and LLMs. Do email me to start a chat.

Email  /  CV  /  Linkedin  /  GitHub  /  Semantic Scholar  /  Google Scholar

profile photo
Research

I am interested in working on meaningful research on natural language processing (NLP) and language model agents. My goal is to build efficient, effective, and faithful systems that work with people on complex real-world tasks.
Recently, I focus on information seeking and planning in long-horizon scenarios.

Some interesting directions I am thinking of:
(1) How agents can learn, maintain memory, and evolve over complex long-horizon tasks;
(2) How models can stay aligned with human motives, goals, and intents, especially under high automation, e.g., by studying multi-turn behaviors and how multi-modal inputs shape model representation spaces;
(3) How agents can proactively help humans seek information, learn, and stay educated amid massive artificial artifacts.

2026
1. Improving Attributed Long-form Question Answering with Intent Awareness
Xinran Zhao, Aakanksha Naik, Jay DeYoung, Joseph Chee Chang, Jena D. Hwang, Tongshuang Wu, and Varsha Kishore
In Proceedings of ICLR 2026.
2. Revela: Dense Retriever Learning via Language Modeling
Fengyu Cai, Tong Chen, Xinran Zhao, Sihao Chen, Hongming Zhang, Tongshuang Wu, Iryna Gurevych, and Heinz Koeppl
In Proceedings of ICLR 2026 (Oral).
Best Paper Award at FrontierIR @ AAAI 2026.
3. Strategic Planning and Rationalizing on Trees Make LLMs Better Debaters
Danqing Wang*, Zhuorui Ye*, Xinran Zhao*, Fei Fang, and Lei Li (*: equal contribution)
In Proceedings of ICLR 2026.
4. DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Rulin Shao, Akari Asai, Shannon Zejiang Shen, Hamish Ivison, Varsha Kishore, Jingming Zhuo, Xinran Zhao, Molly Park, Samuel G. Finlayson, David Sontag, Tyler Murray, Sewon Min, Pradeep Dasigi, Luca Soldaini, Faeze Brahman, Wen-tau Yih, Tongshuang Wu, Luke Zettlemoyer, Yoon Kim, Hannaneh Hajishirzi, and Pang Wei Koh
In Proceedings of ICML 2026 (Spotlight).
5. Reading the Tea Leaves, Again: Can Large Language Models Predict Language Model Training Runs?
Xinran Zhao, Xiang Li, Jinchuan Tian, Zichun Yu, and Tongshuang Wu
In Submission.
6. Knowing When to Stop: The Evidence–Action Gap in Coding Agents
Tianyu Chen*, Xinran Zhao*, Vijay Viswanathan, Shoubin Yu, Mingyuan Zhou, and Tongshuang Wu (*: equal contribution)
In Submission.
7. Pivot-ICL: Adaptive Exemplar Selection for In-Context Learning
Xinran Zhao, Hanie Sedghi, Azade Nova, and Tongshuang Wu
In Submission.
8. The Ramon Llull's Thinking Machine for Automated Ideation
Xinran Zhao, Boyuan Zheng, Chenglei Si, Haofei Yu, Ken Liu, Runlong Zhou, Ruochen Li, Tong Chen, Xiang Li, Yiming Zhang, and Tongshuang Wu
In Submission. Presented at LM4Sci @ COLM 2025.
9. Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation
Chenyang Yang, Xinran Zhao, Tongshuang Wu, and Christian Kästner
In Submission (on arXiv).
10. ReSHAPE: User-Guided Revision of Planning Specifications
Yilin Zhang, Xinran Zhao, and Tongshuang Wu
In Submission.
11. Adaptive-GEPA: Adapting Meta-Harnesses to Heterogeneous Requests
Tianyu Chen, Yasi Zhang, Ruiyi Wang, Xinran Zhao, Taoran Li, and Mingyuan Zhou
In Submission.
12. RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems
Yixiao Zeng, Tianyu Cao, Danqing Wang, Xinran Zhao, Zimeng Qiu, Morteza Ziyadi, Tongshuang Wu, and Lei Li
In Submission (on arXiv).
13. Efficient Test-Time Adaptation through Human-AI Interaction
Zora Zhiruo Wang, Apurva Gandhi, Rulin Shao, Aspen Chen, Jonas Mueller, Zhiqi Liang, Jett Chen, Michael Ryan, Qianou Ma, Luxi He, Zhoujun Cheng, Andre He, Seungone Kim, Jiayi Geng, Mingqian Zheng, Weiwei Sun, Zheyuan Zhang, Xinran Zhao, Yike Wang, Abe Hou, Liwei Jiang, Pang Wei Koh, Diyi Yang, Graham Neubig, and Daniel Fried
arXiv preprint, 2026.
peer-reviewed
1. Improving Large Language Model Planning with Action Sequence Similarity
Xinran Zhao, Hanie Sedghi, Bernd Bohnet, Dale Schuurmans, and Azade Nova
In Proceedings of ICLR 2025.
2. MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers
Jushaan Singh Kalra*, Xinran Zhao*, To Eun Kim, Fengyu Cai, Fernando Diaz, and Tongshuang Wu (*: equal contribution)
In Proceedings of EMNLP 2025.
3. cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree
Yilin Zhang, Xinran Zhao, Zora Zhiruo Wang, Chenyang Yang, Jiayi Wei, and Tongshuang Wu
In Findings of EMNLP 2025. [code] 
4. SPHERE: An Evaluation Card for Human-AI Systems
Qianou Ma*, Dora Zhao*, Xinran Zhao, Chenglei Si, Chenyang Yang, Ryan Louie, Ehud Reiter, Diyi Yang+, and Tongshuang Wu+
In Findings of ACL 2025. [website] 
5. Dense X Retrieval: What Retrieval Granularity Should We Use?
Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu, Kaixin Ma, Xinran Zhao, Hongming Zhang, Dong Yu
In Proceedings of EMNLP 2024
6. MixGR: Enhancing Retriever Generalization for Scientific Domain through Complementary Granularity
Fengyu Cai, Xinran Zhao, Tong Chen, Sihao Chen, Hongming Zhang, Iryna Gurevych, Heinz Koeppl
In Proceedings of EMNLP 2024
7. HiMemFormer: Hierarchical Memory-Aware Transformer for Multi-Agent Action Anticipation
Zirui Wang, Xinran Zhao, Simon Stepputtis, Woojun Kim, Tongshuang Wu, Katia P. Sycara, Yaqi Xie
In Video-Language Models @ NeurIPS 2024
8. Beyond Relevance: Evaluate and Improve Retrievers on Perspective Awareness
Xinran Zhao, Tong Chen, Sihao Chen, Hongming Zhang, Tongshuang Wu
In Proceedings of COLM 2024
9. "Merge Conflicts!" Exploring the Impacts of External Distractors to Parametric Knowledge Graphs
Cheng Qian, Xinran Zhao, Sherry Tongshuang Wu
In Proceedings of COLM 2024
10. Fact-and-Reflection (FaR) Improves Confidence Calibration of Large Language Models
Xinran Zhao, Hongming Zhang, Xiaoman Pan, Wenlin Yao, Dong Yu, Tongshuang Wu, Jianshu Chen
In Findings of ACL 2024
11. GeoHard: Towards Measuring Class-wise Hardness through Modelling Class Semantics
Fengyu Cai, Xinran Zhao, Hongming Zhang, Iryna Gurevych, Heinz Koeppl
In Findings of ACL 2024
12. Thrust: Adaptively Propels Large Language Models with External Knowledge
Xinran Zhao, Hongming Zhang, Xiaoman Pan, Wenlin Yao, Dong Yu, and Jianshu Chen
In Proceedings of NeurIPS 2023
13. Video State-changing Object Segmentation
Xiang Li, Jiangwei Yu, Xinran Zhao, Hongming Zhang, and Yu-Xiong Wang
In Proceedings of ICCV 2023.
14. Towards Reference-free Text Simplification Evaluation with a BERT Siamese Network Architecture
Xinran Zhao, Esin Durmus, and Dit-Yan Yeung
In Findings of ACL 2023.
15. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (BIG-bench)
Aarohi Srivastava et al (100+ authors), Xinran Zhao
In TMLR
16. On Measuring the Intrinsic Few-Shot Hardness of Datasets
Xinran Zhao*, Shikhar Murty*, and Christopher Manning (*: equal contribution)
In Proceedings of EMNLP 2022.
17. Weakly Supervised Text Classification using Supervision Signals from a Language Model
Ziqian Zeng, Weimin Ni, Tianqing Fang, Xiang Li, Xinran Zhao, and Yangqiu Song
In Findings of NAACL 2022. [details] 
18. PCR4ALL: A Comprehensive Evaluation Benchmark for Pronoun Coreference Resolution
Xinran Zhao, Hongming Zhang, and Yangqiu Song
In Proceedings of LREC 2022. [paper] /  [details] 
19. Leveraging Topic Relatedness for Argument Persuasion
Xinran Zhao, Esin Durmus, Hongming Zhang, and Claire Cardie
In Findings of ACL 2021. [details] 
20. Probing Toxic Content in Large Pre-Trained Language Models
Nedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song, and Dit-Yan Yeung
In Proceedings of ACL 2021. [details] 
21. WinoWhy: A Deep Diagnose of Essential Commonsense Knowledge for Answering Winograd Schema Challenge
Hongming Zhang*, Xinran Zhao*, and Yangqiu Song (*: equal contribution)
In Proceedings of ACL 2020.
[paper] /  [code & data] /  [details] 
22. A Brief Survey and Comparative Study of Recent Development of Pronoun Coreference Resolution
Hongming Zhang, Xinran Zhao, and Yangqiu Song
In CRAC @ EMNLP 2021.
[paper] /  [code & data] /  [details] 
23. Learning Contextual Causality Between Daily Events from Time-consecutive Images
Hongming Zhang, Yintong Huo, Xinran Zhao, Yangqiu Song, and Dan Roth
In Causality in Vision @ CVPR 2021. [details] 
24. The Effects of Fear of Missing Out on Social Media Posting Preferences
Yue Xi, Jiale Huo, Xinran Zhao, Yushi Jiang, Qiang Yang
In European Journal of Marketing.
Academic Events
  • Invited to join a panel discussion for the Global Embodied Ecosystem Night at IROS 2026, Fall 2026.
  • Invited to give a talk (topic: Continual Learning for Coding Agents) at Google DeepMind, Summer 2026.
  • Invited to give a talk (topic: Improving Large Language Model Planning with Action Sequence Similarity) for H2 Group at UW, Summer 2025.
  • Invited to give a talk (topic: Towards Building Retrieval Systems in Complex Realistic Scenarios) for Samaya AI, Summer 2024.
  • Invited to serve as an area chair for ACL and EMNLP (2024–present) and NAACL (2025–present).
  • Invited to serve as a reviewer for ACL, EMNLP, NAACL, EACL, AACL, COLM, NeurIPS, ICLR, AAAI, and ECCV.
  • Invited to give a talk (topic: On Task Difficulty of Few-shot Learning) for ICML 2022 Commonsense Tutorial, Summer 2022.
  • Invited to give a talk (topic: NLP as a Tool for Scientific Discovery) for the Department of Marketing @ National University of Singapore, Spring 2022.
Projects
1. Can Pre-trained Language Models Understand Definitions?
Advised by Christopher Manning
CS 224N Final Project.
2. Seek to Embed ASER: A Large-scale Eventuality Knowledge Graph
Advised by Xin Liu and Prof. Yangqiu Song
An acknowledged contributor (paper was accepted by WWW 2020). [paper] /  [details] 
3. An Online Learning Platform with NLP Supported Teaching Assistant
Advised by Prof. Dit-Yan Yeung.
Final Year Project. [details] 
4. A Comparative Study on the Sentimental Characteristics of Chinese and Western Tourists
Advised by Xi Yue and Prof. Yushi Jiang
In Annual Conference of Journal of Marketing Science 2019, China.
[paper (Chinese)] /  [paper (English)] /  [details] 
5. One-shot Pokemon Classification
Advised by Prof. Qifeng Chen.
Coursework. [report]  [details] 
Miscellaneous

Besides researching, I am also keen on other activities:

  • I am an enthusiastic bodybuilder who can do 140 kg deadlift and 140 kg squat (sets).

  • I was a member of the HKUST Mandarin Debate team, where we won the First Prize for the Bay Area Debate Championship and Top 8 for the sixth and seventh International Mandarin Debate Invitational in Singapore.

  • I am a traditional Chinese poem writer who has published a few poems in newspapers.

  • I was an active member of the HKUST Archery Club.


  • This page is forked from Jon Barron's home page.