My research focuses on large language models (LLMs) for astronomy and lunar exploration, with particular interests in multimodal perception, knowledge-augmented generation, and efficient reasoning. My current research centers on developing foundation models and intelligent systems for lunar exploration, with an emphasis on applying advances in artificial intelligence to practical challenges in astronomy and space exploration. I conduct astronomical research under the supervision of Prof. Nan Li at the National Astronomical Observatories (NAOC) and Associate Prof. Ning An at the University of Chinese Academy of Sciences (UCAS).
I am currently a research intern at The Future Laboratory, Tsinghua University, where I work on LLMs and human-computer interaction for lunar exploration. Previously, I interned at the National Astronomical Observatories (NAOC), CAS and the Department of Earth and Space Sciences, SUSTech, where I gained hands-on experience in cutting-edge astronomical and space science projects.

Xin-Yu Xiao, Zhixian He, Shiqi Wang, Ye Tian, Qianchen Xia
Empirical Methods in Natural Language Processing (EMNLP) Industry Track 2026 Accepted
Lunar-R1 is an 8B reasoning model designed for lunar exploration tasks. It introduces Latent Difficulty Perception to estimate task difficulty and adapt the length of its reasoning process. This mechanism concentrates computation on difficult problems while avoiding unnecessary tokens on simpler ones. Experiments show a 38.9% reduction in token usage together with improved accuracy.
Xin-Yu Xiao, Zhixian He, Shiqi Wang, Ye Tian, Qianchen Xia
Empirical Methods in Natural Language Processing (EMNLP) Industry Track 2026 Accepted
Lunar-R1 is an 8B reasoning model designed for lunar exploration tasks. It introduces Latent Difficulty Perception to estimate task difficulty and adapt the length of its reasoning process. This mechanism concentrates computation on difficult problems while avoiding unnecessary tokens on simpler ones. Experiments show a 38.9% reduction in token usage together with improved accuracy.

Xin-Yu Xiao, Ye Tian, Erwei Yin, Zhixian He, Shiqi Wang, Yalei Liu, Qianchen Xia
Annual Meeting of the Association for Computational Linguistics (ACL) 2026 Published
Lunar-Bench is a 3,000-task benchmark for evaluating task-oriented reasoning and decision-making in lunar exploration scenarios. Each task tests whether a model can reason about a lunar situation and produce an appropriate decision. Environmental Scenario Indicators measure safety, efficiency, integrity, and alignment. The benchmark provides a structured basis for comparing the reliability and practical usefulness of lunar reasoning systems.
Xin-Yu Xiao, Ye Tian, Erwei Yin, Zhixian He, Shiqi Wang, Yalei Liu, Qianchen Xia
Annual Meeting of the Association for Computational Linguistics (ACL) 2026 Published
Lunar-Bench is a 3,000-task benchmark for evaluating task-oriented reasoning and decision-making in lunar exploration scenarios. Each task tests whether a model can reason about a lunar situation and produce an appropriate decision. Environmental Scenario Indicators measure safety, efficiency, integrity, and alignment. The benchmark provides a structured basis for comparing the reliability and practical usefulness of lunar reasoning systems.

Xin-Yu Xiao, Yalei Liu, Xiangyu Liu, Zengrui Li, Erwei Yin, Qianchen Xia
Annual Meeting of the Association for Computational Linguistics (ACL) 2025 Published
Lunar Twins introduces domain-specific large language models and data for lunar exploration. The system includes the Chang'e and Yutu models, together with a collaborative multi-agent workflow for generating and solving lunar tasks. It also establishes a specialized lunar dataset integrating information from Chang'e missions. Experiments show that the resulting models outperform comparable general-purpose models on lunar-domain tasks.
Xin-Yu Xiao, Yalei Liu, Xiangyu Liu, Zengrui Li, Erwei Yin, Qianchen Xia
Annual Meeting of the Association for Computational Linguistics (ACL) 2025 Published
Lunar Twins introduces domain-specific large language models and data for lunar exploration. The system includes the Chang'e and Yutu models, together with a collaborative multi-agent workflow for generating and solving lunar tasks. It also establishes a specialized lunar dataset integrating information from Chang'e missions. Experiments show that the resulting models outperform comparable general-purpose models on lunar-domain tasks.