AI4AIR: A Comprehensive Survey on Large Language Models for AI Research
Preprint, under review, 2026
TL;DR This survey introduces AI4AIR, a comprehensive review of large language models as pivotal components within machine learning research pipelines, covering data engineering, model design and optimization, model evaluation, and closed-loop AI research automation.
Abstract Language-mediated automation is beginning to complement the human-centered trial-and-error process in AI research. Among current AI tools, large language models (LLMs) have become a central interface for generation, knowledge synthesis, and reasoning in research workflows. While LLMs are now widely used to support general scientific workflows such as literature review and scientific writing, their specific roles and deeper contributions to the core lifecycle of AI research itself remain insufficiently explored in a systematic manner. To bridge this gap, this survey introduces AI4AIR (short for AI for AI Research), which comprehensively reviews LLMs as pivotal components within machine learning research pipelines. We construct a structured two-dimensional taxonomy. One axis spans major research domains including natural language processing, computer vision, data mining, and general machine learning. The other follows the research pipeline stages, encompassing data engineering, model design and optimization, model evaluation, and the cross-stage closed-loop automation that connects them. Within this framework, we identify five recurring roles of LLMs, namely annotator, synthesizer, optimizer, evaluator, and orchestrator, through which LLMs contribute to AI research workflows. We further discuss bottlenecks such as contamination, hallucination, and reliability under feedback-driven use, and outline future directions for improving both the efficiency and the reliability of AI research and discovery.
Project page: Awesome LLMs for AI Research.
Recommended citation: Ao, Xiang, Junhong Lian, Hanyang Li, Siyi Wang, Yiran Qiao, Yi Qiao, Jiaqi Xu, Qing He, and Xueqi Cheng. "AI4AIR: A Comprehensive Survey on Large Language Models for AI Research." Preprint, under review. 2026.
Show citation
@article{ao2026ai4air,
title = {AI4AIR: A Comprehensive Survey on Large Language Models for AI Research},
author = {Xiang Ao and Junhong Lian and Hanyang Li and Siyi Wang and Yiran Qiao and Yi Qiao and Jiaqi Xu and Qing He and Xueqi Cheng},
year = {2026},
note = {Preprint, under review},
url = { https://ict-find-lab.github.io/Awesome-LLMs-for-AI-Research/ }
}
Leave a Comment