Shervin Minaee, Tomáš Mikolov, Narjes Nikzad, M. Chenaghlu, R. Socher, Xavier Amatriain, Jianfeng Gao
A comprehensive survey paper that summarizes major LLM families, techniques, datasets, and evaluation methods, gaining attention since ChatGPT.
The rapid development of LLMs has led to diverse model types, training methods, and evaluation criteria, making it difficult to grasp the overall picture. Researchers and practitioners need a systematic overview and comparative analysis.
The paper analyzes the characteristics, contributions, and limitations of three major LLM families: GPT, LLaMA, and PaLM. It also overviews LLM construction and augmentation techniques (fine-tuning, prompt engineering, RAG, etc.), surveys datasets for training, fine-tuning, and evaluation, and compares the performance of several popular LLMs on representative benchmarks.
Provides a comprehensive overview of the current state of LLM research, serving as a useful reference for both beginners and experts. It clarifies the strengths and weaknesses of each model, systematizes the status of datasets and evaluation methods, and suggests future research directions (efficiency, reliability, ethics, etc.), guiding subsequent studies.