2Department of Industrial Engineering, Faculty of Engineering, Gazi University, Ankara, 06560, Türkiye
3AltexSoft Inc., Foster City, California, 94404, USA
Abstract
In academic research, systematic literature reviews play a key role in bringing together what is known about a given topic, but carrying out such reviews requires considerable time and effort, especially in fast-moving areas like logistics and transportation. Even the early task of locating and filtering relevant publications from the large body of available work can be slow and labor-intensive. Selection and interpretation of articles also remain prone to personal judgment, which threatens the reproducibility that scientific inquiry depends on. Against this backdrop, this study examines how advanced large language models can be incorporated into the systematic literature review process. To do so, we asked GPT-3.5 to screen research articles using a set of predefined criteria, testing it for both speed and accuracy. The input consisted of titles and abstracts from published work on the Electric Vehicle Routing Problem, a topic of growing importance in sustainable logistics, and each article was run through a carefully prepared prompt. GPT-3.5 scored 96% accuracy and finished the entire task in less than three minutes, whereas a PhD student working on the same task reached 84% accuracy over three working days. These outcomes suggest that large language models can be useful for automating parts or the whole of the literature review process. The workload can be reduced without sacrificing, and in fact even improving, accuracy. The study also provides a starting point for broader use of such models in academic research, illustrating how they may improve the speed and reproducibility of review work across different disciplines.
