https://doi.org/10.1007/s41781-026-00167-6
Research
Predicting Job Turnaround Time in Large-Scale Distributed Computing Environments with Graph Neural Networks
1
Brookhaven National Laboratory, Upton, NY, USA
2
Oak Ridge National Laboratory, Oak Ridge, TN, USA
3
University of Pittsburgh, Pittsburgh, PA, USA
4
Carnegie Mellon University, Pittsburgh, PA, USA
5
University of Massachusetts, Amherst, MA, USA
6
SLAC National Accelerator Laboratory, Menlo Park, CA, USA
a
This email address is being protected from spambots. You need JavaScript enabled to view it.
Received:
22
May
2026
Accepted:
29
May
2026
Published online:
27
June
2026
Abstract
Large-scale research infrastructures increasingly depend on distributed computing platforms to deliver timely data processing and analysis to their scientific communities. In such environments, job turnaround time is a key operational quantity because it affects user-perceived latency, workflow planning, and the efficient use of heterogeneous computing resources across facilities processing millions of jobs per week. Predicting turnaround time is difficult because it depends not only on intrinsic job characteristics but also on dynamic infrastructure conditions such as queue pressure, resource availability, brokerage decisions, and site-specific operating behavior. We study this problem in the Production and Distributed Analysis (PanDA) workload management system, which supports data-intensive scientific computing across grid, cloud, and high-performance computing resources and is used by large scientific collaborations including ATLAS at the Large Hadron Collider and ePIC at the Electron-Ion Collider. We present PanDA-GNN, a graph neural network for job-level turnaround-time prediction that represents jobs together with their local execution context. Using data from five of the most active PanDA computing sites, with approximately 2.1 million training jobs and separate validation and test sets of approximately 0.46 million and 0.40 million jobs, respectively, we show that graph-based context improves predictive performance over strong non-graph baselines. On the held-out test set, PanDA-GNN achieved R2 = 0.94, MAE = 84.72 minutes, MedAE = 39.62 minutes, RMSE = 163.59 minutes, and MAPE = 20.39%. These results show that graph-based workload modeling is a practical approach for turnaround-time prediction in large-scale research computing infrastructures and, more broadly, illustrate how data-driven methods can support informed scheduling, provisioning, and latency-aware workflow management across shared distributed research computing infrastructures.
Key words: Research infrastructures / Distributed scientific computing / Workload management / Job turnaround time prediction / Job scheduling / Graph neural networks / PanDA
© The Author(s) 2026
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.

