The goal of the AuroraGPT Inference Team is to build, maintain, and provide scalable inference services and APIs to enable the AuroraGPT community to access various state-of-the-art LLMs. As modern research increasingly hinges on advanced language models, traditional commercial cloud platforms introduce severe drawbacks regarding operating costs, data isolation, and computational ceilings. To eliminate these barriers, our team built a cluster-agnostic framework that converts institutional supercomputers into an on-premise AI generation engine.
The Inference team had developed an inference framework and service that runs parallel inference jobs on multiple ALCF systems: Minerva (Nvidia B200), Sophia (Nvidia A100), Metis (SambaNova), and soon on Tara (Nvidia GH200). Leveraging Globus Auth and Globus Compute, it provides cloud-like access to diverse AI models via an OpenAI-compliant API and also a web interface. The cluster-agnostic API allows requests to be distributed across federated clusters. The framework supports multiple inference backends (e.g., vLLM), auto-scales resources, maintains “hot” nodes for low-latency execution, and offers both high-throughput batch and interactive modes. It addresses the growing demand for private, secure, and scalable AI inference in scientific workflows, allowing researchers to generate billions of tokens daily on-premise without relying on commercial cloud infrastructure.
Developing and deploying the service across these platforms has given us experience with AI inference on different types of AI hardware and has provided infrastructure for the broader community. A paper describing the inference service was published at the AI4S workshop at SC25.
ALCF continues to actively develop and maintain this service. See https://docs.alcf.anl.gov/services/inference-endpoints/ for the latest on how to access it and the models currently served.
Publication
Aditya Tanikanti, Benoit Cote, Yanfei Guo, Le Chen, Nickolaus Saint, Ryan Chard, Ken Raffenetti, Rajeev Thakur, Thomas Uram, Ian Foster, Michael E. Papka, Venkatram Vishwanath, “FIRST: Federated Inference Resource Scheduling Toolkit for Scientific AI Model Access,” Proceedings of the 6th Workshop on Artificial Intelligence and Machine Learning for Scientific Applications (AI4S) at SC25, November 2025.
Figure 1: Architecture of the Globus-Compute-based inference service that we have developed for serving LLMs on the Sophia GPU cluster at ALCF. Authenticated and authorized users can make requests to run supported LLMs; requests are routed to compute resources on Sophia nodes.
Team Members
Rajeev Thakur serves as the primary Group Leader for the Inference Team. For further information about our inference framework, API access documentation, custom model deployments, or potential research collaborations, please contact him directly.