Publications

Review publications from the AuroraGPT team.

Published

Latent Reasoning Guidance for Parallel Code Translation


Bitan, T., E. Kaplan, R. Bar-Yadin, L. Ghrayeb, L. Chen, S. Jhaveri, N. Hasabnis, and G. Oren, "Latent Reasoning Guidance for Parallel Code Translation," arXiv preprint arXiv:2606.05518, (August/2026). doi: https://doi.org/10.48550/arXiv.2606.05518

Tackling complex coding tasks often requires autonomous agents and iterative repair pipelines. These increasingly rely on large amounts of test-time computation, often spending many decoding and repair steps before discovering whether a program compiles, runs, or validates. Executable parallel-code translation is an effective setting for earlier guidance because success is behavioral rather than textual. However, most guidance methods act only after complete programs or textual traces are decoded. This motivates the question: can latent reasoning provide an earlier intervention point, before the model commits to code? We study a test-time latent guidance method for this setting that trains a smaller Process Reward Model (PRM) over continuous latent prefixes and uses it to select among alternate hidden-state trajectories before final code decoding, separately from but compatible with post-decoding optimization. On a 76-task ParaTrans benchmark evaluation, latent PRM guidance improves mean validation rate from 32.89\% with unguided latent reasoning to 42.1\%, outperforming fine-tuned and vanilla baselines in the same setting. These gains persist under the same three-attempt repair loop. These results provide bounded evidence that useful alternative latent continuations exist and that PRM-scored latent branch selection can improve executable outcomes in this setting without retraining the main generative model.

journal article
Published

SURE: data-efficient safety guardrailing via internal representations and uncertainty-weighted pseudo-labels


Li, H., J. Duan, C. Yuan, J. Diffenderfer, S. Madireddy, B. Kailkhura, and K. Xu, "SURE: data-efficient safety guardrailing via internal representations and uncertainty-weighted pseudo-labels," npj Artificial Intelligence, (July/2026), Nature Portfolio. doi: 10.1038/s44387-026-00141-y

Large language models (LLMs) can produce harmful content when prompted with unsafe inputs, yet inference-time safety interventions often degrade general-purpose performance through over-refusal. We propose SURE (Semi-supervised Uncertainty-weighted Representation-based External detection), a framework that trains a lightweight classifier on the LLM’s hidden state representations to detect harmful inputs entirely outside the inference pipeline, preserving the model’s original capabilities. To overcome annotation scarcity, SURE queries the target LLM to assess unlabeled inputs and weights each pseudo-label by the classifier’s confidence, the LLM’s uncertainty, measured as the perplexity of its single-token safety self-assessment, and their agreement. Experiments across three open-source LLMs (8B–70B parameters) show that SURE achieves average harmonic mean scores of 88–90% with only 80 labeled samples, surpassing both inference-time and representation-based baselines. Performance plateaus at 40 labeled examples, and cross-category generalization confirms over 90% accuracy on held-out categories, indicating that the learned representations capture general safety features.

conference paper
Published

Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation


Chen, L., N. Xu, W. Chen, B. Lei, P-H. Lin, D. Zhou, R. Thakur, C. Ding, A. Jannesari, and C. Liao, "Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation," Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), (July/2026), Association for Computational Linguistics (ACL), pp. 33781-33803. doi: https://doi.org/10.18653/v1/2026.acl-long.1557

Large language models (LLMs) have shown remarkable capabilities in code translation, yet their performance deteriorates in low-resource programming domains such as Fortran and emerging frameworks like CUDA, where high-quality parallel data are scarce. We present an automated dataset generation pipeline featuring a dual-LLM Questioner–Solver design that incorporates external knowledge from compilers and runtime feedback. Beyond traditional source–target code pair datasets, our approach additionally generates (1) verified translations with unit tests for assessing functional consistency, and (2) multi-turn dialogues that capture the reasoning process behind translation refinement. Applied to Fortran→C++ and C++→CUDA, the pipeline yields 3.64k and 3.93k dialogues, respectively. Fine-tuning on this data yields dramatic improvements in functional correctness, boosting unit test success rates by over 56% on the challenging C++-to-CUDA task. We show that the generated data enables a 7B open-weight model to significantly outperform larger proprietary systems on key metrics like compilation success.

conference paper
Published

ParaCodex: A Profiling-Guided Autonomous Coding Agent for Reliable Parallel Code Generation and Translation


Kaplan, E., T. Bitan, L. Ghrayeb, L. Chen, T. Yotam, N. Hasabnis, and G. Oren, "ParaCodex: A Profiling-Guided Autonomous Coding Agent for Reliable Parallel Code Generation and Translation," Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), (July/2026), Association for Computational Linguistics, pp. 16113–16136. doi: https://doi.org/10.18653/v1/2026.acl-long.732

Parallel programming is central to HPC and AI, but producing code that is correct and fast remains challenging, especially for OpenMP GPU offload, where data movement and tuning dominate. Autonomous coding agents can compile, test, and profile on target hardware, but outputs are brittle without domain scaffolding.We present ParaCodex, an HPC-engineer workflow that turns a Codex-based agent into an autonomous OpenMP GPU offload system using staged hotspot analysis, explicit data planning, correctness gating, and profiling-guided refinement. We evaluate translation from serial CPU kernels to OpenMP GPU offload kernels on HeCBench, Rodinia, and NAS. After excluding five kernels, ParaCodex succeeded on all 31 valid kernels. In 27/31 (87%) of these valid cases, the generated kernels improved GPU time over reference implementations, a result that holds independently on both the A100 and RTX 4060. The resulting OpenMP kernels achieve geometric-mean speedups of 3.1 (A100) and 3.6 (RTX 4060) on HeCBench and 1.5 and 1.1 on Rodinia, and outperform a zero-shot Codex baseline on all suites. We also evaluate CUDA -> OpenMP offload translation on ParEval, where ParaCodex maintains high compilation and validation rates in code-only and end-to-end settings.

Published

ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation.


Jhaveri, S., E. Kaplan, T. Yotam, L. Chen, T. Bitan, N. Hasabnis, and G. Oren, "ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation.," (June/2026). doi: https://doi.org/10.13140/RG.2.2.25807.85920

Modern compute-intensive software is written against a rapidly changing ecosystem of accelerators, programming APIs, compiler stacks, and portability layers. As hardware and software platforms evolve, kernels in AI, scientific computing, simulation, graphics, and data-intensive workloads often need to migrate across CUDA, OpenMP, OpenCL, OpenMP target offload, and related parallel APIs. Large language models and autonomous coding agents are increasingly proposed as tools for this migration, but the field still lacks a reliable way to measure whether they can preserve the low-level parallel semantics that make such translations behaviorally valid under declared oracles, including thread indexing, synchronization, memory management, host-device coordination, and API-specific execution structure. We present PARBENCH, a kernel-centric benchmark framework designed to isolate and measure LLM-based parallel API translation under executable, reproducible conditions. Unlike general code benchmarks that are largely sequential, parallel-code benchmarks that emphasize generation from natural-language prompts, or repository-level studies that confound translation with build-system reconstruction, PARBENCH fixes the surrounding build, run, and verification infrastructure through declarative benchmark specifications and asks models to translate only the computational kernels. The benchmark draws on multiple open-source HPC suites to cover representative cross-API translation directions among CUDA, OpenMP, OpenCL, and OpenMP target offload. To probe whether success reflects robust translation rather than surface-form memorization, PARBENCH also includes AST-driven intended behavior-preserving, baseline-validated source augmentation. We further evaluate PARBENCH on state-of-the-art open and proprietary LLMs, showing that it exposes persistent barriers to reliable parallel code translation, including direction asymmetry, multi-file coordination, incomplete API adaptation, and uneven robustness to source-level perturbations.

conference paper
Published

AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine


Siebenschuh, C., Hippe, K., Gokdemir, O., Brace, A., Khan, A., Hossain, K., et al., "AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine," Eighth Conference on Machine Learning and Systems (MLSys, 2025, (April/2026)

Language models for scientific tasks are trained on text from scientific publications, most distributed as PDFs that require parsing. PDF parsing approaches range from inexpensive heuristics (for simple documents) to computationally intensive ML-driven systems (for complex or degraded ones). The choice of the "best" parser for a particular document depends on its computational cost and the accuracy of its output. To address these issues, we introduce an Adaptive Parallel PDF Parsing and Resource Scaling Engine (AdaParse), a data-driven strategy for assigning an appropriate parser to each document. We enlist scientists to select preferred parser outputs and incorporate this information through direct preference optimization (DPO) into AdaParse, thereby aligning its selection process with human judgment. AdaParse then incorporates hardware requirements and predicted accuracy of each parser to orchestrate computational resources efficiently for large-scale parsing campaigns. We demonstrate that AdaParse, when compared to state-of-the-art parsers, improves throughput by 17× while still achieving comparable accuracy (0.2 percent better) on a benchmark set of 1000 scientific documents. AdaParse's combination of high accuracy and parallel scalability makes it feasible to parse large-scale scientific document corpora to support the development of high-quality, trillion-token-scale text datasets.

Published

SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation


Shomee, H. H., R. Chaturvedi, Y. Xie, and T. Mallick, "SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation," arXiv, (February/2026), arXiv

Large language models (LLMs) are increasingly used to support question answering and decision-making in high-stakes, domain-specific settings such as natural hazard response and infrastructure planning, where effective answers must convey fine-grained, decision-critical details. However, existing evaluation frameworks for retrieval-augmented generation (RAG) and open-ended question answering primarily rely on surface-level similarity, factual consistency, or semantic relevance, and often fail to assess whether responses provide the specific information required for domain-sensitive decisions. To address this gap, we propose a multi-dimensional, reference-free evaluation framework that assesses LLM outputs along four complementary dimensions: specificity, robustness to paraphrasing and semantic perturbations, answer relevance, and context utilization. We introduce a curated dataset of 1,412 domain-specific question-answer pairs spanning 40 professional roles and seven natural hazard types to support systematic evaluation. We further conduct human evaluation to assess inter-annotator agreement and alignment between model outputs and human judgments, which highlights the inherent subjectivity of open-ended, domain-specific evaluation. Our results show that no single metric sufficiently captures answer quality in isolation and demonstrate the need for structured, multi-metric evaluation frameworks when deploying LLMs in high-stakes applications.

Published

LSHBloom: Internet-Scale Text Deduplication


Khan, A., R. Underwood, C. Siebenschuh, Y. Babuji, A. Ajith, K. Hippe, O. Gokdemir, A. Brace, K. Chard, and I. Foster, "LSHBloom: Internet-Scale Text Deduplication," PUBLISHED??, (January/2026)

conference paper
Published

FIRST: Federated Inference Resource Scheduling Toolkit for Scientific AI Model Access


Tanikanti, A., B. Côté, Y. Guo, L. Chen, N. Saint, R. Chard, K. Raffenetti, R. Thakur, T. Uram, I. Foster, M. E. Papka, and V. Vishwanath, "FIRST: Federated Inference Resource Scheduling Toolkit for Scientific AI Model Access," Proceedings of the SC '25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, (November/2025), St. Louis, Missouri, USA, Association for Computing Machinery, pp. 52-60. doi: https://doi.org/10.1145/3731599.3767346

We present the Federated Inference Resource Scheduling Toolkit (FIRST), a framework enabling Inference-as-a-Service across distributed High-Performance Computing (HPC) clusters. FIRST provides cloud-like access to diverse AI models, like Large Language Models (LLMs), on existing HPC infrastructure. Leveraging Globus Auth and Globus Compute, the system allows researchers to run parallel inference workloads via an OpenAI-compliant API on private, secure environments. This cluster-agnostic API allows requests to be distributed across federated clusters, targeting numerous hosted models. FIRST supports multiple inference backends (e.g., vLLM), auto-scales resources, maintains "hot" nodes for low-latency execution, and offers both high-throughput batch and interactive modes. The framework addresses the growing demand for private, secure, and scalable AI inference in scientific workflows, allowing researchers to generate billions of tokens daily on-premises without relying on commercial cloud infrastructure.

workshop paper
Published

Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models


Gokdemir, O., Getty, N., Underwood, R., Madireddy, S., Cappello, F., et al. , "Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models," Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC25) Workshops, 2025, (November/2025)

As scientific knowledge grows at an unprecedented pace, evaluation benchmarks must evolve to reflect new discoveries and ensure language models are tested on current, diverse literature. We propose a scalable, modular framework for generating multiple-choice question-answering (MCQA) benchmarks directly from large corpora of scientific papers. Our pipeline automates every stage of MCQA creation, including PDF parsing, semantic chunking, question generation, and model evaluation. As a case study, we generate more than 16,000 MCQs from 22,000 open-access articles in radiation and cancer biology. We then evaluate a suite of small language models (1.1B–14B parameters) on these questions, comparing baseline accuracy with retrieval-augmented generation (RAG) from paper-derived semantic chunks and from reasoning traces distilled from GPT-4.1. We find that reasoning-trace retrieval consistently improves performance on both synthetic and expert-annotated benchmarks, enabling several small models to surpass GPT-4 on the 2023 Astro Radiation and Cancer Biology exam.

conference paper
Published

Beyond End-to-End: Understanding the Limits of LLMs in Scientific Problem Solving


Liu, Y., Di, S., Getty, N., Mallick, T., Underwood, R., and Jin, S., "Beyond End-to-End: Understanding the Limits of LLMs in Scientific Problem Solving," Frontiers in Generative AI for HPC Science and Engineering: Foundations, Challenges, and Opportunities. , (November/2025), Trillion Parameter Consortium at The International Conference for High Performance Computing, Networking, Storage, and Analysis (TPC@SC), 2025

Multimodal large language models (MLLMs) are now widely used across many applications, including scientific question answering that requires combining visual and textual inputs. However, existing benchmarks in this area are mostly end-to-end, making it difficult to pinpoint where models fail. To address this gap, we design an evaluation framework that decomposes scientific question answering into subtasks for fine-grained assessment. We evaluate two MLLMs, Gemini 2.5 Pro and Qwen2.5-VL-32B-Instruct, on questions involving high-resolution visual data. Results show that accurate answers are unattainable without scripting or tool use. Although both models can solve individual subtasks, such as mapping cities to coordinates or computing pixel positions, they often fail to integrate these abilities in end-to-end reasoning, producing large deviations. Our findings highlight the importance of benchmarks that expose reasoning bottlenecks and suggest that agent-based or multi-model approaches may be required to achieve reliable performance on complex scientific tasks.

conference paper
Published

Exploring Distributed Vector Databases Performance on HPC Platforms: A Study with Qdrant


Ockerman, S., Gueroudji, A., Oh, S. Y., Underwood, R., Chia, N., Chard, K., Ross, R., and Venkataraman, S. , "Exploring Distributed Vector Databases Performance on HPC Platforms: A Study with Qdrant," Trillion Parameter Consortium at The International Conference for High Performance Computing, Networking, Storage, and Analysis (TPC@SC), 2025, (November/2025)

Vector databases have rapidly grown in popularity, enabling efficient similarity search over data such as text, images, and video. They now play a central role in modern AI workflows, aiding large language models by grounding model outputs in external literature through retrieval-augmented generation. Despite their importance, little is known about the performance characteristics of vector databases in high-performance computing (HPC) systems that drive large-scale science. This work presents an empirical study of distributed vector database performance on the Polaris supercomputer in the Argonne Leadership Computing Facility. We construct a realistic biological-text workload from BV-BRC and generate embeddings from the peS2o corpus using Qwen3-Embedding-4B. We select Qdrant to evaluate insertion, index construction, and query latency with up to 32 workers. Informed by practical lessons from our experience, this work takes a first step toward characterizing vector database performance on HPC platforms to guide future research and optimization.

journal article
Published

Context Length Alone Hurts LLM Performance Despite Perfect Retrieval


Y. Du, M. Tian, S. Ronanki, S. Rongali, S. Bodapati, A. Galstyan, A. Wells, R. Schwartz, E. Huerta, and H. Peng., "Context Length Alone Hurts LLM Performance Despite Perfect Retrieval," Empirical Methods in Natural Language Processing (EMNLP), (October/2025)

Large language models (LLMs) often fail to scale their performance on long-context tasks performance in line with the context lengths they support. This gap is commonly attributed to retrieval failures -- the models' inability to identify relevant information in the long inputs. Accordingly, recent efforts often focus on evaluating and improving LLMs' retrieval performance: if retrieval is perfect, a model should, in principle, perform just as well on a long input as it does on a short one -- or should it? This paper presents findings that the answer to this question may be negative. Our systematic experiments across 5 open- and closed-source LLMs on math, question answering, and coding tasks reveal that, even when models can perfectly retrieve all relevant information, their performance still degrades substantially (13.9%--85%) as input length increases but remains well within the models' claimed lengths. This failure occurs even when the irrelevant tokens are replaced with minimally distracting whitespace, and, more surprisingly, when they are all masked and the models are forced to attend only to the relevant tokens. A similar performance drop is observed when all relevant evidence is placed immediately before the question. Our findings reveal a previously-unrealized limitation: the sheer length of the input alone can hurt LLM performance, independent of retrieval quality and without any distraction. They motivate our simple, model-agnostic mitigation strategy that transforms a long-context task into a short-context one by prompting the model to recite the retrieved evidence before attempting to solve the problem. On RULER, we observe a consistent improvement of GPT-4o up to 4% on an already strong baseline.

tech report
Published

Probing the Critical Point (CritPt) of AI Reasoning: A Frontier Physics Research Benchmark


Zhu, M., Tian, M., Yang, X., Zhou, T., Yuan, L., Zhu, P., et al., "Probing the Critical Point (CritPt) of AI Reasoning: A Frontier Physics Research Benchmark," arxiv.org, (September/2025)

While large language models (LLMs) with reasoning capabilities are progressing rapidly on high-school math competitions and coding, can they reason effectively through complex, open-ended challenges found in frontier physics research? And crucially, what kinds of reasoning tasks do physicists want LLMs to assist with? To address these questions, we present the CritPt (Complex Research using Integrated Thinking - Physics Test, pronounced "critical point"), the first benchmark designed to test LLMs on unpublished, research-level reasoning tasks that broadly covers modern physics research areas, including condensed matter, quantum physics, atomic, molecular & optical physics, astrophysics, high energy physics, mathematical physics, statistical physics, nuclear physics, nonlinear dynamics, fluid dynamics and biophysics. CritPt consists of 71 composite research challenges designed to simulate full-scale research projects at the entry level, which are also decomposed to 190 simpler checkpoint tasks for more fine-grained insights. All problems are newly created by 50+ active physics researchers based on their own research. Every problem is hand-curated to admit a guess-resistant and machine-verifiable answer and is evaluated by an automated grading pipeline heavily customized for advanced physics-specific output formats. We find that while current state-of-the-art LLMs show early promise on isolated checkpoints, they remain far from being able to reliably solve full research-scale challenges: the best average accuracy among base models is only 5.7%, achieved by GPT-5 (high), moderately rising to around 10% when equipped with coding tools. Through the realistic yet standardized evaluation offered by CritPt, we highlight a large disconnect between current model capabilities and realistic physics research demands, offering a foundation to guide the development of scientifically grounded AI tools.