Jhaveri, S., E. Kaplan, T. Yotam, L. Chen, T. Bitan, N. Hasabnis, and G. Oren, "ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation.," (June/2026). doi: https://doi.org/10.13140/RG.2.2.25807.85920
Modern compute-intensive software is written against a rapidly changing ecosystem of accelerators, programming APIs, compiler stacks, and portability layers. As hardware and software platforms evolve, kernels in AI, scientific computing, simulation, graphics, and data-intensive workloads often need to migrate across CUDA, OpenMP, OpenCL, OpenMP target offload, and related parallel APIs. Large language models and autonomous coding agents are increasingly proposed as tools for this migration, but the field still lacks a reliable way to measure whether they can preserve the low-level parallel semantics that make such translations behaviorally valid under declared oracles, including thread indexing, synchronization, memory management, host-device coordination, and API-specific execution structure. We present PARBENCH, a kernel-centric benchmark framework designed to isolate and measure LLM-based parallel API translation under executable, reproducible conditions. Unlike general code benchmarks that are largely sequential, parallel-code benchmarks that emphasize generation from natural-language prompts, or repository-level studies that confound translation with build-system reconstruction, PARBENCH fixes the surrounding build, run, and verification infrastructure through declarative benchmark specifications and asks models to translate only the computational kernels. The benchmark draws on multiple open-source HPC suites to cover representative cross-API translation directions among CUDA, OpenMP, OpenCL, and OpenMP target offload. To probe whether success reflects robust translation rather than surface-form memorization, PARBENCH also includes AST-driven intended behavior-preserving, baseline-validated source augmentation. We further evaluate PARBENCH on state-of-the-art open and proprietary LLMs, showing that it exposes persistent barriers to reliable parallel code translation, including direction asymmetry, multi-file coordination, incomplete API adaptation, and uneven robustness to source-level perturbations.