The Data Team works on preparing, curating, and optimizing massive natural and synthetic datasets to fuel next-generation scientific AI. To deliver on this mission and handle trillions of tokens of text, the team pioneers high-performance data engineering pipelines that eliminate traditional scaling bottlenecks. Our research directly addresses the heavy computational challenges of transforming raw, complex scientific literature repositories into highly optimized, clean training datasets. Two foundational technologies developed by our team illustrate these efforts:
Key Initiatives & Core Activities
Beyond developing core software frameworks, our team drives foundational research across data synthesis, supercomputing workflows, and multi-institutional partnerships. Our active initiatives include:
- Developing Adaptive Parsing Architectures (AdaParse): We build data-driven software engines that solve the extreme computational cost of accurately extracting text and formulas from PDFs. By using machine learning to intelligently route document workloads, our tools eliminate the trade-off between parsing speed and data quality.
- Engineering Extreme-Scale Deduplication Pipelines (LSHBloom): We design memory-efficient indexing systems that remove duplicate, repetitive text from massive internet and scientific datasets. This software prevents AI models from overfitting, saving millions of hours of expensive supercomputing training time.
- Adapting Complex Domain Data via Data Narratives: We are researching generic techniques to translate dense, highly structured database records into natural language descriptions optimized for AI training and inference. For example, we are currently exploring this framework within the biology domain to convert complex genomic and biological datasets into descriptive text formats that Large Language Models can readily interpret.
- Engineering Scalable Software for AI Agents: We are actively researching and designing the high-throughput infrastructure required to support autonomous AI agents that interact heavily with scientific data. This work ensures that future AI systems can quickly query, process, and reason through massive datasets in real time.
- Supercomputing Workflow Optimization: Our teams work on scaling the AuroraGPT data pipeline on the world's most powerful supercomputers at Argone. Through deep architectural tuning of the Parsl parallel workflow framework, our research focuses on severe memory bottlenecks and engineered custom checkpointing systems, to enable reliable extraction of massive data footprints across leadership-class hardware.
- Strategic National Lab & Academic Partnerships: We maintain a collaborative, multi-institutional research ecosystem to advance the frontier of scientific AI. Current active collaborations include joint research with Temple University on data format trade-offs in AI reasoning traces, the University of Wisconsin-Madison on eliminating scaling bottlenecks in massive vector databases, and the Advanced Photon Source (APS) on generalizing our data narratives framework to new scientific domains.
While these core initiatives guide our long-term research direction, our immediate impact is driven by the production-ready tools we build and release to the open-science community.
To overcome the specific engineering bottlenecks of document ingestion and dataset cleanliness at a trillion-token scale, our team developed two foundational technologies: AdaParse, our adaptive parallel parsing engine, and LSHBloom, our memory-efficient deduplication pipeline. The architecture, performance metrics, and deployments of these two flagship artifacts are detailed below.
References & Academic Citations
If you use our frameworks or research artifacts in your work, please cite our papers:
[1] Siebenschuh, C., Hippe, K., Gokdemir, O., Brace, A., Khan, A. M., Hossain, K., Babuji, Y., Chia, N., Vishwanath, V., Stevens, R., Ramanathan, A., Foster, I., & Underwood, R. (2025). AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine. Proceedings of the Eighth Conference on Machine Learning and Systems (MLSys 2025).
[2] Khan, A., Underwood, R., Siebenschuh, C., Babuji, Y., Ajith, A., Hippe, K., Gokdemir, O., Brace, A., Chard, K., & Foster, I. (2024). LSHBloom: Memory-efficient, Extreme-scale Document Deduplication. arXiv preprint arXiv:2411.04257.
Team Members
Ian Foster and Robert Underwood serve as the leadership contacts for the Data Team. For further information about our research, open-source software, datasets, or potential collaborations, please contact them directly.