REDI — automated pipeline for scientific AI datasets
ArXiv paper introduces REDI, an open-source framework that standardizes the messy path from raw scientific data to AI-ready training sets across climate, materials science, and fusion research.
• Five-stage pipeline: ingest, preprocess, transform, structure, output
• Built-in reproducibility instrumentation and provenance tracking
• Deploys as an agent-callable skill for autonomous workflows
• SetGo companion tool automates FAIR compliance and catalog publication