We are turning data into usable assets in the age of AI

AI is only as effective as the data it’s built on. Scientific data must be structured, curated, and AI-ready from the start. That’s the core mission of the Scientific Data Division: treating data as first-class, usable assets in the age of AI.

Founded in 2021, the Scientific Data Division fosters breakthrough discoveries through the application and development of novel data science methods, technologies, and infrastructures in partnership with science domain experts.

Our tight-knit team works across scientific disciplines in order to develop and apply cutting-edge computational methods and tackle complex problems in cosmology, physics, biosciences, and other scientific domains.

Through collaboration, we gain new insights and expertise

We manage the full scientific data lifecycle. From acquisition and cleaning to analysis, publication, and long-term preservation, SDD supports every stage of the data journey. This work is increasingly essential as scientific data grows in volume, complexity, and variety—across instruments, simulations, and disciplines.

Usability is essential—and that starts with UX/UI. We build systems scientists actually want to use—because they’re designed with them in mind. Human-centered design, intuitive interfaces, and workflow-aware tools reduce friction and make data easier to work with, reuse, and trust.

Statistical thinking runs through everything we do. From data quality and reproducibility to uncertainty quantification and trustworthy AI, statistical rigor underpins our work across the board.

We focus on designing tools and systems that fit how scientists work

Scientific data moves through many hands, systems, and stages. It’s rarely linear, and we build infrastructure that supports that full lifecycle without losing meaning or quality along the way. We use UX and human-centered design to create flexible, relevant tools that meet the needs of our users. Our job is to build platforms that support collaboration, reuse, and scale—so data can move easily across teams, domains, and facilities.

As science becomes more data-driven, we develop methods and tools that support infrastructure for complex, distributed workflows. 

Learn more about how our team works together to accelerate AI innovation.