Skip to main content

Science Data Pipelines

NAS Division experts work with various science teams to provide and operate custom data pipelines that expedite the steps involved in processing the massive amounts of raw data obtained from NASA’s ground- and space-borne observatories. These steps include: cleaning the data to correct errors, omissions, or inconsistencies; and visualizing, modeling, and analyzing the data.

Image of TESS Data Pipeline
Graphic showing how the TESS science data pipeline works. Wendy Stenzel, NASA/Ames

The first and second generations of NAS data pipeline software were developed to support the Kepler and TESS planet-hunting missions. The NAS Division now has over a decade of experience operating pipelines to process the enormous volumes of data produced by these missions. Between them, Kepler and TESS have discovered most of the exoplanets known to science, and the archived data products produced by these pipelines provide a rich dataset for use by researchers all over the world.

The third generation of NAS data pipeline software builds on the heritage of the Kepler and TESS software. Named Ziggy, this software provides a standalone application that can be used to build pipelines for any purpose. For example, Ziggy has been used to build a prototype data pipeline for NASA’s Earth System Observatory's Surface Biology & Geology Mission, slated for launch in the late 2020s, which will collect 2.4 terabytes (TB) of data per day and produce over 40 TB of science data product per day.

Ziggy includes a wide array of improvements over the earlier pipeline applications, and is available as open source software on GitHub, along with extensive documentation and examples. Using Ziggy, researchers can develop a pipeline that can take advantage of the tremendous compute and data storage resources provided by the NAS Division.