SLAC's neural network compresses scientific data 100x without loss
New AI method preserves fine-grained details like X-ray speckles that traditional compression erases, addressing storage crisis at next-gen facilities.
Researchers at SLAC National Accelerator Laboratory have developed an artificial intelligence system that compresses massive scientific datasets by up to 100-fold while retaining subtle details that conventional compression methods erase.
The neural network-based approach addresses a looming data crisis: next-generation experiments like SLAC's upgraded Linac Coherent Light Source (LCLS) will generate nearly one terabyte of data per second—far exceeding current storage and analysis capabilities. The work was published in Nature Machine Intelligence on August 24, 2026.
How the system preserves scientific signal
Conventional compression treats all data uniformly, often eliminating fine-grained features that contain valuable scientific information. In X-ray imaging, for example, tiny speckle patterns reveal how materials are structured and how that structure changes over time.
"Those speckles often reflect the underlying arrangement, disorder, or dynamics of a material," said Yuan Ni, a SLAC research associate who led the work. "If we lose them, we would lose unique scientific insights."
The new method uses wavelet analysis to separate dataset features by scale, then compresses different-scaled features separately through a neural network. This ensures finer details aren't lost in generalized compression. The neural network learns a compact representation that preserves these features while achieving 10- to 100-fold file size reductions depending on the data type and desired fidelity.
Selective decompression cuts retrieval time
Beyond storage efficiency, the system enables researchers to decompress only specific regions of interest rather than entire files—a process that can take minutes to days with conventional methods.
"If you are using a conventional compressor, you would need to decompress the entire file," said Zhantao Chen, now an assistant professor at the University of Texas at Austin, who worked on the method as a SLAC research associate. "This method can decompress only the region of interest rather than the entire dataset, so it's much more efficient."
The team tested the approach on diverse data types including molecular measurements from multiple experimental techniques, solar magnetic field observations, and photographs. The neural network adapted to different data characteristics, learning which features matter for each measurement type.
Why it matters
The LCLS will eventually generate up to one million X-ray pulses per second to capture atomic and molecular motion. Instruments like the X-ray photon fluctuation spectroscopy (XPFS) system—designed to study exotic topological and quantum materials—require novel processing to handle this data volume. Without new compression approaches, storage costs and analysis bottlenecks could limit the scientific return from billion-dollar facilities. This method provides a path to manage data floods while preserving the signal researchers need for discovery.
Joshua Turner, a lead scientist at SLAC and the Stanford Institute for Materials and Energy Sciences who served as principal investigator, noted the method works alongside existing data-reduction techniques rather than replacing them. "There are many applications in science now where data storage and analysis speed are really important problems, and I think this method is a clever way to solve them," Turner said.
The research involved collaborators from UC Davis and Carnegie Mellon University. The team trained neural networks using Perlmutter, a computational resource at Lawrence Berkeley National Laboratory's National Energy Research Scientific Computing Center. Details were first reported by SLAC National Accelerator Laboratory.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
