Skip to main content

Cloudera Adds NVIDIA CUDA-X GPU Acceleration to Apache Spark 4.1

Cloudera said it enabled native GPU acceleration for Apache Spark 4.1 in Cloudera Data Engineering, with support from NVIDIA libraries. The change is intended to shorten processing time for Spark workloads without requiring code changes.

Cloudera linked the effort to challenges it described with large-scale Spark jobs taking hours and increasing cloud compute costs. It also said the approach was designed to support governance and operational consistency in hybrid environments through Cloudera Data Engineering and Cloudera Anywhere Cloud.

The capability uses the NVIDIA CUDA-X library, with the NVIDIA cuDF plug-in for Apache Spark. Cloudera said the plug-in would support Cloudera Anywhere Cloud and provide built-in deployment without manual driver configuration, while using Cloudera Unified Data Fabric for enterprise security and governance.

Cloudera said the GPU acceleration capability for Apache Spark would be available in Cloudera Data Engineering as part of the Cloudera Anywhere Cloud announced at EVOLVE Singapore on August 20, 2026. The company said additional demonstrations and technical sessions were scheduled for NVIDIA GTC Berlin and Cloudera EVOLVE New York later this year.

“For many organizations, AI isn't limited by models. It's limited by how quickly they can turn raw data into trusted, usable insights,” said Leo Brunnick, Chief Product Officer at Cloudera. “Accelerating Spark inside Cloudera Data Engineering helps remove that bottleneck, allowing customers to move from data preparation to analytics and AI faster while keeping governance, security, and operational consistency at the center of their strategy.”

“The fastest path to accelerating AI deployments is the one that aligns with how enterprises already operate today,” said Pat Lee, vice president, Strategic Enterprise Partnerships at NVIDIA. “With NVIDIA AI infrastructure and CUDA-X libraries now native to Cloudera Data Engineering, enterprises can lower costs and dramatically speed up Apache Spark pipelines without changing a single line of PySpark or SQL code, turning business data into a foundation for AI.”

n

Provided by Globe Newswire on behalf of Cloudera. Click to read original content.