Amazon Web Services Announces AWS Glue 6.0: A Major Leap Forward in Serverless Data Processing with Apache Spark 4.1 and Apache Iceberg v3

SEATTLE — In a move set to reshape the economics and capabilities of cloud-scale data engineering, Amazon Web Services (AWS) has officially announced the general availability of AWS Glue 6.0. The latest iteration of the company’s fully serverless data integration and extract, transform, and load (ETL) service introduces a transformative 30% price reduction compared to previous generations, while simultaneously rolling out comprehensive support for Apache Iceberg v3.
Built on a deeply modernized runtime stack featuring Apache Spark 4.1, Python 3.13, and Scala 2.13, AWS Glue 6.0 aims to deliver drastically faster query performance, streamlined semi-structured data management, and real-time streaming capabilities characterized by single-digit millisecond latencies.
According to AWS, this release establishes AWS Glue as the most complete implementation of Apache Iceberg v3 available on any fully managed, serverless Spark service on the market today. The new version is available starting today across all AWS regions where AWS Glue operates.
1. Main Facts: What’s New in AWS Glue 6.0
AWS Glue 6.0 represents one of the most comprehensive modernization efforts in the history of the service. It addresses two of the most pressing challenges facing modern data platforms: the escalating cost of cloud data processing and the operational friction involved in managing complex, evolving semi-structured data formats like JSON, application logs, and clickstream events.
The 30% Price Reduction
For enterprise data teams operating under continuous pressure to optimize cloud expenditures, the headline financial metric of AWS Glue 6.0 is its 30% lower pricing structure relative to prior versions. By passing infrastructure efficiency gains directly to customers, AWS is lowering the barrier to entry for large-scale batch processing, complex ETL pipelines, and continuous data lake maintenance. Billing remains second-based and hourly for crawlers and jobs, while the AWS Glue Data Catalog retains its predictable monthly fee structure, featuring a generous tier of one million free stored objects and one million free accesses.
Modernized Runtime Engine: Spark 4.1, Python 3.13, and Scala 2.13
Under the hood, AWS Glue 6.0 abandons legacy dependencies in favor of an entirely updated execution environment.
- Apache Spark 4.1: Brings the latest execution optimizations, enhanced memory management, and advanced query planning capabilities to serverless pipelines.
- Python 3.13: Empowers data scientists and engineers to write modern, performant PySpark code leveraging the latest language features, security updates, and ecosystem packages.
- Scala 2.13: Provides robust, high-performance compilation options for enterprise applications requiring strict type safety and heavy JVM integration.
Full Apache Iceberg v3 Support and the VARIANT Data Type
Built on top of Iceberg version 1.11.0, AWS Glue 6.0 delivers the complete Apache Iceberg v3 specification. The standout architectural feature in this release is native support for the VARIANT data type, complete with built-in shredding support.
In traditional data lake architectures, semi-structured formats like JSON often required flattening schemas or casting fields into bulky string data types, which frequently resulted in inflated storage footprints, expensive custom parsing scripts, and brittle pipelines that broke whenever an upstream application changed its payload structure.
With VARIANT shredding in AWS Glue 6.0, data engineering teams can ingest, store, and query complex JSON and event data natively. The engine automatically optimizes how semi-structured attributes are read, achieving dramatically faster query performance without requiring duplicate data copies, rigid schema definitions, or custom pre-processing code. This capability fundamentally transforms how organizations manage telemetry, logs, and dynamic document data at petabyte scale.

2. Chronology: The Evolution Leading to Glue 6.0
To understand the strategic significance of AWS Glue 6.0, it is helpful to trace the evolution of AWS’s serverless data integration ecosystem and the broader shift toward open table formats in modern data architectures.
The Rise of Serverless ETL (2017–2020)
When AWS initially launched Glue, the goal was to eliminate the operational overhead of provisioning, configuring, and scaling Apache Spark clusters for ETL workloads. Over its first few years, Glue evolved from a basic cataloging and crawler tool into a robust execution engine capable of running complex Spark and Python scripts without requiring infrastructure management. However, early iterations faced criticisms regarding startup latency and cost relative to managing self-hosted EMR clusters.
Embracing Open Table Formats (2021–2023)
As data lakes matured into "lakehouses," the industry experienced a paradigm shift away from proprietary formats toward open table formats like Apache Hudi, Delta Lake, and Apache Iceberg. Recognizing this trend, AWS steadily integrated native support for Iceberg into Glue, enabling ACID transactions, time travel, and schema evolution directly on Amazon S3. Successive releases optimized catalog integration and query federation via Amazon Athena and Amazon EMR.
The Push for Modern Runtimes and Cost Optimization (2024–2026)
As enterprise data volumes exploded, customers demanded faster execution engines and more cost-effective pricing models to combat cloud cost fatigue. AWS responded by systematically modernizing its analytics portfolio.
- Glue 4.0 and 5.0: Laid the groundwork by introducing faster job startup times and incremental updates to Spark and Python runtimes.
- The Culmination (August 2026): AWS releases Glue 6.0, representing a synchronized leap forward. By coupling an underlying engine upgrade (Spark 4.1) with native Iceberg v3 capabilities and a permanent 30% price cut, AWS has positioned Glue as a foundational pillar for next-generation cloud data architecture.
3. Supporting Data & Technical Architecture
The technical underpinnings of AWS Glue 6.0 are designed to maximize throughput while minimizing operational complexity for data architects.
Performance Metrics and Benchmarks
Internal evaluations and early customer feedback indicate that workloads running on AWS Glue 6.0 benefit significantly from the underlying optimizations in Spark 4.1. Key architectural enhancements include:
- Vectorized Readers: Improved memory layout and vectorized processing pipelines reduce CPU cycles spent on row-by-row data translation.
- Optimized Shuffle Operations: Enhanced network and disk shuffle mechanisms minimize bottlenecks during large-scale join and aggregation operations.
- Streaming Latency: Real-time streaming ETL jobs benefit from optimized micro-batch processing, achieving end-to-end latencies in the single-digit milliseconds when consuming from sources like Amazon Kinesis or Amazon MSK (Managed Streaming for Apache Kafka).
Storage and Cost Efficiency Breakdown
| Component | Traditional Approach (Pre-Glue 6.0) | AWS Glue 6.0 Approach | Operational & Financial Benefit |
|---|---|---|---|
| Semi-Structured Data | String columns or heavily flattened schemas | Native VARIANT data type with automatic shredding |
Eliminates duplicate copies; up to 30-50% faster query read performance. |
| Runtime Engine | Spark 3.x / Older Python & Scala runtimes | Spark 4.1, Python 3.13, Scala 2.13 | Faster execution times, modern language features, lower compute minutes. |
| Pricing Model | Standard legacy tier pricing | Permanent 30% price reduction | Direct reduction in monthly cloud infrastructure spend for ETL pipelines. |
| Table Format Support | Partial or basic Iceberg features | Complete Apache Iceberg v3 specification | Robust transactional guarantees, hidden partitioning, and advanced maintenance. |
4. Official Responses and Industry Perspectives
AWS leadership and engineering teams have emphasized that Glue 6.0 is a direct response to evolving customer needs around cost management and data architecture modernization.
"With AWS Glue 6.0, we wanted to address the two most common refrains we hear from data leaders: ‘How can we process more data without increasing our budget?’ and ‘How can we stop spending weeks writing custom code just to handle changing JSON schemas?’" said a senior AWS product spokesperson. "By delivering a 30% price reduction alongside native Apache Iceberg v3 support and the revolutionary
VARIANTdata type, we are giving our customers an unmatched serverless platform that is faster, cheaper, and profoundly easier to use."
Early industry analysts and data engineering experts have echoed these sentiments, noting that the combination of lower pricing and advanced table format support makes serverless Spark an increasingly compelling alternative to provisioned cluster management.

Data platform engineers have praised the elimination of custom parsing logic for JSON payloads, noting that the new VARIANT type dramatically reduces code maintenance overhead across large enterprise pipelines.
5. Implications for Enterprise Data Strategies
The release of AWS Glue 6.0 carries profound implications for data engineering teams, Chief Data Officers (CDOs), and enterprise cloud architectures.
1. Accelerated Migration to Open Lakehouses
With full Iceberg v3 support baked directly into a fully serverless managed service, organizations no longer need to choose between the operational simplicity of serverless ETL and the open, vendor-neutral governance of modern table formats. Glue 6.0 removes the friction of maintaining custom Iceberg maintenance jobs (such as compaction, snapshot expiration, and manifest file optimization), allowing teams to build robust, ACID-compliant data lakes with minimal operational overhead.
2. Radical Simplification of Event-Driven Pipelines
Organizations ingesting massive volumes of web analytics, IoT telemetry, and application logs can now leverage the VARIANT data type to streamline their ingestion layers. By avoiding schema-on-write constraints and brittle flattening transformations, data teams can achieve higher agility when upstream applications evolve, drastically reducing pipeline failure rates and engineering hours spent on maintenance.
3. Re-evaluating Cloud Budgets
The 30% price reduction changes the cost-benefit calculus for organizations considering migrating legacy on-premises Hadoop clusters or expensive third-party ETL appliances to the cloud. When combined with serverless auto-scaling—where compute resources scale down to zero when jobs are idle—Glue 6.0 offers a highly predictable and cost-effective economic model for enterprise data processing.
Getting Started and Migration Pathways
Adopting AWS Glue 6.0 requires no breaking API changes, allowing organizations to transition smoothly without rewriting their orchestration logic.
- Creation and Updates: Developers can select the new runtime version by specifying the existing
--glue-versionparameter (6.0) via the AWS Command Line Interface (AWS CLI), AWS SDKs, or infrastructure-as-code tools. - AWS Glue Studio: Users navigating the AWS Glue Studio console can select Glue 6.0 – Supports Spark 4.1, Scala 2, Python 3 directly under the Job Details tab.
- Interactive Notebooks: Data scientists using Jupyter-based notebooks or interactive sessions can initialize the environment by setting
%glue_version 6.0in their magic commands. - Automated Upgrades: To assist teams with migration, AWS has provided the Spark upgrade agent within Glue Studio, alongside automated upgrade features designed to transition legacy jobs seamlessly.
Regional availability spans all standard AWS regions where AWS Glue operates. For organizations looking to integrate AI assistants into their development workflows, AWS also supports the new release via the AWS MCP Server and associated plugins, enabling developers to query documentation, check regional availability, and troubleshoot error codes using their preferred AI tools.
As enterprises continue to scale their data operations in the cloud, AWS Glue 6.0 establishes a new benchmark for performance, cost efficiency, and open-format integration in serverless data processing.
