Amazon Web Services Unveils AWS Glue 6.0: A Major Leap Forward in Serverless Data Processing and Cost Reduction

SEATTLE — In a move set to reshape enterprise data architectures, Amazon Web Services (AWS) has officially announced the general availability of AWS Glue 6.0. This latest iteration of the fully serverless data integration and ETL (Extract, Transform, Load) service introduces a modernized runtime stack, robust support for Apache Iceberg v3, and—perhaps most notably for cost-conscious enterprise leaders—a substantial 30% price reduction compared to previous versions.
Built upon the powerful foundations of Apache Spark 4.1, Python 3.13, and Scala 2.13, AWS Glue 6.0 positions itself as the most comprehensive and performant serverless Spark-managed service on the market. By marrying high-end runtime performance with deep Apache Iceberg v3 integration, AWS is directly targeting the growing enterprise demand for efficient, scalable, and cost-effective data lakehouse operations.
Main Facts: What is AWS Glue 6.0?
AWS Glue 6.0 is a comprehensive architectural overhaul of Amazon’s flagship serverless data integration service. Designed to handle modern, massive-scale data workflows, the new release centers on three primary pillars: dramatic cost savings, advanced open-table format support, and next-generation runtime performance.
- 30% Price Reduction: AWS has slashed pricing by 30% across the board for Glue 6.0 workloads, making large-scale data processing significantly more affordable.
- Apache Iceberg v3 Integration: Based on Iceberg 1.11.0, the release provides full compliance with the Iceberg v3 specification, highlighted by the revolutionary
VARIANTdata type with shredding support. - Modernized Runtime Engine: The service now runs natively on Apache Spark 4.1, Python 3.13, and Scala 2.13, ensuring blazing-fast execution speeds and compatibility with the latest language and framework standards.
- Real-Time Capabilities: Enhanced streaming features allow for real-time data ingestion and processing with single-digit millisecond latency.
- Seamless Migration Path: Enterprises can adopt Glue 6.0 with zero API changes required, utilizing automated upgrade agents and simple configuration flags within the AWS Management Console, CLI, or SDKs.
Chronology: The Evolution Leading to Glue 6.0
To understand the significance of AWS Glue 6.0, it is helpful to look at the trajectory of cloud-based data processing and the open-table format revolution over the last several years.
The Rise of Serverless ETL (2017–2020)
When AWS Glue was first introduced, its primary promise was liberating data engineers from the operational overhead of provisioning, configuring, and scaling clusters. Over successive iterations (Glue 1.0, 2.0, and 3.0), AWS steadily improved startup times, introduced support for newer Python and Spark versions, and optimized resource utilization. However, as data volumes exploded and organizations shifted away from proprietary data warehouses toward open data lakes, the demand for open-table formats intensified.
The Open Table Format Boom and Iceberg (2021–2024)
As data lakes matured, transactional capabilities became paramount. Formats like Apache Iceberg emerged as industry standards, allowing users to perform ACID transactions, time travel queries, and schema evolution directly on cloud object storage like Amazon S3. AWS steadily built out support for Iceberg in previous Glue versions, but limitations in semi-structured data handling and parsing overhead remained friction points for enterprise developers.
The Modernization Era: Spark 4.1 and Glue 6.0 (2025–2026)
With the open-source community rallying around Apache Spark 4.1 and Iceberg v3, AWS recognized the need for a holistic architectural refresh. Development focused heavily on modernizing the underlying execution engine to support advanced semi-structured data types natively—culminating in the July/August 2026 release of AWS Glue 6.0. By lowering prices concurrently, AWS has signaled an aggressive push to capture market share from competing data analytics platforms.
Supporting Data and Technical Deep Dive
The architectural enhancements in AWS Glue 6.0 are not merely incremental; they target specific data engineering bottlenecks that have historically plagued large-scale analytics pipelines.
The Power of the VARIANT Data Type and Shredding
One of the most significant technical achievements in Glue 6.0 is its complete implementation of the Apache Iceberg v3 specification, anchored by the 1.11.0 release. The crown jewel of this update is the VARIANT data type equipped with built-in shredding support.

Traditionally, processing semi-structured data—such as nested JSON files, application logs, and event streams—required expensive workarounds:
- Flattening schemas ahead of time, which frequently broke downstream pipelines when schemas evolved.
- Storing data as raw strings, requiring custom parsing logic and resulting in slow query read performance.
- Duplicating data across multiple tables to optimize for different query patterns.
With VARIANT shredding, Glue 6.0 allows data engineers to ingest, store, and query semi-structured payloads natively. The engine automatically "shreds" the internal structure, optimizing physical storage layouts for columnar reads. The result is dramatically faster query performance, reduced storage overhead, and complete resilience against breaking changes when upstream JSON schemas shift.
Runtime Performance: Spark 4.1, Python 3.13, and Scala 2.13
Under the hood, Glue 6.0 leverages Apache Spark 4.1, the latest evolution of the ubiquitous distributed compute engine. Spark 4.1 brings critical performance optimizations, query planner enhancements, and memory management improvements. Combined with Python 3.13—known for its execution speedups and refined typing capabilities—and Scala 2.13, developers can execute complex ETL pipelines with substantially lower resource footprints.
Furthermore, these performance gains extend to real-time data streaming architectures, where Glue 6.0 achieves single-digit millisecond latency. This enables organizations to build unified pipelines that seamlessly handle both batch historical processing and real-time operational analytics without switching tools.
Official Responses and Accessibility
AWS has ensured that transitioning to Glue 6.0 is frictionless for existing users. No core API changes are mandatory; teams can adopt the new runtime simply by updating configuration parameters.
Getting Started and Migration Pathways
Engineers can spin up Glue 6.0 jobs using the standard create-job or update-job APIs via the AWS Command Line Interface (AWS SDKs, or through visual interfaces like AWS Glue Studio, Amazon SageMaker Unified Studio, and traditional Jupyter notebooks).
Within the AWS Glue Studio console, configuring a job is as straightforward as navigating to the Job Details tab and selecting the version labeled:
Glue 6.0 – Supports Spark 4.1, Scala 2, Python 3
For organizations managing large fleets of legacy jobs, AWS has introduced the Spark upgrade agent within Glue Studio, alongside an auto-upgrade feature designed to streamline the migration process safely and systematically.

Pricing and Regional Availability
AWS Glue 6.0 is generally available today across all global AWS regions where AWS Glue operates.
The pricing model remains straightforward and transparent:
- ETL Jobs and Crawlers: Billed by the second on an hourly rate, now at a 30% discount compared to prior generations.
- Data Catalog: Billed via a simplified monthly fee for storing and accessing metadata. To encourage early adoption and small-scale experimentation, AWS continues to offer the first million objects stored and the first million accesses completely free of charge.
Implications for the Data Engineering Industry
The launch of AWS Glue 6.0 carries profound implications for data teams, enterprise architects, and the broader cloud analytics ecosystem.
1. The Democratization of Advanced Data Lakehouses
By pairing advanced Iceberg v3 capabilities with a 30% price reduction, AWS is removing economic and technical barriers to building modern data lakehouses. Historically, maintaining high-performance open-table formats at scale required specialized tuning and incurred steep infrastructure bills. Glue 6.0 abstracts away this complexity, making enterprise-grade data management accessible to smaller engineering teams.
2. Death of the Traditional ETL Pipeline Breakage
Schema drift has long been the bane of data engineers. Whenever a third-party API or application altered a JSON log structure, downstream ETL jobs would crash, demanding urgent manual intervention. The introduction of the VARIANT data type with shredding support effectively neutralizes this pain point. By allowing schema evolution without pipeline failure, Glue 6.0 promises to reclaim countless hours previously lost to maintenance and debugging.
3. Heightened Competitive Pressure in the Analytics Market
As cloud providers and independent data platforms vie for dominance in the data lakehouse space, pricing pressure is mounting. AWS’s aggressive 30% price drop on a modernized Spark 4.1 runtime sets a new benchmark for cost-performance efficiency. Competitors will likely be forced to respond with similar optimizations or pricing adjustments.
Summary
AWS Glue 6.0 represents a watershed moment for serverless data processing. By combining substantial cost savings with deep open-source innovations like Apache Iceberg v3 and Spark 4.1, Amazon Web Services has delivered a toolkit that is faster, cheaper, and fundamentally easier to maintain. For organizations looking to modernize their data architecture without inflating their cloud budgets, Glue 6.0 offers an exceptionally compelling path forward.
