October 1, 2026

Amazon S3 Tables Embraces Apache Iceberg V3: A Paradigm Shift for Petabyte-Scale Analytics and Data Management

amazon-s3-tables-embraces-apache-iceberg-v3-a-paradigm-shift-for-petabyte-scale-analytics-and-data-management

amazon-s3-tables-embraces-apache-iceberg-v3-a-paradigm-shift-for-petabyte-scale-analytics-and-data-management

SEATTLE — In a major development for cloud-native data architecture, Amazon Web Services (AWS) has announced full support for the Apache Iceberg V3 specification across Amazon S3 Tables. This launch marks a turning point for data engineers, enterprise architects, and analytics teams struggling with the operational bottlenecks of previous-generation data lake table formats.

By integrating the advanced capabilities of Iceberg V3—including native deletion vectors, automated row lineage, and powerful new data types like variant, geometry, geography, and nanosecond timestamps—Amazon S3 Tables aims to drastically reduce query latency, slash storage overhead, and streamline complex compliance operations at petabyte scale.


Main Facts: What is Changing in Amazon S3 Tables?

The integration of Apache Iceberg V3 into Amazon S3 Tables is not merely an incremental update; it represents a comprehensive overhaul of how semi-structured, geospatial, and high-frequency time-series data can be managed within open data lakes.

Key highlights of the announcement include:

  • Native V3 Data Types: Amazon S3 Tables now natively supports variant, nanosecond-precision timestamps, geometry, geography, and unknown data types. This eliminates the legacy necessity of encoding complex payloads into cumbersome strings or integers.
  • Deletion Vectors: Replacing the positional delete files of Iceberg V2, V3 introduces compact binary deletion vectors. This addresses one of the most resource-intensive pain points in big data analytics: handling row-level updates and deletes without triggering expensive, performance-degrading table compactions.
  • Built-In Row Lineage: Every record is now automatically injected with _row_id and _last_updated_sequence_number, allowing downstream data pipelines to execute incremental reads and state changes without scanning entire tables.
  • Seamless In-Place Upgrades: Organizations can instantly upgrade existing V2 tables to V3 with a single SQL command (ALTER TABLE ... SET TBLPROPERTIES ('format-version' = '3')) without requiring tedious and costly data rewrites.
  • Zero Additional Cost: The V3 feature set is available immediately in all AWS Regions that support S3 Tables, with no supplementary charges beyond standard S3 Tables storage and operation pricing.

Chronology: The Evolution to Iceberg V3 on AWS

To understand the weight of this release, one must look at the architectural trajectory of modern data lakes.

The Rise of Open Table Formats

For years, organizations building data lakes faced a fundamental dilemma. While cloud object storage like Amazon S3 offered virtually infinite scalability and low-cost storage, it lacked the ACID (Atomicity, Consistency, Isolation, Durability) transaction guarantees and table management features found in traditional relational database management systems (RDBMS).

Apache Iceberg emerged as the open-source community’s definitive answer to this challenge. By abstracting file locations into a structured table format, Iceberg brought enterprise-grade features—such as time travel, schema evolution, and hidden partitioning—directly to open Parquet files sitting on object storage.

The Limitations of Iceberg V2

As analytics workloads expanded into petabyte and exabyte scales, teams utilizing Apache Iceberg V2 began hitting operational limits. When compliance mandates (such as GDPR or CCPA) required the deletion of thousands of user records from massive multi-billion-row tables, Iceberg V2 generated thousands of small positional delete files. These files severely degraded query performance until background compaction jobs could run.

Furthermore, modern applications generating semi-structured JSON payloads, high-precision IoT telemetry with nanosecond timestamps, or complex geospatial mapping coordinates were forced to rely on sub-optimal workarounds. Developers frequently had to serialize complex structures into string columns, forcing every downstream analytical query to incur the heavy computational cost of parsing strings at runtime.

The V3 Breakthrough and AWS Integration

Recognizing these community-wide friction points, the Apache Iceberg project finalized the V3 specification to directly tackle structural inefficiencies, introduce native semi-structured processing, and optimize row-level mutation tracking.

Building upon its robust ecosystem support, AWS moved rapidly to incorporate these enhancements natively. Starting today, Amazon S3 Tables—a purpose-built storage tier designed specifically to automate the maintenance, compaction, and replication of Iceberg tables—fully supports the entire V3 specification.


Supporting Data: Technical Deep Dive and Practical Implementation

To evaluate the operational impact of the upgrade, enterprise analytics teams must examine how Iceberg V3 alters daily database management workflows, query performance, and storage efficiency.

Handling Semi-Structured Data with the Variant Type

In modern retail, fintech, and SaaS applications, event streams vary wildly in structure. A single clickstream table might capture page views, user searches, and financial checkouts, each with entirely distinct attributes.

Under V2, managing this required complex schema evolution or storing payloads as raw JSON strings. With V3 and Amazon S3 Tables, engineers can leverage the variant data type:

CREATE TABLE my_catalog.namespace.clickstream (
  event_id bigint,
  event_time timestamp,
  user_id string,
  payload variant
)
USING iceberg
TBLPROPERTIES ('format-version' = '3');

When data is written to this table, the storage engine automatically shreds the variant data into hidden columns and builds internal statistics. At query time, these statistics enable aggressive file pruning, bypassing the massive I/O overhead traditionally associated with parsing JSON strings on the fly:

SELECT
  event_id,
  user_id,
  variant_get(payload, '$.action', 'string') AS action,
  variant_get(payload, '$.amount', 'double') AS amount
FROM my_catalog.namespace.clickstream
WHERE variant_get(payload, '$.action', 'string') = 'purchase'
  AND variant_get(payload, '$.amount', 'double') > 50.00;

Streamlining Compliance Deletes via Deletion Vectors

Compliance and data privacy regulations require rapid removal of user data. In V2 architectures, deleting 50,000 records from a 2-billion-row table created an avalanche of small positional delete files, crippling query speeds.

Amazon S3 Tables now support all Apache Iceberg V3 data types | Amazon Web Services

Iceberg V3 solves this via merge-on-read configurations and deletion vectors. By altering the table properties:

ALTER TABLE my_catalog.namespace.clickstream
SET TBLPROPERTIES (
  'write.delete.mode' = 'merge-on-read',
  'write.update.mode' = 'merge-on-read',
  'write.merge.mode' = 'merge-on-read'
);

When a compliance delete is executed:

DELETE FROM my_catalog.namespace.clickstream
WHERE user_id = 'user-42';

The engine writes a single, highly compact binary deletion vector file instead of rewriting data files or generating thousands of discrete delete objects. Amazon S3 Tables then handles the eventual merging and compaction during routine, automated maintenance cycles.

Powering Incremental Pipelines with Row Lineage

Traditional data pipelines often rely on full table scans or complex watermark heuristics to identify newly changed records. Iceberg V3 eliminates this friction by automatically appending _row_id and _last_updated_sequence_number to every record.

Downstream ETL pipelines can now query changes instantly:

SELECT *, _row_id, _last_updated_sequence_number
FROM my_catalog.namespace.clickstream
WHERE _last_updated_sequence_number > 42;

This capability transforms ETL efficiency, allowing downstream jobs to checkpoint sequence numbers and process only delta changes rather than consuming compute resources on redundant full-table scans.


Official Responses and Ecosystem Integration

AWS offers the broadest native Apache Iceberg support among major cloud service providers, integrating table specifications across ingestion, storage, cataloging, and analytics engines.

Daniel Abib, representing the AWS engineering team behind the launch, emphasized the seamless interoperability designed into the new release. Through native support for the Iceberg REST Catalog (IRC) API across both Amazon S3 Tables and the AWS Glue Data Catalog, enterprises are not locked into a single compute engine.

Cross-Service Compatibility on AWS

  • Storage & Optimization: Amazon S3 Tables provides the foundational, purpose-built storage layer that automatically executes background compaction, Intelligent-Tiering, and replication.
  • Data Ingestion & Processing: Analytics engineers can write and transform V3 data seamlessly using Amazon EMR Spark.
  • Governance & Cataloging: AWS Glue 6.0 offers full Apache Iceberg V3 support and integrated catalog management.
  • Query & BI Analytics: Enterprise users can execute high-performance queries across V3 tables using Amazon Redshift.

Furthermore, to assist developers and system administrators in navigating the migration to V3, AWS has integrated comprehensive documentation and troubleshooting workflows directly into the AWS MCP Server and associated plugins, allowing AI-assisted tooling to query API references and regional availability in real time.


Implications: What This Means for the Enterprise Data Landscape

The arrival of Apache Iceberg V3 support in Amazon S3 Tables carries profound implications for data-driven organizations.

1. Drastic Reduction in Cloud Storage and Compute Waste

The combination of deletion vectors and variant data types directly addresses hidden cloud bills. By removing the need to over-provision compute clusters just to parse unstructured strings or brute-force through uncompacted positional delete files, organizations will see immediate reductions in both compute spend (such as EMR or Redshift CU hours) and S3 storage request overhead.

2. Accelerated Data Governance and Compliance Agility

Data privacy regulations demand agility. The ability to execute rapid, lightweight deletions without degrading database performance ensures that legal and compliance teams can enforce data retention and right-to-be-forgotten policies without pushing engineering infrastructure to its limits.

3. Maturation of Open Data Lakehouse Architectures

As open table formats like Iceberg continue to absorb features traditionally reserved for proprietary data warehouses (such as atomic mutations, efficient row-level versioning, and rich native types), the case for proprietary data lock-in weakens. Enterprises can maintain absolute ownership of their raw Parquet files in Amazon S3 while enjoying the performance, reliability, and governance of an enterprise-grade database engine.

Migration Considerations

While upgrading to V3 is a seamless, one-way atomic operation via a simple ALTER TABLE property update, enterprise architects must exercise caution. Because the Apache Iceberg specification does not support downgrading from V3 to V2, database administrators must verify that all downstream analytics engines, third-party query tools, and custom ingestion pipelines accessing the catalog fully support the V3 specification before initiating migrations. AWS provides backward compatibility for V2 readers during transitional phases to ensure smooth organizational change management.

Getting Started

Amazon S3 Tables support for all Apache Iceberg V3 features is available today across all AWS regions where S3 Tables operate. Organizations can begin creating V3 table buckets or upgrading existing workloads via the Amazon S3 console, API, or AWS Command Line Interface (CLI) at no extra charge beyond standard S3 pricing.