Delta Lake for Rust and Python has hit a major milestone: deltalake 1.0.0 is now
available. This release delivers a complete DataFusion 55 + Arrow 59 upgrade, column
mapping support compatible with Databricks Unity Catalog, Spark-parity deletion
vector CDF streaming, and buoyant_kernel 0.28 for faster bug fixes and snapshot metrics.
Performance & Scale Fixes
The latest releases of the deltalake crate contain a tremendous number of of performance fixes, especially in the newer Datafusion table providers. Much of this is credit to the diligent work by a number of delta-rs contributors with @ethan-tyler leading the charge on a number of the performance enhancements.
ReduceDeltaScanNext file_id memory
overhead dropped [PR #4519], metadata-only scan batches chunked within session
limits [PR #4730] to prevent byte-array overflow [Issue #4727], batch partition splitting
efficiency improved [PR #4715], and row-order preservation across scans maintained
[PR #4692]. Unbounded in-flight uploads that could cause OOM on slow
object stores are now capped [PR #4708], and full vacuum scans on multi-level
partitions run faster than ever [PR #4635]. Together these fixes ensure smooth
scans and writes, from single-digit GB tables all the way to petabyte tables.
Breaking Changes
For AWS S3 uses the deltalake 1.0 release has a very important behavior
change: the S3DynamoDbLogStore support was removed and instead S3-based
"conditional put" semantics are now the default supported mode for writes on
S3.
S3DynamoDbLogStore allowed coordinated commits with a similarly configured
Delta/Spark writers on the same table. As of the 1.0.0 release of
delta-rs this coordination is no longer needed for the majority of users.
For Rust<->Rust readers and writers we recommend upgrading all writers to 1.0 at the same time to ensure that all potential concurrent writers are cooperating with the newer conditional put behaviors.
From our testing Databricks (DBR) based writers will cooperate correctly with delta-rs 1.0 as the underlying machinery supports these conditional put semantics already.
Concurrent writers using deltalake 1.0 and Apache Spark should upgrade to Delta Lake 4.2 or later which have the library updates needed to interact safely with AWS S3.
More technical details around the change can be found in this GitHub discussion
Column Mapping & Schema Flexibility
Column mapping support is ready for production [PR #4505]. You can now
create tables with arbitrary physical–logical column name mappings, making delta-rs
tables compatible with Databricks Unity Catalog NiFi templates. Physical execution
pipelines [PR #4493] and format fixes now round-trip correctly
with PySpark while preserving partition identities. Schema evolution receives a powerful
DROP NOT NULL constraint management API [PR #4552] that lets users
dynamically lift NOT NULL constraints on existing columns wherever同上、 muscular workflows
demand. Finally, Python users can commit AddAction and RemoveAction in the same
transaction [Issue #4480], simplifying idempotent schema
migrations.
CDF (Change Data Feed) Maturity
Change Data Feed support reaches Spark-parity with deletion vector CDF [PR #4639] mirroring Spark 3.x semantics. CDF scans now prune by partition predicates [PR #4487] and preserve merge row ordinals [PR #4473], enabling streaming bootstraps that are temporal and partition-layered: first snapshot → change feed, matching Spark [Issue #4554]. If you've wanted to build Kafka-to-Delta CDC pipelines that are stream-native without hand-crafted transaction stitching, 1.0.0 delivers that foundation.
buoyant_kernel 0.28
As we said in the initial
buoyant_kernel announcement post, this distribution gives
us an unblock path for curated patches ahead of upstream release cycles, directly
benefiting all delta-rs 1.0.0+ users.
buoyant_kernel 0.28.0 just released with refined snapshot metrics:
LogSegmentLoadSuccess, ProtocolMetadataLoadSuccess, and SnapshotBuildSuccess events now carry load-type +
source fields with crc_versions_behind counts, letting users diagnose slow startup
and out-of-sync recovery paths [delta-kernel-rs PR #2915, PR #2916]. The meatiest improvements, merged into delta-rs
1.0.0 via [PR #4712], stream directly into every scan, giving Delta Lake
applications immediate visibility into initial snapshot contamination, stale metadata
runs, and snapshot Hektar resolution costs.
DataFusion 55 + Arrow 59 Upgrade
We promised upgraded data/tasks infrastructure and are delivering across the board: DataFusion 55.x and Arrow 59.x chains are now live [PR #4655] with full zero-copy parquet uploads reintroduced [PR #4705]. These upgrades give users the latest Arrow 59 features, improved compression/encodings, and DataFusion 55's enhanced SQL planning and execution efficiency. The upgrade fully satisfies the Arrow 58 + DataFusion 53 promise from April's buoyant_kernel post and pushes the stack even further forward.
Thank You
Our sincerest thanks to the 17 contributors who made delta-rs 1.0.0 a reality:
- @adamreeve
- @akashjainn
- @akhramov
- @ethan-tyler
- @fallintoplace
- @hntd187
- @ion-elgreco
- @jx2lee
- @JustinRush80
- @kszucs
- @laverem
- @liamphmurphy
- @mahesh-desu
- @matt-lebl
- @osmant
- @plaindocs
- @Rodrigo-Palma
- @adampolomski
- @ArulJerald
- @seokjin0414
Delta Lake on Rust and Python is growing rapidly. If your organization wants to improve how you move data, react quicker with usable streams, or operate larger real-time data lakes at lower cost with delta-rs 1.0.0, drop me an email and let's chat!
