Quick Facts
- Databricks announced Lakebase Change Data Feed in public preview at Data + AI Summit 2026, designed to eliminate manual CDC pipelines between OLTP and analytical systems.
- Lakebase, built on Neon’s serverless PostgreSQL, reached general availability on Feb. 3, 2026, and has grown adoption at more than twice the pace of Databricks’ core data warehousing product.
- Databricks crossed a $5.4 billion revenue run rate in Q4 with more than 65% year-over-year growth, and holds a $134 billion valuation after raising roughly $5 billion in equity financing.
Databricks used its Data + AI Summit 2026 this week to announce a direct attack on the data pipeline market. The company introduced Lakebase Change Data Feed, now in public preview, alongside deeper integration between its Lakebase operational database and the Lakeflow data engineering product.
The event drew more than 30,000 attendees in person at San Francisco’s Moscone Center from June 15 through 18, with tens of thousands more participating remotely across 150 countries.
The Pipeline Problem
Traditional data architectures require separate pipelines to move records from transactional databases into analytical systems. Teams must configure database connectors, monitor replication states, absorb performance impacts on production databases, and track failures across disconnected tools.
Databricks argues this creates duplicated data, inconsistent permissions, fragmented observability, and continuous performance tuning overhead. The Lakebase Change Data Feed is positioned as the fix.
How Lakebase Change Data Feed Works
The feature pushes application writes into Unity Catalog within one minute. That latency reduction lets machine learning models retrain and score against current application data without a separate ingestion pipeline.
Lakebase serves as the Bronze layer in a medallion architecture. High-velocity updates happen inside Postgres, and the full change history flows into the Lakehouse automatically as SCD Type 2 records. Databricks says this eliminates the custom CDC stack that previously powered ETL, streaming workflows, and audit logs.
Lakebase itself is built on Neon’s serverless PostgreSQL technology. Its architecture separates compute and storage, using Amazon S3 as the primary source of truth. Stateless Postgres compute nodes, Safekeepers that replicate the write-ahead log, and Pageservers that serve data pages from object storage form the core components.
Lakeflow and Open Standards
Databricks also packaged its data pipeline capabilities under the Lakeflow brand. The company is open-sourcing Lakeflow’s pipeline model, positioning it as a new ETL standard built for AI workloads. Lakebase and Lakeflow together connect transactional data in Postgres to analytical data in Delta Lake through Unity Catalog.
Business Scale Behind the Bet
Databricks reported a $5.4 billion annualized revenue run rate as of its fiscal Q4, with more than 65% year-over-year growth. AI revenue alone exceeds a $1.4 billion run rate. Net revenue retention stands above 140%, and more than 800 customers spend at least $1 million annually. More than 70 customers exceed $10 million in annual spend.
The company completed more than $7 billion in financing, including roughly $5 billion in equity at a $134 billion valuation and approximately $2 billion in debt capacity. Databricks generated positive free cash flow over the past 12 months.
More than 20,000 organizations use the platform, including Mastercard, AT&T, Bayer, Unilever, and 70% of the Fortune 500. Lakebase, which entered public preview at the 2025 Data + AI Summit, has seen adoption grow at more than twice the rate of Databricks’ established data warehousing product since its February 2026 general availability date.
What This Means for Software Teams
For SaaS companies running separate operational and analytical stacks, the Databricks pitch is consolidation under a single governance layer. Unity Catalog would manage permissions across application data, analytics, and AI models without separate tooling for each.
The Lakebase Change Data Feed targets engineering teams that currently maintain custom CDC infrastructure. If the one-minute latency claim holds in production, it removes a common bottleneck for real-time ML feature pipelines and agent memory systems that depend on current application state.
Read more: Databricks declares the end of pipelines with a unified platform for operational and analytical data
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
