Vantage Insights

AI-First Salesforce & Databricks Consulting

Row Filtering in Lakeflow Connect: What It Actually Buys Salesforce Teams

Row Filtering in Lakeflow Connect: What It Actually Buys Salesforce Teams

Databricks added row filtering to Lakeflow Connect’s managed ingestion connectors in January 2026, and the release note undersells it: “row filtering for managed ingestion connectors to improve performance and minimize data duplication.” For teams pulling Salesforce data into a lakehouse, that is a real capability. It is also a narrower one than the pitch suggests, and the gap between what people assume and what the connector actually supports is where pipelines get designed wrong.

What the Feature Does

Lakeflow Connect’s Salesforce connector now accepts a row_filter condition in the pipeline spec, applied like a SQL WHERE clause during both the initial load and every incremental sync after it, according to Databricks’ row filtering documentation. Supported operators are AND, OR, =, !=, <, <=, >, and >=, though OR is only available on the Salesforce and Google Analytics connectors; ServiceNow, for example, supports AND only. String literals need quotes, numeric literals do not. That is the whole surface area. No LIKE, no IN, no subqueries.

The immediate use case is straightforward: a sandbox or dev environment does not need five years of closed Opportunity history, and a pipeline pulling only current fiscal year records ingests less, costs less to run, and finishes faster. For orgs with large Salesforce instances, that is a meaningful operational win, not a marketing one.

The Limitation Most Teams Will Miss

Here is the part that will surprise anyone reading “SQL WHERE clause” and assuming they can filter on arbitrary business logic. Per Databricks’ row filtering documentation, Salesforce row filtering is restricted to exactly two columns: the primary key (Id) and a single cursor column, which the connector auto-selects per object from a priority list, SystemModstamp first, then LastModifiedDate, CreatedDate, or LoginTime, whichever is present. You cannot filter Opportunities by StageName, Accounts by RecordTypeId, or Cases by Status, and you cannot combine two different cursor columns in one condition. If the business reason for filtering is “we only want closed-won” or “exclude test accounts,” row filtering as shipped does not do that. You are filtering by when a record changed, not by what it is.

That constraint is not a bug, it follows from how the connector does incremental extraction, but it means row filtering solves the “reduce volume by time window” problem well and the “reduce volume by business segment” problem not at all. Teams that need the second one are still looking at a transformation step downstream, typically a view or a Lakeflow Declarative Pipeline filter applied after ingestion, which defeats part of the cost savings since the full object still lands before it gets trimmed.

Two Gotchas That Change the Design

Two behaviors in the documentation change how you should design around this, and both are easy to miss on a first read.

First, deletes are not retroactive. If a row already ingested no longer matches the filter condition, it is not removed from the target table. The filter only governs what comes in, not what is already there. Second, rows that existed before a filter was added or changed are not automatically picked up. Widening a filter after the fact requires a full refresh of the pipeline to backfill the newly-included rows. Both of these mean row filtering is closer to a one-way valve set at pipeline creation than a dynamic query you can safely loosen and tighten. Changing it mid-life is a full-refresh operation, and that needs to be budgeted for, not discovered during an incident.

The connector’s schema evolution behavior compounds this. New and deleted columns propagate automatically, and column renames are handled as an add plus a drop, but data type changes on existing columns are not automated. Combine a narrow, timestamp-only row filter with a manual data type migration path and you have two places where “it will just work” is the wrong assumption for a healthcare or financial services pipeline carrying compliance obligations.

Where It Genuinely Pays Off

None of this makes row filtering a weak feature, it makes it a specific one. It is well suited to a small set of real problems: keeping non-prod environments smaller than prod, capping historical backfill volume on first load, and running lighter incremental pipelines against high-churn objects like Tasks or Events where most records fall outside a recent time window. Paired with Lakeflow Connect’s existing column selection, which lets you drop columns you never needed in the first place, the two features together do meaningfully cut the footprint of a Salesforce object landing in Unity Catalog, provided the filtering logic stays anchored to time and identity rather than business state.

What We Would Do

When we scope a Salesforce-to-Databricks ingestion pipeline, we treat row filtering as a volume control, not an access control. Business-level filtering (by record type, owner, or status) belongs in a governed view or a downstream transformation with its own tests, not in the ingestion layer where it is invisible to anyone reading the pipeline config six months later. We also document the full-refresh cost of any filter change before the pipeline ships, so widening a time window later is a planned operation instead of a surprise during a data audit. That discipline is consistent with how we build any production pipeline: agentic delivery speeds up the implementation work, but the governance decisions, the filter boundaries, and the compliance review still get made by a person who understands what the data is for.

Row filtering is a good addition to Lakeflow Connect, and worth adopting for the specific problem it solves. It is not a substitute for a governance layer, and treating it as one is the mistake to design against before the first pipeline goes live.

Get In Touch

Tell us where your Salesforce org and your AI ambitions stand today.

Chat On WhatsApp