Share this on:
What You'll Learn
Snowflake and Databricks are Salesforce Data 360’s two flagship zero-copy partners, and the marketing on both sides tends to describe the connection the same way. Instant. Seamless. No data movement required.
What gets left out is that zero copy isn’t one mechanism. It’s three, each with a different cost profile. And there’s a specific set of workloads where the honest answer is still to ingest the data instead of federating it.
What Zero Copy Is, and Is Not
Zero copy is bidirectional federation. Data 360 can read data live from Snowflake or Databricks, and share data back to either platform, with no data movement, copying, or reformatting in either direction.
A metadata layer tracks where each piece of data actually lives and routes a query to it, rather than pulling that data into Salesforce first and creating a copy that then has to be kept in sync by hand.
It is not a data modeling shortcut. A federated table still has to map onto Data 360’s Data Model Objects. It’s the same governed layer every native Salesforce record goes through. Connecting to Snowflake or Databricks changes where the data physically sits. It doesn’t decide what that data means inside the platform. And skipping that modeling step is the single most common reason a federation project stalls.
Query Federation, File Federation, & Cached Acceleration Explained
Underneath the single label “zero copy,” Data 360 actually runs three distinct mechanisms, and they aren’t interchangeable.
Query federation
Sends a live query to Snowflake or Databricks and lets that system’s own compute execute it, returning only the result set. Nothing persists in Data 360 afterward. It’s the right default for reporting and ad hoc analysis against data that’s already reasonably clean. And it’s the mode most demos default to, because it needs the least setup.
File federation
Skips the external system’s compute layer entirely and reads directly from the underlying storage, typically Apache Iceberg tables. This is what makes large lake data, historical archives, semi-structured logs, and document stores usable from Data 360 without pipelining them into a CRM object model. Data that was effectively locked away becomes queryable in place.
Cached acceleration
Sits on top of either mode, materializing a temporary, scheduled-refresh copy of frequently queried data inside Data 360. It trades some freshness for not hitting the external warehouse on every single read. This matters once a query pattern runs often enough that the live-query cost adds up.
None of the three is a universally better choice.
They answer two different questions:
- How fresh this needs to be.
- How often it’s going to get queried.
When to Ingest Instead of Using Zero Copy
Federation earns its reputation in most of the cases it gets used for.
Three failure modes worth naming directly, because none of them are edge cases. They’re the specific, recurring reasons a federation deployment runs into trouble once it’s live.
1
Sub-second reads are the first.
A live query’s speed is entirely a function of the external system’s current load. And a warehouse that answers one analyst instantly can slow down noticeably once a dashboard fires a dozen queries against it at once. Anything that needs a guaranteed response time under real concurrent traffic is a better fit for ingestion or a cached layer. Because federation ties that guarantee to a system Data 360 doesn’t control.
2
Heavy pre-consumption transforms are the second
Federated data still has to resolve into Data 360’s data model. And if the logic required to get it there is substantial, that transformation belongs upstream of the federation layer.
3
Availability-critical paths are the third
Federation has no local fallback. If the external source goes down, the data goes with it. And for any workflow where that’s not acceptable, that single dependency settles the decision before anything else gets evaluated.
Closing the Loop: Sharing the Golden Record Back to Snowflake & Databricks
Most conversations about zero copy stop at pulling data in. The more useful half of the design is what flows back out once Data 360 has already done the work of resolving that data into a single trustworthy record. And this is where treating Snowflake and Databricks as one partner stack actually pays off, because Data 360 closes the loop with both.
On the Snowflake side, a unified profile built inside Data 360 can be shared back into Snowflake in real time – where a data science team uses that enriched, governed data to build and train models directly.
On the Databricks side, the same golden record can be shared into Databricks without a separate export step, and Einstein Studio and Databricks Mosaic AI are built around exactly that handoff. A team trains a model in Databricks using Data 360 data, then brings the trained model back into Salesforce to power personalization, prediction, or an Agentforce agent.
Neither direction requires the underlying data to be copied to make the round trip work. That’s the part of this architecture worth designing for on purpose.
How LumenData Architects Zero Copy Across Salesforce, Snowflake, & Databricks
Picking the right pattern for a given workload, knowing when ingestion is the more honest call, and making sure the golden record gets shared back out to the right place- that’s the work behind LumenData’s engagements.
That means:
Pairing zero-copy federation to both Snowflake and Databricks with multidomain MDM that resolves conflicting records into one governed golden record before federation ever exposes them.
That ordering is deliberate. What gets federated is already trustworthy. And what flows back out to either platform doesn’t create a second copy drifting out of sync with the first.
We are an Approved Delivery Partner with Databricks, a Platinum Partner with Informatica, and a Premier Services Partner with Snowflake, with 75+ certifications across the team.
Considering a zero-copy architecture across your own Snowflake or Databricks environment? Talk to LumenData about which pattern actually fits your workload.
About LumenData
LumenData is a leading provider of Enterprise Data Management, Cloud and Analytics solutions and helps businesses handle data silos, discover their potential, and prepare for end-to-end digital transformation. Founded in 2008, the company is headquartered in Santa Clara, California, with locations in India.
With 150+ Technical and Functional Consultants, LumenData forms strong client partnerships to drive high-quality outcomes. Their work across multiple industries and with prestigious clients like Versant Health, Boston Consulting Group, FDA, Department of Labor, Kroger, Nissan, Autodesk, Bayer, Bausch & Lomb, Citibank, Credit Suisse, Cummins, Gilead, HP, Nintendo, PC Connection, Starbucks, University of Colorado, Weight Watchers, KAO, HealthEdge, Amylyx, Brinks, Clara Analytics, and Royal Caribbean Group, speaks to their capabilities.
For media inquiries, please contact: marketing@lumendata.com.
Considering a zero-copy architecture across your own Snowflake or Databricks environment?
Authors
References
- Reddit Drives Revenue with Agentforce-Powered Advertiser Support - Salesforce
- Snowflake - Data 360 Partner - Salesforce
- Salesforce and Databricks Partnership: Data 360, Agentforce, and AI - Salesforce
- LumenData Achieves Approved Delivery Partner Status with Databricks Professional Services - LumenData
- Informatica MDM | Data-Driven Salesforce Alignment - LumenData
- LumenData + Snowflake Implementation - LumenData


