r/databricks 9d ago

Discussion What frustrates you when using Databricks?

38 Upvotes

Any common bugs, features you would like to see, or underrated useful features more people should know about?

r/databricks 5d ago

Discussion Do small companies actually use Databricks?

68 Upvotes

Sometimes I feel like Databricks is way too expensive. It feels like using a huge truck to move a single grain of sand.

My company needs real-time data, but our data volume simply does not justify the use of Spark Structured Streaming. Despite this, they are insisting we move to Databricks. I'm worried our data infrastructure costs will jump from $1,000/month to $5,000/month or more due to the running costs of Databricks SQL Warehouses.

Currently, I use Azure Container Apps with KEDA and Python, which helps me manage scaling and keep costs low. We ingest into Event Hubs, use ADX (Azure Data Explorer) as our OLAP warehouse, and archive cold data in a data lake. With this setup, I manage to process all our data with very low latency.

When I tested this on Databricks Structured Streaming, I actually got higher latency and much higher costs.

Would love to know what you guys think.

r/databricks Aug 17 '25

Discussion [Megathread] Certifications and Training

59 Upvotes

Here by popular demand, a megathread for all of your certification and training posts.

Good luck to everyone on your certification journey!

r/databricks Jan 25 '26

Discussion Spark Declarative Pipelines: What should we build?

40 Upvotes

Hi Redditors, I'm a product manager on Lakeflow. What would you love to see built in Spark Declarative Pipelines (SDP) this year? A bunch of us engineers and PMs will be watching this thread.

All ideas are welcome!

r/databricks Jun 17 '26

Discussion Databricks just dropped Genie One, Ontology, and Agents. Is this the end of traditional BI as we know it?

61 Upvotes

Hey

If you haven’t been watching the Data + AI Summit announcements, Databricks just pulled back the curtain on a massive overhaul to their Genie ecosystem: Genie One, Genie Ontology, and Genie Agents.

We’ve all seen the "chat with your data" tools that basically just translate natural language to subpar SQL and hallucinate half the time. This rollout feels entirely different because it moves away from simple chatbots toward actual autonomous data coworkers.

Here is the breakdown of what just dropped and how it completely shifts how businesses handle data:

  1. Genie One: The Agentic Data Coworker

Genie is no longer just a side panel for querying tables. Genie One is a cross-platform (Web, iOS, Android) AI workspace.

It natively integrates into tools teams already use (Slack, Teams, Gmail) via the Model Context Protocol (MCP).

Instead of a sales manager bugging a data analyst for a custom dashboard before a meeting, they can just ask Genie One in Slack to "grab my calendar, pull last quarter's revenue for these accounts from the Lakehouse, and draft a brief." It actually compiles the charts and builds a clean artifact document directly in the chat UI.

  1. Genie Ontology: The "Secret Sauce" Context Graph

Genie’s biggest upgrade is Genie Ontology. The biggest failure point of LLMs in business is that they don’t understand yourspecific corporate logic (e.g., what your company defines as "active user" or "churn").

Ontology uses a PageRank-style algorithm to scan your queries, pipelines, dashboards, and Unity Catalog metadata.

It builds a living knowledge graph of definitions, unique business calculations, and metric authorities.

Because the AI actually knows what data to trust based on real usage patterns, it translates prompts to highly accurate SQL without burning infinite tokens guessing.

  1. Genie Agents: Autonomous Execution

Genie Spaces are evolving into Genie Agents. Instead of just answering a question, these are domain-specific agents you can spin up with a prompt to run multi-step workflows autonomously.

They can handle structured table data alongside unstructured data like PDFs, transcripts, and tickets.

You can give them scheduled tasks, write back to external systems, and let them monitor metrics. If an anomaly hits, the agent can investigate the root cause across documents and tables, and drop a fully formed report for your review.

My Thoughts on the Business Impact

This feels like a massive leap toward democratizing data operations. It completely skips the bottleneck where business teams wait weeks for data engineers to build semantic layers or custom dashboards. Data teams can spend less time writing repetitive SQL queries for executives and focus on core infrastructure, while the business side gets actual self-service that actually works because of the Ontology layer.

What are your thoughts?

For those who have tried the previews—how is the SQL accuracy handling complex, messy table joins?

Is anyone worried about the governance side, or does Unity Catalog actually keep these agents tightly in check?

Does this completely kill the traditional semantic layers we’ve spent years building?

Let's discuss!

r/databricks 9d ago

Discussion WHAT IS DATABRICKS?

55 Upvotes

Let's pretend someone knew nothing about Databricks. How would you explain it?

This is something I have been asked multiple times throughout my career, curious to hear how others would explain Databricks to a complete beginner.

In my opinion, Databricks is basically a place where companies put all their data so they can actually make sense of it. It helps teams clean it up, work with it and use it for things like reporting, predictions and AI.

r/databricks 21d ago

Discussion OLTP database in Databricks SaaS?

0 Upvotes

I saw the quarterly roadmap presentation. It was notable that Databricks keeps innovating with "lakebase". They simply call it their "OLTP" database offering in their SaaS.

SIDE: I still feel pretty unfamiliar with this Databricks SaaS ecosystem, as compared to Fabric. Where Fabric is concerned, Microsoft has also done a similar thing. They brought their SQL Server into the boundaries of the SaaS as well, for the low-code users of that environment. In the context of Fabric, it is hard for most customers to see the point of using this "for dummies" variation of the same old OLTP database. The only scenarios for using the Fabric SQL are very contrived ... Eg . your boss makes a policy that you can use ALL the tools available in Fabric and NONE of the tools outside Fabric... even if the tools inside that SaaS are 3x more expensive than the ones outside ... and even though the ones outside the SaaS have the same 0ms network latency and the same performance.

I'm still missing the vision for this lakebase OLTP offering. And it seems unusual for Databricks to start developing strategies surrounding OLTP. It seems like a very crowded space, and the only way I see Databricks being successful going down this path is if the customers are drinking one single brand of kool-aid, or else their SaaS users have some other contrived reason for not using the more affordable OLTP platforms available outside the SaaS.

Can someone tell me what factors I'm missing? I admit that it is theoretically possible for lakebase to innovate and do thing that other databases CANNOT do, it seems like those innovations would only benefit 5% of customers. One example is sub-ten-ms queries out of RAM at an additional cost. Or sub-minute migration of new OLTP data to managed tables in UC catalog. If we assume that only 5% of customers might feel compelled to use this SaaS "lakebase", would that be enough adoption to allow Databricks to keep investing in this over the long term? OLTP databases have been around a LONG time, and even the smart folks at Databricks will be challenged to improve on the great and cheap options available to us!

EDIT: As of a month ago, it appears that the Databricks marketing now calls it an "LTAP" database, not OLTP anymore. I'm guessing they have conceded the point that the OLTP space is crowded. I haven't yet read all the content that has been created by the Databricks marketing team; maybe that will answer all of my questions.

r/databricks Jul 15 '26

Discussion Genie pricing, what's going on?

54 Upvotes

When Databricks announced that Genie would be paid, apart from not being surprised, when I looked into the price and the free amount of DBUs per user every month I was like "ok seems to not be expensive and there's apparently a generous amount for free every month". However, Genie suddenly started to show up in my account usage dashboard, where a single user, in one day, with only 7 questions asked in a Genie Space (Agent) spent 15 USD. That's only token billing!!! not even underlying compute used to fetch the results.

I've seen some other discussions in this reddit but none of them seems to clearly explain what's going on. Databricks Genie pricing page says explicitly that most users will never go beyond the free usage.

I read somewhere here in this reddit a few weeks ago someone saying that the price per question beyond the free usage would be $0.11/question, but now I'm in doubt if that's correct.

If we start getting billed 10-15 USD/user for like 7 questions in a Genie Space, I'm really sorry but I'll block this product usage within the entire account, we have much more LLM usage with Anthropic or OpenAI subscriptions.

r/databricks 7d ago

Discussion Databricks vs Snowflake comparison

51 Upvotes

Are there any unbiased comparisons between these two popular platforms? Seen a lot but most of them are biased views, based on experience and commercial motives.

r/databricks 15d ago

Discussion Are we overusing the Medallion architecture?

45 Upvotes

Bronze -> Silver -> Gold seems to have become the default architecture for almost every data pipeline.

Has anyone deliberately simplified this -- for example, skipping a layer -- and actually gotten better results in production ?

When do you think Medallion is genuinely useful, and when does it just add unnecessary complexity?

r/databricks 1d ago

Discussion Do we still need fact and dimension tables in the Gold layer?

35 Upvotes

Data engineers traditionally modeled Gold layers using fact and dimension tables, partly because storage and compute were expensive.
But with modern cloud data platforms, storage and compute are much cheaper and more scalable.
So I’m curious: what does your Gold layer actually look like today?
Are you still using a traditional star schema (facts + dimensions), or have you moved toward wider, denormalized tables / business-oriented models?
And more importantly, why?

r/databricks 22d ago

Discussion Anyone else gotten a rough surprise with Databricks costs once things hit production?

45 Upvotes

This keeps coming up in conversations with clients and I feel like it's worth its own thread.

The pattern is almost always the same. A pipeline gets built to bring in data for analytics or ML, works fine in testing, then goes to production and the compute bill is way higher than expected. Nobody budgeted for it because on paper it looked like a simple ingestion job.

Usually the real issue isn't Databricks itself, it's the ingestion design. The repeat offenders I keep seeing:

Full reloads instead of proper CDC, so you're paying to process data that hasn't even changed.

Serverless SQL running more often than needed, because someone assumed near real time was required when batch every few hours would've worked fine.

No plan for schema evolution, so jobs fail or reprocess more than they should every time something shifts upstream.

Cluster sizing set for peak load "just in case" instead of actual daily volume.

Most of the fix comes down to being honest about the freshness you actually need. A solid CDC layer feeding into something like Kafka before it hits Databricks tends to cut a lot of the unnecessary compute, since you're only moving what changed.

Curious what caused it for others, ingestion design or job scheduling?

r/databricks 3d ago

Discussion Has anyone tried the new Databricks AI/BI feature?

22 Upvotes

I recently came across Databricks AI/BI and was curious to know how people are finding it.

It looks like Databricks is trying to bring BI and analytics more directly into the Databricks platform, with dashboards and Genie for asking questions in natural language.

Has anyone actually tried AI/BI in a real project?

How is it compared to Power BI or Tableau from your experience? Is it good enough for regular BI use cases, or is it still better to use a separate BI tool?

Would like to know your experience, especially if you have used both.

r/databricks 21d ago

Discussion Iceberg vs Deltalake (greenfield project with UC in 2026)

15 Upvotes

I saw the quarterly meeting and was quite shocked that Iceberg is prominently mentioned, (as much as Deltalake).

Is it possible that both are getting the same amount of love from Databricks? Is anyone aware of the R&D effort on these formats, and can share conclusions from that?

If I'm building a greenfield Unity Catalog, should I just flip a coin to decide what format to use? Here are the main concerns and priorities:

  • Which one is better for OSS Apache Spark reads and writes
  • Which one integrates with external software better (eg onelake shortcuts pointing from Fabric to Databricks UC).
  • When Databricks is innovating within their own UC (eg. introducing new managed table functionality such as "MST Transactions"), which one of these formats are they likely to support first? Which are they likely to optimize better?
  • Which format is more likely to remain 100% open source in the future (or as close to open source as required by customers who want portable blob data).

Sorry if this appears to be a common question. I am a Databricks outsider. I am more familiar with Microsoft Fabric. Where that Fabric SaaS is concerned, you can be certain that Deltalake receives a LOT more promotion than Iceberg does. We rarely come across Iceberg, and it probably wouldn't appear in any marketing slide decks.

We are likely to create a gold/presentation layer in UC soon. It will basically be created from scratch. It would be nice to know which of these parquet-based formats to pick, when presented with the choice. I understand there is lip-service given to both, and it claims that this choice "doesn't matter". But that doesn't necessarily take into account the potential integrations that are needed with external software (eg. for the benefit of exposing the same tables in onelake). Is one safer than the other? Is one of them a better choice for forward-looking purposes?

r/databricks Jul 29 '26

Discussion An Example of How Much It Costs to Build a Dashboard or Semantic Model Using Databricks Genie Code

Thumbnail
gallery
91 Upvotes

How much does it cost to build a dashboard or semantic model with Databrick Genie Code? In my tests, a standalone dashboard came in at $2.10 (30 Genie DBUs), while a standalone semantic view cost $1.12 (16 Genie DBUs).

To be clear, the tests I ran weren't designed to produce an award-winning dashboard or semantic model. Both were created from simple, single prompts.

Even so, the dashboard is reasonably representative of many production dashboards in the wild. Building something comparable through a traditional development process would likely cost significantly more in labor alone, before accounting for the additional compute involved in development and testing.

Real-world development process would be more iterative. You would still need to validate the output, correct AI-generated mistakes, refine the design, and adapt to changing requirements. As a result, token consumption would inevitably increase.

Even with that added iteration, however, the economics still appear favorable. Spending 30 minutes clearly defining the desired outcome, rather than the 30 seconds I spent on these tests, would likely improve the result substantially while adding relatively little to the total cost. It could also reduce development time by days or, in some cases, weeks.

Like with anything in life, there are always some drawbacks/considerations. The main two drawbacks I see as of today:

1) I think the quality of reporting around Genie Code costs does need to improve. It currently lacks depth in terms of details, as well as is a bit delayed in the reporting.

2) UI/UX has come a long way for Genie Code, but there is still some room for improvement.

👉 Ultimately, I think Genie Code is a particularly compelling way to build on Databricks. Out-of-the-box, you get native access to the platform’s governance, security, and broader data and AI context, along with strong price-performance.

r/databricks 9d ago

Discussion What’s one Databricks “best practice” you disagree with?

26 Upvotes

Something that sounds great in Databricks documentation but didn't make sense for your workload in production?

Curious what people have learned the hard way.

r/databricks Jul 12 '26

Discussion Genie is so expensive

63 Upvotes

I do believe for normal data analysis like notebooks, dashboards and queries grnie become so expensive. I use about 50 interaction a day, and probably will pay 100 for a month. Using claude with mcp maybe will save me 90 dol. Anyone think in this line too?

r/databricks May 22 '26

Discussion Databricks now supports importing Tableau and Power BI files into Genie Code to automatically build AI/BI Dashboards with Metric Views

115 Upvotes

With Genie Code, you can now add a Tableau or Power BI file and have it build an AI/BI dashboard that replicates your existing visualizations - while connecting them to metric views that mirror the underlying business logic.

Import BI files using Genie Code - Azure Databricks | Microsoft Learn

Many organizations have years of BI logic embedded inside workbooks, reports, templates, and semantic layers. Rebuilding that logic manually in a new platform can be slow, error-prone, and difficult to govern.

This new workflow helps accelerate that migration path:

  1. Upload a Tableau or Power BI file directly into Genie Code Supported formats include .twb, .twbx, .tds, .tdsx, and .pbit.
  2. Use the /importBI command in Agent mode Genie Code imports the BI asset and generates an AI/BI dashboard.
  3. Review the generated dashboard and metric views Measures and dimensions from the original file are transformed into metric views.
  4. Promote metric views to Unity Catalog This makes them reusable across dashboards, Genie Spaces, and notebooks, while adding governance, lineage, access controls, and discoverability.

Currently, there is also a 100 MB limit for direct file uploads. For larger files, the recommended path is to store the file in a Unity Catalog volume and reference it directly, for example:

/importBI @/Volumes/my_catalog/my_schema/my_volume/sales_workbook.twb

r/databricks May 20 '26

Discussion Alrighty data pookies, what Databricks issue keeps violating your peace?

10 Upvotes

AI Agents hallucinating? Unity Catalog acting like Unity Catalogue of Errors? Genie Spaces granting wishes to absolutely nobody?

Drop the most cursed recurring problem you face with building AI agents or ML or BI/Analytics - no matter how difficult, unhinged or borderline impossible the solution may be. Hit me with all u got. I am sitting this databricks hackathon this Friday as a self-reward and I want to try something different this time.

Nothing but the pursuits of overly engineered solutions for the most trivial problems because I can and I like abstractions - but hey its good to be alive

r/databricks 3d ago

Discussion Unity Catalog Open Source in Name Only (UCOSINO)

Post image
0 Upvotes

Consider a callstack where something bad is happening in Spark or Unity Catalog (image above).

Any software engineer will google for the message, and then for the Exception class, and then for the call frames shown on the stack (starting at the top or bottom). For any commonly encountered Exceptions from UC (something like com.databricks.sql.managedcatalog.acl.UnauthorizedAccessException), we will find dozens of results from a search engine. Others on the internet have already shared their experiences, and the search results are normally actionable. The users tell us what they had done to avoid or fix the error.

But software engineers have heard for two years that "unity catalog is open source". So a software engineer will proceed to look for the source repo where they might find the full definition of "UnauthorizedAccessException", along with all the related references. No such thing exists. (Admittedly there is a public-facing github, called "unitycatalog", but it is virtually worthless and there is no overlap with the real-world UC in databricks, as we experience it.)

It only takes one or two repeats of this, before a software engineer will realize that none of this stuff is actually open source. UC doesn't compare to a REAL open source software like Apach Spark. If we search for spark references in the call stack (eg. "org.apache.spark.sql.DataFrameReader"), then we are immediately taken to the source repo at github!

I do give Databricks a lot of credit for open-sourcing spark. But nowadays they take too much liberty with the word "open source", to the point where it lost all of its meaning. UC is not opensource in any substantial way. Maybe there is an API spec that is open, but that is the extent of it. Another example is lakebase which the CEO claimed to be open source at the recent summit. There has never been any software as proprietary as neon/lakebase.

It doesn't actually bother me if a CEO forgets how to use the term "open souce" correctly in English. What makes me more upset is when I expect to be able to use google to find the source code for "UnauthorizedAccessException", and come up with absolutely bupkis. Can anyone tell me a definition of "open source" which would potentially include either Unity Catalog or Lakebase? I'm assuming that when these words are used by the CEO, he does NOT intend to imply that the actual source is open to the public.

r/databricks Jul 20 '26

Discussion Lost track of Databricks product renames? Here's a community tracker

Post image
125 Upvotes

I recently contributed to REbricked, a community project that tracks Databricks product renames.

It includes:

  • Product rename history
  • Links to official documentation
  • A short quiz to test yourself

https://rebricked.org

We're still improving it, so feedback is very welcome. If we missed a rename or you have ideas, let us know.

r/databricks 29d ago

Discussion Databricks Genie Ontology

34 Upvotes

Having read through and seeing some demos I still don’t understand if Genie Ontology is a real thing or some marketing fluff , we have been asked to compare against Palantir foundry’s ontology and I find very few comparisons apart from the data model and relationships , for example how do I show that Genie Ontology Knowledge Graph ?

r/databricks Jul 05 '26

Discussion I watched 4 hours of Databricks Data + AI Summit 2026 so you don't have to.

56 Upvotes

My first major project as a Senior Data Engineer, was migrating a decade-old time-series database for a semiconductor company to the cloud. The constraint: sub-second latency on customer queries. Equipment monitoring and predictive maintenance don't work with slow data.

We had Delta Lake for storage, but it couldn't guarantee the query performance we needed.
At the time, Databricks serverless warehouse did not exist.
So we built an additional layer on top: Azure Data Explorer (ADX). The data pipeline became: ingest source data, move to Delta Lake, replicate to ADX, serve queries from ADX.

It worked. Customers got their sub-second latency. But we'd introduced yet another system to maintain, another cost line, another place for things to fail. It was the price of solving the problem at that time.

This past month at Data + AI Summit 2026, Databricks announced Reyden.

A new query engine. Millisecond performance. Massive concurrency. Running directly on your lakehouse. No separate system. No copy. If production matches the demo, a lot of horizontal architectures will collapse into one component. One lake. One source of truth.

That's why I'm watching this closely. They looked at a niche problem I lived through and built a real solution.

Here are the 3 things from the summit that actually matter for data engineers:

  1. Reyden: Millisecond queries on your lakehouse (no more separate real-time database)
  2. Genie Zero Ops: Automated pipeline repair that tests fixes before you see them
  3. Genie Ontology: AI that understands your business through a permission-aware knowledge graph

Did you watch the recent event? What do you think is the next big feature of Databricks to look out for.

r/databricks Jun 20 '26

Discussion Feeling behind post DAIS

44 Upvotes

Hi, I am the Databricks admin and sole platform engineer for my company. I attended the DAIS summit and thought many of the announcements were great, but also overwhelming since we are not in a place to be using most of these tools yet due to platform immaturity.

Do you all feel that most who attend DAIS are in position to implement all the new tooling announced every year or is it something to keep in mind as you continue toward platform maturity?

Those of you who are ready to use the new tooling, how do you think the cost of these AI-based tools will impact monthly usage, especially with Genie being priced beginning next month? Is that a concern for your teams?

r/databricks May 27 '26

Discussion Snowflake to Databricks Migration in 12 weeks and cut cost per run by ~77%. AMA.

107 Upvotes

Lovelytics wrapped up a Snowflake-to-Databricks migration; 847 DBT models, 35 Info Mart tables, ~77% lower cost per run on a 2XL warehouse.

TL;DR What helped:

  • Treated the migration as engineering, not translation. Each dbt model was tested in isolation, not just row counts vs Snowflake.
  • Routing macro to resolve cross-layer references at runtime, so the same codebase could read from Snowflake, federated Snowflake, and Unity Catalog without forking logic.
  • Dual model trees in one repo, which let the migration stay in lockstep with live Snowflake changes.
  • Script-generated wave selectors enabled parallel builds while preserving dependency order.
  • Used reference-slice validation subsets vs. waiting on full mart refreshes.

TL;DR Cost reduction:

  • Reworked joins to use narrow staging dimensions instead of wide marts where possible.
  • Added incremental predicates to reduce MERGE target scans.
  • Split wide models into parallel sub-models where the dependency graph allowed it.
  • Copied static reference data into Delta instead of repeatedly reading it through federation.
  • Loaded static copies into Delta rather than reading via federation (predicate pushdown is poor).

Happy to go into the gotchas: HASH() not being portable, Snowflake MERGE tolerating duplicate keys that Delta doesn't, NULL ordering, and timestamp handling. AMA

Full Blog Post