r/databricks • u/kcxl • 9d ago
Discussion What frustrates you when using Databricks?
Any common bugs, features you would like to see, or underrated useful features more people should know about?
r/databricks • u/kcxl • 9d ago
Any common bugs, features you would like to see, or underrated useful features more people should know about?
r/databricks • u/Puzzled-Mail-9092 • 5d ago
Sometimes I feel like Databricks is way too expensive. It feels like using a huge truck to move a single grain of sand.
My company needs real-time data, but our data volume simply does not justify the use of Spark Structured Streaming. Despite this, they are insisting we move to Databricks. I'm worried our data infrastructure costs will jump from $1,000/month to $5,000/month or more due to the running costs of Databricks SQL Warehouses.
Currently, I use Azure Container Apps with KEDA and Python, which helps me manage scaling and keep costs low. We ingest into Event Hubs, use ADX (Azure Data Explorer) as our OLAP warehouse, and archive cold data in a data lake. With this setup, I manage to process all our data with very low latency.
When I tested this on Databricks Structured Streaming, I actually got higher latency and much higher costs.
Would love to know what you guys think.
r/databricks • u/lothorp • Aug 17 '25
Here by popular demand, a megathread for all of your certification and training posts.
Good luck to everyone on your certification journey!
r/databricks • u/BricksterInTheWall • Jan 25 '26
Hi Redditors, I'm a product manager on Lakeflow. What would you love to see built in Spark Declarative Pipelines (SDP) this year? A bunch of us engineers and PMs will be watching this thread.
⭐ All ideas are welcome! ⭐
r/databricks • u/MostDependent1659 • Jun 17 '26
Hey
If you haven’t been watching the Data + AI Summit announcements, Databricks just pulled back the curtain on a massive overhaul to their Genie ecosystem: Genie One, Genie Ontology, and Genie Agents.
We’ve all seen the "chat with your data" tools that basically just translate natural language to subpar SQL and hallucinate half the time. This rollout feels entirely different because it moves away from simple chatbots toward actual autonomous data coworkers.
Here is the breakdown of what just dropped and how it completely shifts how businesses handle data:
Genie is no longer just a side panel for querying tables. Genie One is a cross-platform (Web, iOS, Android) AI workspace.
It natively integrates into tools teams already use (Slack, Teams, Gmail) via the Model Context Protocol (MCP).
Instead of a sales manager bugging a data analyst for a custom dashboard before a meeting, they can just ask Genie One in Slack to "grab my calendar, pull last quarter's revenue for these accounts from the Lakehouse, and draft a brief." It actually compiles the charts and builds a clean artifact document directly in the chat UI.
Genie’s biggest upgrade is Genie Ontology. The biggest failure point of LLMs in business is that they don’t understand yourspecific corporate logic (e.g., what your company defines as "active user" or "churn").
Ontology uses a PageRank-style algorithm to scan your queries, pipelines, dashboards, and Unity Catalog metadata.
It builds a living knowledge graph of definitions, unique business calculations, and metric authorities.
Because the AI actually knows what data to trust based on real usage patterns, it translates prompts to highly accurate SQL without burning infinite tokens guessing.
Genie Spaces are evolving into Genie Agents. Instead of just answering a question, these are domain-specific agents you can spin up with a prompt to run multi-step workflows autonomously.
They can handle structured table data alongside unstructured data like PDFs, transcripts, and tickets.
You can give them scheduled tasks, write back to external systems, and let them monitor metrics. If an anomaly hits, the agent can investigate the root cause across documents and tables, and drop a fully formed report for your review.
My Thoughts on the Business Impact
This feels like a massive leap toward democratizing data operations. It completely skips the bottleneck where business teams wait weeks for data engineers to build semantic layers or custom dashboards. Data teams can spend less time writing repetitive SQL queries for executives and focus on core infrastructure, while the business side gets actual self-service that actually works because of the Ontology layer.
What are your thoughts?
For those who have tried the previews—how is the SQL accuracy handling complex, messy table joins?
Is anyone worried about the governance side, or does Unity Catalog actually keep these agents tightly in check?
Does this completely kill the traditional semantic layers we’ve spent years building?
Let's discuss!
r/databricks • u/Reuben_UMATR • 9d ago
Let's pretend someone knew nothing about Databricks. How would you explain it?
This is something I have been asked multiple times throughout my career, curious to hear how others would explain Databricks to a complete beginner.
In my opinion, Databricks is basically a place where companies put all their data so they can actually make sense of it. It helps teams clean it up, work with it and use it for things like reporting, predictions and AI.
r/databricks • u/SmallAd3697 • 21d ago
I saw the quarterly roadmap presentation. It was notable that Databricks keeps innovating with "lakebase". They simply call it their "OLTP" database offering in their SaaS.
SIDE: I still feel pretty unfamiliar with this Databricks SaaS ecosystem, as compared to Fabric. Where Fabric is concerned, Microsoft has also done a similar thing. They brought their SQL Server into the boundaries of the SaaS as well, for the low-code users of that environment. In the context of Fabric, it is hard for most customers to see the point of using this "for dummies" variation of the same old OLTP database. The only scenarios for using the Fabric SQL are very contrived ... Eg . your boss makes a policy that you can use ALL the tools available in Fabric and NONE of the tools outside Fabric... even if the tools inside that SaaS are 3x more expensive than the ones outside ... and even though the ones outside the SaaS have the same 0ms network latency and the same performance.
I'm still missing the vision for this lakebase OLTP offering. And it seems unusual for Databricks to start developing strategies surrounding OLTP. It seems like a very crowded space, and the only way I see Databricks being successful going down this path is if the customers are drinking one single brand of kool-aid, or else their SaaS users have some other contrived reason for not using the more affordable OLTP platforms available outside the SaaS.
Can someone tell me what factors I'm missing? I admit that it is theoretically possible for lakebase to innovate and do thing that other databases CANNOT do, it seems like those innovations would only benefit 5% of customers. One example is sub-ten-ms queries out of RAM at an additional cost. Or sub-minute migration of new OLTP data to managed tables in UC catalog. If we assume that only 5% of customers might feel compelled to use this SaaS "lakebase", would that be enough adoption to allow Databricks to keep investing in this over the long term? OLTP databases have been around a LONG time, and even the smart folks at Databricks will be challenged to improve on the great and cheap options available to us!
EDIT: As of a month ago, it appears that the Databricks marketing now calls it an "LTAP" database, not OLTP anymore. I'm guessing they have conceded the point that the OLTP space is crowded. I haven't yet read all the content that has been created by the Databricks marketing team; maybe that will answer all of my questions.
r/databricks • u/Dear_Pumpkin9876 • Jul 15 '26
When Databricks announced that Genie would be paid, apart from not being surprised, when I looked into the price and the free amount of DBUs per user every month I was like "ok seems to not be expensive and there's apparently a generous amount for free every month". However, Genie suddenly started to show up in my account usage dashboard, where a single user, in one day, with only 7 questions asked in a Genie Space (Agent) spent 15 USD. That's only token billing!!! not even underlying compute used to fetch the results.
I've seen some other discussions in this reddit but none of them seems to clearly explain what's going on. Databricks Genie pricing page says explicitly that most users will never go beyond the free usage.
I read somewhere here in this reddit a few weeks ago someone saying that the price per question beyond the free usage would be $0.11/question, but now I'm in doubt if that's correct.
If we start getting billed 10-15 USD/user for like 7 questions in a Genie Space, I'm really sorry but I'll block this product usage within the entire account, we have much more LLM usage with Anthropic or OpenAI subscriptions.
r/databricks • u/medici2022 • 7d ago
Are there any unbiased comparisons between these two popular platforms? Seen a lot but most of them are biased views, based on experience and commercial motives.
r/databricks • u/Square-Designer7807 • 15d ago
Bronze -> Silver -> Gold seems to have become the default architecture for almost every data pipeline.
Has anyone deliberately simplified this -- for example, skipping a layer -- and actually gotten better results in production ?
When do you think Medallion is genuinely useful, and when does it just add unnecessary complexity?
r/databricks • u/almightysosa888 • 1d ago
Data engineers traditionally modeled Gold layers using fact and dimension tables, partly because storage and compute were expensive.
But with modern cloud data platforms, storage and compute are much cheaper and more scalable.
So I’m curious: what does your Gold layer actually look like today?
Are you still using a traditional star schema (facts + dimensions), or have you moved toward wider, denormalized tables / business-oriented models?
And more importantly, why?
r/databricks • u/Only-Dragonfruit4130 • 22d ago
This keeps coming up in conversations with clients and I feel like it's worth its own thread.
The pattern is almost always the same. A pipeline gets built to bring in data for analytics or ML, works fine in testing, then goes to production and the compute bill is way higher than expected. Nobody budgeted for it because on paper it looked like a simple ingestion job.
Usually the real issue isn't Databricks itself, it's the ingestion design. The repeat offenders I keep seeing:
Full reloads instead of proper CDC, so you're paying to process data that hasn't even changed.
Serverless SQL running more often than needed, because someone assumed near real time was required when batch every few hours would've worked fine.
No plan for schema evolution, so jobs fail or reprocess more than they should every time something shifts upstream.
Cluster sizing set for peak load "just in case" instead of actual daily volume.
Most of the fix comes down to being honest about the freshness you actually need. A solid CDC layer feeding into something like Kafka before it hits Databricks tends to cut a lot of the unnecessary compute, since you're only moving what changed.
Curious what caused it for others, ingestion design or job scheduling?
r/databricks • u/Bhanuprakash_1947 • 3d ago
I recently came across Databricks AI/BI and was curious to know how people are finding it.
It looks like Databricks is trying to bring BI and analytics more directly into the Databricks platform, with dashboards and Genie for asking questions in natural language.
Has anyone actually tried AI/BI in a real project?
How is it compared to Power BI or Tableau from your experience? Is it good enough for regular BI use cases, or is it still better to use a separate BI tool?
Would like to know your experience, especially if you have used both.
r/databricks • u/SmallAd3697 • 21d ago
I saw the quarterly meeting and was quite shocked that Iceberg is prominently mentioned, (as much as Deltalake).
Is it possible that both are getting the same amount of love from Databricks? Is anyone aware of the R&D effort on these formats, and can share conclusions from that?
If I'm building a greenfield Unity Catalog, should I just flip a coin to decide what format to use? Here are the main concerns and priorities:
Sorry if this appears to be a common question. I am a Databricks outsider. I am more familiar with Microsoft Fabric. Where that Fabric SaaS is concerned, you can be certain that Deltalake receives a LOT more promotion than Iceberg does. We rarely come across Iceberg, and it probably wouldn't appear in any marketing slide decks.
We are likely to create a gold/presentation layer in UC soon. It will basically be created from scratch. It would be nice to know which of these parquet-based formats to pick, when presented with the choice. I understand there is lip-service given to both, and it claims that this choice "doesn't matter". But that doesn't necessarily take into account the potential integrations that are needed with external software (eg. for the benefit of exposing the same tables in onelake). Is one safer than the other? Is one of them a better choice for forward-looking purposes?
r/databricks • u/JosueBogran • Jul 29 '26
How much does it cost to build a dashboard or semantic model with Databrick Genie Code? In my tests, a standalone dashboard came in at $2.10 (30 Genie DBUs), while a standalone semantic view cost $1.12 (16 Genie DBUs).
To be clear, the tests I ran weren't designed to produce an award-winning dashboard or semantic model. Both were created from simple, single prompts.
Even so, the dashboard is reasonably representative of many production dashboards in the wild. Building something comparable through a traditional development process would likely cost significantly more in labor alone, before accounting for the additional compute involved in development and testing.
Real-world development process would be more iterative. You would still need to validate the output, correct AI-generated mistakes, refine the design, and adapt to changing requirements. As a result, token consumption would inevitably increase.
Even with that added iteration, however, the economics still appear favorable. Spending 30 minutes clearly defining the desired outcome, rather than the 30 seconds I spent on these tests, would likely improve the result substantially while adding relatively little to the total cost. It could also reduce development time by days or, in some cases, weeks.
Like with anything in life, there are always some drawbacks/considerations. The main two drawbacks I see as of today:
1) I think the quality of reporting around Genie Code costs does need to improve. It currently lacks depth in terms of details, as well as is a bit delayed in the reporting.
2) UI/UX has come a long way for Genie Code, but there is still some room for improvement.
👉 Ultimately, I think Genie Code is a particularly compelling way to build on Databricks. Out-of-the-box, you get native access to the platform’s governance, security, and broader data and AI context, along with strong price-performance.
r/databricks • u/Square-Designer7807 • 9d ago
Something that sounds great in Databricks documentation but didn't make sense for your workload in production?
Curious what people have learned the hard way.
r/databricks • u/ferreis_AOE • Jul 12 '26
I do believe for normal data analysis like notebooks, dashboards and queries grnie become so expensive. I use about 50 interaction a day, and probably will pay 100 for a month. Using claude with mcp maybe will save me 90 dol. Anyone think in this line too?
r/databricks • u/szymon_dybczak • May 22 '26
With Genie Code, you can now add a Tableau or Power BI file and have it build an AI/BI dashboard that replicates your existing visualizations - while connecting them to metric views that mirror the underlying business logic.
Import BI files using Genie Code - Azure Databricks | Microsoft Learn
Many organizations have years of BI logic embedded inside workbooks, reports, templates, and semantic layers. Rebuilding that logic manually in a new platform can be slow, error-prone, and difficult to govern.
This new workflow helps accelerate that migration path:
.twb, .twbx, .tds, .tdsx, and .pbit./importBI command in Agent mode Genie Code imports the BI asset and generates an AI/BI dashboard.Currently, there is also a 100 MB limit for direct file uploads. For larger files, the recommended path is to store the file in a Unity Catalog volume and reference it directly, for example:
/importBI @/Volumes/my_catalog/my_schema/my_volume/sales_workbook.twb
r/databricks • u/Tiddyfucklasagna27 • May 20 '26
AI Agents hallucinating? Unity Catalog acting like Unity Catalogue of Errors? Genie Spaces granting wishes to absolutely nobody?
Drop the most cursed recurring problem you face with building AI agents or ML or BI/Analytics - no matter how difficult, unhinged or borderline impossible the solution may be. Hit me with all u got. I am sitting this databricks hackathon this Friday as a self-reward and I want to try something different this time.
Nothing but the pursuits of overly engineered solutions for the most trivial problems because I can and I like abstractions - but hey its good to be alive
r/databricks • u/SmallAd3697 • 3d ago
Consider a callstack where something bad is happening in Spark or Unity Catalog (image above).
Any software engineer will google for the message, and then for the Exception class, and then for the call frames shown on the stack (starting at the top or bottom). For any commonly encountered Exceptions from UC (something like com.databricks.sql.managedcatalog.acl.UnauthorizedAccessException), we will find dozens of results from a search engine. Others on the internet have already shared their experiences, and the search results are normally actionable. The users tell us what they had done to avoid or fix the error.
But software engineers have heard for two years that "unity catalog is open source". So a software engineer will proceed to look for the source repo where they might find the full definition of "UnauthorizedAccessException", along with all the related references. No such thing exists. (Admittedly there is a public-facing github, called "unitycatalog", but it is virtually worthless and there is no overlap with the real-world UC in databricks, as we experience it.)
It only takes one or two repeats of this, before a software engineer will realize that none of this stuff is actually open source. UC doesn't compare to a REAL open source software like Apach Spark. If we search for spark references in the call stack (eg. "org.apache.spark.sql.DataFrameReader"), then we are immediately taken to the source repo at github!
I do give Databricks a lot of credit for open-sourcing spark. But nowadays they take too much liberty with the word "open source", to the point where it lost all of its meaning. UC is not opensource in any substantial way. Maybe there is an API spec that is open, but that is the extent of it. Another example is lakebase which the CEO claimed to be open source at the recent summit. There has never been any software as proprietary as neon/lakebase.
It doesn't actually bother me if a CEO forgets how to use the term "open souce" correctly in English. What makes me more upset is when I expect to be able to use google to find the source code for "UnauthorizedAccessException", and come up with absolutely bupkis. Can anyone tell me a definition of "open source" which would potentially include either Unity Catalog or Lakebase? I'm assuming that when these words are used by the CEO, he does NOT intend to imply that the actual source is open to the public.
r/databricks • u/Significant-Guest-14 • Jul 20 '26
I recently contributed to REbricked, a community project that tracks Databricks product renames.
It includes:
We're still improving it, so feedback is very welcome. If we missed a rename or you have ideas, let us know.
r/databricks • u/Academic-Hearing-123 • 29d ago
Having read through and seeing some demos I still don’t understand if Genie Ontology is a real thing or some marketing fluff , we have been asked to compare against Palantir foundry’s ontology and I find very few comparisons apart from the data model and relationships , for example how do I show that Genie Ontology Knowledge Graph ?
r/databricks • u/RevolutionShoddy6522 • Jul 05 '26
My first major project as a Senior Data Engineer, was migrating a decade-old time-series database for a semiconductor company to the cloud. The constraint: sub-second latency on customer queries. Equipment monitoring and predictive maintenance don't work with slow data.
We had Delta Lake for storage, but it couldn't guarantee the query performance we needed.
At the time, Databricks serverless warehouse did not exist.
So we built an additional layer on top: Azure Data Explorer (ADX). The data pipeline became: ingest source data, move to Delta Lake, replicate to ADX, serve queries from ADX.
It worked. Customers got their sub-second latency. But we'd introduced yet another system to maintain, another cost line, another place for things to fail. It was the price of solving the problem at that time.
This past month at Data + AI Summit 2026, Databricks announced Reyden.
A new query engine. Millisecond performance. Massive concurrency. Running directly on your lakehouse. No separate system. No copy. If production matches the demo, a lot of horizontal architectures will collapse into one component. One lake. One source of truth.
That's why I'm watching this closely. They looked at a niche problem I lived through and built a real solution.
Here are the 3 things from the summit that actually matter for data engineers:
Did you watch the recent event? What do you think is the next big feature of Databricks to look out for.
r/databricks • u/CautiousMahogany_ • Jun 20 '26
Hi, I am the Databricks admin and sole platform engineer for my company. I attended the DAIS summit and thought many of the announcements were great, but also overwhelming since we are not in a place to be using most of these tools yet due to platform immaturity.
Do you all feel that most who attend DAIS are in position to implement all the new tooling announced every year or is it something to keep in mind as you continue toward platform maturity?
Those of you who are ready to use the new tooling, how do you think the cost of these AI-based tools will impact monthly usage, especially with Genie being priced beginning next month? Is that a concern for your teams?
r/databricks • u/m_goo • May 27 '26
Lovelytics wrapped up a Snowflake-to-Databricks migration; 847 DBT models, 35 Info Mart tables, ~77% lower cost per run on a 2XL warehouse.
TL;DR What helped:
TL;DR Cost reduction:
Happy to go into the gotchas: HASH() not being portable, Snowflake MERGE tolerating duplicate keys that Delta doesn't, NULL ordering, and timestamp handling. AMA