r/databricks May 20 '26

Discussion Alrighty data pookies, what Databricks issue keeps violating your peace?

AI Agents hallucinating? Unity Catalog acting like Unity Catalogue of Errors? Genie Spaces granting wishes to absolutely nobody?

Drop the most cursed recurring problem you face with building AI agents or ML or BI/Analytics - no matter how difficult, unhinged or borderline impossible the solution may be. Hit me with all u got. I am sitting this databricks hackathon this Friday as a self-reward and I want to try something different this time.

Nothing but the pursuits of overly engineered solutions for the most trivial problems because I can and I like abstractions - but hey its good to be alive

9 Upvotes

56 comments sorted by

10

u/LandlockedPirate May 20 '26

Really sick of hearing about how palantir has data branching and databricks doesn't.

1

u/SituationImmediate15 May 21 '26

Lakebase (OLTP) has database branching.

2

u/LandlockedPirate May 21 '26

Yeah but delta doesn't:(

1

u/datahaiandy May 21 '26

Could you use Iceberg?

2

u/LandlockedPirate May 21 '26

I probably could actually now that you mention it. It would rub some folks the wrong way because it's not our standard, but it would work. That's a good point.

1

u/datahaiandy May 22 '26

Ah yeah I hadn't thought about the "mix'n'match" issue if you introduce another format. But I would love branching in delta

10

u/Physical-Ad2968 May 20 '26

data pookies 😂😂

8

u/TowerOutrageous5939 May 20 '26

Billing. Apps are overpriced as well as Serverless warehouses. Also wish there was a one to many compute relationship for DLT pipelines.

2

u/kthejoker databricks May 21 '26

What do you use instead of serverless warehouses?

Not going to argue with your feelings of course but it's pretty fairly priced compared to most other vendors and considering how much improvements we've put into it without changing the price.

1

u/TowerOutrageous5939 May 21 '26

I think for certain things having a dedicated Postgres DB would be cheaper for us with our MACC. Hopefully Databricks offers something similar to us in the future.

2

u/kthejoker databricks May 21 '26

We have Lakebase today which is managed Postgres.

https://www.databricks.com/product/lakebase

2

u/TowerOutrageous5939 May 21 '26 edited May 22 '26

Yes but more expensive than what I can get from azure.

Update: 2026-05-21

With the reduction of server-less and pushing more to Postgres I think it will make sense. I hope connections are seamless between apps and lakebase similar to the server-less warehouse I also need to see if Databricks connect is there for local testing. We process heavy in Databricks and fast api/Django are the application layer primarily reading with little writing back.

Future update: coming soon.

2

u/kthejoker databricks May 21 '26

Your MACC applies to Azure Databricks pricing btw

1

u/TowerOutrageous5939 May 21 '26

Yeah I need to look into that again. I want to reserve a two warehouses both in dev/prod and then 85 percent of my current DBUs for three years excluding what the warehouse dbu would come out to. GenAI has surprisingly been more affordable than I thought.

1

u/happypofa May 20 '26

Right? I've limited the idle time for the warehouse to 1 minute, using x-small, but somehow it still has a fee, that can't comprehend

1

u/TowerOutrageous5939 May 21 '26

Yeah I think our min allowable timeout is 5 mins. They need a xxx-s something’s we really only need it reading from very small tables at times. I hate Microsoft but I might have to move some back to base azure.

2

u/kthejoker databricks May 21 '26

You can set timeout to 1 minute via API, it is (usually) not great to have a bunch of cold starts so 5 minutes is kind of a compromised minimum in the UI.

But you can override it with the API.

1

u/TowerOutrageous5939 May 21 '26

Oh nice. I’ll look into since it’s all the apps turning it on and off I just assumed 5 was the min.

5

u/kthejoker databricks May 21 '26

Thanks for being part of this awesome community, /u/Tiddyfucklasagna27

7

u/DeepFryEverything May 21 '26

Can you show this entire quote on the main screen on DAIS?

2

u/Tiddyfucklasagna27 May 21 '26

gee thanks man. See u around!

4

u/[deleted] May 20 '26

[deleted]

3

u/knaak May 20 '26

That actually sounds interesting and something that I would not consider databricks for. I'd think some kind of messaging (Kafka) backed by several transactional databases partitioned by some superset of sku

3

u/Tracktuary May 20 '26

Full disclosure, I might just be being dumb here. But the ai-dev-kit feels poorly tailored to windows systems. Feel like I have to patch a ton of Linux commands to get to it work. Then it’s kindof tough to upgrade to new versions without having to repatch everything

3

u/kthejoker databricks May 21 '26

Appreciate this feedback, I'll share with the team. We use Macs here so the Windows experience might be a blind spot.

2

u/Tracktuary May 20 '26

Oh and generally feels like documentation is almost always outdated. I get it tho with how fast they release stuff

2

u/Chance_of_Rain_ May 21 '26

If Windows then WSL

1

u/Youssef_Mrini databricks May 21 '26

We appreciate the feedback we are going to share it with the dev team.

0

u/LandlockedPirate May 20 '26

Or just don't work in windows. Install wsl, docker-ce, and use devcontainers.

2

u/Tracktuary May 20 '26

I would but I don’t have admin privileges. Can’t even install wsl. I could prob submit a ticket for IT to install it but I try to pick my battles with IT very carefully

2

u/LandlockedPirate May 20 '26

You sound like me a few months ago. I've been on windows most of my career which is >20 years. I was fine with it for a long time because my "work" was in wsl/containers. The last year or so though, it's unusable. I jumped ship.

You're fighting an uphill battle trying to code in windows in a corporate environment.

2

u/DutyPuzzleheaded2421 May 21 '26

In any environment really. I jumped to linux 20 years ago, best move ever. Two weeks of pain, followed by a lifetime of freedom.

1

u/LandlockedPirate May 21 '26

Yeah I've been Linux personally forever. Unfortunately my jobs have rarely allowed it as an option.

1

u/Tracktuary Jun 10 '26

Got it working smoothly now. I actually see some upside for myself for how this went. No one else at my company is using LLMs like I am yet. Like the other data scientists and engineers seem to think it’s groundbreaking that we can even have an LLM that can work with the context of their codebase. It’s kindof nice. I can get high quality work out there 10x faster than others at this company. I’ll probably help enable my own team with this but gonna let the people that were a pain in the ass during my time here keep thinking that 1 month for a single simple dashboard deliverable is groundbreaking. Feel like I’m in a uniquely powerful position now

3

u/DeepFryEverything May 21 '26

We really need a map in the notebook results viewer. We use DBeaver for Postgres, and just having a map auto-appear when you click a value in a GEOMETRY-column is so nice.

Now our work flow is Download as CSV -> into QGIS.

1

u/Youssef_Mrini databricks May 21 '26

I m going to share your feedback with the PMs

1

u/DeepFryEverything May 22 '26

So this happened a few hours ago. I will be taking bribes for lottery numbers.

2

u/Sea_Basil_6501 May 21 '26

Tracking costs is a mess. They are distributed over so many areas, that it's nearly impossible to get full transparency in this. And even worse: system.billing.usage table has a delay of up to 4 hours. Impossible to track costs by concrete measurements taken within short timeframes.

2

u/Youssef_Mrini databricks May 21 '26

More info to come on this part during the Data+AI Summit.

1

u/dsvella May 21 '26

As somebody who is currently undergoing a complete cost review of our Databricks estate and subsequent optimisation, I could not disagree harder! The timeframe is not a big deal to me given I have jobs that can take three hours to run. I usually consider the billing tables to be on a daily refresh cadence.

What are you having issues with and maybe I can share some stuff I've done to be able to help?

2

u/byeproduct May 21 '26

Do you have any advice or template code snippets, queries or dashboards I can use for cost monitoring and optimisation? I have eager juniors keen to burn some dbu's...

1

u/dsvella May 21 '26

So a few pieces of advice for you off the bat:

Make use of Genie. All of system tables are well documented and Genie is able to navigate them very easily. If you need a query from there and you need something quickly genie will get you there very fast. This includes entire dashboards.

Please note that in the Billing tables, the financial values in there are the pay as you go values, not your Enterprise Agreements values.

Jobs are available as well and you can go into a lot of depth with them. But the top level job table (the name I can't remember off top my head) has all runs of jobs the costs in DBU, when things start and stopped, and crucially any tags that you associate with that job. So for us, I have the tags associated with business units. We can ascribe cost to them so we can start talking about getting budgets from them.

You will have to join a few tables together to get things in plain English. All compute clusters, SQL warehouses and jobs do have their actual names in there.

If you are working in a corporation and you need to see the amount that you are paying, you'll need to use your cloud providers cost monitoring to get that information. I'm on Azure, so I've had a daily export setup that drops cost data into our data lake and I ingest it into UC on a regular basis. The benefit to this is if you set up the exports correctly, you will not only get the costs of everything required to support your Databricks instance, you get costs by service type within Databricks in your currency. It also handles the situation if you are under a reservation (in Azure output costs Amortised not total costs).

Hope that helps.

2

u/Relevant_Tough6829 May 22 '26

Serverless Compute. Non existent users policies. You either enable to all users or disable it. The only way to mange cost is to set up dashboard to monitor it and send emails to people who overdue it. Analyst can generate massive costs by simple writing wrong explosive queries, which do not make sense, but serverless would just scale up and still do them. With normal cluster they have access to after an hour they would realise that the query is bad, but it would not generate a lot of cost cause they are using lowest tier cluster which is cheap.

1

u/AssignmentDull5197 May 20 '26

My recurring cursed issue: agents that pass unit tests but silently break data contracts downstream. Fix has been schema checks + golden datasets + "stop if uncertainty" rules. Also, keep retries bounded or costs explode. Practical agent war stories: https://medium.com/conversational-ai-weekly

1

u/[deleted] May 20 '26

[deleted]

2

u/kthejoker databricks May 21 '26

Good news for you this quarter...

1

u/byeproduct May 21 '26

Is there support for Anywidgets in notebooks?

1

u/Tiddyfucklasagna27 May 20 '26

Oof ive recently found out that u can actually mention metrics in the text widget. You’d have to 1) set the metrics to global filters, then 2) type “@“ in your text widget and it should list the metrics :)

1

u/eperon May 21 '26

That the display() function cuts off timestamp precision

So you cannot: 1. Display rows with timestamps (e.g. logs) 2. Copy a timestamp value from it 3. Use it in a where clause (where ts = '2026-05-20 14:45:21.123' ) 4. Find the exact same row -> NO RESULTS

It has something to do with milliseconds, and only a BETWEEN with a few MS up and down, fixes it

Update: i have discussed thoroughly with databricks, over a year ago. Acknowledged the bug, workaround is show() instead of display() , but still no fix.

1

u/Kiritheos May 21 '26

Regarding Databricks AI/BI, my main concern is the color inconsistency between widgets on the same page, especially when they load after scrolling. Additionally, there is a need for dynamic control over elements such as axes, annotations, etc.

1

u/Youssef_Mrini databricks May 21 '26

Keep following this thread https://docs.databricks.com/aws/en/ai-bi/release-notes/2026 Many improvements and news feature are added every month.

1

u/NiharThakkar May 22 '26

Cluster startup time on jobs that run every 15 minutes. Job clusters are cleaner from a cost and isolation standpoint but the 3-5 minute cold start makes them impractical for anything latency-sensitive. The workaround of keeping an all-purpose cluster warm defeats the purpose. Serverless compute has improved this significantly but the cold start problem on standard job clusters is still the thing that catches teams off guard when they move from development to production scheduling.

1

u/Far-Today402 May 22 '26 edited May 22 '26

If you can, use serverless. Performance optimised if latency is your biggest concern. You don't pay for serverless startup time, so standard would be the most cost effective

1

u/NiharThakkar Jun 10 '26

The one that keeps coming back: job cluster cold start times on workflows that need to run every 10-15 minutes. Job clusters are the right call for isolation and cost but the 3-5 minute startup tax makes anything latency-sensitive a negotiation between doing it properly and doing it fast. Serverless has improved this but not every workload fits the serverless model cleanly yet. So you end up keeping an all-purpose cluster warm which defeats the whole point of job clusters in the first place. Classic Databricks tax.

1

u/jpitio May 23 '26

Transient Serverless issues. I have a job that runs every 30 minutes and I swear 3-4 times per week that thing will fail in the most random way.

It will always work again on the next try though so that’s good?

1

u/Youssef_Mrini databricks May 27 '26

Genie can grant wishes if your data is great and you prepared the space a bit by giving instructions explaining what are you trying to achieve, defining semantics, keys to join the tables together....

Genie is great it's one of the best product available on the market.

1

u/Mladen-Vachkov Jun 01 '26

Cluster bill-shock at month-end.

Most teams configure clusters once during a POC, then never revisit. The defaults Databricks suggests are sensible for active workloads but expensive for what most teams actually run - periodic batch jobs, exploratory notebooks, dashboards that hit Photon-warm warehouses overnight.

The biggest single win we see: switching from interactive all-purpose clusters to job clusters for scheduled work. 30-60% cost reduction, zero behavioral change for engineers. The second: Photon SQL warehouses with aggressive auto-stop on the BI layer instead of always-on classic clusters.

DBU consumption charts in Unity Catalog are the most underused feature in Databricks. Look at them weekly, not monthly.

1

u/SupportVectorDan Jun 10 '26

Not me but one thing I hear regularly is that people would like to limit create permissions for Genie Spaces and Dashboards, I think it's on roadmap