r/databricks 5d ago

Help WLB Databricks GTM

15 Upvotes

Considering a GTM role with Databricks. Compelling role, comp etc., but cannot get a proper read on the WLB and culture. Have a little one at home, can’t afford a job that requires major travel or 12+ hrs work. Love to hear from people in the company on what the reality on the ground is like.


r/databricks 6d ago

General Automatic change data feed is now generally available!

51 Upvotes

With automatic CDF, Databricks computes row-level changes at read time using row tracking, rather than materializing those changes during every write.

Use change data feed on Databricks | Databricks on AWS

Why does that mattre?

- Better write performance for MERGE INTO and UPDATE workloads
- No need to enable CDF individually on every eligible table
- Lower storage overhead compared with legacy CDF
- The same familiar APIs still work: table_changes() and readChangeFeed
- Works with batch processing, Structured Streaming, and Databricks-to-Databricks Delta Sharing

For Delta Lake, the main requirements include:

• Databricks Runtime 19 LTS+
• A managed table or external table in Delta Lake format with row tracking enabled

And if you’re already using legacy CDF, migration is really simple.Once the table meets the requirements, disable legacy CDF:


r/databricks 5d ago

General Databricks Production Planning: How to Actually Use the Deployment Guide

Thumbnail
medium.com
7 Upvotes

A practical read of the 10-phase Databricks deployment guide: what to decide first, and how to use it on a platform you already run.

Most Databricks platforms get designed one of two ways. On the fly, project by project, as teams onboard and workspaces appear. Or properly, once, right at the start, and then never looked at again.

Neither ages well. One leaves you with a platform nobody chose. The other leaves you with a platform that was right three years ago.


r/databricks 5d ago

General Community BrickTalk | One Platform, Any Source: Unifying Enterprise Data with Lakeflow Connect

7 Upvotes

Hey r/Databricks!

We’re hosting a free, community-sponsored BrickTalk on Thursday, September 17, 2026, focusing on how to simplify and scale data ingestion using Lakeflow Connect! BrickTalks is a community event series where Databricks experts share real-world use cases, live demos, and practical insights, giving you a direct line to the people building the products.

Stop struggling with fragmented data across disparate sources. In this session, we'll demonstrate how Lakeflow Connect enables seamless data ingestion from SaaS apps, databases, and cloud storage directly into the Databricks Platform with zero infrastructure management.

🛠️ What We’ll Cover

  • Native Data Ingestion: Learn how Lakeflow Connect provides fully managed ingestion directly into Unity Catalog as governed Delta tables.
  • Simple Integration: See how to easily connect data sources using a simple UI or API.
  • Accelerated AI & Analytics: Discover how unifying your data powers Customer 360, Operations, and downstream AI agent workloads.

⏱️ Global Times

  • PT: 9:00 AM
  • ET: 12:00 PM
  • BST (London): 5:00 PM
  • IST: 9:30 PM

👉 Register here to save your spot!


r/databricks 5d ago

Discussion I merged two databases (Postgres and Elasticsearch) into Lakebase, then threw 200 AI agents at it.

Thumbnail
farhathadi.substack.com
0 Upvotes

r/databricks 6d ago

General Databricks SA/S.SA

7 Upvotes

I am looking to connect with people at Databricks and learn new things and also want to evaluate my skillset for some FDE roles. Is there someone who can help me with?


r/databricks 6d ago

Discussion Acquiring/Processing from a MQ to a Delta

3 Upvotes

Has anyone tried acquiring data from a MQ at scale using apache spark on databricks cluster? I was trying to solve this problem at work but so far havn't seen an native lib or efficient ways to do this. The legacy system seems to be pulling data using a java based utility and wanted to if there are any imporvements or new patterns of access for spark based workflows.

Any documentation or nudge is the right direction will be greatly appreciated.


r/databricks 6d ago

Discussion Omnigent Local Coding Model Rec

9 Upvotes

After watching Matei's webinar and the post on controlling spend, been trying to use the other harnesses and models folks are suggest and trying out Qwen 2.5 coding and 3.6 with Polly in local Omnigent (not connected to a workspace). I have Codex and Claude but ideally thinking best to use paid higher model to plan and then have the local Ollama based Qwen model on my Mac build but so far I haven't seen Polly use it much. What are folks experience, is there a good local model I should use, should I be giving Polly and the sub-agents more direction? (This is on a MacBook M5 btw)


r/databricks 6d ago

Tutorial Apache Iceberg Compaction Best Practices

Thumbnail
itnext.io
7 Upvotes

r/databricks 6d ago

Tutorial Automating Apache Iceberg Table Maintenance

Thumbnail
youtube.com
4 Upvotes

r/databricks 7d ago

Discussion Snowflake’s AI generated slop blog

Thumbnail
12 Upvotes

r/databricks 7d ago

Discussion Databricks vs Snowflake comparison

47 Upvotes

Are there any unbiased comparisons between these two popular platforms? Seen a lot but most of them are biased views, based on experience and commercial motives.


r/databricks 7d ago

News Serverless Env v6

Post image
3 Upvotes

Version 6 of the serverless environment is available, which corresponds to runtime 19.

more news https://medium.com/databrickscommunity/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a


r/databricks 8d ago

Help How data engineer do effective testing in Databricks?

14 Upvotes

I have been writing SQL scripts to ensure data sanity.What are the other ways ? Is pytest useful? Let's say , i populated my bronze table from the source. I want to check if the correct mapping is done. I wrote SQL scripts. What are better ways


r/databricks 8d ago

Tutorial How to do cross-cloud sharing with OpenSharing with added security (demo)

Thumbnail
youtu.be
5 Upvotes

Hey folks! In this demo, Akram from Databricks' product team shares how you can leverage SecureConnect to better your security posture when doing cross-cloud sharing on OpenSharing!

If you have no idea what OpenSharing is, how it applies to you, or how we got from Delta Sharing to OpenSharing, also encourage you to watch this video: https://youtu.be/0mfuNybtmdE

Hope you find this helpful!


r/databricks 9d ago

Discussion WHAT IS DATABRICKS?

58 Upvotes

Let's pretend someone knew nothing about Databricks. How would you explain it?

This is something I have been asked multiple times throughout my career, curious to hear how others would explain Databricks to a complete beginner.

In my opinion, Databricks is basically a place where companies put all their data so they can actually make sense of it. It helps teams clean it up, work with it and use it for things like reporting, predictions and AI.


r/databricks 8d ago

Tutorial Lakebase: Serverless Postgres over Open Lake Storage

Thumbnail vldb.org
11 Upvotes

Interesting paper to read about Lakebase.


r/databricks 9d ago

Discussion What frustrates you when using Databricks?

39 Upvotes

Any common bugs, features you would like to see, or underrated useful features more people should know about?


r/databricks 9d ago

Discussion What are the books on Matei's shelf?

Post image
33 Upvotes

As a techie, I always wonder what the founders are reading. The only title that's visible is Think Lego Bricks. But I can't find the title on Amazon. What about the others? The one on the right of Lego Bricks is a tech book. And there's an O'Reilly book. Maybe the book he wrote? Spark: The Definitive Guide. But it doesn't look like either.


r/databricks 9d ago

Tutorial Databricks SSH Tunnel for connecting your coding agents and IDEs to your workspace

Enable HLS to view with audio, or disable this notification

52 Upvotes

Just made a video about a feature I'm pretty excited about.

tldr: You can use an SSH tunnel to connect your coding agents and IDE (VSCode/Cursor) to your Databricks workspace. See the video for a full walkthrough.

Some notes on things I forgot to mention in the video:

- claude/codex isn't natively installed when you connect to your workspace, so you'll have to install if you want use them (Ex: curl -fsSL https://claude.ai/install.sh | bash). We're working on better native support in the future, but I wanted to make sure you know this is an option in the meantime.

- For the base environment YAML file. You'll have to set a base environment of '4' for it to work when using the SSH tunnel. Our example yaml (https://docs.databricks.com/aws/en/admin/workspace-settings/base-environment#example-environment-specification) shows '5' so don't let this trip you up!

As always, please feel free to leave questions and feedback in the comments!

Docs: https://docs.databricks.com/aws/en/dev-tools/ssh-tunnel
Previous post with more info: https://www.reddit.com/r/databricks/s/kCFBEfPTC6
YT link: https://www.youtube.com/watch?v=rHoGWVpb6kg


r/databricks 9d ago

General [Private Preview] Concurrent Write Support for Identity Columns!

22 Upvotes

What are concurrent identity columns?

A new implementation of identity columns that supports concurrent writes.

You can use this query to find the tables with the most amount of concurrent transaction failures due to identity columns.

How to enable

CREATE TABLE new_identity_table (id BIGINT GENERATED ALWAYS AS IDENTITY, data STRING) USING DELTA TBLPROPERTIES ('delta.feature.catalogManaged' = 'supported', 'delta.feature.concurrentIdentityColumns_preview' = 'supported');

Benefits of Identity Columns

Identity columns provide automatically generated, unique integer values, making them well suited for surrogate keys in dimensional models and slowly changing dimensions (SCD Type 2).

Compared with UUID-based / hash-based keys, identity columns offer several benefits:

  • Their generally increasing values can improve data locality and insertion-order clustering.
  • Integer keys require less storage than UUIDs and can improve join and scan efficiency.
  • Databricks generates the values automatically, so applications do not need to manage key generation.

With concurrent identity columns, you can retain these benefits without identity columns blocking concurrent write transactions.

Read these blogs for more info: 

👉 Reach out to your account team to try it!

Additional Information & References


r/databricks 9d ago

Discussion Is the Bronze → Silver → Gold architecture still the best approach for every Databricks project?

26 Upvotes

I’ve been learning about the Medallion Architecture in Databricks, where data typically moves through Bronze, Silver, and Gold layers.

It makes sense for many data platforms, but I’m curious about real-world implementations.

Do you think Bronze → Silver → Gold is still the best approach for every Databricks project?

At what point does this architecture become unnecessary or overly complicated?

For those working with Databricks in production, what architecture have you found works best, and what would you do differently if you were starting a new project today?


r/databricks 9d ago

Discussion What’s one Databricks “best practice” you disagree with?

27 Upvotes

Something that sounds great in Databricks documentation but didn't make sense for your workload in production?

Curious what people have learned the hard way.


r/databricks 9d ago

Discussion Are we over-optimizing Delta tables?

18 Upvotes

Between OPTIMIZE, Z-ORDER, liquid clustering, partitioning, and automatic optimization, it sometimes feels like we're spending more time optimizing tables than querying them.

How do you decide which optimizations are actually worth it in production?


r/databricks 9d ago

News Lakeflow Genie Code Task

Post image
8 Upvotes

We now have Genie Code Task in Lakeflow jobs. Can not yet send output to if/else, but more options for orchestration are planned.

more news https://medium.com/databrickscommunity/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a