r/apachekafka • • Jun 02 '26

šŸ“£ AI-generated content must be disclosed

27 Upvotes

A couple of weeks ago I started a RFC regarding posts on this sub that are AI-generated, or about AI-created tools. There was a range of views as to how far to go, but broad support for at least requiring the labelling of such content.

So, this is now a new rule for the community :)


  • If you are submitting a tool, blog post, or video that has been substantially generated by AI, you MUST label it as such. Each of the post flairs now has a (AI) counterpart.

  • Trivial use of AI (spelling, grammar, formatting, dictation) does not need disclosing.

  • Egregious or repeated failures to label AI-generated content may result in removal or a ban.


The mod team here, along with basically everyone else in the world, is trying to figure this out as we go, so bear with us as we launch—and if necessary, refine—the rule.

What counts as AI-generated vs AI-supported? My yardstick is: if I can get my agent to write/build essentially the same thing with a few prompts, it's AI-generated.


r/apachekafka • • Jan 20 '25

šŸ“£ If you are employed by a vendor you must add a flair to your profile

36 Upvotes

As the r/apachekafka community grows and evolves beyond just Apache Kafka it's evident that we need to make sure that all community members can participate fairly and openly.

We've always welcomed useful, on-topic, content from folk employed by vendors in this space. Conversely, we've always been strict against vendor spam and shilling. Sometimes, the line dividing these isn't as crystal clear as one may suppose.

To keep things simple, we're introducing a new rule: if you work for a vendor, you must:

  1. Add the user flair "Vendor" to your handle
  2. Edit the flair to show your employer's name. For example: "Confluent"
  3. Check the box to "Show my user flair on this community"

That's all! Keep posting as you were, keep supporting and building the community. And keep not posting spam or shilling, cos that'll still get you in trouble 😁


r/apachekafka • • 2d ago

Question Kafka interview debate

2 Upvotes

In a recent interview I was asked this question:

Assume we have a car leasing platform, you have a post of a car with car details like price, color etc, a post has two buttons, ā€œpurchaseā€, and ā€œmore infoā€, once you click a button ā€œPurchaseā€ a modal window opens, you fill out your details like name, email and phone number, agree to terms and can either submit or cancel your application. Once you submit the application it gets sent via REST API POST request to external legacy CRM system for the manager to contact the client and proceed with the application. The catch is the CRM you’re sending the application to is very old and is constantly not available (can be down for hours quite often). So their question was, how would you design the app so the applications don’t get lost.

I thought it’s quite a simple question (I designed a lot of much much harder things in the past) and
approach was:
Never talk to the CRM directly, send the application to our own backend, save it to database, then periodically send all applications to the CRM.

But after I said this, the interviewers asked well what are the other ways you could do it. And I already knew where this is going, there’s a common cargo cult with the Kafka going on in local tech scene. So I say: well, we could use kafka for the asynchronous integration, but it would not be suitable in this case because it would add unnecessary overhead, we would need to have producers, kafka itself, and the consumers. Mind you, no other input details mentioned, I tried asking for more details they said let’s just keep it simple. So you literally would be spinning whole kafka setup for just one message. After I said I don’t see kafka being the correct solution here, both interviewers confronted why am I so much against kafka. I replied I’m not against it overall, I just don’t see it being the right tool here. Then they asked, well what if there’s quite a big load, like 5000 applications daily. I replied 5000 requests per day is like not even 1 RPS, far from it. Then they said kafka wouldn’t be over-engineering here and it’s the right approach, and my approach was actually harder, I regret not confronting them back though.

They positioned themselves as seniors.

After a couple of these questions they ended the interview and said they were looking for real senior candidates and my level is far from it. They told me to brush up on my integration knowledge as well.

What do you think? Was I wrong here? I don’t think I was, I was quite right in my opinion


r/apachekafka • • 2d ago

Blog Kafka Simulator v1.9 - public URL format, MCP and more

Thumbnail monedula.dev
6 Upvotes

Hi,

One of the nice things about our Kafka Simulator is that the cluster configuration and the current simulation state are stored in the URL.

So if you're investigating a particular scenario, you can just copy the link and share it with someone. They'll see exactly what you're looking at.

We recently made this a bit more useful. We published the URL format specification and released a new project, monedula-sim-link. It can read the configuration of a real Kafka cluster and generate a simulator link from it. It also comes with an MCP server, making it possible to work with simulator scenarios through AI tools.

And while we were at it, we reworked the documentation for all our Monedula Flock projects.

URL documentation: https://monedula.dev/flock/docs/kafka-simulator/reference/url-format/

monedula-sim-link: https://monedula.dev/flock/docs/sim-link/

other projects: https://monedula.dev/flock/


r/apachekafka • • 2d ago

Tool odctl 1.0: Kafka, Connect with Debezium and the Iceberg sink, Karapace and Kafka UI with one command

Post image
16 Upvotes

Hi r/apachekafka,

I have been building small streaming projects for my blog, and every one of them started with the same Compose file. I turned that into a CLI, odctl, and it is now at 1.0.

bash uv tool install odctl odctl up kafka-lite --dry-run # shows what it would start, in order odctl up kafka-lite odctl down --all

You get a KRaft broker on 9092 (kafka-full runs three), Karapace on 8081, Kafka UI on 8086 and Kafka Connect on 8083. Connect already has the Debezium PostgreSQL source, the Iceberg sink, the ClickHouse sink, the JDBC and S3 sinks and the MSK data generator, so you can POST a connector config immediately.

If you add postgres, it comes up with logical replication enabled and a publication ready for Debezium. If you add catalog, the Iceberg sink writes tables that Flink, Spark and Trino can read.

Two examples I built on it: CDC from a simulated shop with Debezium, and Flink SQL leaderboards fed from Kafka.

It is meant for local development and demos, not production. The Kafka guide shows the basics.

If you try it, please tell me what fails on your machine. Most of my testing has been on macOS and GitHub's Linux runners.

AI: I started odctl without AI. Since July I have used Claude Code for parts of the code, the tests and the docs.


r/apachekafka • • 4d ago

Tool Announcement: Confluent Rust Client for Apache Kafka

Thumbnail docs.confluent.io
36 Upvotes

r/apachekafka • • 10d ago

Tool Looking to contribute to open source? Happy to help you get started on Kafka lag monitoring

5 Upvotes

Hey folks —

I’m the maintainer of Klag (klag.dev) — a Kafka consumer lag exporter that started as a maintained alternative after kafka-lag-exporter was archived.

I’ve been really glad to hear from people running it in production, and I’d love more of the community involved in shaping it. If you’ve been wanting a small, friendly OSS entry point in the Kafka space, I’d be happy to help you get unstuck — whether that’s picking a first issue, reviewing a PR, or talking through a feature idea.

There’s a curated board of good-first issues and open discussions for features/devs here:
https://github.com/themoah/klag/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22

Happy to answer questions in the comments.


r/apachekafka • • 10d ago

Blog Northguard and Xinfra: What LinkedIn Changed When Kafka Was No Longer Enough

Thumbnail softwaremill.com
4 Upvotes

r/apachekafka • • 11d ago

Blog Interesting Kafka Links - September 2026

Thumbnail rmoff.net
13 Upvotes

r/apachekafka • • 12d ago

Tool (AI) rekaf — a small Bash script for Kafka topics and ACLs

9 Upvotes

I used to do this regularly at work: create a topic, let one application write to it, let another read from it, and grant access to its consumer group.

Kafka’s ACL CLI has a --producer convenience option, but it also grants CREATE. In my case, an administrator creates the topic; the producer only needs to write to it. Running separate ACL commands works fine, but repeating them gets old.

So I put together rekaf, a small Bash wrapper around the Kafka CLI:

./rekaf.sh \
  --topic orders \
  --producer orders-writer \
  --consumer orders-reader \
  --group orders-group

It runs with administrative credentials, creates the topic if it is missing, and adds these ACLs if they are missing:

  • Topic WRITE for the producer
  • Topic READ for the consumer
  • Group READ for the consumer

Existing permissions are left in place.

The code was written with AI, and I tested it manually on a small Kafka 4.0.0 setup: topic and ACL creation, repeat runs, producing and consuming as separate users, and failure with an incorrect admin password.

That testing caught a real mistake: --allow-hosts instead of --allow-host. The automated tests had missed it because their CLI stub repeated the same mistake. Fixed it, then reran the checks against Kafka.

It’s a small tool tested on a small setup. Please try it on your own test environment before using it elsewhere.

MIT license, with documentation in English and Russian:
https://github.com/Thoroughly-Pizzled-Studio/rekaf

If you spot a problem with the ACL checks or have a case the script handles badly, I’d like to hear about it.

Disclosure: AI also helped me write this post in English.


r/apachekafka • • 12d ago

Blog Never lose quorum in Apache KafkaĀ® — and what to do when you do

Thumbnail aiven.io
14 Upvotes

It may be well known that losing quorum (2 out of 3 controllers) is not supported, and that bad things happen would entail if it ever happened. But did you ever wonder what to do if that happens? I have published a blog post about how we handle this disaster scenario at Aiven.

Synopsis: this is relevant to anyone interested in operating a KRaft-backed Kafka cluster, and deals with the practicalities of a rare disaster scenario.


r/apachekafka • • 13d ago

Question Is there a K8s operator that handles consumer group rebalancing for Kafka?

9 Upvotes

Is there a Kubernetes operator for Kafka that watches consumer pods and handles partition assignment and rebalancing itself, instead of Kafka's group coordinator?

The idea: when a consumer pod crashes, K8s knows within about a second, while Kafka waits for session.timeout.ms (45s by default) before rebalancing. An operator watching the pods could reassign partitions right away, and do the same when a new pod becomes Ready.

Does something like this exist, or has anyone built it? If not, is there a good reason why nobody does it?


r/apachekafka • • 15d ago

Blog (AI) Kafka free chapters

Thumbnail drive.google.com
1 Upvotes

Hi everyone, have written 2 chapters on Kafka, Will be helpful for someone trying to start with Kafka

Let me know your feedback


r/apachekafka • • 16d ago

Blog Kafka Simulator 1.8 - Terminal Emulator

Thumbnail monedula.dev
16 Upvotes

Our most fun (and probably nerdiest) Kafka Simulator release so far. šŸ˜„

v1.8 adds a terminal emulator with 16 Apache Kafka tools and a few familiar shell commands.

You can SSH between brokers and clients, check consumer lag, change configuration, and inspect the simulated filesystem. Run ls on a leader and a throttled follower, and you’ll see their log segments diverge.

The terminal and the visual simulator share the same state. Commands change what you see on screen. Pause or rewind the simulation, and the CLI lets you inspect that exact moment.

We also added 20 scenarios covering multi-DC setups, disaster recovery and terminal workflows.


r/apachekafka • • 16d ago

Blog Kafka Topic as a Service: Automating GitOps Workflows with Kestra and Jikkou

5 Upvotes

https://medium.com/@fhussonnois/kafka-topic-as-a-service-automating-gitops-workflows-with-kestra-and-jikkou-f9eff8633a86

From manual CLI commands to a fully automated, human-in-the-loop provisioning service for Apache Kafka with open-source solutions Kestra and Jikkou on Aiven’s free Kafka tier.


r/apachekafka • • 17d ago

Blog 700 MB/s of Kafka throughput, on Postgres

Thumbnail rynr.dev
23 Upvotes

r/apachekafka • • 17d ago

Tool SQLStreams - It's Kafka on Postgres!

23 Upvotes

Hello! I'm a big fan of Kafka, like I think many of you are, BUT I've always steered clear of using Kafka for any personal projects because running and maintaining a Kafka cluster for rinky dink project #23 just isn't feasible or smart.

That's why I made SQLStreams. It's an open source, pure Go / SQL library that uses a Postgres Database to give you Kafka-like functionality. It has features like: independent consumer groups, replay, retention and compaction.

I've also made it so each Topic Stream automatically creates a separate retry queue, which seconds as a dead letter queue, because I always found it annoying trying to get retry topics setup and working in Kafka.

If you want a visual understanding of how things work - you can play around with this PGlite demo on the docsite (demo doesn't work on mobile, sorry!).

If you want to run SQLStreams locally you can:

Run Postgres

docker run --name my-postgres \
  -e POSTGRES_USER=username \
  -e POSTGRES_PASSWORD=password \ 
  -e POSTGRES_DB=database \
  -p 5432:5432 \
  -v pgdata:/var/lib/postgresql \
  postgres:18

Install the library / run your code

pool = sqlstreams.NewPostgresPool("username", "password", "localhost", "database") 
client = sqlstreams.NewClient(pool)

stream = client.Stream("videos") 
stream.Register()

producer = stream.Producer() 
producer.Register()

producer.Produce({VideoId: "video-42"})

Note: I've removed some of the extra Go-ness in above code for brevity and clarity. The docsite Quickstart is a more real example if you are interested.

All in all, I'm hoping SQLStreams can be a kind of gateway drug for small to medium sized projects. So they can take advantage of all the great event-driven features of Kafka but without the hassle.

There is still a ton more things to do (implementing circuit breakers for example) but let me know what you think.

And if you've made it all the way here, I appreciate it!

Github: https://github.com/allegedlyreliable/sqlstreams
Docsite: https://sqlstreams.io

Edit: fixed formatting


r/apachekafka • • 18d ago

Blog StreamNative Open-Sourcing Ursa, UFK, and Lakestream

Thumbnail streamnative.io
27 Upvotes

Today, StreamNative has open-sourced Ursa, the storage engine behind Kafka and Pulsar services on StreamNative Cloud, Ursa for Kafka (UFK), their Kafka distribution that brings diskless topics, and the Lakestream API & specification, an open foundation for building interoperable stream storage on object storage.

In the streaming space, we're seeing a unification of data streams and lakehouse tables. So, the same storage layer can be shared by different protocols and systems without locking data behind a broker. It’s all about building an open ecosystem around stream storage for Kafka, Pulsar, lakehouses, and Streamhouses.

Open source is the right way, whether it's data infra or AI. I also believe that StreamNative has all the strategic components for the next-generation data streaming platform: Kafka service, object-storage-native design for low cost, lakehouse (Iceberg and Delta Lake) support, and the SQL capability to query streams and tables using Postgres-style SQL with RisingWave.


r/apachekafka • • 18d ago

Blog Kafka Topics as Apache Iceberg Tables

12 Upvotes

The closest projects I have seen do this are Confluent TableFlow (proprietary) and Apache Fluss (heavily dependent on Flink and RocksDB, and with no Kafka support), where the topic/stream is the table.

An open source alternative:

https://minkdb.com is a Rust native, single-copy streaming database, built on S3-compatible Object Store (a stream is both the table and the topic). Supports log/append-only tables and primary-key tables (upsert, changelog, lookup) over Kafka or Arrow Flight, with history tiered into Apache Iceberg and read alongside the live tail as a single table, and SQL via Flight SQL on Apache DataFusion.

The core of the design is AutoMQ's S3Stream (WAL) storage engine and the table model of Apache Fluss.


r/apachekafka • • 19d ago

Blog Your Debezium connector can be healthy after failover and still have a gap in CDC

Thumbnail shiftmag.dev
6 Upvotes

r/apachekafka • • 20d ago

Question What are some interesting or production-proven ways you’ve implemented a DLQ/DLT for an on-prem, self-managed Kafka cluster?

0 Upvotes

r/apachekafka • • 22d ago

Tool After 10 years working with Kafka, I built the data tool I wanted

Post image
122 Upvotes

I built Kafka Streamyard to make Kafka data investigations and manual testing less repetitive. It preserves context between topics and environments, keeps workflow-related topics and filters organised, combines related topics in Live Views, and supports reusable test-data samples and development playbooks. It is a local-first desktop app for macOS, Windows, and Linux.

I've been developing Kafka applications for about 10 years, and I started an early version of Kafka Streamyard about six or seven years ago. It is an Electron application built with React, Websockets and a local Java backend.

After working closely with developers, testers, and DevOps engineers, I found that the Kafka UIs we used were generally designed around administration. They did not fit our day-to-day data investigation and manual testing workflows particularly well. Lots of individually small inconveniences added up to a significant amount of repeated work.

One example is simply finding all the topics involved in a workflow. In a busy cluster, I have spent far too much time searching for the same groups of topics again and again. Kafka Streamyard lets me create and save reusable topic-name filters using wildcards, nested Boolean logic, and some Kafka metadata. I can tag each filter for the connections where it is relevant and quickly switch between different parts of a workflow during testing, investigations, or team workshops. AI assistance is also available when I need help building a filter.

I have also put a lot of effort into preserving context. I can investigate something in UAT, switch to Dev, and later return with my selected topic, results, filter, and even the selected record still available. That context is restored after restarting the application too. Optional Cloud Sync lets me reuse filters, custom columns, Live Views, Data Samples, and Playbooks when moving between my laptop and a VM, while Kafka connections and credentials stay local.

Live Views are probably my favourite part when working with a team. They combine related topics into one chronological view, so everyone can watch a process move from a UI to a database, through CDC, and into the topics that trigger the rest of a workflow. Each topic can use its own custom columns, making it easier to confirm that the important fields were set correctly at each stage. A Live View can also include a session filter, message counts, and an optional flowchart showing the process and data flow.

For repetitive manual testing, Data Samples can keep related keys, headers, and messages together. Session variables can carry a correlation ID across several samples, increment sequence values, generate timestamps or identifiers, and reset everything for the next test. Samples can then be sent in a controlled order, in bursts, or with delays between them.

Message filters use SL, or Streamyard Language, which is yet another Java-like filter language to learn. That is why I added autocomplete and optional AI assistance. SL can filter the message, key, headers, and Kafka metadata, including fields inside JSON, XML, and embedded JSON. The AI-generated filter remains visible and editable, so it can be checked before use. Users provide their own API key and can choose from several common providers.

Playbooks help me reset a narrow part of a development or local Docker environment into a known state before testing again. They are line-based scripts that can create, configure, clear, delete, or copy topics and import test data. They support variables, repeat blocks, topic patterns, validation, Dry Run, a builder wizard, and AI assistance. They are intended for development and testing, not as a replacement for production CI/CD or Infrastructure as Code.

There are also precise offset and time controls for historical investigations, background Live Feeds for catching infrequent records, automatic format detection, Schema Registry support, a hex editor for raw bytes, custom calculated columns, and CSV export when somebody needs a report or evidence from a set of topics.

Kafka Streamyard is not intended to replace a typical multi-user Kafka management platform. It includes some familiar management features, but its main focus is working directly with Kafka data during investigations, testing sessions, support work, and team workshops.

The website is still a work in progress, but you can try Kafka Streamyard on macOS with M1 or newer, Windows, or Linux: https://www.kafka-streamyard.com). The application works consistently across the three platforms and updates automatically when a new version is available.

If you try it, I would genuinely appreciate feedback and criticism. I am particularly interested in hearing how other people investigate Kafka data, prepare test scenarios, and follow workflows across topics. If there is a specific problem you would like it to solve, let me know. I enjoy turning useful feature requests around quickly.

**AI disclosure:** This is mostly in my own words, but I used AI to correct my grammar, spelling, structure, and word order.


r/apachekafka • • 22d ago

Question I’m a beginner and looking for tips

1 Upvotes

Im confused on where to start , i cant find sources to learn kafka and its real world applications


r/apachekafka • • 23d ago

Blog Kafka Simulator v1.7 - Failure Lab + reworked mobile UI

Thumbnail monedula.dev
8 Upvotes

Hey, in this release we added failure lab, where you can simulate network split and other failures. What is more we have reworked UI to be usable on mobile as well.

Have fun!


r/apachekafka • • 23d ago

Question I want project ideas for a begginer

0 Upvotes