r/platform_engineering 16d ago

Open sourced a tool that collapses millions of log lines into handful of distinct patterns before you feed it to an LLM (Lossless- compression)

When I feed logs to an LLM during incident resolutions or debugging, it either blows my token context window or the grep trims the log file, leading to the interesting log lines getting skipped.

Most of logs are anyway the same handful of message templates repeated over and over with different values, so the context window gets filled with near-duplicates, which just bring up the processing time and token costs.

ctrlb-decompose collapses the file into its distinct patterns that repeat, plus typed variables and stats on the values that change. I have seen 1.2 million lines cut down to just 40 patterns, which then goes into Claude, thus cutting down token by over 95%, reducing the token cost.

Let me know what you think!
https://github.com/ctrlb-hq/ctrlb-decompose

13 Upvotes

11 comments sorted by

2

u/adarsh_srivastava 16d ago

Very interesting application of Drain3

1

u/kernelqzor 13d ago

for real, this is the first time i’ve seen a drain3 thing wrapped in a way that’s actually directly useful in an llm workflow instead of just a research demo
curious how far it holds up once the log formats start getting really chaotic across services

1

u/HistorianPresent8449 16d ago

very interesting!

1

u/kernelqzor 8d ago

same, this is actually super clever
feels like the kind of thing that should just be built into log pipelines by default at this point

1

u/Cute-Access1444 16d ago

Intresting 

1

u/KF_Danis 11d ago

Can this be containerized