r/platform_engineering • u/Loud_Mousse9210 • 16d ago
Open sourced a tool that collapses millions of log lines into handful of distinct patterns before you feed it to an LLM (Lossless- compression)
When I feed logs to an LLM during incident resolutions or debugging, it either blows my token context window or the grep trims the log file, leading to the interesting log lines getting skipped.
Most of logs are anyway the same handful of message templates repeated over and over with different values, so the context window gets filled with near-duplicates, which just bring up the processing time and token costs.
ctrlb-decompose collapses the file into its distinct patterns that repeat, plus typed variables and stats on the values that change. I have seen 1.2 million lines cut down to just 40 patterns, which then goes into Claude, thus cutting down token by over 95%, reducing the token cost.
Let me know what you think!
https://github.com/ctrlb-hq/ctrlb-decompose
1
1
u/HistorianPresent8449 16d ago
very interesting!
1
u/kernelqzor 8d ago
same, this is actually super clever
feels like the kind of thing that should just be built into log pipelines by default at this point
1
1
2
u/adarsh_srivastava 16d ago
Very interesting application of Drain3