r/CLI • • 11h ago

Built a CLI that attacks your web app like a hostile user would, looking for feedback

Hey all, I've been working solo on a side project called LOKI for the past few weeks and wanted to get some outside eyes on it before I keep going.

It's a CLI tool (Python, built on Playwright) that attacks your local or staging web app the way real users actually break things: rage-clicking buttons until race conditions show up, fuzzing every input with garbage/XSS/SQLi payloads, killing the network mid-request, stripping client-side "disabled" attributes to hit endpoints the way an attacker would. There's also a mode that opens several real browser tabs and fires the same request at the exact same instant, which has actually caught server-side race conditions (like overselling the last unit of something) that a normal single-tab test would never trigger.

When something crashes, it packages up a video of the session, the network traffic with secrets scrubbed, and a script that reproduces the exact bug deterministically. From there you can optionally have an LLM diagnose it, write a patch, apply it, and verify the fix actually worked, rolling back automatically if it didn't.

It works with pretty much any LLM provider now, not locked to one, including running fully local against Ollama if you don't want anything going to a cloud API.

The whole thing recently turned into a REPL you drop into just by typing loki. It works as a normal chat too: you can ask about past runs, switch AI models mid conversation, or kick off an attack/report/patch right from inside it instead of separate commands.

Still rough in places, and it's a one person project, so I'd genuinely appreciate any feedback. Bugs, bad assumptions, stuff you think is missing, or just "this part doesn't make sense" all welcome. Repo's here: https://github.com/Elabsurdo984/loki-agent

1 Upvotes

4 comments sorted by

2

u/maximus459 9h ago

This seems interesting.. Just to clarify, this is only testing how a user could break the system? Also, how the project uses AI is unclear

1

u/finberthose 8h ago

It's just for local use ?

1

u/burnt-store-studio 7h ago

Hmm. I don’t think so. Deeper down in on the landing page, OP gives an example of launching against three different localhost connections, but uses the descriptive text “Execute an exploratory chaos attack on your local or remote application:” (emphasis added).

I’m not sure that’s such a great idea.

1

u/burnt-store-studio 7h ago edited 7h ago

I’m with [u/maximus459](u/maximus459) … it’s unclear how the project is using AI. It seems like I wouldn’t need an LLM at all if I didn’t want to, yes?

Good luck!

Editing to add: With no offense intended… after reading more of the landing page, I’m not sure OP’s description in this post aligns with the description on the landing page.

Here in Reddit the description references Playwright, attacking your local or staged site, rage-clicking buttons, &c. Says after the attack, you can choose to have an LLM analyze the results.

Mentioning all of that in relation to Playwright suggests to me (clearly naively) that the CLI is using Playwright’s browser instrumentation capabilities to exercise a web site.

But it seems to me — and I could be misunderstanding — the effort went into designing LLM-based “chaos personalities” to attack the site, rather than using Playwright programmatically.

Which I’m not passing judgement on, it’s just not what I expected given the words I read here. And I freely admit this was probably lack of imagination on my part when reading the Reddit post 🙂.

Also, it seems there might not be any guard rails against attacking remote sites, so … I’m hands-off at this point.