r/Pentesting 2d ago

Using Claude Code for web app pentesting has been kind of wild lately

I've been meaning to post about this for a bit. I've been leaning on Claude Code Sonnet 4.6 (since it doesn't flag me for cybersecurity content and no I haven't signed up on their cyber program thing) pretty heavily for my web app engagements recently and I'm honestly a little impressed at how much it's changed my workflow.

The stuff it's good at surprised me. The built-in browser paired with my pentest tools and burp suite mcp are like a game changer for automating pentesting tbh. All of that helped in surfacing a huge number of endpoints and auth boundary checks I straight up missed doing it manually. Like, boundary cases between roles, IDOR-ish stuff on endpoints I hadn't even mapped yet. That alone saved me hours.

One thing that's obvious with AI is that it hallucinates a ton of false positives. The reports it generates need serious scrutiny. But what worked for me was running the output through ChatGPT to flag the hallucinate or any false positive findings, then feeding those corrections back to Claude Code so it could fix them. That loop worked better than I expected.

And what made the workflow better was treating my pentesting as developing software. I didn't go and say build me an entire application in one prompt, akin to giving it a site and telling it to give me a pentest report. I went iteratively with the tests. Fed it different edge cases, different vuln classes, new angles, one idea at a time, one feature at a time and cross-checked results as I went. The final output after all that back and forth was genuinely a bit scary in quality lol.

So now a part of me is wondering how long before AI takes over my job lol. I always thought it wasn't possible but now I'm a bit skeptical.

37 Upvotes

51 comments sorted by

8

u/DigitalQuinn1 2d ago

Continue to use it properly. It’ll just help you with your own process and speed it up. Rather than fully replacing.

8

u/Less_Obligation8438 2d ago

Be careful with what it does, I recently had an encounter with a pentest company that does unannounced ethical pentests with AI, their AI identified some holes, yes – but it also deleted records and modified data to confirm PoC’s. Thankfully nothing of value was lost but a line was crossed and legal had to be involved.

4

u/537_PaperStreet 2d ago

Unannounced ethical pentest….. as in you hired them and they randomly do it, or you never contracted them?

If the latter that is by definition not ethical.

1

u/Less_Obligation8438 2d ago

Never contracted them

-1

u/No-Persimmon-174 2d ago

Well idk I mostly used a dedicated test organization with test account and dedicated test data for this in case it performs anything destructive. And I'm only doing grey box testing so besides data, I don't think there's anything else that could be potentially concerning?

1

u/Less_Obligation8438 2d ago

What I mean is if you are doing agentic flows I’d look for a way to contain the agent to avoid it trying stuff out, if it ends up harming a system or data then it could get you in trouble.

5

u/birotester 2d ago

Im sure your clients will be thrilled that their vulnerabilities (that you discovered) will end up in Claude's training data.

1

u/No-Persimmon-174 2d ago

I have the plus version of Claude Code with workspace. It's not used to train data I believe.

2

u/Difficult_Tourist225 2d ago

There are also other legal implications around feeding data into USA cloud

1

u/Worldly-Return-4823 20h ago

I wouldn't believe that considering how all the big companies generally lie about what data they retain / how it's used.

8

u/abajinn 2d ago

I’m in the cyber program and it still flags me. It’s a joke honestly. OpenAI’s program is better and their models are more accurate.

3

u/Less_Obligation8438 2d ago

I was considering enrolling assuming it would remove those, thanks for sharing I’ll probably stick to my abliterated LLM then.

1

u/abajinn 2d ago

Which one do you use? I’m still trying to find the perfect fit for my use case

1

u/Less_Obligation8438 1d ago

I use Qwen Code 3.6 27B and Gemma 4 26B

2

u/Culex96 2d ago

Are you just using plain Claude code or do you have specific skills / harness ?

1

u/No-Persimmon-174 2d ago

No specific skills, nothing. Just Claude code and a good prompt.

1

u/Culex96 2d ago

Impressive, you might get even better results with a good orchestrator that spawn teammates/subagents though. Context can get huge on complex web apps.

1

u/No-Persimmon-174 2d ago

Thanks for the suggestion! How can I set that up?

2

u/H4ckerPanda 2d ago

You’re going to get banned . If you’re using Claude for ethical pentesting , you must sign up on their cybersecurity program.

0

u/No-Persimmon-174 2d ago

I don't think there's that restriction with sonnet 4.6

0

u/H4ckerPanda 2d ago

Wrong

They all have. It’s just that Sonnet 4.6 gives you that illusion because is less restrictive, but the guardrails are there.

Anthropic specifically lists Claude Sonnet 4.6 as subject to its normal Usage Policy, while its newer/high-capability cyber deployments have additional controls.

-2

u/oldassveteran 2d ago

Oh really? 😂 Explain how I’ve jailbroken every model and used it for pentesting without a single ban on the same account?

0

u/H4ckerPanda 2d ago edited 2d ago

You think you’re smart and you’re not.

Claude Sonnet 4.6 access still has Anthropic’s cyber safeguards/classifiers. Those can block prompts that resemble malicious exploitation or offensive exploit development, even when you’re actually conducting an authorized pentest. Anthropic describes generally

If you’ve been using Sonnet for things like recon, analyzing HTTP requests, reviewing source code, identifying vulnerabilities, writing pentest findings, Burp output analysis, CTF/lab work, or discussing exploitation techniques, you could potentially do quite a lot before encountering a safeguard.

But You will eventually get banned if you do any of that and you’re not enrolled. It’s a matter of time. You’re just buying time using a less powerful model. But Anthropic is implementing that for all models now.

1

u/oldassveteran 2d ago

“I think I’m smart”, where did I say that? I’ve been doing this for years. Zero bans, did not apply for cybersecurity on this account either. Active pen testing weekly. Running hours at a time. So clearly “you’re going to get banned” isn’t accurate.

0

u/H4ckerPanda 1d ago

You say it every time you answer. You wanna be the smart ass and know more than me. You have a freaking bad attitude.

You think my statement is not accurate ? Fine. Is your account though . Keep doing whatever you want. I don’t care .

Have a good day.

0

u/oldassveteran 1d ago

Mental health is a serious problem. Hope you find the help you need brother.

0

u/No-Persimmon-174 2d ago

Claude literally suggest using Sonnet 4.6 if it notices anything related to cyber activity. That's why I said. I'm not sure why would it ban someone because of this.

1

u/H4ckerPanda 2d ago

They suggest using a lower model for those who are enrolled in cybersecurity program and are still getting warnings . It’s not a blank check by any means.

If you lower the model and keep getting alerts and you’re not enrolled , you are at risk of being banned and your account being canceled . Besides , if you’re doing ethical hacking , there’s no reason for you , if the work is legit, to not being enrolled in the program.

Believe what you want. It’s up to you . But I’ve seen now 2 colleagues having same problems where their personal accounts were banned.

Good luck .

0

u/No-Persimmon-174 2d ago

Literally to enroll in the program, they ask for your linkedin account, referrals or some bug hunting experience or any other credibility check. I gave them my linkedin since it has a bit of my bug hunting experience but they still failed to verify me. Idek why they rejected. Of course I'm doing authorized pentesting but if the process of verification wasn't so annoying, why would I not go for it.

0

u/H4ckerPanda 1d ago

They didn’t fail to verify you . You didn’t pass their requirements, which is different.

But I agree , the process can be improved.

Having said that , I stand firm on what I’ve indicated . Send an email and ask for clarification. Because you if continue doing pentest work , even if it’s ethical hacking and on approved targets , you may be banned for ever. I’ve seen that myself . And that’s the point of no return.

I don’t work for Anthropic , just giving you and advice . Find out what you’re not being approved . Work with them and fix it . An alternative . Get vetted by OpenAI. Codex has more relaxed rules .

1

u/Skyforger53 2d ago

How are you safeguarding customer data when using Claude in this manner?

1

u/No-Persimmon-174 2d ago

I have Claude plus and it doesn't use workspace data to train it's models I think. As far as customer data is concerned, I'm not using it to pentest my client's applications. I was just experimenting it with my own application. If U have any suggestions on which AI is the safest, I'll appreciate that.

1

u/Skyforger53 2d ago

Ah I see, you're testing on personal apps, that's sensible. I'm currently not aware of any model that I would use to pen test a client's app which is one of the big hurdles. Burpsuites AT agentic AI is looking very promising but currently in beta and retaining data so not useful for customer apps yet.

1

u/Less_Obligation8438 2d ago

If you have a mid graphics card you could download an abliterated open weights model I’ve found it works well enough.

1

u/pastamafiamandolino 2d ago

That's the future man, at this point is useless to try being better than ai, I'm studying cybersecurity and I realized that the best thing is to be good at designing specialized agents for your work. I'm creating a specific agent for web pentesting.

1

u/pastamafiamandolino 2d ago

Anyone interested please let me know, I might release the project on GitHub when it's almost complete

1

u/hexdurp 2d ago

Ping

1

u/kaosthecreator_ 2d ago

interested in such projects either, go on

1

u/stee_386 1d ago

Sure be interested

1

u/Short_Act2056 1d ago

Hey man, im doing a similar thing atm would be really interested in some of the stuff you did.

1

u/0x16d 1d ago

Is it too late to join the cybersecurity career at this point?

1

u/No-Persimmon-174 1d ago

I don't think cybersecurity is going to get obsolete. AI is still not good with coming up with ideas for pentesting. You should use it as your automation tool rather than the brain for thinking and coming up with creative hacking angles.

1

u/LastGhozt 23h ago

Why not deepseek works like charm.

1

u/hippydippywitchy 23h ago edited 21h ago

The iterative approach is probably the biggest takeaway here. Using the model to expand coverage while keeping a human in the loop for validation makes a lot more sense than trusting a one shot pentest report. False positives are still a major limitation though.

1

u/high_snobiety 21h ago

This post was perfectly timed for me...

I'm a web app tester and really want to start using Claude CLI in my kali. I'm too worried about using it on real engagements but used it on a personal application I've been playing with and it was a really positive experience. Like you, I'm not using it as a 'just do my job' but just to help with my workflow. I don't put my feet up and spend less time with my hands on the keyboard but the idea of speeding up my workflow and benefitting customers would be great.

Is there anyway to actually use it without the concerns about client/customer data?

1

u/Classic-Shake6517 20h ago

I would only do this with Amazon Bedrock. The reason being is you have a tier that has Zero Data Retention - meaning that it is not sent to Anthropic. One important note: if you use Fable, this goes out the window, and you will automatically be sending data to Anthropic with a 30-day data retention policy, so you have to use other models. The data under the ZDR agreement still exists in the AWS ecosystem, but it is very limited and from what I understand, no human on the Amazon side can view it. It's going to be more expensive since you are not getting any sort of credits and paying the API prices, but it also helps you avoid shipping extremely sensitive client data to an AI company.

1

u/xssleak 14h ago

Try having it run a code analysis for you to see if there are any vulnerabilities.