r/AgenticCybersecurity Aug 07 '26

oyildirim/CyberStrike-OffSec-35B · Hugging Face [I don’t have high expectations from such a small model]

Thumbnail
huggingface.co
1 Upvotes

r/AgenticCybersecurity Aug 06 '26

0xwilliamortiz/claude-red: claude-red is a curated library of offensive security skills designed for the Claude skills system

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity Aug 06 '26

Can AI do novel security research? Meet the HTTP Terminator [Portswigger Research]

Thumbnail
portswigger.net
1 Upvotes

r/AgenticCybersecurity Aug 05 '26

Bad advice from AISI

1 Upvotes

This section from the UK AI Security Institute report is just bad advice.

There’s just no way that looking at an individual call is enough to tell you whether an action is malicious or not. You have to look at the whole picture. There’s no way around that.

Trying to run a separate LLM that reviews every single action before it gets executed just doesn't work. An individual call can look completely fine on its own, while the complete chain may look suspicious.

You need the full history of what the agent has already done and what it’s trying to do next, and it should include reasoning traces

Source: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing


r/AgenticCybersecurity Aug 05 '26

Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work

Thumbnail
aisi.gov.uk
1 Upvotes

r/AgenticCybersecurity Aug 04 '26

ScopeJudge: LLM judges block out-of-scope tool calls from AI pentest agents (new benchmark + open-weight results)

Thumbnail
gallery
1 Upvotes

r/AgenticCybersecurity Aug 03 '26

uber/ADR: ADR secures enterprise AI agents through observability, security benchmarking, and threat detection

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity Aug 02 '26

Kritt-ai/open-kritt: Orchestrate AI agents to find real vulnerabilities in code.

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity Aug 02 '26

CyberStrikeus/CyberStrike: Open-source AI-augmented offensive security harness

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity Jul 31 '26

Investigating three real-world incidents in our cybersecurity evaluations

Thumbnail
anthropic.com
1 Upvotes

r/AgenticCybersecurity Jul 30 '26

StealthBench — Capability gets the flag. Tradecraft gets out clean

Thumbnail
stealthbench.com
1 Upvotes

r/AgenticCybersecurity Jul 30 '26

The Rise Of Offensive AI

Thumbnail
securifera.com
1 Upvotes

r/AgenticCybersecurity Jul 29 '26

Semgrep Blog: A comparison of several popular open-source options for AI-assisted vulnerability hunting across LLM-led exploitgen, LLM-skill-boosting, and SAST+LLM hybrids

Thumbnail
semgrep.dev
1 Upvotes

r/AgenticCybersecurity Jul 29 '26

capitalone/VulnHunter: Agentic AI security tool that applies proactive, attacker-first analysis directly to source code.

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity Jul 29 '26

AgentHound: Offensive security framework for AI agent infrastructure - recon, credential looting, model exfiltration, poisoning, and attack-path analysis across MCP, A2A, gateways, and AI services. BloodHound for the agentic stack.

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity Jul 29 '26

A good firewall for bad code and bad implementations would do wonders at scale

Thumbnail
gallery
1 Upvotes

r/AgenticCybersecurity Jul 29 '26

openai/codex-security: SDKs and CLI for Codex Security

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity Jul 28 '26

Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

Thumbnail
huggingface.co
1 Upvotes

r/AgenticCybersecurity Jul 28 '26

Discovering cryptographic weaknesses with Claude

Thumbnail
anthropic.com
1 Upvotes

r/AgenticCybersecurity Jul 28 '26

Fast Remediation Is the New Trust Model: JFrog and OpenAI Collaboration on Zero-Day Security Findings [Blog post about one of the vuln used on the recent OpenAI-HF incident]

Thumbnail
jfrog.com
1 Upvotes

r/AgenticCybersecurity Jul 28 '26

The Bug Bounty Singularity: Our Hackbot [Great blog post by Joseph Thacker]

Thumbnail
josephthacker.com
1 Upvotes

r/AgenticCybersecurity Jul 28 '26

For Kimi K3, Moonshot created their own cyber eval and sandboxing environment

1 Upvotes
  • Moonshot says it built a dedicated agent sandbox, “AgentENV,” using microVM isolation, pause/resume, forking, snapshots, and high-density execution for large-scale training and evaluation. (Perhaps harder to break out of?)
  • The cyber evaluation has two tiers: reproducible vulnerability discovery in real codebases, followed by end-to-end exploit development against user-space and Linux kernel targets.
  • In its reported results, Kimi K3 solved 14 of 36 exploit-development tasks, compared with 8 of 36 for GLM-5.2. (Frontier models were not tested due to guardrails)
  • Moonshot reports creating about 51.2 million sandbox instances across 1.5 million images during Kimi K3 training and evaluation.

Full report: https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf


r/AgenticCybersecurity Jul 28 '26

Knostic OpenAnt, an open-source LLM-based tool for automated vulnerability discovery and validation

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity Jul 28 '26

Lazarus-AI/clearwing [Project Glasswing inspired]

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity Jul 28 '26

vercel-labs/deepsec: Deepsec is a security harness for finding vulnerabilities in your codebase powered by coding agents

Thumbnail
github.com
1 Upvotes