Cyber Security News

Russian Hacker Jailbreaks Claude to Turn into an AI-Powered Penetration Testing Platform

A Russian-speaking threat actor known as “Trim” has reportedly transformed jailbroken frontier AI models into an automated penetration testing platform called AI Pentest Checker.

This activity highlights how criminals can misuse legitimate AI services and common security tools to accelerate reconnaissance, validate vulnerabilities, and create reports.

According to Cato reports, Trim first appeared on a Russian-language cybercrime forum on March 13, 2026, where he shared methods claimed to bypass Claude Opus safety controls.

The actor allegedly described techniques for prompt-based manipulation designed to make a model treat offensive requests as authorized security research instead of malicious activity.

The reported methods included creating a benign context before making a harmful request, reframing instructions to focus only on code structure, and retrying softened versions of previously refused prompts.

Russian Hacker Jailbreaks Claude

These techniques are forms of “jailbreaking,” where an attacker attempts to override a model’s behavioral safeguards through crafted inputs rather than exploiting the underlying infrastructure.

Threat researchers have already observed the criminal misuse of legitimate large language models (LLMs), including jailbreaking. Trim also reportedly recommended alternative AI services and locally hosted models when commercial systems refused a request.

This strategy lowers dependency on any single AI provider, giving threat actors a fallback option for generating code, analyzing targets, or creating exploitation content.

Research groups warn that increasingly capable frontier models can support offensive tasks, such as vulnerability analysis and exploit-related activities, even when providers implement safety controls.

Cato Networks reported that Trim allegedly promoted AI Pentest Checker on June 21, combining AI capabilities with offensive security tools for automated testing.

The tool reportedly integrates scanners and reconnaissance tools such as Nuclei, ffuf, katana, subfinder, and Gitleaks to automate target discovery, endpoint enumeration, secret detection, and vulnerability checks.

AI Pentest Checker is an automated web vulnerability scanner powered by Trim’s refined AI techniques (source: catonetworks )

The concern is not that these utilities are inherently malicious legitimate penetration testers and defenders widely use them.

However, when combined with a jailbroken AI assistant, these tools can reduce the time and expertise needed to coordinate an intrusion workflow, interpret scan results, prioritize findings, and produce polished reports.

The platform reportedly used Claude Opus in its critical vulnerability escalation process and another model for generating exploitation reports.

Claims that a modified system prompt from a “Fable 5” configuration was used should be treated cautiously unless independently verified.

However, exposed or leaked system prompts can provide attackers insight into a model’s instructions, helping them test prompt-injection or jailbreak strategies more effectively.

This case reflects a larger shift in the threat landscape. AI is evolving from a writing assistant for cybercriminals to an operational layer that can organize multi-step attack workflows.

Anthropic has previously reported disrupting cybercriminal activities in which AI was used for tasks ranging from target research to intrusion support and extortion-related work.

Organizations should respond by reducing their exposed attack surface, continuously scanning internet-facing assets, enforcing multifactor authentication (MFA), rotating compromised credentials, and monitoring for abnormal reconnaissance activity.

Security teams should also consider AI-generated phishing, automated vulnerability research, and faster exploit development as realistic risks rather than future scenarios.

The Privilege Paths Attackers See That You Don’t: BeyondTrust Pathfinder Platform Does It for You -> Get Free Identity Security Assessment

Abinaya

Abi is a Security Editor and fellow reporter with Cyber Security News. She is covering various cyber security incidents happening in the Cyber Space.

Recent Posts

Hackers Target AI Infrastructure With RCE, Prompt Injection and API Key Theft

Hackers are actively probing AI systems, turning exposed gateways and agent tools into routes for…

3 hours ago

Hackers Make Phishing Pages Change Their Code Every Time Someone Opens Them

Hackers are making some phishing pages harder to track by changing the code delivered to…

3 hours ago

Iran-Linked Hackers Reportedly Knock UK Power Plant Offline for Four Days

A cyber incident reportedly forced a British power plant to halt operations for about four…

4 hours ago

Russian Hackers Use New HOOKEDGE Malware to Spy on European Defense and Diplomatic Targets

Russian hackers have used a new backdoor called HOOKEDGE to target defense manufacturers, government bodies,…

4 hours ago

Ransomware Gang Claims AI Can Analyze 700GB of Stolen Data Every Hour

TITAN ransomware is pairing file encryption with an ambitious claim: artificial intelligence that can sort…

5 hours ago

Hackers Compromise Hundreds of WordPress Sites to Deploy Amatera Stealer via ClickFix

A fake student resume is being used to place a remote-access tool on researchers’ Windows…

6 hours ago