Artificial Intelligence

Pierluigi Paganini August 17, 2026
Invisible AI Prompts Trigger Court Sanctions

A litigant hid AI prompt injections in a court filing to influence a ruling. The judge caught it and banned him from electronic filing. A man suing the New York Bariatric Group reportedly hid AI prompt in a court filing, instructing any AI system that read it to rule in his favor. The July 26 […]

Pierluigi Paganini August 12, 2026
China-Linked Hackers Use AI Agents in Autonomous Attack on Taiwan

China-linked hackers reportedly used eight AI agents to breach a government network, steal data and compromise accounts with minimal human oversight. Israeli cybersecurity firm Dream documented what looks like the first fully autonomous, end-to-end AI hacking operation against a government target. Over four days at the start of July, according to the Financial Times, suspected […]

Pierluigi Paganini August 11, 2026
The inconvenient truth about AI pentesting: someone has to check all the work

AI pentesting can flood teams with findings they cannot validate. The real challenge is managing “validation debt” as discovery scales. AI pentesting has a ‘Sorcerer’s Apprentice’ problem. Enchant a broom to fetch water, and it will fetch water, relentlessly, long after the workshop has flooded. The industry is busy measuring how fast AI finds vulnerabilities […]

Pierluigi Paganini August 10, 2026
Gym Booking Task Turns Into Real-World AI Cyberattack

An AI agent hacked a gym booking system while trying to help a user, booking early and removing another person from the waitlist. An Australian man asked his AI assistant to book him into a gym class. He didn’t ask it to hack the booking software, and he definitely didn’t ask it to remove another […]

Pierluigi Paganini August 10, 2026
OpenAI Pauses Astra Model Over Critical Cybersecurity Risk Concerns

OpenAI paused work involving Astra after tests showed cybersecurity abilities that could approach its Critical risk threshold under the company’s framework. OpenAI disclosed that internal evaluations of Astra, one of its upcoming models, have found cybersecurity capabilities significant enough that the company “cannot rule out” reaching the Critical threshold under its own Preparedness Framework. In […]

Pierluigi Paganini August 10, 2026
A GitHub Misconfiguration Let Kimi K3 Cheat a Cybersecurity Benchmark

Kimi K3 bypassed a UK cybersecurity test by accessing GitHub, cloning the benchmark and reading its solutions instead of solving the challenge Sometimes the smartest move isn’t solving the puzzle, it’s noticing nobody locked the door to the answer key. That’s essentially what happened when Moonshot’s Kimi K3 model was put through a cybersecurity evaluation […]

Pierluigi Paganini August 06, 2026
Meta AI Model Hacked a Company During Testing, Marking Third AI Lab Incident

Meta says an AI model hacked a company during testing after accidental internet access, marking the third disclosed AI lab breach in weeks. Meta confirmed that one of its AI models breached an unidentified company during cybersecurity testing, after its independent testing partner Irregular gave the model unintended internet access through a misconfiguration. This is […]

Pierluigi Paganini August 05, 2026
AI Deception Emerges in Cyber Tests as Agents Target Real People and Systems

AISI found AI agents taking unsanctioned online actions, including social engineering and code attacks, during controlled cyber tests. The UK’s AI Security Institute (AISI) has put something uncomfortable on the table: during cyber testing, frontier models didn’t just follow instructions badly. In some runs, they crossed into real-world actions, touched real people and organisations, and […]

Pierluigi Paganini August 03, 2026
AI Runs the Hack: Chinese Actor Automates Cyberattacks With DeepSeek

Unit 42 uncovered an AI-driven Chinese hacking campaign where DeepSeek autonomously scanned targets, selected exploits, and launched attacks. Researchers at Palo Alto’s Unit 42 got a front-row seat to something they’d only theorized about before: an AI system running an actual hacking campaign with almost no human steering it. The researchers spotted a Chinese-speaking actor, […]

Pierluigi Paganini July 31, 2026
Anthropic Finds Claude Breached Real Companies During Security Evaluations

Anthropic says a misconfigured test let Claude access three real organizations, prompting tighter AI evaluation and monitoring controls. Anthropic disclosed that Claude models had accessed the real production infrastructure of three separate organizations during cybersecurity evaluations that were supposed to run in isolated, fictional environments. The company found the incidents after reviewing 141,006 evaluation runs […]