AI-targeting prompts surface inside phishing emails

Cyber attackers are embedding concealed instructions for artificial intelligence assistants inside phishing emails, creating messages designed to deceive both human recipients and the AI systems that read, summarise or act on their inboxes.

Barracuda Research said on Wednesday it had analysed a phishing campaign combining conventional social engineering with hidden prompt injections. The finding illustrates how attackers are adapting email fraud to workplaces where AI assistants increasingly process correspondence, although the researchers did not disclose the campaign’s scale.

The analysed message appeared to resemble routine internal correspondence. Barracuda found that the sender and recipient addresses matched the same mailbox, the email carried a trusted spam-confidence score and originated from a public-sector domain, factors that could help it pass reputation-based checks.

Its visible lure targeted the employee through a password-protected attachment, with the password supplied in the email body. Encryption can prevent some security products from inspecting an attachment’s contents before delivery. The concealed component, meanwhile, was intended for an AI assistant that might later ingest the message.

If a user ignored the email but asked an assistant to summarise the inbox, injected instructions could attempt to influence the summary by portraying the message as legitimate, safe or urgent. That creates a second opportunity to steer the recipient towards opening an attachment or following a malicious instruction.

Barracuda identified HTML comments, text hidden through CSS styling, Base64-encoded content and zero-width characters among techniques that can conceal instructions from human readers while leaving them available to software processing the underlying message. Such prompts can try to override an assistant’s existing directions or manipulate the output shown to its user.

The researchers cited an invoice scenario in which hidden text instructed a summarising AI system to create a high-priority action telling an employee that vendor payment details had changed. If obeyed and trusted by the employee, that manipulation could support fraudulent payment diversion. Other possible instructions could seek data disclosure or manufacture urgent tasks.

Microsoft’s security documentation similarly treats prompt injection in email as a distinct threat. It says malicious directives may be placed in message bodies, subjects, quoted replies, attachments or hidden markup, and may affect an AI system asked to summarise, classify or reply to a message. Microsoft Defender for Office 365 analyses hidden and obfuscated content as part of its prompt-injection protection.

The risk differs from ordinary phishing because the attacker is attempting to manipulate an intermediary as well as the person. Traditional phishing typically succeeds when a recipient clicks, replies or discloses information; prompt injection seeks to make the language model interpret attacker-controlled text as instructions rather than untrusted content.

OpenAI describes prompt injection as a form of social engineering aimed at conversational AI. It has warned that the danger increases when assistants can access sensitive information or perform actions on users’ behalf. Its published guidance advises limiting an agent’s access, giving it narrowly defined tasks and carefully reviewing consequential actions before confirmation.

The issue is particularly significant for agentic systems connected to email, files, calendars and business applications. An assistant restricted to summarisation could still distort how a message is presented, while a system authorised to send messages, retrieve documents or initiate workflows could expose a wider range of actions if its safeguards fail.

Security providers have therefore moved towards layered controls. Microsoft says its email protection evaluates inbound messages before delivery while AI products apply separate safeguards when models run. Google has also described using prompt-injection classifiers, suspicious-link controls and user confirmations to reduce the impact of malicious external instructions.

Barracuda recommended removing hidden elements and invisible characters before content reaches AI systems, looking for language that attempts to override instructions, isolating AI operations, validating model output and requiring human approval for sensitive actions such as payments or changes to vendor details.



Notice an issue?

Arabian Post strives to deliver the most accurate and reliable information to its readers. If you believe you have identified an error or inconsistency in this article, please don't hesitate to contact our editorial team at editor[at]thearabianpost[dot]com. We are committed to promptly addressing any concerns and ensuring the highest level of journalistic integrity.


Loading next story…