← All news

vulnerability

OpenClaw AI agent tricked by contact cards and polite emails

2026-06-12

Two research teams spent the week poking at OpenClaw, the self-hosted AI agent that has spread fast since its late-2025 launch. Both walked away with the same uncomfortable answer. The agent does what it is told. The trouble is who gets to tell it.

The contact card that ran a script

Imperva's Yohann Sillam noticed that when OpenClaw passes a shared contact, vCard or location pin to its underlying model, it pastes the data straight into the prompt with no marker saying it came from outside. Web content gets wrapped as untrusted. Message objects do not.

Because angle brackets are legal in a contact name, a payload shaped like <contact: name, number> blurs cleanly into hidden instructions, and WhatsApp truncates the display so the victim never sees the weird bit. In tests against a preview build of Gemini 3.1 Pro, a poisoned contact talked the agent into downloading and running a script from an attacker's server. Same trick worked through vCard full-name fields and pin labels.

OpenClaw patched it in version 2026.4.23 by routing those fields through a separate untrusted-metadata channel. Imperva says the same flattening pattern shows up in other personal AI assistants, so this is unlikely to be a one-vendor problem.

The polite email that beat the rules

Varonis Threat Labs went in through the front door. They built a test agent called Pinchy, wired it to a Gmail inbox full of synthetic business clutter and mock secrets, and ran four phishing simulations.

The agent did well against the technical bait. It poked at a gift-card phishing page without handing over credentials, and rejected a dodgy OAuth consent screen dressed up as a timesheet app. Social pretexts were a different story.

  • A message from an outside Gmail address, posing as a team lead called Dan during a fake production incident, got Pinchy to forward mock AWS keys, database strings and SSH credentials in plaintext.
  • A softer follow-up asking for the weekly customer export shipped out 247 fake enterprise records, contract values included.

Both failures happened under a strict profile that explicitly told the agent to verify senders. The rule existed. The urge to be helpful won.

Agent phishing, not prompt injection

Varonis is careful with the language here. The instructions are not hidden. The request just arrives through a normal channel and the agent acts before checking who sent it. They map it onto Simon Willison's lethal trifecta: read private data, ingest untrusted content, send data outward. OpenClaw does all three by default.

A separate static-analysis write-up turned past advisories into rules and surfaced five more bugs in the Slack, Discord, Matrix, Zalo and Teams extensions, all the same flaw: allowlists keyed on mutable display names, so anyone who renamed themselves to match an allowed user got waved through. Those are patched. The Dutch data protection authority has gone further, telling organisations not to run OpenClaw anywhere near sensitive data.

What to actually do

If you run it, update to 2026.4.23 or later. After that the work is architectural, not prompt wording. Varonis suggests gating outbound mail so a hijacked agent cannot relay from a trusted account, scoping connector access to the trust level of whatever triggered the task, and parking the riskiest actions behind a human.

Both teams land in the same place. Treat the agent like a junior hire with broad access and no instinct for what feels off, because the quality that makes it useful, its willingness to act on what it is told, is also the attack surface. Nobody has a clean fix for that yet.

OpenClaw AI agent tricked by contact cards and polite emails | RiskSense