Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

OpenAI updates ChatGPT Atlas to counter manipulation attacks

OpenAI has released a security update for ChatGPT Atlas’s browser agent to address new prompt injection attack risks.

ChatGPT Atlas allows users to automate browser actions, increasing its exposure to adversarial attacks.

OpenAI said the update includes an adversarially trained model and enhanced safeguards, prompted by internal automated red teaming that uncovered new attack methods.

Prompt injection attacks embed malicious instructions into content, aiming to manipulate AI agents to perform unintended actions.

OpenAI uses reinforcement learning with large language models to simulate attacks and identify vulnerabilities before they can be exploited.

The company said prompt injection remains an open security challenge, and expects ongoing work on defenses.

OpenAI advises users to limit logged-in access, carefully review agent confirmation prompts, and give specific instructions to reduce risks when using agent features.

🔗 Source: OpenAI

🧠 Food for thought

Implications, context, and why it matters.

OpenAI’s reinforcement learning red teaming finds multi-step attacks

  • OpenAI built a large language model (LLM)-based automated attacker trained with reinforcement learning for automated red teaming (simulated adversary testing). It found new tactics 1.
  • The bot runs attack trials in simulation and watches AI replies. It steered AI browser agents into harmful workflows spanning tens or hundreds of steps 1.
  • In one demo, the attacker slipped a malicious email into a user’s inbox. The AI agent followed hidden instructions and sent a resignation instead of an out-of-office reply 1.
  • In a demo after security updates, ChatGPT Atlas’s browser agent mode detected prompt injection and flagged it to the user, the company said 1. The UK’s National Cyber Security Centre said such attacks “may never be totally mitigated” 1.

Opportunities for third-party security testing as agentic browser vulnerabilities persist

  • The prompt injection problem leaves room for security startups and testing providers to build automated red teaming tools for agentic browsers (AI-controlled browser agents) 2. Frameworks like Garak (an open-source LLM security scanner) and PyRIT (an open-source adversarial testing toolkit) target general LLM issues 2.
  • Enterprise security teams can stand out with continuous monitoring that mixes automated attack simulation with expert analysis, since some testing tools produce false positives that require tuning plus interpretation 2.
  • Investors should watch the risk value trade-off for agentic browsers. Experts say they do not yet deliver enough value to justify access to email or payment data 1. Security tools may mature before browser agent use.

Recent OpenAI developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.