🧔♂️ A friendly human may check it before it goes live. More news here
OpenAI unveils GPT-5.3 Codex-Spark for real-time coding
OpenAI has announced a research preview of GPT-5.3-Codex-Spark, a smaller, fast inference version of GPT-5.3-Codex designed for real-time coding tasks.
The model, developed in partnership with Cerebras, runs on Cerebras’ Wafer Scale Engine 3 hardware, enabling high-speed responses of over 1,000 tokens per second.
It features a 128k context window and is currently text only.
The model aims to support interactive coding workflows, allowing users to make targeted edits and see immediate results.
OpenAI reports improvements in latency and response streaming, reducing overhead by up to 80%.
Codex-Spark is available to ChatGPT Pro users through the Codex app, CLI, and VS Code extension, with usage governed by separate rate limits.
The company plans to expand access and introduce additional capabilities, including larger models and multimodal inputs, in future updates.
🔗 Source: OpenAI
🧠 Food for thought
Implications, context, and why it matters.
OpenAI’s speed boost goes beyond a new chip
- OpenAI and Cerebras signed a multi-year deal that OpenAI said last month is worth over $10 billion, so this is not a small trial 1.
- Codex-Spark clears 1,000 tokens per second by pairing Cerebras’ architecture with the Wafer Scale Engine 3 (WSE-3) for low-latency inference (running an AI model to generate outputs) 2.
- That pace comes with lower benchmark results, since OpenAI says Codex-Spark trails GPT-5.3-Codex on SWE-Bench Pro and Terminal-Bench 2.0 (coding-focused evaluation benchmarks) even while finishing tasks in a fraction of the time 3.
- Software work also mattered, since OpenAI cut client-server roundtrip overhead by 80% and time-to-first-token by 50% using infrastructure changes like a persistent WebSocket connection (a way to keep an app and server continuously connected) 4.
Specialized AI hardware and faster interactions are gaining ground
- OpenAI is spreading its hardware bets, using task-specific chips for low-latency inference while GPUs still carry training and other workloads 5.
- This leaves room for specialized chip companies to challenge Nvidia by winning narrow use cases where fast responses matter most 5.
- Real-time feedback can shift expectations, moving products away from prompt-and-wait agents toward interactive and collaborative AI experiences 1.
- Developers can split work across models, using fast systems for quick steps while handing longer or harder jobs to stronger background models 1.
Recent OpenAI developments
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.




