Open AI rolls out new model that enables voice chats with ChatGPT

OpenAI CEO Sam Altman / Photo credit: Y Combinator
Open AI has unveiled GPT-4 omni (GPT-4o), the first version of its advanced AI model.
GPT-4o isn’t confined to text, the AI research lab says. It can accept input in any combination of text, audio, and image, and can churn out any combination of these media.
With this new model, users can have a conversation with ChatGPT directly. “The new voice (and video) mode is the best computer interface I’ve ever used,” OpenAI CEO Sam Altman said in a blog post. “Getting to human-level response times and expressiveness turns out to be a big change.”
With an average response time of 320 milliseconds, GPT-4o is designed to match human response times in conversations. It outperforms its predecessor, GPT-4 Turbo, with non-English content, but matches it in performance on text in English. GPT-4o is also 50% cheaper to deploy.
Safety was a priority in GPT-4o’s development, which included post-training model refinement and training data filtering, OpenAI said. Over 70 external experts reviewed the model as well.
GPT-4o can be accessed on the ChatGPT platform beginning today. Developers can also access GPT-4o in the API as a text and vision model, with its new audio and video capabilities slated to launch later to a group of trusted partners.
The new model’s launch comes after OpenAI rolled out Sora, a model that can create videos from text, in February.
See also: What I learned about genAI from failing to scale my AI startup
Editing by Eileen C. Ang
(And yes, we’re serious about ethics and transparency. More information here.)
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.




