OpenAI is moving ChatGPT's voice capabilities beyond conversation and into tasks that require multiple online steps. On Wednesday, the company will bring its more advanced GPT-Live voice system to ChatGPT's web and mobile apps, allowing users to issue spoken instructions for tasks involving browsers and multiple applications. The update expands OpenAI's presence in the personal AI assistant market and supports its longer-term hardware ambitions.
GPT-Live Connects Voice With Task Execution
The launch connects OpenAI's voice model, GPT-Live, with an AI model capable of operating a computer or browser. GPT-Live can listen to users and respond in real time. It was previously used in ChatGPT's standard conversational voice mode and has already been linked to task-execution models in the desktop application.
Atty Eleti, who leads ChatGPT's voice product, demonstrated how the system works. Using only voice commands, he asked it to analyze his DoorDash spending, pay $5 to reserve a pickleball court and reorganize his calendar. The sequence involved spending data, a payment and personal scheduling, showing that the voice feature is designed not only to answer questions but also to complete a series of online actions for the user.
Eleti said OpenAI wants people to interact with AI primarily through voice in the future, making speech one of the main interfaces between users and machines. He described that as a long-term direction for the team.
More Than 150 Million Users Have Used Voice and Dictation
Eleti said more than 150 million people currently use ChatGPT's dictation and voice-mode tools. Voice input can be faster and more natural than typing, allowing users to capture ideas without picking up a device or handle tasks while commuting or walking.
OpenAI also stressed that voice mode does not require every input and response to be spoken. Users can make a request verbally, while ChatGPT may respond with text, audio or a combination of both depending on the task. Eleti said the team considered a practical difference when designing the experience: speaking is generally faster than typing, while reading text is often faster than listening to a complete spoken response.
Personal Assistant Competition Moves to Mobile Use
Personal AI assistants have drawn increasing attention in Silicon Valley. Products such as Instinct and Meta's Muse are also being developed to handle routine online tasks on behalf of users. ChatGPT already offers many of these capabilities, but OpenAI has not launched a separate personal assistant product with a narrower focus. The voice update places its existing task-execution tools behind an interface better suited to mobile use.
In the near term, users can hand off some online tasks while walking, commuting or unable to operate a device with both hands. Voice mode may also be useful for activities that require speaking, such as language learning and speech practice. Members of OpenAI's product team have used the feature while commuting and walking.
Device Constraints and User Habits Remain Factors
Voice assistants still face practical limits. The smart-speaker market has shown that voice interaction can retain a user base, but most office workflows remain built around desktop computers and keyboards. Voice input may be convenient without being suitable for every task involving extensive data checks, sensitive information or multiple confirmation steps.
OpenAI's hardware plans are also pushing its voice capabilities toward standalone devices. The company plans to launch a smart speaker in 2027. If that plan goes ahead, the combination of GPT-Live and task-execution models could become a key interaction method for the hardware. For now, however, users mainly access these capabilities through ChatGPT's web, mobile and desktop apps. The scope of payments, account authorization and cross-application actions will continue to depend on product permissions and the user's operating environment.