第一财经

OpenAI Updates Its Speech Model, Bringing the User Experience on the Consumer Side in Line with Google and DouBao

原文:OpenAI更新语音模型,C端交互体验向谷歌、豆包看齐

Summary of Key Points

OpenAI has introduced a new generation of voice model, GPT-Live, which addresses the issues of high latency and mechanical disconnect in previous voice interactions. With its full-duplex architecture, it enables natural conversation where both listening and speaking occur simultaneously, and it can also leverage GPT-5.5 to handle complex tasks. Along with Google’s Gemini Live and ByteDance’s DouBao voice products, this indicates that voice has evolved from a “nice-to-have” feature of AI to a core means of attracting and retaining users. The user experience on the consumer (C) side has become more accessible (for example, even the elderly can use it), while the business (B) side offers potential new revenue growth through intelligent agents.

1. GPT-Live Finally Solves OpenAI’s Voice Problems

Previously, OpenAI’s voice functionality worked in a three-step process: first converting your speech to text (ASR), then having the AI generate a response (LLM), and finally converting the text back to speech (TTS). This approach was not only slow (with long delays) but also missed out on many details—such as the tone, emotion of your voice, or even background noises. As a result, the responses were always cold and robotic, resembling conversations with a robot reading from a script.

Later, GPT-4o tried to combine these steps into one model, which reduced latency but resulted in an unstable experience. With GPT-Live, the architecture has been completely restructured to use full-duplex technology, allowing the AI to listen to you speak and respond at the same time. For instance, it might say “Hmm hmm” to show it’s listening, or wait quietly while you think, making the conversation feel more like a conversation with a real person and providing more accurate emotional recognition.

2. Two Key Technical Advancements: Full-Duplex Architecture + GPT-5.5 Support

The main highlights of GPT-Live are:

1. Full-duplex Architecture: This breaks the traditional pattern of waiting for you to finish speaking before responding. For example, if you ask “How’s the weather today? Also, recommend a nearby café,” the AI starts processing the question about the weather as soon as it hears it. By the time you finish asking about the café, it may already have prepared the answer and can seamlessly provide recommendations without any delay.

2. Backend Support with GPT-5.5: For complex tasks (like analyzing the market trends for new energy vehicles), the AI can quietly call on the more powerful GPT-5.5 to gather information and perform advanced reasoning, ensuring the conversation remains smooth without any interruption.

3. Comparison with Google and ByteDance: Similar Experiences, Different Technical Approaches

The main players in real-time voice interaction technology are three companies:

  • Google Gemini Live: Uses a unified multimodal model (handling audio, video, and text together) and is accelerated by its own TPU chips. The advantage is strong multimodal integration, but it lacks specialized optimization for voice interactions.
  • ByteDance DouBao Voice: Utilizes its proprietary “Seeduplex” voice model designed specifically for voice interactions and was one of the first to implement full-duplex technology.
  • OpenAI GPT-Live: Now also includes full-duplex functionality and combines it with the powerful reasoning capabilities of GPT-5.5.

Although all three offer real-time conversations, their technical approaches differ, indicating that the “optimal” solution for voice interaction has not yet been found. However, there’s consensus that voice is no longer a dispensable feature but a crucial means to attract and retain users.

4. Voice Becomes the “Traffic Password” of AI: Great Future Potential

Currently, 150 million people use ChatGPT’s voice features, and one of the team members mentioned chatting with the AI for 30-40 minutes (similar to chatting with a friend while walking). In the future, voice could become the primary mode of interaction. For example, you might manage complex tasks (like scheduling your week or monitoring smart home devices) without having to type; you could simply converse with the AI.

This is significant for AI companies:

  • User Growth: Voice makes technology more accessible to a wider range of users, including the elderly and children.
  • Hardware Implementation: Smart speakers, robots, and in-vehicle AI systems all require voice interaction.
  • Commercialization: Enterprises can use voice-based intelligent agents for customer service, office tasks, and other applications, creating new revenue opportunities.

5. More Accessible on the C Side, with Business (B) Side Intelligent Agents Being the Next Frontier

User feedback indicates that GPT-Live’s experience on the consumer side is very user-friendly, on par with ByteDance DouBao and Google Gemini. The real challenge lies on the business side: can GPT-Live integrate voice interactions effectively with intelligent agents for long-term, complex tasks (such as customer service or data analysis)? If it can, this could become a new revenue source for OpenAI, as business customers are willing to pay for more professional services.

In summary, the launch of GPT-Live marks a shift from “usable” to “user-friendly” AI voice interaction, suggesting that future conversations with AI will be as natural as talking to people around us.