Deep Dive
Voice versus dictation
The demo opens with a user asking ChatGPT to help with to-dos while music plays, establishing the conversational tone. OpenAI clarifies the core distinction: dictation converts speech into a prompt you send, while Voice mode maintains an ongoing conversation that persists as you work. Voice can see your screen when you request it, talk through projects with you in real time, and keep processing in the background while you switch between applications. This transforms the interaction from transactional to collaborative.
Cross-app context and background work
The user brainstorms Dev Day activation ideas with the AI, which then creates a pitch deck without the user leaving the conversation. When the user mentions a DX team off-site, Voice proactively reads the calendar to provide dates and times. The system combines context from Slack, Gmail, and calendar in parallel — it pulls the off-site request, checks availability on specific dates, and understands the full picture without the user manually opening each app. When the user asks about flights to Paris, Voice can stage bookings in a travel tool without executing them, keeping the human in final-decision mode while the assistant handles the legwork.
Maintaining flow state
A key feature is asynchronous notification — when work completes, Voice brings the user back with a voice notification rather than forcing them to check status. The user can ask ChatGPT to open the pitch deck in Chrome, then continue talking while the browser loads in the background. Once the deck is ready, the AI announces it vocally. The user can then decide to share it with teammates, review it further, or move to the next task. This model lets professionals stay in a single conversation thread and continuous mental flow instead of context-switching between windows and tools, which OpenAI positions as the fundamental advantage over older voice interfaces.