OpenAI
OpenAIJul 29
Tech

Using Voice in ChatGPT Work

3 min video3 key momentsWatch original
TL;DR

ChatGPT Voice lets you talk to the AI while working across apps, access screen context, and handle multiple tasks without leaving your flow.

Key Insights

1

Voice differs fundamentally from dictation — it's a persistent assistant that stays with you across apps and can see your screen on demand, not just a text input method.

2

Connected apps contextChatGPT Voice can access multiple connected tools like Slack, Gmail, and calendar simultaneously to gather context without forcing you to navigate between apps manually.

3

Background executionTasks can run in the background while you switch between apps — the AI brings you back with voice when something's ready, letting you stay in flow state.

Want this for every new video OpenAI posts? Brevyd summarizes each upload automatically, the morning it drops.

Deep Dive

Voice versus dictation

The demo opens with a user asking ChatGPT to help with to-dos while music plays, establishing the conversational tone. OpenAI clarifies the core distinction: dictation converts speech into a prompt you send, while Voice mode maintains an ongoing conversation that persists as you work. Voice can see your screen when you request it, talk through projects with you in real time, and keep processing in the background while you switch between applications. This transforms the interaction from transactional to collaborative.

Cross-app context and background work

The user brainstorms Dev Day activation ideas with the AI, which then creates a pitch deck without the user leaving the conversation. When the user mentions a DX team off-site, Voice proactively reads the calendar to provide dates and times. The system combines context from Slack, Gmail, and calendar in parallel — it pulls the off-site request, checks availability on specific dates, and understands the full picture without the user manually opening each app. When the user asks about flights to Paris, Voice can stage bookings in a travel tool without executing them, keeping the human in final-decision mode while the assistant handles the legwork.

Maintaining flow state

A key feature is asynchronous notification — when work completes, Voice brings the user back with a voice notification rather than forcing them to check status. The user can ask ChatGPT to open the pitch deck in Chrome, then continue talking while the browser loads in the background. Once the deck is ready, the AI announces it vocally. The user can then decide to share it with teammates, review it further, or move to the next task. This model lets professionals stay in a single conversation thread and continuous mental flow instead of context-switching between windows and tools, which OpenAI positions as the fundamental advantage over older voice interfaces.

Takeaways

  • Use Voice mode for multi-step projects where you need to reference existing data — calendar, email, Slack — without manually opening each app.
  • Stage high-stakes decisions like bookings or drafts in Voice without auto-executing; let the AI prepare them for your final click.
  • Start with Voice when you're already in flow on a primary task and need a thinking partner that won't break your focus.

Key moments

0:45Voice versus dictation defined

Dictation turns what I say into a prompt. Voice is an ongoing conversation that stays with me while I work.

1:30Cross-app context gathering

Voice can use its own on-screen context, connected apps, and can gather the other context it needs for tasks in the background.

2:40Async notifications keep flow

When something is ready, a voice brings it back, and I decide whether to redirect it, share it, or take it to the next step.

You just read one. Brevyd does this for every upload.

Follow OpenAI and every new video comes back as a summary like this, in your morning briefing. No watching required.