OpenAI
OpenAIJul 28
Tech

Introducing gpt-transcribe and gpt-live-transcribe

2 min video2 key momentsWatch original
TL;DR

OpenAI released gpt-transcribe and gpt-live-transcribe, two new models that handle 57 languages, multi-language switching, and domain-specific terms with better accuracy on accents, names, and background noise.

Key Insights

1

Both models support mid-session language switching without losing accuracy — you can code-switch between English and Spanish in the same conversation.

2

Vocabulary prompting for precisionDevelopers can inject domain-specific vocabulary (phishing, ARR, A1C) so the models prioritize correct terminology over common words.

3

Background noise filteringgpt-live-transcribe processes background noise and side conversations intelligently, filtering them out so only the primary speaker's words appear in the transcript.

Want this for every new video OpenAI posts? Brevyd summarizes each upload automatically, the morning it drops.

Deep Dive

Two models, two use cases

OpenAI introduced gpt-transcribe for batch processing complete audio files and gpt-live-transcribe for real-time streaming with an open connection. The batch model is built for call archives, podcasts, and large jobs where a 30-minute file processes in under a minute and accuracy matters more than speed. The streaming model targets captions, voice dictation, and live interfaces where latency directly impacts user experience. Both support 57 languages out of the box.

Accuracy where it usually fails

Both models handle notoriously difficult transcription cases better than predecessors: accents, code-switched speech, short answers, proper nouns, and numbers. Developers can provide domain-specific vocabulary lists to ensure the models prioritize industry terms like phishing, ARR, or A1C. The demo showed real-time switching between English and Spanish in the same session, with the model following both directions without degradation. Background noise handling is also improved, so cafes, conference halls, and loud offices won't pollute the final transcript.

Takeaways

  • Use gpt-live-transcribe for voice interfaces and real-time captions, gpt-transcribe for archival work and podcasts where you can wait 60 seconds.
  • Feed the models your domain vocabulary up front (terms from finance, healthcare, tech) so they weight those correctly over acoustic similarities.

Key moments

0:21Spanish code-switch mid-session

Y también puedo cambiar al español en la misma sesión. El modelo sigue transcribiendo en tiempo real y ahora vuelvo al inglés.

0:4330-minute file in under a minute

The model will take a little less than a minute to process it, and then the transcript is ready for whatever comes next.

You just read one. Brevyd does this for every upload.

Follow OpenAI and every new video comes back as a summary like this, in your morning briefing. No watching required.