
OpenAI announced on X today (July 29) that it is launching two transcription models, GPT-Live-Transcribe and GPT-Transcribe, with support for API access.

Pricing
GPT-Live-Transcribe costs $0.017 per minute for transcription (note: approximately 0.12 yuan at the current exchange rate)
GPT-Transcribe costs $0.0045 per minute for transcription (approximately 0.03 yuan at the current exchange rate)
OpenAI stated that compared with the company's previous models, these two transcription models can better understand context and provide more accurate transcription for real-world audio in various accents and languages, including phrases, numbers, technical terms, and speech with loud background noise.
GPT-Live-Transcribe is built specifically for low-latency, real-time transcription, while GPT-Transcribe is optimized for asynchronous transcription of completed audio files and batch workloads.

On the Context Aware ASR benchmark, GPT-Transcribe's semantic accuracy increased from 41.6% without free-form context to 45.2% with context.

Across 22 languages in Common Voice, GPT-Transcribe had a transcription error rate of 19.27%, compared with 40.37% for Whisper (whisper-1). Across nine languages in a real-world audio recording benchmark, it achieved a transcription error rate of 8.98%, compared with 15.21% for Whisper (whisper-1).
