VentureBeat
Follow
Microsoft AI’s MAI-Transcribe-2 undercuts OpenAI, Google and ElevenLabs on price and speed
Microsoft has launched MAI-Transcribe-2, a new speech-recognition model that is faster, more accurate, and significantly cheaper than existing offerings from competitors like OpenAI and Google. This advanced model, priced at just ten cents per hour of audio, represents a substantial cost reduction from its predecessor. The company's strategy involves developing its own high-quality AI models, designed to replace those previously licensed from OpenAI. MAI-Transcribe-2 supports 60 languages and is built for the complexities of real-world business audio, including background noise and overlapping speech. Key features like speaker diarization, word-level timestamps, and keyword biasing are included at no extra charge. The model also offers configurable output styles and handles code switching between languages within a single conversation. Microsoft claims top rankings on benchmarks like FLEURS and Artificial Analysis, highlighting its performance in accuracy and speed. The rapid release cadence of three models in five months demonstrates Microsoft's accelerated development in this area. This push to build in-house AI capabilities is driven by a desire for independence and improved profit margins. By transcribing audio at a lower cost, Microsoft can integrate these features more affordably into its own products like Teams, Word, and Excel. This move positions Microsoft to compete more effectively in the AI market, reducing reliance on external partners. The company is directly challenging specialized transcription vendors by bundling essential features at an unprecedented price point. While Alibaba is noted as a strong competitor, Microsoft aims to offer a compelling alternative for enterprises.