Apple's new SpeechAnalyzer API achieved a word error rate (WER) of 2.12% on clean speech and 4.56% on noisy speech in benchmark tests against Whisper models and Apple's previous SFSpeechRecognizer. The tests, conducted on 5,559 standard utterances, showed SpeechAnalyzer running about three times faster than Whisper Small while delivering higher accuracy, according to get-inscribe.com.
The benchmark compared Apple's SpeechAnalyzer, Whisper Small, Whisper Base, Whisper Tiny, and the legacy SFSpeechRecognizer using the LibriSpeech dataset. SpeechAnalyzer outperformed all Whisper models, including Whisper Small, which had a WER of 3.74% on clean speech and 7.95% on noisy speech. The legacy SFSpeechRecognizer ranked last with a 9.02% WER on clean speech. All engines ran fully on-device on an Apple M2 Pro with macOS 26.5.1, ensuring consistent testing conditions, get-inscribe.com reported.
The results highlight Apple's advancement in on-device speech recognition technology, delivering both improved accuracy and speed compared to open-source Whisper models. The SpeechAnalyzer's smaller model size and faster processing could benefit applications requiring real-time transcription and voice commands. This development positions Apple ahead in the competitive speech recognition space, where accuracy and efficiency are critical for user experience and privacy.
The benchmark data and transcripts for all 5,559 test utterances were released publicly by get-inscribe.com on July 13, providing transparency and a resource for further research and development in speech recognition technologies.