
AI Models
Microsoft VibeVoice Speech-to-Text Model Tested
Microsoft's VibeVoice offers speech-to-text transcription with speaker diarization under an MIT license. Developer Simon Willison ran it on a Mac using a quantized MLX version, processing an hour of podcast audio in under nine minutes. The model outputs detailed JSON with timestamps and speaker labels, though it limits input to one hour.
Apr 283 minNeura News