Skip to main content
Transcribe audio to text using Whisper and other speech recognition models.
WhollyAPI hosts Whisper and other speech recognition models. Given an audio file, they produce transcribed text with per-sentence timestamps. Browse all speech recognition models.

Models

  • openai/whisper-large — best accuracy
  • openai/whisper-medium, openai/whisper-small, openai/whisper-base — faster, lighter
  • openai/whisper-timestamped-medium — per-word timestamp segmentation

Example

Supported audio formats

  • mp3
  • wav

Response

Additional parameters

Each model exposes different parameters (language, task, etc.). Check the model’s API documentation page for details.