Category field guide

Voice, audio & transcription

Synthetic voice, transcription, and persona tools used for narration, prototypes, accessibility, and production experiments.

Tools covered
4
Full reviews
1
Field notes
3

My position

How I decide what belongs in this layer

ElevenLabs, PlayHT, Whisper, and Resemble AI cover generation and transcription rather than one universal audio job. I have used this layer for multi-persona prototypes, narration planning, and turning recorded material into editable text.

The operational issues are as important as vocal quality: consent, voice rights, disclosure, pronunciation, speaker separation, cost at production volume, and whether the final audio still sounds appropriate for the audience.

The working set

4 tools, with the evidence level visible.

A field-note label is not a downgrade or a placeholder. It is the honest boundary between useful experience and a completed repeatable review.

Full reviewTested integration

ElevenLabs

High-quality synthetic speech for narration, differentiated personas, and localization experiments.

Convincing audio with real production value; consent and pronunciation review are non-negotiable.Read the profile →
Field notesTested and compared

PlayHT

A synthetic-voice option evaluated for narration, voice variety, and comparison with other text-to-speech platforms.

A credible voice-generation alternative; the right decision comes from the actual script, voice rights, pronunciation, and production volume.Read the profile →
Field notesUsed in workflow design

Whisper

A speech-to-text model used as a transcription layer for turning recordings into searchable, editable, and automatable text.

A strong transcription foundation when the workflow preserves the recording and reviews domain-specific errors.Read the profile →
Field notesPrototype evaluation

Resemble AI

A synthetic-voice and voice-cloning platform evaluated for custom voices, real-time use, and persona-oriented application work.

Interesting for controlled custom-voice applications; consent, provenance, abuse controls, and operational latency belong in the core design.Read the profile →

Selection rules

The rules I use before adding another tool.

Understandable, rights-aware audio that survives a real editorial review.

  1. 01Confirm rights before generation.
  2. 02Write for speech, not the page.
  3. 03Review every proper noun and command.
  4. 04Keep the original recording when transcription matters.