IMG 20210102 170742 scaled

From macOS say to Azure Neural Voices: Picking a Text-to-Speech Voice for Video Overlays

· ·

I needed a voiceover for a short video overlay this week. Nothing fancy, just a few lines of narration to sit on top of a screen recording. My first instinct was the say command that ships with every Mac, because it’s already there and it takes about four seconds to get something talking.

That instinct was half right. say is brilliant for quick experiments, but the moment you want a nicer voice, and especially the moment you want to publish the result, things get more interesting. Here’s what I learned, in the order I learned it.

Choosing a voice with say

The say command takes a -v flag for the voice:

say -v Daniel "Hello there"

Daniel is the British English male voice and a decent default if you’re in the UK. To find out what else you have installed, pass a question mark as the voice name:

say -v '?'

That prints every installed voice along with its language code and a sample sentence. It’s worth scrolling through, because macOS ships with a lot of them, and some are surprisingly characterful.

Getting better voices

The voices installed by default are the compact versions. Apple offers much higher quality Enhanced and Premium versions for many of them, and they’re free to download:

  1. Open System Settings
  2. Go to Accessibility > Spoken Content
  3. Click the info button next to System Voice, then Manage Voices
  4. Tick the voices you want and let them download

Once downloaded, they appear in say -v '?' and you can call them by their full name, quotes included:

say -v "Serena (Premium)" "This sounds a lot better."

The difference between a compact voice and its Premium version is big enough that I’d always grab the Premium one if you’re going to listen to more than a sentence.

Other useful flags

A few flags turn say from a toy into something genuinely useful:

  • -r 180 sets the speaking rate in words per minute. The default is a little brisk for narration, so I found something between 160 and 180 easier to follow.
  • -o out.aiff writes the audio to a file instead of playing it through your speakers. This is the one you want for video work.
  • -f script.txt reads the text from a file, which is much easier than escaping a long script on the command line.

Put together, a usable narration command looks like this:

say -v "Daniel (Enhanced)" -r 170 -f script.txt -o voiceover.aiff

If you leave out -v entirely, say uses whatever is set as the System Voice in that Spoken Content panel. That detail turns out to matter for the next part.

What about the Siri voices?

The Siri voices are the best sounding voices on a Mac, so naturally that’s what I wanted for the overlay. The catch is that they don’t show up in say -v '?', and you can’t select them by name with -v.

There is a workaround. Because say falls back to the System Voice when you don’t pass -v, you can make a Siri voice the System Voice and then call say without the flag:

  1. Go to System Settings > Accessibility > Spoken Content
  2. Set System Voice to one of the Siri voices (it downloads the first time you pick it)
  3. Run say with no -v:
say -o voiceover.aiff "Your overlay script here"

On some macOS versions that’s all you need. On others, the Siri voices will happily speak out loud but refuse to render to a file, and you end up with silence or a fallback voice in your .aiff.

If that happens, the reliable fix is to capture the audio as it plays:

  1. Install BlackHole, a free virtual audio driver
  2. Set BlackHole as your output device
  3. Run say "your text" without -o
  4. Record the BlackHole input in QuickTime or Audacity

It’s clunky, but it works, and you get a clean recording with no room noise.

The licensing catch

This is the bit I nearly skipped past, and it’s the most important part of the whole post.

The macOS software licence covers the built-in system voices, Siri voices included. It allows you to use them while running macOS and to create content for your own personal, non-commercial use. It explicitly rules out recording, publishing or redistributing those voices in a commercial or public context, and that includes non-profit use.

So if your video is a personal project that never leaves your machine, carry on. If it’s going on a company website, a product demo, a YouTube channel, a client deliverable, or anywhere public, the Siri voice and every other built-in macOS voice is off the table. That ruled them out for me, because the overlay was for work.

The good news is that commercial text-to-speech has become very good, and in some cases very cheap.

Azure neural voices

I already run infrastructure on Azure, so Azure AI Speech was the obvious place to look. Its neural voices are close to the Siri voices in quality, there’s a wide choice of British English voices, and the output is licensed for exactly this kind of use.

Does it cost anything?

For a video overlay, almost certainly not. The free (F0) tier of Azure AI Speech includes 500,000 characters of neural text-to-speech per month. A few minutes of narration is a few thousand characters, so you could voice a lot of videos before paying a penny.

Beyond the free tier, standard neural voices are billed per character, at roughly $16 per million characters at the time of writing. Check the official pricing page for your region before relying on that figure.

Two things to watch:

  • HD voices, Custom Neural Voice and Personal Voice aren’t covered by the free allowance. Stick to the standard neural voices if you want to stay at zero.
  • You can only have one F0 Speech resource per subscription, so if you’ve created one before, reuse it.

The no-code route: Speech Studio

You don’t need to write any code to get an audio file out of Azure:

  1. In the Azure portal, create a Speech resource and choose the Free F0 pricing tier
  2. Open Speech Studio and pick Audio Content Creation
  3. Paste in your script
  4. Choose a voice; en-GB-SoniaNeural and en-GB-RyanNeural are good British starting points
  5. Adjust pacing, pauses and pronunciation in the editor
  6. Export as MP3 or WAV and drop it into your video editor

The editor is the real win here. You can slow a single phrase, add a pause before a key point, or fix how a product name is pronounced, all without regenerating the whole thing.

The command line route

If you’d rather stay in the terminal, as I usually would, the REST API is simple enough to call with curl. Grab the key and region from your Speech resource, then:

curl -X POST "https://uksouth.tts.speech.microsoft.com/cognitiveservices/v1" \
  -H "Ocp-Apim-Subscription-Key: $AZURE_SPEECH_KEY" \
  -H "Content-Type: application/ssml+xml" \
  -H "X-Microsoft-OutputFormat: audio-24khz-48kbitrate-mono-mp3" \
  -H "User-Agent: voiceover-script" \
  -d '<speak version="1.0" xml:lang="en-GB">
        <voice name="en-GB-RyanNeural">
          <prosody rate="-5%">Welcome to the demo. Let me show you how this works.</prosody>
        </voice>
      </speak>' \
  -o voiceover.mp3

Swap uksouth for whichever region your resource lives in. The <prosody> tag is optional; I found a small negative rate made narration feel less rushed, much like lowering -r with say.

SSML gives you the same control as the Speech Studio editor, just in text form. <break time="500ms"/> adds a pause, and <emphasis> leans on a word. For a longer script, I’d keep the SSML in a file and pass it with -d @script.ssml.

Which should you use?

Here’s where I landed:

  • Quick test, or just for you: say with a Premium or Enhanced voice. It’s instant and needs nothing installed.
  • Personal project that needs the nicest voice: a Siri voice set as the System Voice, with BlackHole as a fallback if file output misbehaves.
  • Anything published, commercial or client facing: Azure neural voices on the free tier. Similar quality, properly licensed, and very likely free at video overlay volumes.

Other commercial services such as ElevenLabs and Google Cloud Text-to-Speech fill the same role and are worth a look if you’re not already on Azure. For me, having it sit in a subscription I already manage made the decision easy.

The lesson I’ll take away is that the licence deserves as much attention as the audio quality. The Siri voices sound lovely, and it would have been very easy to drop one into a work video without thinking twice. A couple of minutes reading the terms saved me from that, and the alternative turned out to cost nothing anyway.


Leave a Reply