Add a key only for the services you use; their models then appear in the toolbar picker. Your audio is sent to the service you pick.

Only filled-in keys are saved.

Speech recognition that runs on this Mac: your audio stays here.

Speech models

How each speech model has done on your own dictations.

ModelWait per audio minuteDictionary correctionsRunsAudio
Wait: how long transcription took after you stopped, per minute of audio; fast mode uploads while you record, and processing is timed separately. Corrections: dictionary replacements per 100 words, in dictations where the dictionary ran. Neither is an accuracy score.

Corrections & formatting

What the chosen decision model did after speech recognition, across all dictations and every speech model. It counts work done, not accuracy.

Heard as What it is Scope
Pinned · every speech model, protected from suggestions Learned · learned for this model only; the list follows the toolbar picker

How the dictionary works

A short guide to fixing the words your speech model gets wrong.

What it's for

Every speech model has words it keeps getting wrong: names, products, technical terms. Your dictionary remembers those mistakes and fixes them in each new dictation, before the text reaches you.

How an entry works

An entry pairs what the speech model writes with what you might have meant. Say it often writes “cloud” when you said “Claude”. But sometimes you really mean a cloud. So the entry keeps both meanings:

cloud → Claude, an AI assistant you ask to write or code
cloud → cloud, internet storage and servers

In “Ask cloud to draft the reply”, Entune writes Claude. In “The backups are in the cloud”, it leaves cloud alone. It chooses from the surrounding words, which is why each meaning has a short description: write what the thing is and what it goes with.

One entry can list several spellings the model writes (“cloud”, “clawed”) and several meanings. Each heard spelling is linked to the meanings it can stand for.

Getting suggestions

You don't have to build the dictionary by hand. Entune can read your dictations and suggest entries, using the dictionary model you choose (its key is in Settings).

Choose Get suggestions. What it looks for depends on your dictionary:

  • With no entries yet for this speech model, it reads your recent transcripts and suggests first entries for the mistakes it finds.
  • Once you have entries, pinned or learned, it also compares what the speech model wrote with what the dictionary changed. It suggests missing entries and fixes: a clearer description, a wider or narrower entry, or removing a learned entry that does more harm than good.
  • Learning from audio, in the ⋯ menu, transcribes recordings again with your current speech model before suggesting entries. Use it when the recordings were made with another model, or come from another dictation app or a folder. It takes longer and uses more tokens.

Each speech model makes its own mistakes, so suggestions are kept for the speech model selected in the toolbar.

Each run reads up to 300 of that model's newest transcripts that no earlier run has read, so running it again moves on to newer dictations. To read transcripts again, turn on Include transcripts already used in the suggestions panel or the ⋯ menu; that costs tokens again and can repeat earlier suggestions.

Learning from audio

Choose a source in the ⋯ menu: your Entune recordings made with other speech models, the recordings another dictation app kept on this Mac (only the audio is copied, never its transcripts), or a folder of audio files. For Entune recordings you can include the ones this speech model already transcribed.

After you pick a source, the timeline shows all its recorded audio, oldest on the left and newest on the right. It measures recorded time, so days without recordings take no space; the dates under the handles show where you are.

Everything is selected at first. Drag the left handle right to keep only your most recent audio, or move either handle to choose any stretch. You'll see exactly how much audio and how many recordings are included, and you can list and play them. Each edge snaps to whole recordings.

Then choose Transcribe and suggest. Entune transcribes the selection with your speech model, then the dictionary model reads the new text. Transcribing again can take a while and your providers may charge for both steps. You can stop at any time; finished transcripts are kept for a retry.

Reviewing suggestions

Suggestions open in a panel beside your dictionary, grouped into new entries, changes and removals. You can edit any of them, remove the ones you don't want with ×, and then apply the rest in one step. Nothing is saved until you apply. While suggestions are open, dictation pauses, so apply or discard them when you're done.

Pinned and learned entries

Learned entries are used only with the speech model they were learned for. Pin an entry to use it with every speech model; pinned entries are yours and suggestions can't remove them. Pinning doesn't make an entry win: Entune still reads the sentence.

Adding an entry yourself

Choose Add and fill in one line: how it's written, what the speech model hears (several, separated by commas) and what it is. For several meanings, a spelling kept as written, or always replacing, open the full editor from the link under that line. Edit on an entry opens the same editor.

Edit entry

Several meanings, keep as written, links to other entries, always replace.

Get suggestions

Suggestions from your recent transcripts. Nothing changes until you review them.

Dictionary generation can take minutes to hours for a large history. It processes batches in sequence; the model and size of the growing dictionary affect the time.

Reads transcripts an earlier run already read. It costs tokens again and can repeat earlier suggestions.

Nothing changes in your dictionary until you apply. While a run or review is open, dictation and dictionary editing are paused.

Shortcuts

Hold-to-talk
One key. Records while held, release to stop.
Hands-free
A combination. Press to start, again to stop.
Cancel dictation
Discard the active shortcut recording.

Appearance

Theme
Follow the system, or fix one.
Text size
Scales the whole window. ⌘+ ⌘− step it, ⌘0 resets.

After speech recognition, Entune can fix known mistakes and tidy the text. A decision model makes each choice: it answers questions about the text and never writes any.

Decision model
Needed for the steps below. Choose where the decisions are made.
The original transcript is always kept in History.
Advanced: waiting and retries
All three steps share the maximum wait, including retries.
Suggestion model
Writes dictionary suggestions when you choose Get suggestions on the Dictionary page. It reads the transcripts it learns from. Dictation and editing work without it.
No model

Local API

Apps and agents you dictate to can send corrections you have confirmed. Each one is added to your dictionary as a pinned entry.
Endpoint
HTTP on this machine only. Nothing else is exposed.
Example request
Send only what the person confirmed, whole words or phrases, never guesses.

Received corrections

Newest first. Each one is already in Pinned; remove it there if it's wrong.

What Entune keeps on this Mac, and what it sends elsewhere.

What leaves this Mac

Speech recognition
Audio goes to the cloud service you pick. Models on this Mac keep it here.
Corrections & formatting
When on, transcript text, the words in question and relevant dictionary meanings go to TypeSafe if Jev is your decision model. With Laya they stay on this Mac.
Dictionary suggestions
When you ask, transcripts and your dictionary go to your suggestion model. Transcripts made only for a suggestion run are not saved by Entune.
Each service's own retention rules apply.

Your data stays yours

No expiry
Entune never automatically expires or deletes your saved audio or transcripts. This includes audio imported to build dictionaries. Saved files are kept locally.
Original audio
A ZIP of all recordings and imported dictionary audio. Large exports can take a while to prepare.
Export audio
Transcripts
JSON with saved text, raw transcripts, models, dates and every attempt. Transcripts used only to build a dictionary from imported audio are not saved.
Export transcripts
Exports are created on this Mac. Saved API keys and settings are excluded.

Start over

Delete all Entune data
Recordings, transcripts, the dictionary, settings, API keys and downloaded models on this Mac. You can export first.

How corrections and formatting work

What happens to a transcript after speech recognition.

The steps

Each transcript goes through the steps you turned on, in this order: your dictionary, repeated fillers, then paragraphs and bullets. The original transcript from the speech model is always kept in History, and you can copy it from there.

Your dictionary

For each word your speech model often gets wrong, the decision model reads the sentence around it and picks the meaning that fits: “Ask cloud to draft it” becomes “Ask Claude to draft it”, while “the backups are in the cloud” stays. Entries set to always apply skip that choice. When no meaning can be chosen, the words stay as they were.

Fillers and formatting

Repeated fillers such as “um um” become one when the decision model hears hesitation; quoted text and code are left alone. Paragraph breaks and bullets are added only where the dictation clearly has them, and every word is kept.

Waiting and failures

The steps share one time limit, including retries; you can change it under Advanced. If the dictionary step fails, you get the original transcript. If fillers or formatting fail, that step is skipped and the text before it is kept.

The decision model

A decision model answers questions about the text and never writes any. For fillers and formatting it receives the whole transcript; for your dictionary, the passage around each word in question with the meanings that could apply, including your personal usage notes. It never receives audio.

Two are available. Jev, from TypeSafe, runs in the cloud on your own TypeSafe key, so the text goes to TypeSafe even when speech recognition runs on this Mac. Laya, an open-weight model from Convai Innovations, runs on this Mac for English, so the text stays here. Its engine is installed once in Terminal, and Entune runs it only while Laya is chosen and a step is on. Laya reads a limited amount of text per question, so in a long dictation its filler and paragraph decisions see only part of it.

How your data works

What stays on this Mac, and what goes where.

On this Mac

Entune keeps its data in one folder on this Mac: your recordings, every transcript (the speech model's original and each processed version), your dictionary, settings, and the API keys you enter. Keys are stored in Entune's database, not in the macOS keychain. Entune has no accounts and sends no usage data.

Audio you import from another dictation app or a folder is copied into that folder; your original files are not changed or moved.

Speech recognition

With a cloud speech model, such as AssemblyAI, Groq or Soniox, each recording's audio is sent to that service to be transcribed. With fast mode, audio goes to AssemblyAI while you are still speaking. Transcribing again with another model sends the audio to that model's service.

With a model on this Mac, such as Whisper or Parakeet, audio stays here. Downloading one fetches the model from Hugging Face.

Corrections and formatting

When Apply your dictionary, Reduce repeated fillers, or Paragraphs and bullets is on, transcript text goes to your decision model: the whole transcript for fillers and formatting, and for your dictionary the passage around each word in question with the meanings that could apply. Audio is never sent.

With Jev, the text goes to TypeSafe, even when speech recognition runs on this Mac. With Laya, it stays on this Mac; its first start fetches the model from Hugging Face.

Dictionary suggestions

Only when you choose Get suggestions, the transcripts it reads and your dictionary are sent to your suggestion model (Anthropic, OpenAI through an API key or a ChatGPT subscription, Google Gemini, Groq or Mistral). With Learn from audio, the chosen audio is first transcribed by your speech model as above. Transcripts made only for a suggestion run are not saved.

Keeping and deleting

Entune never deletes recordings or transcripts on its own. Exports are created on this Mac and leave out your keys and settings.

What a cloud service keeps, and for how long, is up to that service and your account with it. Deleting data in Entune can't remove copies a service kept, exports you saved, or copies in Time Machine or other backups of this Mac.

Delete all Entune data

This can't be undone.

Everything Entune keeps in its data folder is deleted:

  • Recordings and their audio
  • Transcripts: originals, processed versions and every attempt
  • Audio copied in from other dictation apps or folders
  • Your dictionary, its automatic copies and suggestion progress
  • Corrections received from other apps
  • Settings, shortcuts and saved API keys
  • Downloaded models, for speech and for Laya
  • The backups folder, and the log (emptied)

Not affected: your original audio files, exports you saved, anything a cloud service kept, Time Machine or other backups, and the Entune app itself. Files are deleted the normal way; this is not a secure erase.

Want a copy first? Export audio Export transcripts