All posts

Turn a voice note into text without sending it away

I use voice notes when the thought is moving faster than my hands. The annoying part usually comes next: find the recording, upload it somewhere, wait for a service to process it, then wonder how long that audio is going to sit on somebody else's server.

Voice Transcribe is my answer to that gap inside Droppy. It records from the Mac or accepts an existing audio file, turns the speech into text locally, and gives you the transcript at the Shelf. The audio and the transcription stay on your Mac during processing.

That privacy claim has a specific boundary. Droppy downloads the Voice Transcribe runtime and whichever model you choose before the first transcription. Live mode has its own Parakeet streaming model download. Those downloads need an internet connection. Once the required files are installed, the actual recording and speech-to-text work run locally.

Start from the Shelf, menu bar, or a shortcut

Voice Transcribe can live as a widget in the Shelf. Start a recording there and the Shelf becomes the recording surface, with a waveform, elapsed time, and a stop button. When you stop, the same area moves into processing and then the result.

Voice Transcribe recording in Droppy's Shelf with a waveform, elapsed time, and stop button
A recording stays attached to the Shelf instead of opening a separate recorder window.

You can also enable its menu bar icon. The menu offers Quick Record, Invisi-record, and Upload Audio File. Quick Record follows the normal recording flow. Invisi-record starts without opening a separate window, which is useful when you want to capture a thought without moving away from the app in front of you. Click the menu bar item while either mode is recording to stop.

Both recording modes can have their own global keyboard shortcut. I like that split because the visible mode is reassuring for a longer note, while the invisible mode makes a short capture feel almost instant. The shortcuts are user-defined in Voice Transcribe settings rather than fixed combinations I chose for everyone.

Tap a step

Start a recording from the Shelf, the menu bar, or a global shortcut.

  1. Press Record, Quick Record, or Invisi-record.
  2. Watch the waveform and elapsed time, then stop.
  3. Whisper or Parakeet transcribes the speech locally.

Send an existing audio file through the same local pipeline.

  1. Choose Upload Audio File from the menu bar.
  2. Pick an MP3, WAV, M4A, or AIFF in the Mac picker.
  3. Only that file enters transcription, still on your Mac.

Decide how long each recording stays once its transcript is ready.

  1. Set retention to Don't Save, 1 day, 7 days, 30 days, or Forever.
  2. Don't Save clears the audio once a transcript succeeds.
  3. Copy the result, or turn on Auto-Copy Result to skip that step.

The same local flow, from first capture to what you keep.

Choose Whisper or Parakeet

The regular transcription path offers two local backends. Whisper has Tiny, Base, Small, Medium, and Large v3 models, ranging from roughly 75 MB to 3 GB. The smaller models favor speed; the larger ones ask more of the Mac in exchange for accuracy. Parakeet uses one multilingual v3 Core ML model described in the app as a 25-language model.

Whisper also exposes an explicit language picker. The current choices are Auto Detect, English, Dutch, German, French, Spanish, Italian, Portuguese, Polish, Russian, Chinese, Japanese, Korean, Arabic, Hindi, Turkish, Ukrainian, Swedish, Danish, Norwegian, and Finnish. Parakeet handles multilingual detection itself, so that picker is disabled when Parakeet is selected.

I do not want model setup hidden behind the first recording. Voice Transcribe shows the runtime and model state in settings, then gives you a clear install or download action. A normal recording only becomes ready after the shared runtime and selected batch model are present.

Live transcription is a separate Parakeet mode

Parakeet can optionally show words while you speak. This streaming path is English only and uses its own on-device Core ML bundle. It does not use the Whisper models, and it is not available when Whisper is selected.

The live model is another deliberate download. You enable Live Transcription in settings and download that model there before starting. Droppy does not silently begin a large download when you press Record. During a live session, finalized phrases and the current in-progress phrase appear in the Shelf as the model hears them.

The result is ready to use

After a batch recording stops, the Shelf shows transcription progress. When the text is ready, the result view shows the transcript and its word count, with a Copy button that puts the full text on the clipboard.

Voice Transcribe result in Droppy's Shelf with transcribed text and a Copy button
The finished transcript is readable in place and can be copied straight to the clipboard.

If every recording ends with the same copy action, turn on Auto-Copy Result. Droppy then skips the result window, copies the completed transcription immediately, and leaves it ready to paste into Notes, Messages, an email, or the document you were already writing.

Upload Audio File uses the same local pipeline for an existing MP3, WAV, M4A, AIFF, or other audio file accepted by the Mac picker. That makes Voice Transcribe useful for more than fresh microphone recordings without changing the privacy boundary.

You decide what happens to the recording

Local transcription does not automatically mean every recording should live forever. The retention setting offers Don't Save, 1 Day, 7 Days, 30 Days, and Forever. Don't Save removes the captured audio as soon as a successful transcript is produced. The timed choices clean up older recordings after their window, and Forever leaves removal to you.

Voice Transcribe settings showing locally retained recordings with dates and durations
Saved recordings remain visible in settings with their date and duration, subject to the retention policy you choose.

Changing to a shorter retention period applies the policy to recordings already on disk. A failed or empty transcription is treated carefully: the recording is kept for retry instead of being thrown away before you can recover it. When a successful result exists under Don't Save, the audio is removed and Save or Retry is no longer offered for a file that no longer exists.

Microphone access stays explicit

The first microphone recording uses the normal macOS permission prompt. If access was denied earlier, Voice Transcribe cannot work around that decision. It explains that microphone access is required and points you back to System Settings, where macOS keeps control of the permission.

Importing an existing audio file is different because Droppy is not listening to the microphone. You choose the file through the Mac file picker, and only that selected file enters the local transcription pipeline.

A voice note should end as your text

The whole flow is intentionally small: press a shortcut or Record, say the thing, stop, and copy the result. You can keep the controls visible in the Shelf, reduce the interaction to the menu bar, or use Invisi-record when even that feels like too much interruption.

Most importantly, the useful output is plain text on your clipboard, while the original recording follows a retention rule you chose. The model files may have to arrive from the internet once, but the voice note itself does not have to leave your Mac to become something you can use.