Skip to content

The latest SayIt release is live for Windows and Linux.

Download SayIt
Local push-to-talk dictation · Windows & Linux

Your voice. Your words. Typed anywhere.

SayIt is a local push-to-talk voice-typing utility for Windows and Linux. Hold a hotkey, speak, release — the transcribed text is inserted at your cursor.

Windows and Linux installers are available in the download center from the latest published release. The source and any release artifacts live on GitHub.

Active window: code editorProfile: Developer

Open the GitHub Actions log for the Next.js build and run getUserById on it.

  • Technical correctiongit hub → GitHub
  • Custom vocabularynext js → Next.js
  • Developer modeget user by id camel case → getUserById

Inserted at your cursor — Esc cancels at any point.

One utterance, start to finish. Every fix shown is a real, toggleable rule — not an artist’s impression.

Why say it

Speak at the speed of thought

Typing tops out at your fingers. Speaking doesn't — and SayIt keeps the loop tight: hold, speak, release, done.

On the keyboard

Ship the release notes before Friday.

Said with SayIt

Ship the release notes before Friday.

An illustration, not a benchmark — transcription time depends on your hardware, utterance length, and model choice.

How it works

Four steps to your words on the page

No accounts, no setup wizard beyond a microphone and a model — just a key and your voice.

  1. Hold the hotkey

    Hold your configurable push-to-talk hotkey and speak naturally.

    Hotkey
  2. Say what you mean

    Dictate prose, code identifiers, URLs — or spoken edits like "scratch that".

  3. Release to transcribe

    Recognition runs locally on your device, then deterministic formatting rules tidy the text.

  4. Text at your cursor

    The result is inserted into the app you were typing in. Changed your mind? Esc cancels.

    Esc

Context awareness

Adapts to where you’re typing

SayIt reads the active application’s metadata — never its contents — and picks a matching formatting profile. A manual override always wins.

Active window: code editorProfile: Developer

You say

“open the ci slash cd logs and check the git hub runner”

Inserted

Open the CI/CD logs and check the GitHub runner.

  • Technical correction
  • Formatting profiles

The Developer profile runs the local technical-correction rules: spoken tech terms come out formatted, while ordinary prose is left alone.

Active window: email clientProfile: Email

You say

“i'll arrive at five actually six new paragraph kind regards”

Inserted

I'll arrive at six.Kind regards

  • Voice edit
  • Structure commands

“Actually six” supersedes “five”, and “new paragraph” becomes a real break — deterministic local rules, no LLM in the loop.

Active window: chat appProfile: Chat

You say

“let's sync tomorrow scratch that let's sync on thursday”

Inserted

Let's sync on Thursday.

  • Voice edit — “scratch that”

“Scratch that” drops what came before it in the utterance, so the message lands as if the false start never happened.

Personalization

A vocabulary that’s actually yours

Rules, snippets, and corrections are visible, editable, and yours to switch off — nothing learns without you.

Custom vocabulary

Spoken-form → written-form mappings you create yourself — case-insensitive, multi-word, longest-match first. Entries exist only because you added them; each one can be edited or switched off.

Snippets

Say a trigger to insert a saved block of text. Snippets are text only — content like git pull is inserted as literal text and never executed.

Voice edit

Spoken corrections — “scratch that”, “replace five with six”, “delete the last sentence” — are applied by deterministic local rules. Ambiguous phrases stay ordinary text.

Developer mode

Casing commands on demand — camel, snake, kebab, pascal — plus a spoken code-block wrapper, for talking to your editor in its own dialect.

Safe zones

Designate protected applications — password managers, banking — where recording is blocked before any audio is captured.

Remember corrections

Fix a transcript and SayIt can offer to remember it. Only an explicit “Remember” creates a rule — there is no silent learning, ever.

Switch off anything

Every local rule can be disabled, and identical input, context, and settings always produce identical output. The behavior is a choice, not a black box.

Your words stay on your machine

Recognition runs on your device. Context detection reads application metadata only — no screenshots, no clipboard scraping, no keylogging. Optional cloud enhancement is off by default and clearly separated from core dictation.

Read the privacy notes
  • Transcription via Sherpa-ONNX, on your device
  • No account and no API key for core dictation
  • Safe zones block recording before audio capture
  • Every local rule is deterministic and can be switched off

Platforms

Where it runs today

Support reflects what has actually been validated — not marketing.

  • WindowsSupported

    Primary target. The automated test suite runs on Windows; packaged-install validation is still in progress.

  • LinuxPartially supported

    AppImage builds exist but are less validated. Global hotkeys may require X11 and may not work under Wayland; clipboard insertion needs xclip.

  • macOSNot supported

    No macOS build exists today. If that changes, it will appear here alongside real download artifacts.

FAQ

Fair questions, straight answers

If a question is missing, the issue tracker takes it.

Does my audio leave my computer?

Core dictation runs on your device via Sherpa-ONNX — no account and no API key. Optional LLM enhancement is a separate, off-by-default feature; nothing is uploaded by core dictation.

Is it streaming dictation?

No. SayIt records while you hold the hotkey and transcribes once, after you release. Nothing is transcribed mid-word while you speak.

Which platforms does it run on?

Windows is the primary target. Linux builds exist as AppImages but are less validated — global hotkeys may require X11 and may not work under Wayland. There is no macOS build.

How large are the speech models?

Models are separate downloads, roughly 100 MB to 1 GB depending on the model you pick.

Does an AI rewrite my words?

No. Corrections are deterministic local rules — no LLM, no network. Identical input, context, and settings produce identical output, and every rule can be turned off.

What if text insertion fails?

The transcript stays recoverable instead of being lost, and Esc cancels while recording or transcribing.

Can I see the source code?

Yes — SayIt is developed in the open under the MIT License, with the full source on GitHub.

Where do downloads come from?

GitHub Releases, with published checksums. The download center reflects the latest published release — the download center says exactly what exists today.

Hold a key. Say the thing.

SayIt for Windows and Linux — hold, speak, release, and the words are typed where your cursor is.

Released under the MIT License — no account, no API key for core dictation.