Skip to content

The latest SayIt release is live for Windows and Linux.

Download SayIt

Features

What SayIt does today

Every item below corresponds to functionality in the current application. Anything not implemented yet simply isn't listed.

Core

Voice dictation

The foundation: hold a key, speak, release. Recognition runs on your machine, and the result lands where you were typing.

Push-to-talk recording

Hold a configurable global hotkey to record; release to transcribe. Press Esc to cancel mid-recording or while transcription is running.

CtrlSpace

Local speech recognition

Transcription runs on your device via Sherpa-ONNX once a model is installed. Decoding happens once, after you release — SayIt is not a streaming dictation system.

Text at your cursor

The transcript is inserted into the app you were typing in, at the cursor. If insertion fails, the text stays recoverable instead of being lost.

Awareness

Context awareness

SayIt looks at which application is active — using process and window metadata only — and picks a formatting profile.

Formatting profiles

Profiles adjust which local correction rules run: Developer, Email, Chat, Notes, Prompt, or Normal. A profile never rewrites your words, and a manual override always wins.

Metadata only

Context detection reads no screenshots, no clipboard contents, no keystrokes, and no page or editor contents — just the application's own metadata.

Deterministic

Local intelligence

A local-first intelligence layer on top of dictation. It uses no LLM and no network — identical input, context, and settings always produce the same output, and every feature can be turned off.

Technical correction

Normalizes common developer terms ("git hub" → GitHub, "ci slash cd" → CI/CD) and formats clearly-dictated URLs, file paths, email addresses, and API/version structures. Conservative by design — ordinary prose is left alone — and it can be turned off.

Voice edit

"I'll arrive at five. Actually six." becomes "I'll arrive at six." Also "scratch that", "replace five with six", and "delete the last sentence". Ambiguous phrases stay ordinary text.

Snippets

Say a trigger like "my github" to insert a saved block of text. Snippets are text only — content like git pull is inserted as literal text and never executed.

Developer mode

"get user by id, camel case" becomes getUserById — plus snake, kebab, and pascal casing, and a spoken code-block wrapper for longer passages.

Safe zones

Designate protected applications — password managers, banking — where recording is blocked before any audio is captured.

Remember corrections

When you fix a transcript, SayIt can offer to remember the fix. Only an explicit "Remember" creates a rule; there is no silent learning.

Structure commands

"New line" and "new paragraph" become real breaks, and clean ordinal dictation ("first … second … third …") becomes a proper list.

The local intelligence layer uses no LLM, no network, no screenshots, no clipboard reading, and no keylogging. Identical input, context, and settings always produce the same output — and every feature can be turned off.

Yours

Personalization

You control the words SayIt knows — nothing is learned without you.

Custom vocabulary

Spoken-form → written-form mappings you create yourself: "next js" → Next.js, "open ai" → OpenAI. Case-insensitive, multi-word, longest-match-first; entries can be edited, categorized, and individually disabled.

Examples

  • “next js”Next.js
  • “open ai”OpenAI
  • “my github”→ your saved snippet

Entries are user-created only, stored locally, and never transmitted.

Honesty

Where the limits are

Knowing what a tool doesn't do is as useful as knowing what it does.

  • Not streaming: transcription happens once, after you release the hotkey — partial words are not transcribed while you speak.
  • Insertion depends on a focused editable text field; paste success cannot be conclusively verified, and your previous clipboard contents may not be restored.
  • Linux is partially supported and less validated; global hotkeys may require X11 and may not work under Wayland.
  • Speech models are separate downloads from the application — roughly 100 MB to 1 GB depending on the model.
  • Speed varies with your hardware, utterance length, and model choice — no fixed numbers are claimed.