Features
What SayIt does today
Every item below corresponds to functionality in the current application. Anything not implemented yet simply isn't listed.
Core
Voice dictation
The foundation: hold a key, speak, release. Recognition runs on your machine, and the result lands where you were typing.
Push-to-talk recording
Hold a configurable global hotkey to record; release to transcribe. Press Esc to cancel mid-recording or while transcription is running.
CtrlSpaceLocal speech recognition
Transcription runs on your device via Sherpa-ONNX once a model is installed. Decoding happens once, after you release — SayIt is not a streaming dictation system.
Text at your cursor
The transcript is inserted into the app you were typing in, at the cursor. If insertion fails, the text stays recoverable instead of being lost.
Awareness
Context awareness
SayIt looks at which application is active — using process and window metadata only — and picks a formatting profile.
Formatting profiles
Profiles adjust which local correction rules run: Developer, Email, Chat, Notes, Prompt, or Normal. A profile never rewrites your words, and a manual override always wins.
Metadata only
Context detection reads no screenshots, no clipboard contents, no keystrokes, and no page or editor contents — just the application's own metadata.
Deterministic
Local intelligence
A local-first intelligence layer on top of dictation. It uses no LLM and no network — identical input, context, and settings always produce the same output, and every feature can be turned off.
Technical correction
Normalizes common developer terms ("git hub" → GitHub, "ci slash cd" → CI/CD) and formats clearly-dictated URLs, file paths, email addresses, and API/version structures. Conservative by design — ordinary prose is left alone — and it can be turned off.
Voice edit
"I'll arrive at five. Actually six." becomes "I'll arrive at six." Also "scratch that", "replace five with six", and "delete the last sentence". Ambiguous phrases stay ordinary text.
Snippets
Say a trigger like "my github" to insert a saved block of text. Snippets are text only — content like git pull is inserted as literal text and never executed.
Developer mode
"get user by id, camel case" becomes getUserById — plus snake, kebab, and pascal casing, and a spoken code-block wrapper for longer passages.
Safe zones
Designate protected applications — password managers, banking — where recording is blocked before any audio is captured.
Remember corrections
When you fix a transcript, SayIt can offer to remember the fix. Only an explicit "Remember" creates a rule; there is no silent learning.
Structure commands
"New line" and "new paragraph" become real breaks, and clean ordinal dictation ("first … second … third …") becomes a proper list.
Yours
Personalization
You control the words SayIt knows — nothing is learned without you.
Custom vocabulary
Spoken-form → written-form mappings you create yourself: "next js" → Next.js, "open ai" → OpenAI. Case-insensitive, multi-word, longest-match-first; entries can be edited, categorized, and individually disabled.
Examples
- “next js”Next.js
- “open ai”OpenAI
- “my github”→ your saved snippet
Entries are user-created only, stored locally, and never transmitted.
Honesty
Where the limits are
Knowing what a tool doesn't do is as useful as knowing what it does.
- Not streaming: transcription happens once, after you release the hotkey — partial words are not transcribed while you speak.
- Insertion depends on a focused editable text field; paste success cannot be conclusively verified, and your previous clipboard contents may not be restored.
- Linux is partially supported and less validated; global hotkeys may require X11 and may not work under Wayland.
- Speech models are separate downloads from the application — roughly 100 MB to 1 GB depending on the model.
- Speed varies with your hardware, utterance length, and model choice — no fixed numbers are claimed.