On-device dictation · v0.1 beta

Speak. It types. Nothing leaves your machine.

Inkvox turns your voice into text in any app, transcribed by Whisper running on your GPU. Your audio never reaches a server, and there's no account to create.

sound in · words out · nothing leaves your machine

Works where you type

Your favorite apps, now voice-powered.

Claude Notion Cursor Gmail GitHub Perplexity Telegram Figma
WhatsApp Linear Discord Obsidian VS Code X Chrome + every other app with a text field

On-device · measured

Faster than your fingers.

Sub-second transcription, and a private log of everything you've dictated. You can search it whenever you want, and it never leaves your machine.

inkvox · history on-device
  • refactor the recorder so it streams samples instead of buffering the whole take just now · 13 w
  • Hi Sarah, the beta build is ready, I'll send the download link this afternoon 3 min · 14 w
  • Standup: beta ships Thursday, QA gets one extra day, Léa owns the changelog 11 min · 13 w
  • Running ten minutes late, start without me and I'll demo Inkvox when I get there 24 min · 15 w
  • Meeting notes: scope the on-device history search, ship the privacy page first 38 min · 12 w
Stored on this machine only. Nothing is ever synced.

0

words dictated
this week

0h

saved vs typing1

0s

avg transcription2

1 Illustrative activity, based on ~40 wpm typed vs ~150 wpm spoken. 2 11 s of speech on a mid-range GPU (RTX 3070), Whisper large-v3-turbo.

How it works

You tap, you talk, it's typed.

Press whichever key you like: the pill opens, you talk, and the text appears straight into the app you're already working in.

claude · ~/dev/inkvox
Welcome to Claude Code!

model: claude-fable-5 · cwd: ~/dev/inkvox

Found 3 files to refactor. Ready when you are.

>
Launch plan · Google Docs

Launch plan

File Edit View Insert Format Tools Help

Share

Stylized illustrations. Product names and logos belong to their respective owners.

100+ languages

Speak whichever language you like.

Whisper detects the language as you speak. You can switch mid-sentence, and there's nothing to set up.

Privacy

The cloud never hears you.

Cloud dictation streams your voice to someone else's servers: every meeting note, every half-formed idea, everything you mutter near your microphone.

Inkvox runs the open Whisper model on your own GPU. Audio goes from your microphone to your screen and nowhere else: it's never written to disk, and it's gone the moment it's transcribed. There's no account to create, and it works just as well on a plane.

Features

A small app that does one thing well.

The whole product fits in one loop: you speak, it transcribes, it's typed.

You speak

100+ languages. Whisper detects them for you, and you can switch mid-sentence.

Inkvox transcribes

Whisper on your GPU: Vulkan on NVIDIA, AMD and Intel, with a CPU fallback if needed. Offline after one ~800 MB model download.

audio uploaded → 0 bytes

The text appears in the focused app

If you can type there, you can dictate there. Your clipboard is put back exactly as it was.

Works with Claude Notion Cursor Gmail WhatsApp Obsidian VS Code Discord Figma Telegram Linear Perplexity Chrome GitHub X + every app with a text field

Inkvox Pro · in the works

Say it messy.
Pro cleans it up.

The Pro rewrite layer turns raw speech into clean, structured text. It's a small language model running on the same GPU, not in the cloud.

you said

okay so um what I want is like a python script that uh goes through a folder, you know, and finds all the images, the duplicates I mean, and uh deletes them, well actually no, moves them somewhere, like a trash folder or something like that

inkvox pro · local llm

Write a Python script that scans a folder for duplicate images and moves them into a trash subfolder.

Every "um" is a token you pay for.

Paste raw dictation into Claude, ChatGPT or Cursor and the filler words bill you twice: they cost tokens, and they blur what you're asking. Pro strips them before they ever reach the model.

raw dictation
≈ 65 tok
with pro
≈ 24 tok

0%tokens for the same request2

2 This very example, with counts approximated at ~1.3 tokens per word.

Better answers, first try

A clear prompt saves you two round trips. Pro reshapes your ramble into the question you actually meant to ask, so the model stops guessing.

Your AI bill, trimmed

Filler words are tokens. Strip them locally, for free, and every paid model call you make after that gets cheaper.

Dictation stays free

Raw dictation will never move behind the paywall. Pro adds the cleanup, unlocked with a license key, and it runs on your GPU too.

The difference

The same job, without the subscription.

Cloud dictation apps charge you monthly to run your voice through their servers. Inkvox does the same job on hardware you already own. You pay nothing, and there's nothing to leak.

Cloud dictation apps

Your voice is processed

on their servers, every word

Works offline

no internet, no dictation

Account

email, login, sync required

Who can replay your audio

whoever their policy allows

A year of dictation

$144–$180

Inkvox

Your voice is processed

on your GPU, it never leaves

Works offline

plane, train, dead zone, same speed

Account

none, install and talk

Who can replay your audio

no one, never stored, never sent

A year of dictation

$0free while in beta, Pro will be a one-time license, not a subscription

Typical cloud dictation pricing, 2026: $12–15 / month.

FAQ

Questions we get asked.

Is it free?

Dictation is free and will stay free. The upcoming AI rewrite will be the paid Pro feature, unlocked with a license key. There's no subscription wall in front of the basics.

What do I need to run it?

Windows 10/11 or macOS. Any Vulkan-capable GPU (NVIDIA, AMD or Intel, which covers most machines from the last decade) gets you sub-second dictation. Without a GPU, Inkvox falls back to CPU with a lighter model.

Where does my audio go?

From your microphone to your GPU, then to your screen. It's processed in memory, never written to disk, and discarded the moment the text is inserted.

Which model does it use?

Whisper large-v3-turbo (quantized, ~800 MB) by default, downloaded on first launch with a progress bar. On a more modest machine, Inkvox picks a lighter model automatically.

What about macOS?

Yes, macOS is part of the beta. It's the same codebase, with a Metal backend. Windows and macOS run the exact same Inkvox.

Is the demo on this page real?

It's a faithful recreation of the real pill, rebuilt in HTML: same look, same states. The actual app is native, and faster than the animation.

Stop typing.
Keep your voice to yourself.

Free beta · one email when it opens · no spam, ever

macOS Windows No account, just download and talk.