Talk to your Mac.
It actually does the work.
Hold Control Option, say what you want, let go. sideOS reads your screen, decides what to do, opens apps, clicks, types — and tells you when it’s done.
Scroll — watch it work on this Mac
Works in every app on your Mac
You hold
Control + Option wakes the notch.
Listening
macOS transcribes on-device. No API call.
Thinking
Reads the live screen in about 50 ms.
Acting
Opens, clicks, types, scripts — for real.
Speaking
Says what it did, out loud.
It doesn’t just answer.
It acts.
Chatbots hand you a paragraph and leave the work to you. sideOS has hands — a toolbox that reaches straight into the apps already open on your screen.
It reads the screen, not a screenshot
sideOS pulls the accessibility tree macOS already maintains for every window: real buttons, real labels, real text fields. Same understanding as a screenshot, a fraction of the time and tokens. Images are the fallback, not the default.
~50 ms
to read the screen
9 tools
in the agent's hands
The toolbox
read_screenReads the live state of the screenopen_appLaunches or focuses an appclick_elementHits buttons, links, menu itemsfocus_elementPuts the caret in a text fieldtype_textTypes into whatever is focusedpress_keysSends any keyboard shortcutrun_applescriptDrives Apple apps directlytake_screenshotVisual fallback for canvas & videowaitWaits for the UI to settle
Things people actually say
One request runs up to 14 steps. Longer jobs belong in background tasks — press Esc at any moment and everything stops instantly.
Point at anything
on your screen.
No more describing what you're looking at. Frame it, ask, and sideOS sees exactly what you see.
Drag while you talk
Holding the keys turns the cursor into a crosshair. Frame a chart, an error, a paragraph — anywhere on screen.
Pixels and elements, together
The accessibility tree can't see charts or video. An image misreads long numbers and code. sideOS sends both, so they cover each other.
Cropped at native resolution
Only the region travels, at full sharpness — clearer than a whole screenshot and a lot cheaper.
Dictation that costs nothing.
Hold Control fn and talk. What you say is typed exactly where your cursor is — a note, an address bar, a form field, a terminal. It never touches the model.
Four gestures
Hold, speak, release
Push-to-talk. The text lands the moment you let go.
Tap once (under 400 ms)
Hands-free. Let go of the keys — the island keeps listening.
Tap again
Wraps up and types everything you said.
Press Esc
Cancels. Nothing gets typed.
Dictation on
On-deviceShip the beta on Friday, tell the team we are moving the review to Monday
- Zero API requests
- Works without an API key
- 62 languages, including yours
No API, no bill
Dictation never reaches Claude. It runs on macOS's own engines and works without an API key.
Your clipboard survives
Text is pasted at the caret, then whatever you had copied is put right back.
Nothing gets lost
No editable field in focus? The text waits on your clipboard and the island tells you so.
Two engines, picked for you
Automatic by default — the newer engine when your language is supported, the classic one otherwise.
Classic
SFSpeechRecognizer- Languages
- 62 languages
- Length
- ~1 minute per take
Available on every macOS version.
New
SpeechAnalyzer · macOS 26- Languages
- 30 languages
- Length
- No time limit
Fully on-device. We fed it 7 min 10 s of audio in one pass — about 50× faster than real time.
A whole workspace,
folded into a notch.
The voice loop is the front door. Behind it sits everything you'd expect from a real assistant — and it all stays on your Mac.
Hand it a job and walk away
Background tasks run in their own loop — up to 60 steps — while you get on with something else. They can search the web, read and write files, run shell commands and render a finished PDF. Anything with teeth asks you first.
Running · step 12 of 60
web_searchQ3 competitor pricingread_file~/Notes/pricing.mdwrite_filereport.mdrender_pdfQ3-summary.pdf
run_shell needs your approval
Meetings, handled
It notices when a call window opens and offers to record. Afterwards it transcribes on-device, writes the summary, works out who said what — and turns every action item into a reminder.
It remembers
How you write, what you're working on, who matters. Ask what it knows, edit it, or just say “forget that”. Switch memory off and nothing is stored or sent.
Reminders that nudge
“Remind me to send this at five.” It lands in the notch at five — overdue ones stay marked until you deal with them.
Or just type
Double-tap Control and a chat field opens in the island; answers stream down as bubbles. Drop in images, PDFs or text files and it reads them.
Bring your own tools
Connect any MCP server — HTTP or stdio, OAuth included — and those tools join the toolbox the assistant already has.
Searchable history
Every run is logged locally, so “what did I do about that invoice last Tuesday?” is a question it can actually answer.
Prefer the keyboard? Control Control opens the chat field. Move the pointer to the notch and the island unfolds into a full control menu.
Your Mac stays yours.
An assistant that can see your screen and press your keys has to earn that. So the design starts from the other end: keep everything local, and send only what the request cannot do without.
- Speech recognition and dictation run on macOS's own engines, on your Mac
- Your Claude API key lives in the Keychain — never written to a file
- Meeting recordings, transcripts, memory and history stay in local storage
- No account, no telemetry, no server of ours in the middle
- Nothing you say or show is used to train anything
- Screen data is read only for the request you just made
Your key, your bill
The only paid part is the Claude API, billed to you by Anthropic directly. We never see it.
Screenshots are the exception
The accessibility tree covers almost everything. Images are captured only when a request truly needs one.
Or go fully offline
Point sideOS at a local MLX model — Qwen3 4B or 8B — and not a single byte leaves your Mac.
The app is free. The key is yours.
There is no subscription and no middle layer. Speech, actions and voice output all use what macOS already ships with — the only thing that ever costs money is the model.
Dictation
No key, no account, no limit.
- Push-to-talk and hands-free modes
- Runs on macOS's own speech engines
- 62 languages on the classic engine
- Types wherever your caret is
Agent
Most usedThe app is free. You bring a Claude API key and pay Anthropic directly.
- Reads the screen, clicks, types, scripts
- Region select, background tasks, meetings
- Memory, reminders, MCP servers
- Turn thinking depth down, or run Haiku 4.5 for roughly a fifth of the cost
Offline
Swap Claude for a local MLX model and disconnect entirely.
- Qwen3 4B and 8B, quantised for Apple Silicon
- Vision-capable options for screenshots
- Nothing leaves the machine
- Downloads once, then works on a plane
A typical command — read the screen, think, act, answer — lands around two to five cents on your own Anthropic bill. Dictation costs nothing either way.
Questions, answered
Everything people ask before they hand an assistant the keyboard.
Only for the agent. Dictation runs entirely on macOS's own speech engines, so it works the moment you install the app — no key, no account, no cost. For the agent you paste a Claude API key from the Anthropic console; it is stored in your Keychain.
Hold two keys.
Start talking.
Control Option for the agent, Control fn for dictation. Free to install, free to dictate with, and about the price of a coffee a week once you put it to work.
- Apple Silicon
- macOS 14 or later
- Accessibility · Mic · Speech