Skip to content

Voice and dictation ​

Users talk to Everycloud more than they type. Agents must handle speech input and give answers that work when heard.

How voice reaches you ​

GestureModeWhat the agent gets
Hold fnDictationNothing: text is typed into the user's app. Cleanup is done by the dictation pipeline, not by you.
Double-tap fn or fn+SpaceHands-free dictationSame as above; the user stops with fn (and the Stop button with PR A).
Hold Control+OptionVoice AskThe transcribed question. Your answer appears in the notch; Continue opens it in Chats.
"add this to my notes …" while dictatingVoice noteFiled into Notes by the app.
Hold to talk in chatVoice reply in chat (PR E)The transcribed message; optional spoken answer.

Replying to voice (rule R7) ​

  • Lead with the answer. At most two short sentences, about 40 words.
  • No markdown, tables, code, URLs, IDs or JSON in what will be read or shown in the pill.
  • Speak numbers and times naturally: "three tasks", "at half past four".
  • List at most three items, then "and two more in chat".
  • If the answer is long, give the gist and offer it in chat.

Good: "You have three tasks today: call Ana, pay rent and book the train. The train is overdue."

Bad: "Here are your tasks:\n| Task | Status |…"

Dictation quirks (rule R8) ​

Speech-to-text makes predictable mistakes. Before acting:

  1. Match the dictionary. The user's dictionary words and their sounds like aliases tell you what a misheard word should be ("every cloud" → "Everycloud", "Jeff" → "Jev" if Jev is a dictionary word).
  2. Match context names. Project, task, note, agent and people names in your context are likely targets ("the spanish exam project" → project Pass the Spanish exam).
  3. Drop filler and self-corrections. "um, Tuesday, no wait, Thursday" means Thursday.
  4. Ask when it matters. If a guess would change what an action does (which project, which note, a recipient's name), ask one short question.
  5. Keep the user's meaning. Never "fix" content into something they didn't say.

With PR G, when the user corrects a word right after dictating, Everycloud learns it into the dictionary (alias = the misheard word), so the next dictation and your context use the right spelling.

Everycloud for Mac. Draft docs: items marked “Coming soon” are not shipped yet.