Appearance
Voice and dictation
Users talk to Everycloud more than they type. Agents must handle speech input and give answers that work when heard.
How voice reaches you
| Gesture | Mode | What the agent gets |
|---|---|---|
| Hold fn | Dictation | Nothing: text is typed into the user's app. Cleanup is done by the dictation pipeline, not by you. |
| Double-tap fn or fn+Space | Hands-free dictation | Same as above; the user stops with fn (and the Stop button with PR A). |
| Hold Control+Option | Voice Ask | The transcribed question. Your answer appears in the notch; Continue opens it in Chats. |
| "add this to my notes …" while dictating | Voice note | Filed into Notes by the app. |
| Hold to talk in chat | Voice reply in chat (PR E) | The transcribed message; optional spoken answer. |
Replying to voice (rule R7)
- Lead with the answer. At most two short sentences, about 40 words.
- No markdown, tables, code, URLs, IDs or JSON in what will be read or shown in the pill.
- Speak numbers and times naturally: "three tasks", "at half past four".
- List at most three items, then "and two more in chat".
- If the answer is long, give the gist and offer it in chat.
Good: "You have three tasks today: call Ana, pay rent and book the train. The train is overdue."
Bad: "Here are your tasks:\n| Task | Status |…"
Dictation quirks (rule R8)
Speech-to-text makes predictable mistakes. Before acting:
- Match the dictionary. The user's dictionary words and their sounds like aliases tell you what a misheard word should be ("every cloud" → "Everycloud", "Jeff" → "Jev" if Jev is a dictionary word).
- Match context names. Project, task, note, agent and people names in your context are likely targets ("the spanish exam project" → project Pass the Spanish exam).
- Drop filler and self-corrections. "um, Tuesday, no wait, Thursday" means Thursday.
- Ask when it matters. If a guess would change what an action does (which project, which note, a recipient's name), ask one short question.
- Keep the user's meaning. Never "fix" content into something they didn't say.
With PR G, when the user corrects a word right after dictating, Everycloud learns it into the dictionary (alias = the misheard word), so the next dictation and your context use the right spelling.