FIT5152 Chap.8 Voice and Multimodal Interaction
Voice and Multimodal Interaction
Week 8 moves beyond the screen. Voice user interfaces let people act through speech, which suits users with physical impairments, people whose eyes are busy and anyone multitasking, but they lack visual signifiers, so they need their own design guidance.
The seminar dissects a voice dialogue into its parts: the invocation, made of a trigger phrase, an invocation name and an optional invocation phrase; the response and prompt that keep the conversation going; intents with their many sample utterances; slots that behave like variables with allowed values; and earcons, the sound equivalent of icons.
It also warns that natural language processing is not the same as intelligence.
A survey of five popular voice guideline sets by Stacy Branham identifies five themes: make conversations human, personal, efficient and relational, and leave users feeling in control. The chapter pairs those with the limitations of voice and their mitigations, from read-backs and undo to visible listening status.
The second half defines multimodal interfaces, which combine several input or output methods, and explains the human–machine loop in which fusion combines input signals into one command and fission spreads output across modalities. It closes with Sharon Oviatt’s myths about multimodal interaction and with gesture, haptic, biometric, olfactory and brain–computer interfaces.
What this chapter covers
- 01
Voice user interfaces and who they help
- 02
Invocation, responses, prompts and intents
- 03
Slots, utterances and earcons
- 04
Five themes from voice guidelines
- 05
Limitations of voice and their mitigations
- 06
Multimodal interfaces, fusion and fission
Worked example · free
Repairing an error turn in a pharmacy voice assistant
- 1Take responsibility instead of blaming the user: “I found two blood pressure prescriptions.”
- 1Offer an action-based choice rather than teaching a command: “Do you want the morning tablet or the evening tablet?”
- 1Keep the user in control and the list short, offering at most two options and allowing “cancel” at any point.
Key terms
- Voice User Interface
- An interface that lets people interact with a system through voice or speech commands.
- Invocation Name
- The part of a spoken request that identifies which app or action should respond.
- Intent
- A single task a user can request from a voice interface, supported by many sample utterances.
- Earcon
- A short sound effect that adds information or personality to an interaction, as an icon does visually.
- Multimodal Interface
- An interface that provides several methods of input or output and uses them in an integrated way.
- Fission
- The splitting of a system’s output across several modalities, such as speech and visual cues.
Voice and Multimodal Interaction FAQ
Who benefits most from voice user interfaces?
The seminar names people with physical impairments, people with limited visual attention such as drivers, and anyone operating a device remotely or while multitasking, though blind users report that long voice dialogues can be tiring.
What is the difference between fusion and fission?
Fusion happens on the input side, combining several signals such as speech and a gesture into one command. Fission happens on the output side, dividing one result across several modalities such as sound and visuals.
What are the five themes in voice design guidelines?
Stacy Branham’s survey found guidelines agree on making conversations human, personal, efficient and relational, and on giving the user a sense of control, for example by letting them stop or cancel at any time.
Why avoid phrases like “I think” in a voice assistant?
Personification raises unrealistic expectations about how intelligent the system is. Plain, transparent responses help users judge what the assistant can actually do and recover when it gets something wrong.
Assessment move
Pick a simple task such as checking a bus time and write a full voice dialogue for it, labelling the trigger phrase, invocation name, optional phrase, intent, slot and one error path. Then score your dialogue against the five guideline themes, rewriting any prompt that teaches commands instead of offering actions.
For multimodal interaction, redraw the human–machine loop from memory, marking where fusion and fission happen, and describe one everyday product that uses each. Finally, read Oviatt’s ten myths and, for three of them, write one sentence on why a designer might wrongly believe it; this turns a list you might otherwise memorise into reasoning you can apply.
Working through Voice and Multimodal Interaction in FIT5152? Sia is AskSia’s AI Information Technology tutor — ask any FIT5152 Voice and Multimodal Interaction question and get a clear, step-by-step explanation grounded in how FIT5152 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.