AI agents that switch between voice, text, and visuals
Sierra's multimodal agents bring voice, text, and visuals into the same customer conversation. Voice is for explaining what you need, a visual for comparing options side by side, and text for referencing something later. Instead of picking just one, the agent automatically shifts between modes as the conversation needs.
No reviews yetBe the first to leave a review for Sierra
Hunter
š
Sierra's multimodal agents bring voice, text, and visuals into the same customer conversation, so people get the best of each medium instead of being stuck with one.
Voice is for explaining what you need, a visual for comparing options side by side, text for referencing something later. Agents built on Sierra anticipate what each moment of the conversation needs and automatically shift modes, without making you restart or repeat yourself.
It's built on Sierra's MCP UI integration, so businesses design and host their own interactive components (product cards, comparison tables, calendars, forms) and drop them into any conversation. Build a component once, and it works everywhere the agent lives, no rebuilding per channel, no separate versions to maintain. Updates reflect everywhere instantly. Components can also expand full-screen for anything that needs more room.
Key features:
Automatic mode-switching between voice, text, and visuals based on conversation context
The detail I keep coming back to is the agent deciding when to switch, rather than the user picking a channel up front. I work on voice AI at Callie Care, where the person on the call is often in their 80s, and the hard part was never the speech model, it was knowing when voice stops being the right surface. A lot of our failure cases are moments where someone needed something they could look at again later. How does the agent decide to break into a visual? Is it intent based, or are you reading signals like repeated clarification and long pauses? And when it hands back to voice, does it carry the same context or restart the turn?
the "automatically shifts modes based on context" part is the piece I'd want to see fail gracefully - what happens when the agent decides a visual is needed but the customer is on a phone call with no screen in view, or picks voice for someone in a quiet office who can't talk back? is there a way for the customer to override the agent's mode choice mid-conversation, or is that decision entirely agent-side?
Sierra's multimodal agents bring voice, text, and visuals into the same customer conversation, so people get the best of each medium instead of being stuck with one.
Voice is for explaining what you need, a visual for comparing options side by side, text for referencing something later. Agents built on Sierra anticipate what each moment of the conversation needs and automatically shift modes, without making you restart or repeat yourself.
It's built on Sierra's MCP UI integration, so businesses design and host their own interactive components (product cards, comparison tables, calendars, forms) and drop them into any conversation. Build a component once, and it works everywhere the agent lives, no rebuilding per channel, no separate versions to maintain. Updates reflect everywhere instantly. Components can also expand full-screen for anything that needs more room.
Key features:
Automatic mode-switching between voice, text, and visuals based on conversation context
MCP UI integration for embedding custom interactive components (product cards, comparison tables, calendars, forms)
Build-once components that work across every channel the agent is deployed on
Full-screen expansion for components that need more room
Instant updates that reflect everywhere without redeploying
Try it at sierra.ai Ā· multimodal agents
P.S. I hunt the latest and greatest launches in tech, SaaS and AI, follow to be notified ā @rohanrecommends
Refocus
The detail I keep coming back to is the agent deciding when to switch, rather than the user picking a channel up front. I work on voice AI at Callie Care, where the person on the call is often in their 80s, and the hard part was never the speech model, it was knowing when voice stops being the right surface. A lot of our failure cases are moments where someone needed something they could look at again later. How does the agent decide to break into a visual? Is it intent based, or are you reading signals like repeated clarification and long pauses? And when it hands back to voice, does it carry the same context or restart the turn?
Dial
the "automatically shifts modes based on context" part is the piece I'd want to see fail gracefully - what happens when the agent decides a visual is needed but the customer is on a phone call with no screen in view, or picks voice for someone in a quiet office who can't talk back? is there a way for the customer to override the agent's mode choice mid-conversation, or is that decision entirely agent-side?