Voice & Multimodal Experience
Add local/private and premium cloud speech plus richer multimodal interaction to the same AI workspace.
Interactive platform map
Architecture in context
The focused blueprint, its required foundation, and declared recommendations.
Voice & Multimodal Experience
Add local/private and premium cloud speech plus richer multimodal interaction to the same AI workspace.
What this replaces
Voice and media features often require separate SaaS tools and lock an application to a single provider.
What your team gains
A hybrid experience layer that can choose local or cloud speech and plug image/voice capabilities into the same AI interface.
What is inside the blueprint
How it fits the platform
AI response/media request -> provider selection -> Kokoro local or Gemini TTS/image provider -> Open WebUI/mobile/channel output.
Kokoro TTS, gemini-tts-proxy, Open WebUI, LiteLLM where applicable
- AICORTEX Core PlatformSpeech services deploy as Core-managed containers.
From prerequisites to operation
- Audio-capable client
- Optional Gemini credentials
- Voice/model selection policy
- Deploy local TTS service
- Configure Gemini proxy
- Register speech endpoints in UI/app
- Define local-vs-cloud routing
- Voice/language
- Performance style
- Provider preference
- Quality/cost/privacy policy
- Audio limits
- Latency
- Memory usage
- Provider availability
- Voice quality tests
- Request volume/cost
What this unlocks with other layers
Multi-Channel Intelligence
The intelligence layer is reusable across interfaces rather than tied to one chat surface.