- Events
- Workshop - Own Your Voice AI Stack: Production at Scale with AWS and NVIDIA
Workshop - Own Your Voice AI Stack: Production at Scale with AWS and NVIDIA
-
-
PRESENCIAL
English
300: avanzado, 400: experto
Voice AI is one of the fastest-moving model categories in 2026 — clinical scribes, real-time customer service, transcription at scale. But the ownership decision isn't unique to voice. Any team relying on third-party APIs inherits someone else's accuracy ceiling, their latency floor, and their pricing curve. The teams pulling ahead are the ones running their own models and fine-tuning for their domain.
We use the voice stack as the teaching vehicle because it exercises every hard trade-off at once: latency, accuracy, streaming, and inference cost. The patterns transfer directly to LLM fine-tuning, embedding models, and multimodal workloads.
What You’ll Learn:
• The full NVIDIA voice model stack — cascaded and unified architectures, and when to choose each
• A repeatable fine-tuning evaluation loop (benchmark → adapt → validate) that works for any open-weight model, not just ASR
• Deployment and runtime optimisation on Amazon EKS with G6 (L4 GPU) instances using Triton and ONNX
• How to reason about the cost, latency, and accuracy trade-offs of self-hosting vs. API
What You’ll Leave With:
- A mental model of model ownership — architecture selection through to runtime optimisation — demonstrated through voice but directly transferable to LLM and embedding workloads
- A production-ready ASR deployment on AWS Containers
- Confidence to apply the same deploy → tune → optimise → scale pattern to TTS, LLM, or multimodal serving
Featuring: A real-world case study from Heidi Health — production ASR fine-tuned for Australian clinical speech on AWS.
Agenda
11:00 PM UTC
Arrival and networking
11:30 PM UTC
Voice AI Fundamentals
The session opens with Heidi Health's Path to Domain-Tuned Transcription, a 15-minute talk from Taha Ansari, AI Engineer at Heidi Health. This is followed by a 45-minute deep dive into Voice AI Fundamentals with Siddhartha Banerjee, AI Technologist at NVIDIA. Next, Vincent Wang, Specialist SA for GenAI at AWS, presents Building Voice AI for Scale with EKS in a 15-minute session. The event closes with a 15-minute Q&A.
1:00 AM UTC
Break
1:15 AM UTC
Lab 1: Deploy ASR on EKS (90 mins)
2:45 AM UTC
Lunch
3:30 AM UTC
Lab 2: Fine-tune ASR model + discuss best practices for runtime optimisation (2 hours)
5:30 AM UTC
Networking drinks