Passer au contenu principalAWS Startups
    1. Events
    2. Building and Scaling GenAI workload with Amazon EKS

    Building and Scaling GenAI workload with Amazon EKS

    Conteneurs

    Développeurs

    Jour:

    -

    Heure:

    -

    Type:

    EN LIGNE

    Intervenants:

    Dima Breydo | Pr. Solutions Architect, AWS, Majid Shokrolahi | Sr. Solutions Architect, AWS

    Langue:

    English

    Niveau(x):

    300 – Avancé, 400 – Expert

    Détails de l’événement

    -

    -

    EN LIGNE

    Intervenants

    Join us for an immersive virtual hands-on workshop exploring how to build and scale production-ready Generative AI deployments on Amazon EKS using NVIDIA GPUs. Learn how to deploy and manage production-grade LLM workloads, implement advanced RAG patterns with vector databases, and leverage open-source frameworks for data-backed LLM applications. Through hands-on labs, you'll gain best practices for deploying, scaling, and observing Gen AI inference workloads on Kubernetes. Whether you're deploying your first language model or scaling existing workloads, this workshop provides the tools and experience you need. 

    What You Will Learn 

    • How to set up an Amazon EKS cluster optimized for NVIDIA GPU workloads 
    • Infrastructure options for running AI/ML workloads on Amazon EKS 
    • Efficient model serving and scaling using vLLM 
    • Distributed inference architecture implementation with Ray 
    • Observability for vLLM and Ray using Prometheus and Grafana 
    • Best practices for production Gen AI deployments on Kubernetes 

    Agenda

    GenAI on EKS Presentation and Workshop Agenda

    8:00 AM UTC

    Introduction & Presentation — Gen AI on EKS architecture, NVIDIA GPU integration

    10:00–10:20 | Introduction & Presentation — Gen AI on EKS architecture, NVIDIA GPU integration

    8:20 AM UTC

    Hands-on workshop

    10:20–10:30 | Getting Started — Environment setup and cluster exploration 10:30–10:45 | Lab 1: GPU Infrastructure Setup 10:45–11:05 | Lab 2: Deploy the Model 11:05–11:15 | Lab 3: Observability End-to-end 11:15–11:25 | Lab 4: Agentic App Building 11:25–11:50 | Lab 5: Inference Benchmarking 11:50–12:15 | Lab 6: KV Cache Offloading 12:15–12:45 | Lab 7: Fine-Tuning on EKS 12:45–13:00 | Wrap-up & Q&A