Workday's engineering team tackled the challenge of scaling machine learning inference for numerous customers by devising a "bin packed shards" strategy on Kubernetes. This approach, detailed in their Medium article from January 2022, involves grouping multiple tenants' ML models into shared units called shards, aiming for efficient resource usage, particularly memory. Kubernetes handles the deployment and scaling of these shards, while Istio's Virtual Services manage the routing of tenant-specific requests. The strategy offers benefits like cost reduction and independent model management but also presents complexities in initial design and ongoing operation, focusing on a balance between efficiency and manageability.
Fler avsnitt av Rapid Synthesis: Delivered under 30 mins..ish, or it's on me!
Visa alla avsnitt av Rapid Synthesis: Delivered under 30 mins..ish, or it's on me!Rapid Synthesis: Delivered under 30 mins..ish, or it's on me! med Benjamin Alloul 🗪 🅽🅾🆃🅴🅱🅾🅾🅺🅻🅼 finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
