Hardware, attention types, and K/V cache sizing

Igor urges analyzing GPU specs, model attention types, and KV cache needs to choose nodes and memory sizing.

Play episode from 43:44

chevron_right

Transcript

chevron_right

Transcript

Episode notes

Roundtable CAST AI episode: Serving LLMs in Production: Performance, Cost & Scale.

Join the Community:

https://go.mlops.community/YTJoinIn

Get the newsletter: https://go.mlops.community/YTNewsletter

MLOps GPU Guide:

https://go.mlops.community/gpuguide

// Abstract

Experimenting with LLMs is easy. Running them reliably and cost-effectively in production is where things break.

Most AI teams never make it past demos and proofs of concept. A smaller group is pushing real workloads to production—and running into very real challenges around infrastructure efficiency, runaway cloud costs, and reliability at scale.

This session is for engineers and platform teams moving beyond experimentation and building AI systems that actually hold up in production.

// Bio

Ioana Apetrei

Ioana is a Senior Product Manager at CAST AI, leading the AI Enabler product, an AI Gateway platform for cost-effective LLM infrastructure deployment. She brings 12 years of experience building B2C and B2B products reaching over 10 million users. Outside of work, she enjoys assembling puzzles and LEGOs and watching motorsports.

Igor Šušić

Igor is a founding Machine Learning Engineer at CAST AI’s AI Enabler, where he focuses on optimizing inference and training at scale. With a strong background in Natural Language Processing (NLP) and Recommender Systems, Igor has been tackling the challenges of large-scale model optimization long before transformers became mainstream. Prior to CAST AI, he worked at industry leaders like Bloomreach and Infobip, where he contributed to the development and deployment of large-scale AI and personalization systems from the early days of the field.

// Related Links

Website: https://cast.ai/

~~~~~~~~ ✌️Connect With Us ✌️ ~~~~~~~

Catch all episodes, blogs, newsletters, and more: https://go.mlops.community/TYExplore

Join our Slack community [https://go.mlops.community/slack]

Follow us on X/Twitter [@mlopscommunity](https://x.com/mlopscommunity) or [LinkedIn](https://go.mlops.community/linkedin)]

MLOps Swag/Merch: [https://shop.mlops.community/]

Connect with Demetrios on LinkedIn: /dpbrinkm

Connect with Ioana on LinkedIn: /ioanaapetrei/

Connect with Igor on LinkedIn: /igor-%C5%A1u%C5%A1i%C4%87/

The AI-powered Podcast Player

Save insights by tapping your headphones, chat with episodes, discover the best highlights - and more!

Get the app

Home Top podcasts Popular guests Top books