Serving LLMs in Production: Performance, Cost & Scale // CAST AI Roundtable

44 snips

Feb 19, 2026

Igor Šušić, founding ML engineer focused on large-scale inference and performance tuning. Ioana Apetrei, senior product manager building accessible, cost-effective LLM deployment. They debate why deployments fail at scale. They cover model routing and cost vs accuracy. They explain time-sharing GPUs, quantization, prefill vs decode separation, and when self-hosting or managed endpoints make sense.

Ask episode

AI Snips

Chapters

Transcript

Episode notes

INSIGHT

Self-Hosting Is Harder Than It Looks

Self-hosting is more complex than expected due to quotas, capacity, and expensive idle GPUs.
Proper orchestration, autoscaling and domain knowledge can yield 2-3x effective capacity gains.

ANECDOTE

Timeshare Setup Cut Costs For A Customer

One customer migrated from AWS SageMaker to a timeshare setup and saved 40% with equal or better latency.
The cost savings allowed that team to expand faster.

ADVICE

Profile First, Exotic Tricks Later

Profile and benchmark your workload before chasing exotic optimizations.
Continuous batching, quantization and bottleneck analysis will solve most performance problems.

Get the Snipd Podcast app to discover more snips from this episode

Get the app

Roundtable CAST AI episode: Serving LLMs in Production: Performance, Cost & Scale.

Join the Community:

https://go.mlops.community/YTJoinIn

Get the newsletter: https://go.mlops.community/YTNewsletter

MLOps GPU Guide:

https://go.mlops.community/gpuguide

// Abstract

Experimenting with LLMs is easy. Running them reliably and cost-effectively in production is where things break.

Most AI teams never make it past demos and proofs of concept. A smaller group is pushing real workloads to production—and running into very real challenges around infrastructure efficiency, runaway cloud costs, and reliability at scale.

This session is for engineers and platform teams moving beyond experimentation and building AI systems that actually hold up in production.

// Bio

Ioana Apetrei

Ioana is a Senior Product Manager at CAST AI, leading the AI Enabler product, an AI Gateway platform for cost-effective LLM infrastructure deployment. She brings 12 years of experience building B2C and B2B products reaching over 10 million users. Outside of work, she enjoys assembling puzzles and LEGOs and watching motorsports.

Igor Šušić

Igor is a founding Machine Learning Engineer at CAST AI’s AI Enabler, where he focuses on optimizing inference and training at scale. With a strong background in Natural Language Processing (NLP) and Recommender Systems, Igor has been tackling the challenges of large-scale model optimization long before transformers became mainstream. Prior to CAST AI, he worked at industry leaders like Bloomreach and Infobip, where he contributed to the development and deployment of large-scale AI and personalization systems from the early days of the field.

// Related Links

Website: https://cast.ai/

~~~~~~~~ ✌️Connect With Us ✌️ ~~~~~~~

Catch all episodes, blogs, newsletters, and more: https://go.mlops.community/TYExplore

Join our Slack community [https://go.mlops.community/slack]

Follow us on X/Twitter [@mlopscommunity](https://x.com/mlopscommunity) or [LinkedIn](https://go.mlops.community/linkedin)]

MLOps Swag/Merch: [https://shop.mlops.community/]

Connect with Demetrios on LinkedIn: /dpbrinkm

Connect with Ioana on LinkedIn: /ioanaapetrei/

Connect with Igor on LinkedIn: /igor-%C5%A1u%C5%A1i%C4%87/