From GPUs to Workloads: Flex AI’s Blueprint for Fast, Cost‑Efficient AI

14 snips

Sep 28, 2025

Brijesh Tripathi, CEO of Flex AI and a former architect at Intel, NVIDIA, Apple, and Tesla, discusses transforming AI workflows by implementing 'workload as a service'. He highlights the importance of minimizing DevOps burdens to enhance productivity, revealing how inconsistent Kubernetes layers create challenges for AI teams. Brijesh elaborates on optimizing training and inference processes and emphasizes Flex AI's focus on easing the complexity of heterogeneous compute while ensuring cost efficiency. His vision aims to empower teams, enabling them to innovate without infrastructure hassles.

Ask episode

AI Snips

Chapters

Transcript

Episode notes

INSIGHT

Match Architecture To Workflow Stage

Different workflow stages (pretraining, fine-tuning, inference) benefit from different architectures.
Flex AI routes stages to suitable hardware so workflows stay unchanged while underlying silicon varies.

ADVICE

Use Workload-As-A-Service

Treat workloads as a managed service: submit your job and let the platform handle orchestration, libs, and networking.
Use workload-as-a-service to cut days or weeks of setup time and accelerate experiments.

ADVICE

Smooth Peaks With Mixed Capacity

Smooth cost by mixing on-prem capacity with burstable external cloud and multi-tenancy to avoid idle reserved GPUs.
Run training and inference side-by-side with preemption and fractional GPUs to raise utilization and cut spend.

Get the Snipd Podcast app to discover more snips from this episode

Get the app

Summary
In this episode of the AI Engineering Podcast Brijesh Tripathi, CEO of Flex AI, talks about revolutionizing AI engineering by removing DevOps burdens through "workload as a service". Brijesh shares his expertise from leading AI/HPC architecture at Intel and deploying supercomputers like Aurora, highlighting how access friction and idle infrastructure slow progress. He discusses Flex AI's innovative approach to simplifying heterogeneous compute, standardizing on consistent Kubernetes layers, and abstracting inference across various accelerators, allowing teams to iterate faster without wrestling with drivers, libraries, or cloud-by-cloud differences. Brijesh also shares insights into Flex AI's strategies for lifting utilization, protecting real-time workloads, and spanning the full lifecycle from fine-tuning to autoscaled inference, all while keeping complexity at bay.

Announcements

Hello and welcome to the AI Engineering Podcast, your guide to the fast-moving world of building scalable and maintainable AI systems
When ML teams try to run complex workflows through traditional orchestration tools, they hit walls. Cash App discovered this with their fraud detection models - they needed flexible compute, isolated environments, and seamless data exchange between workflows, but their existing tools couldn't deliver. That's why Cash App rely on Prefect. Now their ML workflows run on whatever infrastructure each model needs across Google Cloud, AWS, and Databricks. Custom packages stay isolated. Model outputs flow seamlessly between workflows. Companies like Whoop and 1Password also trust Prefect for their critical workflows. But Prefect didn't stop there. They just launched FastMCP - production-ready infrastructure for AI tools. You get Prefect's orchestration plus instant OAuth, serverless scaling, and blazing-fast Python execution. Deploy your AI tools once, connect to Claude, Cursor, or any MCP client. No more building auth flows or managing servers. Prefect orchestrates your ML pipeline. FastMCP handles your AI tool infrastructure. See what Prefect and Fast MCP can do for your AI workflows at aiengineeringpodcast.com/prefect today.
Your host is Tobias Macey and today I'm interviewing Brijesh Tripathi about FlexAI, a platform offering a service-oriented abstraction for AI workloads

Interview

Introduction
How did you get involved in machine learning?
Can you describe what FlexAI is and the story behind it?
What are some examples of the ways that infrastructure challenges contribute to friction in developing and operating AI applications?
- How do those challenges contribute to issues when scaling new applications/businesses that are founded on AI?
There are numerous managed services and deployable operational elements for operationalizing AI systems. What are some of the main pitfalls that teams need to be aware of when determining how much of that infrastructure to own themselves?
Orchestration is a key element of managing the data and model lifecycles of these applications. How does your approach of "workload as a service" help to mitigate some of the complexities in the overall maintenance of that workload?
Can you describe the design and architecture of the FlexAI platform?
- How has the implementation evolved from when you first started working on it?
For someone who is going to build on top of FlexAI, what are the primary interfaces and concepts that they need to be aware of?
Can you describe the workflow of going from problem to deployment for an AI workload using FlexAI?
One of the perennial challenges of making a well-integrated platform is that there are inevitably pre-existing workloads that don't map cleanly onto the assumptions of the vendor. What are the affordances and escape hatches that you have built in to allow partial/incremental adoption of your service?
What are the elements of AI workloads and applications that you are explicitly not trying to solve for?
What are the most interesting, innovative, or unexpected ways that you have seen FlexAI used?
What are the most interesting, unexpected, or challenging lessons that you have learned while working on FlexAI?
When is FlexAI the wrong choice?
What do you have planned for the future of FlexAI?

Contact Info

Parting Question

From your perspective, what are the biggest gaps in tooling, technology, or training for AI systems today?

Links

The intro and outro music is from Hitman's Lovesong feat. Paola Graziano by The Freak Fandango Orchestra/CC BY-SA 3.0