Suresh Michael
All workshops
LLM Engineering

Large Language Model Inferencing in Production

A technical session on serving LLMs for real: usage patterns, managed services versus self-hosting, vector stores, LLMOps, and running inference on Kubernetes with Kubeflow.

Delivered
October 12, 2025
Format
Technical talk
Audience
Platform & ML engineers

The slide deck is locked

Sign in below to open the deck and the rest of the write-up.

Slides

What we covered

Getting a model to answer is a weekend. Getting it to answer at volume, at a cost you can defend, is the work. This session is about the second one.

  • Usage patterns, and how the shape of your traffic decides everything downstream.
  • Managed AI services and Model Garden: what you give up and what you stop maintaining.
  • Vector stores in the cloud, and choosing one against your retrieval pattern rather than a benchmark.
  • MLOps and LLMOps: what carries over from the former, and what genuinely doesn't.
  • Where AI meets Kubernetes, and the Kubeflow architecture underneath it.
  • Serving LLMs on Kubeflow, end to end.
  • AI agents, and what they add to the inference bill.

Members only

Keep reading

The slide deck and the rest of this write-up are free, sign in with Google and they stay unlocked on this device.

No newsletter, no spam. Your email is used to keep you signed in.

Share this workshop
Copied