TECHNICAL DEEP DIVE

AI Inference Infrastructure Explained for Product Managers

By Institute of AI PM·14 min read·Oct 9, 2026

TL;DR

Inference infrastructure is everything that happens between "user sends a prompt" and "user sees a response." It determines your product's latency, throughput, cost, and reliability at scale. PMs who understand inference can make better decisions about model choice, feature scoping, and cost targets. This guide covers the full stack: GPUs, serving frameworks, KV caches, batching, and the latency vs throughput tradeoff that shapes every AI product in production.

The AI PM Minute

One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.

No fluff. Unsubscribe anytime.

What the Inference Stack Actually Is

Training a model is expensive and happens once. Inference is what runs every time a user interacts with your product. For most AI PMs, inference is the system you actually own and operate day to day, and its cost can exceed training costs within weeks of launch.

The inference stack has four layers, each affecting your product differently:

Hardware layer

What it is: GPUs or specialized accelerators (TPUs, Trainium, Groq LPUs) that physically run matrix multiplications. The hardware choice determines raw speed and cost per token.

PM implication: When you pick a model provider, you're implicitly picking their hardware tier. Providers using older A100s will be cheaper but slower than those on H100s or H200s. Price differences of 5x between providers often trace back to this.

Serving framework

What it is: Software that manages how prompts move from API request to GPU computation. Common frameworks include vLLM, TensorRT-LLM, and SGLang. Model providers run these internally; you see only the API.

PM implication: Framework choice affects throughput (requests per second) and latency. vLLM's PagedAttention optimization, for example, can double throughput by managing GPU memory more efficiently. This is why the same model can be 2x faster on one provider vs. another.

Orchestration layer

What it is: Load balancers, request queues, autoscaling logic, and fallback routing. Handles traffic spikes, routes requests to available GPU capacity, and retries failed requests.

PM implication: This is where your SLA is made or broken. A model that takes 300ms for p50 requests may still hit 10s for p99 if the orchestration layer queues requests during spikes. Always ask your ML team for p95 and p99 latency, not just average.

Caching layer

What it is: Semantic caches, KV caches, and prompt prefix caches that avoid re-running computation for repeated or similar inputs. KV cache is the most impactful: it stores intermediate attention computations so they don't need to be recomputed for each new token.

PM implication: Prompt caching can reduce latency and cost dramatically for products with long, repeated system prompts. If your product has a 5,000-token system prompt sent on every request, enabling prefix caching can cut costs by 60% or more.

The GPU Landscape: What PMs Need to Know

GPU model names appear in provider documentation and infrastructure discussions. You don't need to know chip specifications, but you do need to understand the generational hierarchy and what it means for your choices.

Frontier: H100 / H200

Nvidia's current flagship datacenter GPUs. H100s have 80GB of HBM3 memory; H200s have 141GB. High memory is critical for running large models without splitting them across GPUs. Most frontier model providers (Anthropic, OpenAI, Google) run on H100/H200 or equivalent. Expect premium pricing.

Mid-tier: A100

Previous flagship, still widely deployed. 40GB or 80GB variants. Slower memory bandwidth than H100s, meaning higher latency per token at the same batch size. Many cost-competitive inference providers and smaller hosted model services run on A100s. Typically 30 to 50% cheaper per million tokens vs H100 providers.

Specialized: TPUs, Trainium, Groq

Google's TPU v5, AWS Trainium, and Groq's LPU are purpose-built for AI workloads. Groq LPUs in particular achieve extremely low latency (sub-100ms for full responses) by eliminating memory bandwidth bottlenecks. Trade-off: they support fewer model architectures.

Consumer / Edge: RTX 4090, M3 Ultra

For on-device or low-cost inference of smaller models (7B to 13B parameters). Relevant if your product deploys locally rather than through a cloud provider. Lower cost, complete privacy, no API dependency, but limited model capability ceiling.

The PM rule on GPU tiers

Don't pick providers based on GPU brand names alone. What matters to you is the resulting latency SLA, cost per million tokens, and reliability track record. Provider benchmarks are more useful than hardware specs when scoping your product's cost model.

The Latency vs Throughput Tradeoff

This is the most important infrastructure concept for PM decision-making. Latency and throughput are in tension: optimizing for one generally hurts the other, and the right tradeoff depends entirely on your product's use case.

Latency

What it is: Time from request submission to first token received (TTFT, time to first token) and time to complete the full response (total generation time).

When it matters: Chat interfaces, real-time voice, interactive coding assistants, anything where a human waits and perceives the delay. If your p95 TTFT exceeds 2 seconds, users disengage.

The tradeoff: To minimize latency, you run requests one at a time or in tiny batches, leaving GPU utilization low. This is expensive per request because you're paying for mostly idle compute.

Throughput

What it is: Number of tokens generated per second across all concurrent requests. High throughput means you serve more users with the same GPU capacity.

When it matters: Batch processing, document analysis, background jobs, async features. If you're summarizing 10,000 support tickets overnight, throughput matters more than TTFT.

The tradeoff: To maximize throughput, you batch requests together so the GPU runs full. This adds latency, because a request may wait in queue until the batch is full before processing starts.

Continuous batching (a technique used in vLLM and similar frameworks) partially resolves this tension by dynamically inserting new requests into in-progress batches rather than waiting for a full batch to complete. It significantly improves throughput without the full latency penalty of static batching. Ask your ML team whether your serving stack uses continuous batching.

Go Deeper in the AI PM Masterclass

The masterclass covers how infrastructure decisions translate directly into product scope and cost: taught live by a Salesforce Sr. Director PM.

KV Cache and Context Length: Why Long Contexts Cost More

Every token the model processes generates Key and Value matrices used by the attention mechanism. These matrices are stored in the KV (Key-Value) cache during generation so the model doesn't recompute them for every new output token. Understanding KV cache explains why long contexts are expensive and why context window size matters for both performance and cost.

Why KV cache size matters

KV cache grows linearly with context length: a 128K token context requires 16x more GPU memory for the cache than a 8K context. This memory competes with batch capacity. When GPU memory fills up with KV cache for one long context, you serve fewer parallel requests, reducing throughput.

Prefix caching

If many requests share the same system prompt prefix, providers can cache the KV state for that prefix and skip recomputing it. For products with long system prompts, this can reduce cost by 50 to 80% and cut TTFT on the first token significantly. Most major providers support it; check your API documentation.

The cold start problem

When your product first receives a request after an idle period, the KV cache is empty and the model must load fully into GPU memory. This can add 5 to 20 seconds to response time. For low-traffic or burst-use products, cold starts are a real user experience problem. Solutions: keep-alive requests, minimum replica counts, or async pre-warming.

Context length vs model capability

Longer context windows don't automatically mean better performance. Empirically, information at the middle of a very long context is attended to less reliably than information at the start or end. Design your product to put critical information at the top of the context, not buried in the middle.

What PMs Should Ask the ML Team

Infrastructure knowledge turns into better product decisions only when you know which questions to ask. These are the five questions that surface the most important tradeoffs before a feature ships.

1

"What are our p50, p95, and p99 latency numbers, not just average?"

Why ask it: Averages hide the tail. A 400ms average with a 8s p99 means 1 in 100 users experiences an 8-second wait. That user may be your largest enterprise customer in the middle of a critical workflow.

2

"Are we using prefix/prompt caching, and how much are we saving from it?"

Why ask it: If your system prompt is long and consistent across requests, caching can cut your monthly model cost materially. If nobody has turned it on, you're leaving money on the table.

3

"What happens to latency when traffic doubles? Is autoscaling configured correctly?"

Why ask it: Many products are tuned for average load. A viral feature launch or press mention can spike traffic 5x in minutes. Know the autoscaling threshold and whether it's tested.

4

"Is the model running at batch size 1, or are we batching requests?"

Why ask it: Batch size 1 means every request gets a dedicated GPU computation, minimizing latency but maximizing cost. For async or background features, batching 8 to 32 requests together can cut GPU costs 70% with minimal user impact.

5

"What is our fallback when the primary provider has an outage?"

Why ask it: All inference providers have incidents. If your product has no fallback, a provider outage is a full product outage. Knowing the fallback plan and its performance characteristics is a PM responsibility.

Turn Infrastructure Knowledge Into Sharper Product Decisions

The AI PM Masterclass teaches you to reason about latency, cost, and infrastructure tradeoffs at the level that makes you a credible technical partner to your ML team.

Before you go: get the AI PM Minute

One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.

No fluff. Unsubscribe anytime.