Splitting an LLM Across GPUs: Tensor vs Pipeline
A model too big for one GPU has to be split. How tensor and pipeline parallelism divide an LLM, what each costs, and when to reach for which.
A model too big for one GPU has to be split. How tensor and pipeline parallelism divide an LLM, what each costs, and when to reach for which.
Add a cache node with plain modulo hashing and almost every key moves. Consistent hashing moves only one arc. How the hash ring and virtual nodes work.
Pick a GPU for an LLM and it 'fits' on paper, then OOMs under load. A quick way to estimate the VRAM a model needs: weights, KV cache, and overhead.
Counting unique items exactly costs memory that grows with your data. HyperLogLog estimates the cardinality of billions in about 12KB. Here's how it works.
How rotary position embeddings encode token order by rotating query and key vectors, and how position interpolation and YaRN stretch a model's context window.
Roaring bitmaps store integer sets in three container types so filters stay small and fast. How they work, why OpenSearch leans on them, and the tradeoffs.
One async request scatters a dozen log lines through concurrent traffic. Thread a correlation ID through FastAPI with contextvars and trace it in OpenSearch.
Mixture of Experts routes each token to a couple of expert layers, so the model runs cheaper per token yet still needs every expert sitting in GPU memory.
Float32 embeddings make a vector index expensive to keep in RAM. Binary quantization cuts them 32x; Hamming search plus rescoring keeps recall high.
Most prompts hitting an LLM app are simple. Route those to a small model and escalate only the hard ones to a large model to cut cost and keep quality.
Kubernetes native sidecars are init containers that keep running. They start before your app, restart on their own, and get shut down last. Here is how.
An LLM streams tokens and React repaints your whole chat history on every one. How to keep the message list still while a single reply streams in.
Exact md5 hashing misses near-duplicate documents; comparing every pair is O(n²). How MinHash and LSH find and cluster them across a RAG corpus.
Run RAG retrieval inside the OpenSearch cluster you already operate. Set up a knn_vector index, pick faiss vs lucene, and avoid the memory traps.
How to make a RAG system cite its sources: force the model to reference chunks by ID, then verify each citation against the source text before trusting it.
FastAPI validates request bodies with Pydantic v2 before your handler runs. Models, Field constraints, custom validators, the 422 response, and failure modes.
Cosine, dot product, and Euclidean rank vector search results the same on normalized embeddings, differently otherwise. How to pick and configure one.
SSE streams one direction only. When an LLM chat needs the browser to interrupt a running response, FastAPI WebSockets give you a full-duplex channel.
Naive RAG retrieves on every turn, even when it shouldn't. How agentic RAG lets the model route, grade its own context, and re-retrieve, plus the tradeoffs.
Brute-force vector search compares every embedding. How an IVF index partitions vectors into cells, probes only the nearest, and where the nprobe knob bites.
Small chunks search precisely but read like fragments. Parent document retrieval matches on child chunks, then feeds the whole parent to the model.
Top-k vector search keeps returning near-duplicate chunks that fill the context window. How Maximal Marginal Relevance reranks for relevant, diverse results.
LLM structured output isn't valid JSON until the last token. Here's how to complete and parse a streaming buffer so you can render fields as they arrive.
A cluster upgrade drains a node and evicts every replica of your service at once. How a PodDisruptionBudget caps voluntary disruptions and keeps you online.
How to parse PDFs for a RAG pipeline: detect born-digital vs scanned files, pull out tables without flattening them, and reach for OCR only when you must.
Reciprocal Rank Fusion merges two ranked lists without tuning score weights. The RRF formula, why the k constant matters, and how to run it in OpenSearch.
Log indexes grow until a data node runs out of disk. How OpenSearch ISM rolls indexes over by size, ages them through hot and warm, then deletes on schedule.
A network retry can submit the same request twice. How idempotency keys in FastAPI make POST endpoints safe to retry without duplicate side effects.
Prompt caching reuses a stable prompt prefix to cut Claude API cost and latency. How breakpoints, TTLs, and cache reads work, and what quietly breaks a hit.
A naive LLM server holds finished slots idle until the slowest request in the batch drains. Continuous batching refills them every step. How it works.
Post-filtering vector search silently drops good results. Here is pre-filter vs filtered ANN for RAG, why filtered HNSW is hard, and how to choose.
The Kubernetes HPA scales pods from a metric, but its defaults thrash and lag under real load. How the control loop works, how to tune it, and where it breaks.
Vector search misses chunks that lost their document context. Contextual Retrieval prepends a short LLM-written summary to each chunk before you embed it.
When an LLM provider degrades, retries make it worse. A practical guide to adding a circuit breaker in Python: the three states, tuning, and failure modes.
A GPU node won't run your Pods until a device plugin advertises it. How Kubernetes discovers GPUs, how to share one across Pods, and where scheduling breaks.
Speculative decoding uses a small draft model to guess tokens a big model verifies in one parallel pass, cutting LLM latency with no change to output.
Trace a RAG pipeline with OpenTelemetry: instrument FastAPI, put each retrieval and model step in its own span, and read the latency waterfall in OpenSearch.
The first token is slow, the rest stream fast. Why LLM inference splits into a compute-bound prefill and a memory-bound decode, and what it costs you.
Retrieval metrics say the right docs came back, not that the answer is right. Build an LLM-as-a-judge to score RAG answers for faithfulness and quality.
Retrieved documents are untrusted input. A practical guide to defending a RAG copilot against direct and indirect prompt injection, and why filters alone fail.
Vector search fails when a short question looks nothing like its answer. HyDE has an LLM draft a fake answer, embeds that, and retrieves against it instead.
Semantic caching for LLM apps: cache answers by embedding similarity in FastAPI, tune the cutoff, and avoid false cache hits. Working code and failure modes.
HNSW makes vector search fast, but the embeddings still fill your RAM. Product quantization compresses them ~32x with a small recall hit. Here is how it works.
Awaiting retrieval and LLM calls one by one wastes seconds per request. Here's how to fan them out with asyncio.gather, bound it, and handle partial failures.
Grouped-query attention shares key/value heads across query heads to cut the KV cache. How GQA sits between MHA and MQA, with the memory math and code.
A self-hosted LLM server wastes most of its GPU memory to KV cache fragmentation. Here is how PagedAttention in vLLM pages the cache like an OS.
Matryoshka embeddings front-load meaning into a vector's first dimensions, so you can truncate them for cheaper, faster RAG search without losing much recall.
Change an OpenSearch mapping without dropping writes or serving stale data. A step-by-step reindex with aliases, the Reindex API, and the failure modes.
Add per-user rate limiting to a FastAPI backend with the token bucket algorithm: an in-process version, an atomic Redis script, 429s, and the failure modes.
Byte-pair encoding turns text into the tokens an LLM bills and reasons over. How BPE merges are learned, why token counts drive cost, and where it breaks.
Your async FastAPI app stalls under load, then times out. Here is how SQLAlchemy pool_size and max_overflow really work, and how to size them for Postgres.
Attention is memory-bound, not compute-bound. Here's how FlashAttention uses tiling and online softmax to skip the N×N matrix and run exact attention faster.
ArgoCD ApplicationSets generate one Application per cluster or environment from a single template. How the generators work, a Git example, and where they bite.
Run finite and scheduled work on Kubernetes with Jobs and CronJobs: completions, parallelism, backoffLimit, cleanup, and the failure modes that bite.
A bi-encoder averages token detail away; a cross-encoder is too slow to rank a corpus. Late interaction with ColBERT sits between them. Here is how it works.
A practical guide to FastAPI dependency injection: how Depends resolves a graph, yield setup and teardown, per-request caching, and where it leaks.
Chunking a document for RAG strips each piece of its context. Contextual retrieval adds an LLM-written note to every chunk before you index it.
Store and query RAG embeddings inside Postgres with pgvector: HNSW indexing, distance operators, metadata filtering, hybrid search, and the tradeoffs I hit.
FastAPI BackgroundTasks run inside your web process and disappear on restart. When that is fine, when you need a real task queue, and how to move over.
Adding a metadata filter to a vector search can silently return fewer results or wreck recall. How post-filter, pre-filter, and filterable HNSW actually differ.
A rolling deploy sends SIGTERM and kills your FastAPI pod mid-request, dropping live SSE streams. How to catch it, drain connections, and shut down cleanly.
Temperature, top-p, and top-k are the three knobs that shape how an LLM picks each token. How each one works, when to reach for it, and how they interact.
An LLM agent that runs long enough fills its context window and starts to slow or fail. How to prune, compact, and offload context so agents keep going.
A CPU-based HPA can't see a backed-up queue, so workers fall behind. How KEDA autoscales Kubernetes workers on queue depth, and where it breaks.
The weights don't fit on the GPU you have. How LLM quantization shrinks them to INT8 or INT4, why GPTQ and AWQ beat naive rounding, and where it breaks.
Prompt caching reuses a request's prefix to cut LLM cost and latency. How the prefix match works, where to put the breakpoint, and the silent cache misses.
The embedding model sets the ceiling on RAG retrieval quality. How to choose one by task fit, sequence length, dimensions, and domain, plus the silent bugs.
Short, vague, follow-up questions don't match how your docs are written. How query rewriting, multi-query expansion, and HyDE fix retrieval before it runs.
Speculative decoding uses a small draft model to guess tokens a big model verifies in one pass, cutting LLM latency 2-3x without changing the output.
Static batching leaves the GPU idle when requests finish at different steps. How continuous batching schedules LLM inference per token to raise throughput.
An LLM agent request hides where the time and tokens went behind one flat log. Trace it with OpenTelemetry spans, the GenAI conventions, and where it breaks.
Exact-match caching misses paraphrases, so LLM bills stay high. Here is how to build a semantic cache with embeddings, a similarity threshold, and its traps.
A RAG copilot reads tickets and logs, so whoever writes them can plant instructions in the prompt. How indirect prompt injection works and how to contain it.
Human review does not scale for grading LLM answers. How to use an LLM as a judge: write a rubric, score with structured output, and control the biases.
Your pod died with OOMKilled and exit 137. What Kubernetes memory requests and limits actually do, what triggers the kill, and how to stop it happening.
A user hits stop or switches chats and the old LLM stream keeps writing tokens and running up cost. How to cancel a streaming fetch in React the right way.
MCP standardizes how LLM agents reach your tools and data. A hands-on guide to building an MCP server in Python, picking a transport, and where it breaks.
A long prompt or a long agent run can hit CUDA out of memory, and the KV cache is usually why. How it grows, the per-token math, and how to shrink it.
LLM tokens arrive one at a time, and re-parsing Markdown on every token flickers and drags. How to render streaming Markdown in React without the jank.
Getting an LLM to return JSON is easy; getting valid JSON every time is not. How to use JSON Schema, constrained decoding, and validation to make it reliable.
Your LLM backend returns 429s the moment traffic bursts. How to retry with backoff and jitter, respect Retry-After, and pace fan-out to stay under the limit.
Every RAG stack leans on HNSW but treats it as a black box. Here is how the layered graph index finds nearest neighbors fast, and the knobs that matter.
One blocking call in a FastAPI route stalls every other request, including live SSE streams. Here is how the event loop breaks, and how to keep it free.
Retrieval puts the right chunk at rank 8, but the generator only reads the top few. How a cross-encoder reranker reorders RAG candidates, and where it fails.
Kubernetes has three health probes and mixing them up causes outages. How liveness, readiness, and startup probes work, and the failure modes to avoid.
An LLM that writes and runs code needs real isolation, not a try/except. How to sandbox AI-generated code with E2B microVMs, and the failure modes.
How to chunk documents for a RAG pipeline: why fixed-size splitting fails, structure-aware splitting, size and overlap tradeoffs, and the failure modes.
Vector search alone misses exact IDs and error codes. Here's how to combine BM25 keyword search with dense retrieval, fuse the rankings with RRF, and rerank.
ArgoCD applies your whole app at once, so a migration races the pods that need it. How sync waves and hooks order a Kubernetes rollout, and where they stall.
Changed your embeddings or added a reranker? Measure it. Build a golden set and score retrieval with recall@k, MRR, and nDCG before you trust the change.
Stream LLM output token by token from a FastAPI backend with Server-Sent Events: working code, the EventSource client, proxy buffering, and failure modes.
A practical look at the agent loop behind LLM tools: how the model asks to call a tool, your code runs it, and the result feeds back until the answer is done.
Build an OpenSearch dashboard that watches a job pipeline: structured logs, an index mapping, the queries behind each panel, and alerts that fire early.
How to keep a RAG index fresh without full rebuilds: detect changed files with checksums, re-embed only what changed, and handle deletions safely.