Skip to main content
All articles

Dynamic LoRA swapping enables multi-tenant LLM serving without cold-starts

Serving many fine-tuned models from one resident base model by hot-swapping adapters, and the memory and storage trade-offs.

Sahil BansalConnect

4 min readOriginally on Medium

The architecture of LLM serving is shifting from monolithic deployment to dynamic, request-level adapter swapping. This transition is finally making multi-tenant GPU utilization viable for specialized workloads.

Why This Matters

Operating specialized LLMs in production has traditionally been a resource nightmare. If you have 50 different fine-tuned models for 50 different customers, the old-school approach was to spin up 50 inference endpoints.

This leads to two unacceptable outcomes: either you bleed money keeping idle GPUs warm, or you force users to endure 30-second cold starts while a container pulls a 20GB model from a registry. Dynamic LoRA (Low-Rank Adaptation) swapping in engines like vLLM solves this by keeping one base model resident in VRAM and hot-swapping tiny adapter weights on the fly.

Architecture / System Design

In a dynamic LoRA setup, the serving engine maintains a single instance of the base model (e.g., Llama-3–70B) in GPU memory. The infrastructure no longer treats the model as a static artifact; instead, it treats the adapter as a request-scoped parameter.

The Adapter Pool

vLLM implements a LoRA cache in both CPU and GPU memory. When a request arrives with a specific adapter_id, the engine checks if the weights are already in the GPU cache. If not, it fetches them from CPU RAM—or in a worst-case scenario, from a local SSD or S3 bucket.

Memory Management

The tension here is between the KV cache and the adapter pool. Every megabyte dedicated to storing LoRA adapters is a megabyte taken away from the KV cache, which directly impacts the maximum concurrency (batch size) the engine can handle. In production, you aren’t just managing model weights; you are managing a multi-tiered memory hierarchy where the ‘active’ set of adapters competes for space with the tokens being generated.

Implementation

Setting up a dynamic LoRA server requires explicit configuration of the adapter capacity. You cannot simply ‘enable’ it; you have to budget for it.

To start a vLLM instance with LoRA support, the entry point looks like this:

python -m vllm.entrypoints.openai.api_server \
    --model /models/llama-3-base \
    --enable-lora \
    --max-loras 20 \
    --max-lora-rank 16 \
    --lora-modules \
        sql-adapter=/models/adapters/sql-finetune \
        email-adapter=/models/adapters/email-finetune

At the application level, the client specifies the adapter in the standard OpenAI-compatible request body:

import openai
response = openai.ChatCompletion.create(
    model="sql-adapter",
    messages=[{"role": "user", "content": "Write a query for..."}],
    extra_body={"prompt_logprobs": 1}
)

From an infrastructure perspective, you need to mount a high-speed NVMe volume to /models/adapters. If you are pulling these from S3 dynamically, your bottleneck isn't the GPU—it's the PCIe bandwidth or the network throughput between the storage layer and the compute node.

Operational Realities

Dynamic swapping introduces ‘jitter’ that doesn’t exist in static deployments. The first request for a cold adapter will always incur a Time-To-First-Token (TTFT) penalty while the weights are paged into VRAM.

  • Swap Latency: Moving a 100MB adapter from CPU RAM to GPU memory is fast, but doing it under heavy load while the engine is busy with a 128-request batch can introduce millisecond-level stalls.
  • VRAM Fragmentation: vLLM manages its own memory pool. If you configure max_loras too high, you risk starving the KV cache, leading to aggressive request preemption and 'context-switching' overhead at the inference level.
  • Observability: You need metrics beyond just GPU utilization. You must track adapter_cache_hit_rate and swap_in_latency. If your hit rate is low, your multi-tenancy strategy is actually just a slow version of serverless cold starts.

Failure Modes / Trade-offs

The ‘Thundering Herd’ of Adapters

When a new tenant goes live and floods the system with requests for a cold adapter, the engine might struggle to fetch and apply weights while maintaining the throughput of existing streams. This can lead to a global latency spike affecting all tenants on that GPU.

Version Mismatch

LoRA adapters are tied to the specific architecture of the base model. If your platform team updates the base model version but doesn’t re-train or validate the 500 adapters in the library, the inference will either fail or, worse, produce silent, high-perplexity garbage.

Cost Implication

While you save on GPU idle time, you increase the operational complexity of the storage layer. You need a highly available, low-latency file system (like Amazon FSx for Lustre or a tuned EFS) to back the adapter library, or your ‘dynamic’ swapping will be throttled by S3 GET requests.

Lessons Learned / Best Practices

  • Pin Popular Adapters: Use the --lora-modules flag to pre-load the top 5 most-used adapters at boot time to ensure they never leave the GPU cache.
  • Rank Consistency: Standardize on a single LoRA rank (e.g., 8 or 16) across all teams. Mixing ranks complicates memory allocation and can lead to inefficient padding in the engine’s internal tensors.
  • Tiered Storage: Keep the ‘active’ adapter set on local NVMe and sync with S3 in the background. Don’t make the inference engine wait for the network.

TL;DR

  • Consolidation: Move from one-GPU-per-model to one-GPU-per-N-tenants by swapping LoRA adapters at request time.
  • Memory Pressure: Adapters compete with the KV cache for VRAM; over-provisioning adapters will kill your max batch size.
  • Latency Jitter: The first request to a cold adapter will have a higher TTFT; monitor your cache hit rates religiously.
  • Infrastructure: Your storage throughput (disk to GPU) becomes a primary bottleneck for model availability.
  • When to use: Use when you have many low-traffic specialized models; avoid if you have a single model with consistent, high-volume traffic.

Final Thoughts

Dynamic LoRA swapping turns LLM serving into a caching problem rather than a capacity problem. It’s the only way to scale specialized AI features without the unit economics falling apart, but it requires a much deeper understanding of the GPU memory hierarchy than traditional model serving.

  • LLM serving
  • GPUs
  • Multi-tenancy

Written by Sahil Bansal

DevOps and platform engineer. I write about the infrastructure decisions I have had to live with.

Connect