Overmind: Cut ML Model Loading from 15s to 0.2s

Community Article
Published August 24, 2026

This is a condensed summary. The full write-up with code walkthroughs for every edge case is on the Meshy Research blog.

TL;DR Loading an ML model with from_pretrained is slow even when the OS page cache is fully warm — the bottleneck isn't disk I/O, it's memory copying. We built Overmind, an open-source caching library that reconstructs cached transformers/diffusers pipelines via zero-copy shared memory, needing no changes to inference code. After the first load, subsequent loads only pay the unavoidable cost of moving weights to the GPU.

Code: github.com/meshy-dev/overmind

Why "warm cache" isn't actually fast

At Meshy's inference scale, we run several distinct model checkpoints that don't all fit in VRAM together, so pipelines get swapped in and out constantly. Naively reloading with from_pretrained every time is far too slow — even with the Linux page cache fully warmed, loading can still take tens of seconds. The overhead isn't reading bytes off disk; it's the repeated memory copies involved in deserializing gigabytes of tensor data into fresh Python/PyTorch objects.

The core idea

A PyTorch tensor is just three things: a type, some metadata, and an UntypedStorage object holding the actual bytes. Overmind monkey-patches from_pretrained calls, and instead of the standard pickle round-trip, it serializes the storage objects by reference into anonymous shared memory (via memfd_create), and reconstructs them on the other side directly from a memoryview, via a small C++ helper that constructs an UntypedStorage directly over the buffer — with no copy in between.

This avoids the two real limitations of PyTorch's built-in tensor-sharing (torch.multiprocessing.reductions): it uses POSIX shm, which is capped (Docker defaults to just 64 MiB), allocates a dedicated shm object per tensor, and deallocates memory after the first unpickle — none of which work for a cache meant to be read many times by many processes.

The hard part: making every model family cooperate

The general mechanism gets you most of the way, but real-world model loading kept breaking it in specific ways, each requiring a targeted fix:

  • bitsandbytes quantized models need custom pickle reduction functions, since their quantized parameter types weren't built to be picklable, and quantization itself has to happen in a throwaway subprocess to avoid leaving stray CUDA state in the server process.
  • diffusers' dynamically generated Python modules get imported at runtime into a diffusers_modules namespace — but the client process doesn't have that namespace in its sys.path, so unpickling breaks unless it's explicitly re-imported first.
  • stable-fast compiled artifacts can't be pickled directly; Overmind routes them through torch.jit.save/load instead, with a custom zero-copy loader and a patch so the "flattening" step preserves the actual class type instead of dropping it.
  • Model arguments that reference other cached models (e.g. passing an already-loaded vae into another from_pretrained call) break the assumption that arguments are simple hashable values — Overmind assigns cache IDs to loaded objects so they can be recovered by reference instead of re-pickled in full.
  • Along the way, some nested (non-top-level) functions turned out to not be picklable at all with the standard library, which forced a switch to dill — slower to serialize, but only paid once per model, not on every subsequent load.

Benchmarking

Testing a typical multi-component Stable Diffusion pipeline (VAE + two ControlNets) on an Intel i9-11900K + RTX 4090:

Scenario Total time
Without Overmind (every load) ~5.6–6.3 s
With Overmind, first load (cache cold) 24.2 s
With Overmind, subsequent loads ~1.1 s

The first Overmind load is slower — it pays the cost of populating shared memory. But every load after that drops to about 1.1 seconds total, of which roughly 0.9 seconds is the unavoidable .to('cuda') transfer — meaning the actual model reconstruction itself is on the order of 0.2 seconds.

An unexpected win: memory, not just speed

Beyond faster iteration, deploying Overmind across multiple GPU-bound application instances on the same node measurably cut system memory usage, since shared checkpoints are no longer duplicated in memory per process. And for pipeline/algorithm developers doing frequent modify-verify loops, saving 10–20 seconds per iteration adds up fast — and matters as much for keeping people in flow as for raw throughput.

Full write-up (every code snippet, every edge case): meshy.ai/blog/fast-ml-model-loading-overmind Code: github.com/meshy-dev/overmind

Community

Sign up or log in to comment