← All tools

GPU Container

Org-only cli Python

Model-aware inference memory-placement planner for single-GPU rigs: profile hardware + model, generate explicit VRAM/RAM/NVMe placement plans across runtimes (llama.cpp/vLLM/...), and prove them with a measured receipt. Not VRAM overflow - declared placement.

Language Python
Last updated Sep 8, 2026
Not in registry Install cmd No releases Documented
cudagpuinferencellama-cppllmmoeoffloadvramwsl2