2026-09-13 06:51 UTC
If you're running local LLMs on a 16-24GB CUDA GPU this might interest you: "Raymond's fork" but with hot-swappable MTP meaning when there's room for it you get faster token generation, and when the VRAM is better used for the context it's ejected without a trace.
I used an LLM to make an LLM faster which I believe means I should start working at Cyberdyne Systems.
https://github.com/troed/llama.cpp-adaptive-kv-streaming
#LLM
Replies (0)
No replies.