Final Unicourse'tan Çalış, Yüksek Notu Garantile!
Vizesine Unicourse'tan Çalış, Yüksek Notu Garantile!
Inferno - a 16GB 5060Ti focused fork of the Blackwell-inference engine NInfern
-
The 5090Ti owners are living a life of luxury with an inference engine (NInfern) that maximizes CUDA 120a performance when running Qwen models. Purely for my own curiosity's sake I've worked with an LLM to make a fork of this project but with my 5060Ti (16GB) as target instead. The upstream author is not interested in taking on us VRAM-poor thus the name change.
My main interest lies in trying to get Qwen 3.8 27B even the slightest bit faster than possible with llama.cpp, but I have no idea if I will succeed.
Where I'm currently at is that the ffn layer offloading (the main trick from Stainless-Bacon that makes 27B possible to run at Q4 quant at all) has been implemented, as well as attention rotation for the KV cache quants and FP6 as the first of the new <8 bit KV cache quant implementations to come. Looking at adding KVarN now.
With Qwen 3.8 27B, MTP and 96000 context size at FP6 KV cache I'm currently seeing around 500 tps PP and 8-10 tps TG. This is slightly below the numbers I have in llama.cpp, although there I'm also using ngram-mod and a 5.0/4.1 KV cache quant so it's not apples to apples.
Posting about it here in case anyone else is as curious as I.

Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Kayıt Ol Giriş