2026-09-08 00:35 UTC
The efficiency of Chinese models really is impressive. I generated sooo much code yesterday with Qwen3.6 35B-A3B running on an RTX 5060 Ti 16GB (+ a little CPU offloading). It got the jobs done at ~50 tokens/sec.
(It’s not super complex code, just some scripts that I would not have taken to time to write manually.)
I’d love to upgrade to something with more VRAM, but even my current card has doubled in price since I bought it last year 😬
Replies (1)
-
@isVeryLoud@lemmy.ca 2026-09-08 01:21
Show me your set up!