2026-09-07 15:37 UTC
They’re oversold though, especially prompt caching and the parameter count war
The US model is that they think more training compute and parameters will result in the winning model, while the Chinese are focusing on RL and parameters efficiency due to compute limits.
Replies (1)
-
@percent@infosec.pub 2026-09-08 00:35
The efficiency of Chinese models really is impressive. I generated sooo much code yesterday with Qwen3.6 35B-A3B running on an RTX 5060 Ti 16GB (+ a little CPU offloading). It got the jobs done at ~50 tokens/sec. (It’s not super complex code, just some scripts that I would not have taken to time to write manually.) I’d love to upgrade to something with more VRAM, but even my current card has doubled in price since I bought it last year 😬