@BarbecueCowboy@lemmy.dbzer0.com
2026-09-07 15:26 UTC
It’s kinda surprising,
I know specifically where one of the big ones hosts its models and its not there, but I guess they could have infrastructure in there.
Replies (1)
-
@panda_abyss@lemmy.ca 2026-09-07 15:37
They’re oversold though, especially prompt caching and the parameter count war The US model is that they think more training compute and parameters will result in the winning model, while the Chinese are focusing on RL and parameters efficiency due to compute limits.