Elektrine lite

← Feed

@GreatBigTable@mastodon.social

2026-09-06 08:47 UTC

Oof. Not the take I would have expected from you. The fact that a well crafted prompt can pull near identical data from the training source shows this is not the case. Touch grass, @eff@mastodon.social, I think you have spent too much time online.

Replies (1)

  • @troed@swecyb.com 2026-09-06 12:01

    @GreatBigTable@mastodon.social This is simply not true in the general case*. The proof is simple and based in computer science and maths: A great LLM to host locally is Qwen 3.8 27B. Quantized to 4 bit weights the size of the model is ~14GB on disk. All of human knowledge cannot be compressed down to 14GB and then decompressed again (see Shannon's theorem). Thus LLMs do not work by storing training data. *) When training has _failed_ and cause what's known as "overfitting" too much training data is stored close to verbatim in the model. This is unwanted (wastes model space) and are extreme outliers model creators work actively and successfully to make sure doesn't happen. @eff@mastodon.social

    Open ##4644725