Elektrine lite

← Feed

@giacomo@snac.tesio.it

2026-08-23 07:49 UTC

You are either misunderstanding the #Shannon theorem, #LLM working or both. In fact, LLM's weights are lossy compressions of the source data ("training data", in #AI parlance). They mimic intelligence by chaining statistically related fragments of such source data that the users have a negligible probability to recognize as quotations (basically because they haven't read the original sources). The low probability lets AI companies (and their addicted users who "want to believe") to dismiss and/or gaslight evidences as anecdotes, but in low frequency areas of the source data, where decompression artifacts are rare enough, they become quite evident (see for example the evidences produced in NYT vs OpenAI). Or, much more simply, see #GitHub #Copilot distributing #GPLv3 code with wrong attribution and under a wrong license.

Replies (1)

  • @troed@swecyb.com 2026-08-23 07:53

    @giacomo You've misunderstood the paper. LLMs _can_ be used as compressors, but that does not mean that LLM models _are_ compressing source data retrievably. The reason Shannon's theorem proves this is simple. The model I used for the task in this thread is a 12B parameter model. It's much (much!) too small to store (compress) any relevant amount of source data so that it can be retrieved (decompressed) in the way you believe. What you're referring to is known as "overfitting" and is something no LLM company wants. It means model parameters are wasted instead of used for the higher level concepts of _how_ to do things. If you're actually interested in the topic we can continue, but considering you seem to believe you "know better" I'm not certain you're discussing in good faith here.

    Open ##4528811