Elektrine lite

← Feed

@troed@swecyb.com

2026-08-22 14:07 UTC

@giacomo Shannon's theorem is enough to show that LLMs can not be explained from them storing the training data.

Replies (1)

  • @giacomo@snac.tesio.it 2026-08-23 07:49

    You are either misunderstanding the #Shannon theorem, #LLM working or both. In fact, LLM's weights are lossy compressions of the source data ("training data", in #AI parlance). They mimic intelligence by chaining statistically related fragments of such source data that the users have a negligible probability to recognize as quotations (basically because they haven't read the original sources). The low probability lets AI companies (and their addicted users who "want to believe") to dismiss and/or gaslight evidences as anecdotes, but in low frequency areas of the source data, where decompression artifacts are rare enough, they become quite evident (see for example the evidences produced in NYT vs OpenAI). Or, much more simply, see #GitHub #Copilot distributing #GPLv3 code with wrong attribution and under a wrong license.

    Open ##4528812