Elektrine lite

← Feed

@fj@mastodon.social

2026-09-03 07:34 UTC

The "breakthrough" from the OpenAI Astra models are the implementation of the concepts from this 2015 paper recurrent neural network paper by Jürgen Schmidhuber from Università della Svizzera italiana. https://arxiv.org/pdf/1511.09249 "The new technique OpenAI is using, known as recurrent depth or looped transformer, allows an AI model to improve its answers by processing the same text multiple times.” https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns

Replies (1)

  • @fj@mastodon.social I see it as more of a 2016 paper but the point the same. A decade old, openly reproduced architecture is getting hyped as a secret breakthrough. I'll write it up. Graves 2016, Adaptive Computation Time for Recurrent Neural Networks (DeepMind; the halting mechanism): https://arxiv.org/abs/1603.08983

    Open ##4610661