The risks specific to large language models
Corpus size, misleading fluency and concentration of power: the critiques that have shaped the field.
ObjectiveExplain the main risks raised by large models and the corresponding guardrails.
In 2021, the paper “On the Dangers of Stochastic Parrots” warned against the race for size. The authors highlight several risks. Huge corpora are impossible to document fully and carry biased or toxic content. Fluent text inspires trust even when it is wrong. The cost of these models concentrates power in the hands of a few players.
- Misleading fluency: require sources and check the key facts.
- Opaque corpora: prefer suppliers that document their data and their limitations.
- Dependency: keep the option of switching models and of running AI on your own servers.
References
- Bender, Gebru, McMillan-Major, Shmitchell (2021). On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. Proceedings of ACM FAccT 2021, p. 610–623. doi.org/10.1145/3442188.3445922