The British National Corpus 1994 to 2026

Martin Wynne, Megan Bushnell


Abstract
The British National Corpus (BNC) is a 100 million word collection of samples of written and spoken language from a wide range of sources, designed to represent a wide cross-section of British English from the later part of the 20th century, both spoken and written. It is one of the first generation of monolingual, synchronic, general, representative corpora of its size, and led the way for other national corpora. It was created by a consortium of academic partners and publishers, with funding from the Department of Trade and Industry in the UK. This posters reflects on a number of lessons learning in more than thirty years, in terms of corpus representativeness, modes of access to the corpus, licensing, and managing the transition from a contemporary synchronic corpus to a historical corpus.
Anthology ID:
2026.cmlc-1.11
Volume:
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Piotr Bański, Dawn Knight, Marc Kupietz, Andreas Witt, Alina Wróblewska
Venues:
CMLC | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
76–77
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-cmlc-11
DOI:
10.63317/4yzqikwnh4b8
Bibkey:
Cite (ACL):
Martin Wynne and Megan Bushnell. 2026. The British National Corpus 1994 to 2026. In Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora, pages 76–77, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
The British National Corpus 1994 to 2026 (Wynne & Bushnell, CMLC 2026)
Copy Citation: