ArticleslgStudy

science

List of text corpora

List of text corpora is a science topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand List of text corpora rather than just read about it. In short: Text corpora (singular: text corpus) are large and structured sets of texts, which have been systematically collected. Text corpora are used by both AI developers to train large language models and corpus linguists and within other branches of linguistics for statistical analysis, hypothesis testing, finding patterns of language use, investigating language change and variation, and teaching language proficiency.

Key takeaways

  • List of text corpora belongs to science; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect List of text corpora to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of List of text corpora from memory before moving on to harder problems.

Reference excerpt

Text corpora (singular: text corpus) are large and structured sets of texts, which have been systematically collected. Text corpora are used by both AI developers to train large language models and corpus linguists and within other branches of linguistics for statistical analysis, hypothesis testing, finding patterns of language use, investigating language change and variation, and teaching language proficiency.

English language American National Corpus Bank of English BookCorpus British National Corpus Bergen Corpus of London Teenage Language (COLT) Brown Corpus, forming part of the "Brown Family" of corpora, together with LOB, Frown and F-LOB COCA: see below at English-Corpora.org COHA: see below at English-Corpora.org Corpus Resource Database (CoRD), more than 80 English language corpora. Coruña Corpus, a corpus of late Modern English scientific writing covering the period 1700–1900, developed by the Muste research group at the University of A Coruña DBLP Discovery Dataset (D3), a corpus of computer science publications with sentient metadata. English-Corpora.org, which contains (among others): iWeb, the Intelligent Web-based Corpus: 14 billion words, 6 countries, 2017 COCA, the Corpus of Contemporary American English: 1.0 billion words, American, 1990-2019 COHA, the Corpus of Historical American English: 475 million words, American, 1820-2019 NOW, News on the Web: 23.2 billion+ words, 20 countries, 2010-present English Trends, a large English monitor corpus of news articles gathered from RSS feeds, 86+ billion words, 2014–present GUM corpus, the open source Georgetown University Multilayer corpus, with very many annotation layers Google Books Ngram Corpus International Corpus of English Oxford English Corpus RE3D (Relationship and Entity Extraction Evaluation Dataset) Santa Barbara Corpus of Spoken American English Scottish Corpus of Texts & Speech Strathy Corpus of Canadian English

European languages CETENFolha Basque: The Corpus of Electronic Texts Corpus Inscriptionum Insularum Celticarum (CIIC), covering Primitive Irish inscriptions in Ogham Google Books Ngram Corpus The Georgian Language Corpus Thesaurus Linguae Graecae (Ancient Greek) Eastern Armenian National Corpus (EANC) 110 million words. Freely searchable online. Spanish text corpus by Molino de Ideas, which contains 660 million words. CorALit: the Corpus of Academic Lithuanian Academic texts published in 1999–2009 (approx. 9 million words). Compiled at the University of Vilnius, Lithuania Reference Corpus of Contemporary Portuguese (CRPC) Turkish National Corpus CoRoLa - The Reference Corpus of the Contemporary Romanian Language (Corpus reprezentativ al limbii române contemporane ) TS Corpus - A large set of Turkish corpora. TS Corpus is a Free&Independent Project that aims to build Turkish corpora, NLP tools and linguistic datasets... MacMorpho - an annotated corpus of Brazilian Portuguese text

Slavic

East Slavic Belarusian N-korpus Russian National Corpus General Internet Corpus of Russian General Regionally Annotated Corpus of Ukrainian Ukrainian Language Corpus on the Mova.info Linguistic Portal Ukrainian Language Corpus Araneum Russicum Russian Corpus of Biographical Texts RuTweetCorp RusAge: Corpus for Age-Based Text Classification

South Slavic Bulgarian National Corpus Macedonian Electronic Corpus Croatian Language Corpus Croatian National Corpus Slovenian National Corpus

West Slavic Czech National Corpus National Corpus of Polish Slovak National Corpora

German German Reference Corpus (DeReKo) More than 4 billion words of contemporary written German. Free corpus of German mistakes from people with dyslexia

Middle Eastern Languages Corpus Inscriptionum Semiticarum Kanaanäische und Aramäische Inschriften Hamshahri Corpus (Persian) Persian in MULTEXT-EAST corpus (Persian) Amarna letters (for Akkadian, Egyptian, Sumerogram's, etc.) TEP: Tehran English-Persian Parallel Corpus PTC: Persian Today Corpus: The Most Frequent Words of Today Persian, based on a one-million-word corpus (in Persian: Vāže-hā-ye Porkārbord-e Fārsi-ye Emrūz), Hamid Hassani, Tehran, Iran Language Institute (ILI), 2005, 322 pp. ISBN 964-8699-32-1 Kurdish-corpus.uok.ac.ir (Kurdish-corpus Sorani dialect) University of Kurdistan, Department of English Language and Linguistics Bijankhan Corpus A Contemporary Persian Corpus for NLP researches, University of Tehran, 2012 Neo-Assyrian Text Corpus Project Quranic Arabic Corpus (Classical Arabic) Electronic Text Corpus of Sumerian Literature Open Richly Annotated Cuneiform Corpus Asosoft text corpus – Central Kurdish (Sorani) Thesaurus Linguae Aegyptiae (ancient Egyptian, Afro-Asiatic)

Turkic languages Uzbek national corpus (20 million words)

Devanagari Nepali Text Corpus (90+ million running words/6.5+ million sentences)

East Asian Languages Kotonoha Japanese language corpus LIVAC Synchronous Corpus (Chinese)

South Asian Languages Hindi: SinMin dataset (Sinhala)

African languages Amharic: Creole (Gulf of Guinea): Hausa: Igbo: Oromo: Yoruba: Zulu:

… excerpt ends here. Continue reading the full article.

Worked examples

Example 1 — a first encounter with List of text corpora

Start with the simplest possible case. Write down what List of text corpora claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In science, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to List of text corpora before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about List of text corpora ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of List of text corpora

In research
List of text corpora appears in science research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses List of text corpora in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
List of text corpora is common in secondary-school and first-year university syllabi. It links to neighbouring topics Corpus linguistics, Linguistics lists, Natural language processing, so understanding it makes those chapters shorter.
In everyday life
Look for List of text corpora outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.
Ask Teacher Smith questions about this articleOpens your AI tutor with a question about “List of text corpora” →

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study List of text corpora in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what List of text corpora means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain List of text corpora out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is List of text corpora in simple terms?

Text corpora (singular: text corpus) are large and structured sets of texts, which have been systematically collected. Text corpora are used by both AI developers to train large language models and corpus linguists and within other branches of linguistics for statistical analysis, hypothesis testin…

Why does List of text corpora matter?

Because it connects several science ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study List of text corpora?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on List of text corpora.

Tags

  • Corpus linguistics
  • Linguistics lists
  • Natural language processing

Keep exploring