ArticleslgStudy

science

Spoken English Corpus

Spoken English Corpus is a science topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Spoken English Corpus rather than just read about it. In short: The Spoken English Corpus (SEC) is a speech corpus collection of recordings of spoken British English compiled during 1984–1987. The corpus manual can be found on ICAME.

Spoken English Corpus — main illustration
Spoken English Corpus — illustration

Key takeaways

  • Spoken English Corpus belongs to science; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Spoken English Corpus to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Spoken English Corpus from memory before moving on to harder problems.

Reference excerpt

The Spoken English Corpus (SEC) is a speech corpus collection of recordings of spoken British English compiled during 1984–1987. The corpus manual can be found on ICAME.

History The Spoken English Corpus (SEC) project was supported jointly in 1984-5 by the Humanities Research Fund at Lancaster University and by IBM (UK) Ltd, and subsequently by IBM UK Ltd. The project was supported by Geoffrey Leech at Lancaster and Geoffrey Kaye at IBM. The project was a collaboration, funded by IBM, between the Unit for Computer Research on the English Language (UCREL) at the University of Lancaster and the IBM Scientific Centre in Winchester.

Compilation SEC comprises 53 recorded passages, mainly from the BBC, spoken in the accent usually referred to as Received Pronunciation, or RP. The collection covers categories such as commentary, news broadcast, lecture, dialogue, poetry and propaganda. The corpus contains 52,637 words, totalling 339 minutes. The compilation of the corpus is described by Lita Taylor in her 1996 article "The Compilation of the Spoken English Corpus."

Transcription

A system was devised for transcription of the intonation of the material in the recordings. Two transcribers, Gerry Knowles and Briony Williams, both supported by Lita Taylor, analysed the entire corpus. The transcription system is explained by Williams, and an experiment was conducted by Brian Pickering to assess the degree of agreement between the two transcribers on a section of the Corpus containing around 1000 tone-units which was transcribed by both transcribers. Good agreement was found.

The whole transcription in print was made in its present form by Peter Alderson, who later took over as Speech Research Manager at IBM. The volume was later entitled "A Corpus of Formal British English Speech: The Lancaster/IBM Spoken English Corpus", and was first published by Longman in 1996, later by Routledge in 2013. The book is currently available from online bookstores including Routledge and Book Depository, or in electronic format from Google Play Books.

Other analyses Grammatical tagging of each word, based on the CLAWS1 tagset, was added to the text of the SEC by an automatic process. The fact that this tagging was in machine-readable form made it possible to relate grammatical and prosodic information in the texts. Subsequent work used probabilistic models to develop further the grammatical tagging and to produce automatic parsing techniques. Anne Wichmann published her research on SEC intonation, "Intonation in Text and Discourse: Beginnings, middles, and ends" in 2000.

Machine-Readable Spoken English Corpus (MARSEC) Although the text and its associated tagging existed in machine-readable form, the recordings themselves existed only as tape-recordings. A collaboration, funded by the Economic and Social Research Council in 1992–4, between speech scientists at the Universities of Lancaster and Leeds in the United Kingdom set out to produce a version of the corpus which contained the recordings in digital form, time-linked to the text. The principal researchers were Gerry Knowles and Tamas Varadi (Lancaster) and Peter Roach and Simon Arnfield (Leeds). The outline of the project is set out in Knowles, and the automatic time-alignment is described by Roach and Arnfield. The digitized recordings were recorded on CD-ROM. It was subsequently made available for downloading for research purposes from Leeds University, though this facility is no longer supported.

Aix-MARSEC The work on MARSEC in Lancaster and Leeds finished around 1995, but the corpus has subsequently been the object of a considerable amount of further development at the University of Aix-en-Provence, France, under the direction of Daniel Hirst. The database consists of two major components: the digitalized recordings from MARSEC and the annotations. Annotations have so far been undertaken at nine levels, including phonemes, syllables, words, stress feet, rhythm units and minor and major turn units. Two supplementary levels, the grammatical annotation by CLAWS and a Property Grammar system developed at Aix-en-Provence, are to be integrated soon. A possible disadvantage of this treatment is that the corpus can only be searched using specially written scripts. The database, together with tools, is available under GNU GPL licensing at the Aix-MARSEC project site.

References

Illustrations

Spoken English Corpus illustration

Worked examples

Example 1 — a first encounter with Spoken English Corpus

Start with the simplest possible case. Write down what Spoken English Corpus claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In science, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Spoken English Corpus before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Spoken English Corpus ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Spoken English Corpus

In research
Spoken English Corpus appears in science research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Spoken English Corpus in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Spoken English Corpus is common in secondary-school and first-year university syllabi. It links to neighbouring topics Applied linguistics, Corpora, Dialectology, so understanding it makes those chapters shorter.
In everyday life
Look for Spoken English Corpus outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Spoken English Corpus in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Spoken English Corpus means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Spoken English Corpus out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Spoken English Corpus in simple terms?

The Spoken English Corpus (SEC) is a speech corpus collection of recordings of spoken British English compiled during 1984–1987. The corpus manual can be found on ICAME.

Why does Spoken English Corpus matter?

Because it connects several science ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Spoken English Corpus?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Spoken English Corpus.

Tags

  • Applied linguistics
  • Corpora
  • Dialectology
  • English corpora
  • Linguistic research
  • Phonetics

Keep exploring