ArticleslgStudy

science

Moby Project

Moby Project is a science topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Moby Project rather than just read about it. In short: The Moby Project is a collection of public-domain lexical resources created by Grady Ward. The resources were dedicated to the public domain, and are now mirrored at Project Gutenberg.

Key takeaways

  • Moby Project belongs to science; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Moby Project to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Moby Project from memory before moving on to harder problems.

Reference excerpt

The Moby Project is a collection of public-domain lexical resources created by Grady Ward. The resources were dedicated to the public domain, and are now mirrored at Project Gutenberg. As of 2007, it contains the largest free phonetic database, with 177,267 words and corresponding pronunciations.

Hyphenator The Moby Hyphenator II contains hyphenations of 187,175 words and phrases (including 9,752 entries where no hyphenations are given, such as through and avoir). The character encoding appears to be MacRoman, and hyphenation is indicated by a bullet (⟨•⟩, character value 165 decimal, or A5 hexadecimal). Some entries, however, have a combination of actual hyphens and character 165, such as "bar•ber-sur•geon". There is little to no documentation of the hyphenation choices made; the following examples might give some flavour of the style of hyphenation used: at•mos•phere; at•tend•ant; ca•pac•i•ty; un•col•or•a•ble.

Languages Moby Language II contains wordlists of five languages: French, German, Italian, Japanese, and Spanish. Their statistics are:

However, some of the lists are contaminated: for example, the Japanese list contains English words such as abnormal and non-words such as abcdefgh and m,./. There are also unusual peculiarities in the sorting of these lists, as the French list contains a straight alphabetical listing, while the German list contains the alphabetical listing of traditionally capitalized words and then the alphabetical listing of traditionally lower-cased words. The list of Italian words, however, contains no capitalized words whatsoever. The lists do not use accented characters, so "e^tre" is how a user would look up the French word être ("to be").

Part-of-Speech Moby Part-of-Speech contains 233,356 words fully described by part(s) of speech, listed in priority order. The format of the file is word\parts-of-speech, with the following parts of speech being identified:

Pronunciator The Moby Pronunciator II contains 177,267 entries with corresponding pronunciations. Most of the entries describe a single word, but approximately 79,000 contain hyphenated or multiple word phrases, names, or lexemes. The Project Gutenberg distribution also contains a copy of the cmudict v0.3. The file contains lines of the format word[/part-of-speech] pronunciation. Each line is ended with the ASCII carriage return character (CR, '\r', 0x0D, 13 in decimal). The word field can include apostrophes (e.g. isn't), hyphens (e.g. able-bodied), and multiple words separated by underscores (e.g. monkey_wrench). Non-English words are generally rendered, as stated in the documentation, without accents or other diacritical marks. However, in 36 entries (e.g. São_Miguel), some non-ASCII accented characters remain, represented using Mac OS Roman encoding. The part-of-speech field is used to disambiguate 770 of the words which have differing pronunciations depending on their part-of-speech. For example, for the words spelled close, the verb has the pronunciation , whereas the adjective is . The parts-of-speech have been assigned the following codes:

Following this is the pronunciation. Several special symbols are present:

The rest of the symbols are used to represent IPA characters. The pronunciations are generally consistent with a General American dialect of English, that exhibits father-bother merger, hurry-furry merger and lot-cloth split, but does not exhibit cot-caught merger or wine-whine merger. Each phoneme is represented by a sequence of one or more characters. Some of the sequences are delimited with a slash character "/", as shown in the following table, but note that the sequence for is delimited by two slash characters at either end:

To this collection are added a number of extra sequences representing phonemes found in several other languages. These are used to encode the non-English words, phrases and names that are included in the database. The following table contains these extra phonemes, but note that the extent to which some of these may exist due to encoding errors is not clear.

Shakespeare Moby Shakespeare contains the complete unabridged works of Shakespeare. This specific resource is not available from Project Gutenberg, but it is available in a 1993 version on the web.

Thesaurus The Moby Thesaurus II contains 30,260 root words, with 2,520,264 synonyms and related terms – an average of 83.3 per root word. Each line consists of a list of comma-separated values, with the first term being the root word, and all following words being related terms. Grady Ward placed this thesaurus in the public domain in 1996. It is also available as a Debian package although the package has been discontinued starting with Bullseye.

Words Moby Words II is the largest wordlist in the world. The distribution consists of the following 16 files:

References

External links Former Moby Project site (icon.shef.ac.uk/Moby/) – No longer accessible. View a copy made by the Wayback Machine, as it was on 30 September 2017. ("Last modified: October 24, 2000") working download site. Project Gutenberg downloads Searching for Rhymes with Perl; corresponding code Wiktionary:Appendix:Moby Thesaurus II http://digital.library.upenn.edu/webbin/gutbook/lookup?num=3201

Worked examples

Example 1 — a first encounter with Moby Project

Start with the simplest possible case. Write down what Moby Project claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In science, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Moby Project before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Moby Project ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Moby Project

In research
Moby Project appears in science research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Moby Project in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Moby Project is common in secondary-school and first-year university syllabi. It links to neighbouring topics Corpora, Linguistic research, Public domain databases, so understanding it makes those chapters shorter.
In everyday life
Look for Moby Project outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.
Ask Teacher Smith questions about this articleOpens your AI tutor with a question about “Moby Project” →

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Moby Project in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Moby Project means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Moby Project out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Moby Project in simple terms?

The Moby Project is a collection of public-domain lexical resources created by Grady Ward. The resources were dedicated to the public domain, and are now mirrored at Project Gutenberg.

Why does Moby Project matter?

Because it connects several science ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Moby Project?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Moby Project.

Tags

  • Corpora
  • Linguistic research
  • Public domain databases

Keep exploring