ArticleslgStudy

computer science

Voice computing

Voice computing is a computer science topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Voice computing rather than just read about it. In short: Voice computing is the discipline that develops hardware or software to process voice inputs. It spans many other fields including human-computer interaction, conversational computing, linguistics, natural language processing, automatic speech recognition, speech synthesis, audio engineering, digital signal processing, cloud computing, data science, ethics, law, and information security.

Voice computing — main illustration
Voice computing — illustration

Key takeaways

  • Voice computing belongs to computer science; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Voice computing to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Voice computing from memory before moving on to harder problems.

Reference excerpt

Voice computing is the discipline that develops hardware or software to process voice inputs. It spans many other fields including human-computer interaction, conversational computing, linguistics, natural language processing, automatic speech recognition, speech synthesis, audio engineering, digital signal processing, cloud computing, data science, ethics, law, and information security. Voice computing has become increasingly significant in modern times, especially with the advent of smart speakers like the Amazon Echo and Google Assistant, a shift towards serverless computing, and improved accuracy of speech recognition and text-to-speech models.

History Voice computing has a rich history. First, scientists like Wolfgang Kempelen started to build speech machines to produce the earliest synthetic speech sounds. This led to further work by Thomas Edison to record audio with dictation machines and play it back in corporate settings. In the 1950s-1960s there were primitive attempts to build automated speech recognition systems by Bell Labs, IBM, and others. However, it was not until the 1980s that Hidden Markov Models were used to recognize up to 1,000 words that speech recognition systems became relevant.

Around 2011, Siri emerged on Apple iPhones as the first voice assistant accessible to consumers. This innovation led to a dramatic shift to building voice-first computing architectures. PS4 was released by Sony in North America in 2013 (70+ million devices), Amazon released the Amazon Echo in 2014 (30+ million devices), Microsoft released Cortana (2015 - 400 million Windows 10 users), Google released Google Assistant (2016 - 2 billion active monthly users on Android phones), and Apple released HomePod (2018 - 500,000 devices sold and 1 billion devices active with iOS/Siri). These shifts, along with advancements in cloud infrastructure (e.g. Amazon Web Services) and codecs, have solidified the voice computing field and made it widely relevant to the public at large.

Hardware A voice computer is assembled hardware and software to process voice inputs. Note that voice computers do not necessarily need a screen, such as in the traditional Amazon Echo. In other embodiments, traditional laptop computers or mobile phones could be used as voice computers. Moreover, there has become increasingly more interfaces for voice computers with the advent of IoT-enabled devices, such as within cars or televisions. As of September 2018, there are currently over 20,000 types of devices compatible with Amazon Alexa.

Software Voice computing software can read/write, record, clean, encrypt/decrypt, playback, transcode, transcribe, compress, publish, featurize, model, and visualize voice files. Here are some popular software packages related to voice computing:

Applications Voice computing applications span many industries including voice assistants, healthcare, e-Commerce, finance, supply chain, agriculture, text-to-speech, security, marketing, customer support, recruiting, cloud computing, microphones, speakers, and podcasting. Voice technology is projected to grow at a CAGR of 19-25% by 2025, making it an attractive industry for startups and investors alike.

Legal considerations In the United States, the states have varying telephone call recording laws. In some states, it is legal to record a conversation with the consent of only one party, in others the consent of all parties is required. Moreover, COPPA is a significant law to protect minors using the Internet. With an increasing number of minors interacting with voice computing devices (e.g. the Amazon Alexa), on October 23, 2017 the Federal Trade Commission relaxed the COPAA rule so that children can issue voice searches and commands. Lastly, GDPR is a new European law that governs the right to be forgotten and many other clauses for EU citizens. GDPR also is clear that companies need to outline clear measures to obtain consent if audio recordings are made and define the purpose and scope as to how these recordings will be used, e.g., for training purposes. The bar for valid consent has been raised under the GDPR. Consents must be freely given, specific, informed, and unambiguous; tacit consent is no longer sufficient.

Research conferences There are many research conferences that relate to voice computing. Some of these include:

International Conference on Acoustics, Speech, and Signal Processing Interspeech AVEC IEEE Int'l Conf. on Automatic Face and Gesture Recognition ACII2019 The 8th Int'l Conf. on Affective Computing and Intelligent Interaction

Developer community Google Assistant has roughly 2,000 actions as of January 2018. There are over 50,000 Alexa skills worldwide as of September 2018. In June 2017, Google released AudioSet, a large-scale collection of human-labeled 10-second sound clips drawn from YouTube videos. It contains 1,010,480 videos of human speech files, or 2,793.5 hours in total. It was released as part of the IEEE ICASSP 2017 Conference. In November 2017, Mozilla Foundation released the Common Voice Project, a collection of speech files to help contribute to the larger open source machine learning community. The voicebank is currently 12GB in size, with more than 500 hours of English-language voice data that have been collected from 112 countries since the project's inception in June 2017. This dataset has already resulted in creative projects like the DeepSpeech model, an open source transcription model.

See also Speech recognition Natural language processing Voice user interface Audio codec Ubiquitous computing Hands-free computing

References

Illustrations

Voice computing: The Amazon Echo, an example of a voice computer
The Amazon Echo, an example of a voice computer

Worked examples

Example 1 — a first encounter with Voice computing

Start with the simplest possible case. Write down what Voice computing claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In computer science, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Voice computing before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Voice computing ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Voice computing

In research
Voice computing appears in computer science research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Voice computing in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Voice computing is common in secondary-school and first-year university syllabi. It links to neighbouring topics Computational fields of study, Computational linguistics, History of human–computer interaction, so understanding it makes those chapters shorter.
In everyday life
Look for Voice computing outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.
Ask Teacher Smith questions about this articleOpens your AI tutor with a question about “Voice computing” →

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Voice computing in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Voice computing means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Voice computing out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Voice computing in simple terms?

Voice computing is the discipline that develops hardware or software to process voice inputs. It spans many other fields including human-computer interaction, conversational computing, linguistics, natural language processing, automatic speech recognition, speech synthesis, audio engineering, digit…

Why does Voice computing matter?

Because it connects several computer science ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Voice computing?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Voice computing.

Tags

  • Computational fields of study
  • Computational linguistics
  • History of human–computer interaction
  • Natural language processing
  • Speech recognition
  • Voice technology

Keep exploring