ArticleslgStudy

science

Mel-frequency cepstrum

Mel-frequency cepstrum is a science topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Mel-frequency cepstrum rather than just read about it. In short: In sound processing, the mel-frequency cepstrum (MFC) is a representation of the short-term power spectrum of a sound, based on a linear cosine transform of a log power spectrum on a nonlinear mel scale of frequency. Mel-frequency cepstral coefficients (MFCCs) are coefficients that collectively make up an MFC.

Key takeaways

  • Mel-frequency cepstrum belongs to science; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Mel-frequency cepstrum to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Mel-frequency cepstrum from memory before moving on to harder problems.

Reference excerpt

In sound processing, the mel-frequency cepstrum (MFC) is a representation of the short-term power spectrum of a sound, based on a linear cosine transform of a log power spectrum on a nonlinear mel scale of frequency. Mel-frequency cepstral coefficients (MFCCs) are coefficients that collectively make up an MFC. They are derived from a type of cepstral representation of the audio clip (a nonlinear "spectrum-of-a-spectrum"). The difference between the cepstrum and the mel-frequency cepstrum is that in the MFC, the frequency bands are equally spaced on the mel scale, which approximates the human auditory system's response more closely than the linearly-spaced frequency bands used in the normal spectrum. This frequency warping can allow for better representation of sound, for example, in audio compression that might potentially reduce the transmission bandwidth and the storage requirements of audio signals. MFCCs are commonly derived as follows:

Take the Fourier transform of (a windowed excerpt of) a signal. Map the powers of the spectrum obtained above onto the mel scale, using triangular overlapping windows or alternatively, cosine overlapping windows. Take the logs of the powers at each of the mel frequencies. Take the discrete cosine transform of the list of mel log powers, as if it were a signal. The MFCCs are the amplitudes of the resulting spectrum. There can be variations on this process, for example: differences in the shape or spacing of the windows used to map the scale, or addition of dynamics features such as "delta" and "delta-delta" (first- and second-order frame-to-frame difference) coefficients. The European Telecommunications Standards Institute in the early 2000s defined a standardised MFCC algorithm to be used in mobile phones.

Applications MFCCs are commonly used as features in speech recognition systems, such as the systems which can automatically recognize numbers spoken into a telephone. MFCCs are also increasingly finding uses in music information retrieval applications such as genre classification, audio similarity measures, etc.

MFCC for speaker recognition

Since Mel-frequency bands are distributed evenly in MFCC, and they are very similar to the voice system of a human, MFCC can efficiently be used to characterize speakers. For instance, it can be used to recognize the speaker's cell phone model characteristics, and further the details of the speaker's voice. This type of mobile device recognition is possible because the production of electronic components in a phone have tolerances, because different electronic circuit realizations do not have exact same transfer functions. The dissimilarities in the transfer function from one realization to another becomes more prominent if the task performing circuits are from different manufacturers. Hence, each cell phone introduces a convolutional distortion on input speech that leaves its unique impact on the recordings from the cell phone. Therefore, a particular phone can be identified from the recorded speech by multiplying the original frequency spectrum with further multiplications of transfer functions specific to each phone followed by signal processing techniques. Thus, by using MFCC one can characterize cell phone recordings to identify the brand and model of the phone. Considering recording section of a cellphone as Linear time-invariant (LTI) filter: Impulse response- h(n), recorded speech signal y(n) as output of filter in response to input x(n). Hence, y ( n ) = x ( n ) ∗ h ( n ) {\displaystyle y(n)=x(n)*h(n)} (convolution) As speech is not stationary signal, it is divided into overlapped frames within which the signal is assumed to be stationary. So, the p t h {\displaystyle p^{th}} short-term segment (frame) of recorded input speech is:

y p w ( n ) = [ x ( n ) w ( p W − n ) ] ∗ h ( n ) {\displaystyle y_{p}w(n)=[x(n)w(pW-n)]*h(n)} , where w(n): windowed function of length W. Hence, as specified the footprint of mobile phone of the recorded speech is the convolution distortion that helps to identify the recording phone. The embedded identity of the cell phone requires a conversion to a better identifiable form, hence, taking short-time Fourier transform:

Y p w ( f ) = X p w ( f ) H ( f ) {\displaystyle Y_{p}w(f)=X_{p}w(f)H(f)}

H ( f ) {\displaystyle H(f)} can be considered as a concatenated transfer function that produced input speech, and the recorded speech Y p w ( f ) {\displaystyle Y_{p}w(f)} can be perceived as original speech from cell phone. So, equivalent transfer function of vocal tract and cell phone recorder is considered as original source of recorded speech. Therefore,

X p w ( f ) = X e p w ( f ) X v ( f ) , H ′ ( f ) = H ( f ) X v ( f ) , {\displaystyle X_{p}w(f)=Xe_{p}w(f)X_{v}(f),H'(f)=H(f)X_{v}(f),}

… excerpt ends here. Continue reading the full article.

Worked examples

Example 1 — a first encounter with Mel-frequency cepstrum

Start with the simplest possible case. Write down what Mel-frequency cepstrum claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In science, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Mel-frequency cepstrum before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Mel-frequency cepstrum ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Mel-frequency cepstrum

In research
Mel-frequency cepstrum appears in science research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Mel-frequency cepstrum in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Mel-frequency cepstrum is common in secondary-school and first-year university syllabi. It links to neighbouring topics Music information retrieval, Signal processing, so understanding it makes those chapters shorter.
In everyday life
Look for Mel-frequency cepstrum outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Mel-frequency cepstrum in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Mel-frequency cepstrum means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Mel-frequency cepstrum out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Mel-frequency cepstrum in simple terms?

In sound processing, the mel-frequency cepstrum (MFC) is a representation of the short-term power spectrum of a sound, based on a linear cosine transform of a log power spectrum on a nonlinear mel scale of frequency. Mel-frequency cepstral coefficients (MFCCs) are coefficients that collectively mak…

Why does Mel-frequency cepstrum matter?

Because it connects several science ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Mel-frequency cepstrum?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Mel-frequency cepstrum.

Tags

  • Music information retrieval
  • Signal processing

Keep exploring