Vocaloid is a singing voice synthesizer software product. Its signal processing part was developed through a joint research project between Yamaha Corporation and the Music Technology Group at Pompeu Fabra University, Barcelona. The software was ultimately developed into the commercial product "Vocaloid" that was released in 2004. The software enables users to synthesize "singing" by typing in lyrics and melody and also "speech" by typing in the script of the required words. It uses synthesizing technology with specially recorded vocals of voice actors or singers. To create a song, the user must input the melody and lyrics. A piano roll type interface is used to input the melody and the lyrics can be entered on each note. The software can change the stress of the pronunciations, add effects such as vibrato, or change the dynamics and tone of the voice. Various voice banks have been released for use with the Vocaloid synthesizer technology. Each is sold as "a singer in a box" designed to act as a replacement for an actual singer. As such, they are often released under a moe anthropomorph avatar, however, there are also voice banks released without an assigned avatar. These avatars are also referred to as Vocaloids, and are often marketed as virtual idols; some have gone on to perform at live concerts as an on-stage projection. The software was originally only available in English starting with the first Vocaloids Leon, Lola and Miriam by Zero-G, and Japanese with Meiko and Kaito made by Yamaha and sold by Crypton Future Media. Vocaloid 3 has added support for Spanish for the Vocaloids Bruno, Clara and Maika; Chinese for Luo Tianyi, Yuezheng Ling, Xin Hua, and Yanhe; and Korean for SeeU. The software is intended for professional musicians as well as casual computer music users. Japanese musical groups such as Livetune of Toy's Factory and Supercell of Sony Music Entertainment Japan have released their songs featuring Vocaloid as vocals. Japanese record label Exit Tunes of Quake Inc. also have released compilation albums featuring Vocaloids.
Technology
Vocaloid's singing synthesis technology is generally categorized into the concatenative synthesis in the frequency domain, which splices and processes the vocal fragments extracted from human singing voices, in the forms of time-frequency representation. The Vocaloid system can produce the realistic voices by adding vocal expressions like the vibrato on the score information. Initially, Vocaloid's synthesis technology was called "frequency-domain singing articulation splicing and shaping" (周波数ドメイン歌唱アーティキュレーション接続法, shūhasū-domein kashō ātikyurēshon setsuzoku-hō) on the release of Vocaloid in 2004, although this name is no longer used since the release of Vocaloid 2 in 2007. "Singing articulation" is explained as "vocal expressions" such as vibrato and vocal fragments necessary for singing. The Vocaloid and Vocaloid 2 synthesis engines are designed for singing, not reading text aloud, though software such as Vocaloid-flex and Voiceroid have been developed for that. They cannot naturally replicate singing expressions like hoarse voices or shouts.
System architecture
The main parts of the Vocaloid 2 system are the score editor (Vocaloid 2 editor), the singer library, and the synthesis engine. The synthesis engine receives score information from the score editor, selects appropriate samples from the singer library, and concatenates them to output synthesized voices. There is basically no difference in the score editor and the synthesis engine provided by Yamaha among different Vocaloid 2 products. If a Vocaloid 2 product is already installed, the user can enable another Vocaloid 2 product by adding its library. The system supports three languages, Japanese, Korean, and English, although other languages may be optional in the future. It works standalone (playback and export to WAV) and as a ReWire application or a Virtual Studio Technology instrument (VSTi) accessible from a digital audio workstation (DAW).
Score Editor
The score editor is a piano roll-style editor to input notes, lyrics, and some expressions. When entering lyrics, the editor automatically converts them into Vocaloid phonetic symbols using the built-in pronunciation dictionary. The user can directly edit the phonetic symbols of unregistered words. The score editor offers various parameters to add expressions to singing voices. The user is supposed to optimize these parameters that best fit the synthesized tune when creating voices. This editor supports ReWire and can be synchronized with DAW. Real-time "playback" of songs with predefined lyrics using a MIDI keyboard is also supported.
Singer library Each Vocaloid license develops the singer library, or a database of vocal fragments sampled from real people. The database must have all possible combinations of phonemes of the target language, including diphones (a chain of two different phonemes) and sustained vowels, as well as polyphones with more than two phonemes if necessary. For example, the voice corresponding to the word "sing" ([sIN]) can be synthesized by concatenating the sequence of diphones "#-s, s-I, I-N, N-#" (# indicating a voiceless phoneme) with the sustained vowel ī. The Vocaloid system changes the pitch of these fragments so that it fits the melody. In order to get more natural sounds, three or four different pitch ranges are required to be stored into the library. Japanese requires 500 diphones per pitch, whereas English requires 2,500. Japanese has fewer diphones because it has fewer phonemes and most syllabic sounds are open syllables ending in a vowel. In Japanese, there are three patterns of diphones containing a consonant: voiceless-consonant, vowel-consonant, and consonant-vowel. On the other hand, English has many closed syllables ending in a consonant, and consonant-consonant and consonant-voiceless diphones as well. Thus, more diphones need to be recorded into an English library than into a Japanese one. Due to this linguistic difference, a Japanese library is not suitable for singing in eloquent English.
Synthesis engine
… excerpt ends here. Continue reading the full article.

![Vocaloid: Voice model developed before the Vocaloid, excitation plus resonances (EpR) model,[11] is a combination of: Spectral modeling synthesis (SMS)Source–filter model The model was developed in 2001 as a source–filter model for voice synthesis,[12] but was only implemented on top of the concatenative synthesis model in the final product[citation needed] as a method of avoiding spectral shape discontinuities at the segment boundaries of concatenation.[13](based on Fig.1 on Bonada et al. 2001)](https://upload.wikimedia.org/wikipedia/commons/thumb/2/22/Excitation_plus_Resonances_%28EpR%29_voice_model_%28Bonada_et_al._2001%2C_Fig.1%29.svg/500px-Excitation_plus_Resonances_%28EpR%29_voice_model_%28Bonada_et_al._2001%2C_Fig.1%29.svg.png?utm_source=en.wikipedia.org&utm_campaign=parser&utm_content=thumbnail)


![Vocaloid: Vocaloid synthesis engine[23]](https://upload.wikimedia.org/wikipedia/commons/thumb/6/65/Vocaloid_Synthesis_Engine_-_en.jpg/1280px-Vocaloid_Synthesis_Engine_-_en.jpg?utm_source=en.wikipedia.org&utm_campaign=parser&utm_content=thumbnail)

