ArticleslgStudy

science

Unicode and email

Unicode and email is a science topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Unicode and email rather than just read about it. In short: Many email clients now offer some support for Unicode. Some clients will automatically choose between a legacy encoding and Unicode depending on the mail's content, either automatically or when the user requests it.

Key takeaways

  • Unicode and email belongs to science; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Unicode and email to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Unicode and email from memory before moving on to harder problems.

Reference excerpt

Many email clients now offer some support for Unicode. Some clients will automatically choose between a legacy encoding and Unicode depending on the mail's content, either automatically or when the user requests it. Technical requirements for sending messages containing non-ASCII characters by email include:

encoding of certain header fields (subject, sender's and recipient's names, sender's organization and reply-to name) and, optionally, body in a content-transfer encoding encoding of non-ASCII characters in one of the Unicode transforms negotiating the use of UTF-8 encoding in email addresses and reply codes (SMTPUTF8) sending the information about the content-transfer encoding and the Unicode transform used so that the message can be correctly displayed by the recipient (see Mojibake) If the sender's or recipient's email address contains non-ASCII characters, sending a message also requires encoding these characters in a format that can be understood by mail servers.

Support in protocols RFC 6531 provides a mechanism for allowing non-ASCII email addresses encoded as UTF-8 in the SMTP and LMTP protocols.

Message headers To use Unicode in certain email header fields, e.g. subject lines, sender and recipient names, the Unicode text has to be encoded using a MIME "Encoded-Word" with a Unicode encoding as the charset. To use Unicode in the domain part of email addresses, IDNA encoding must traditionally be used. Alternatively, SMTPUTF8 allows the use of UTF-8 encoding in email addresses (both in a local part and in domain name) as well as in a mail header section. Various standards had been created to retrofit the handling of non-ASCII data to the originally ASCII-only email protocol:

RFC 2047 provides support for encoding non-ASCII values such as real names and subject lines in email headers. RFC 5890 provides support for encoding non-ASCII domain names in the Domain Name System. RFC 6532 allows the use of UTF-8 in a mail header section.

Message bodies As with all encodings apart from US-ASCII, when using Unicode text in email, MIME must be used to specify that a Unicode transformation format is being used for the text. UTF-7, an obsolete encoding, had an advantage on obsolete non–8-bit clean networks over Unicode encodings in that it does not require a transfer encoding to fit within the 7-bit limits of legacy Internet mail servers. On the other hand, UTF-16 must be transfer encoded to fit the data format of SMTP. Although not strictly required, UTF-8 is usually also transfer encoded to avoid problems across 7-bit mail servers. MIME transfer encoding of UTF-8 makes it either unreadable as plain text (in the case of base64) or, for some languages and types of text, heavily size-inefficient (in the case of quoted-printable). Some document formats, such as HTML, PostScript and Rich Text Format, have their own 7-bit encoding schemes for non-ASCII characters and can thus be sent without using any special email encodings. HTML email can use HTML entities to use characters from anywhere in Unicode even if the HTML source text for the email is in a legacy encoding (e.g. 7-bit ASCII).

See also Comparison of email clients Email address internationalization International email Unicode and HTML

References

External links SIL's freeware fonts, editors and documentation

Worked examples

Example 1 — a first encounter with Unicode and email

Start with the simplest possible case. Write down what Unicode and email claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In science, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Unicode and email before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Unicode and email ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Unicode and email

In research
Unicode and email appears in science research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Unicode and email in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Unicode and email is common in secondary-school and first-year university syllabi. It links to neighbouring topics Email, Email clients, Unicode, so understanding it makes those chapters shorter.
In everyday life
Look for Unicode and email outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.
Ask Teacher Smith questions about this articleOpens your AI tutor with a question about “Unicode and email” →

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Unicode and email in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Unicode and email means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Unicode and email out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Unicode and email in simple terms?

Many email clients now offer some support for Unicode. Some clients will automatically choose between a legacy encoding and Unicode depending on the mail's content, either automatically or when the user requests it.

Why does Unicode and email matter?

Because it connects several science ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Unicode and email?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Unicode and email.

Tags

  • Email
  • Email clients
  • Unicode

Keep exploring