Character Encoding

COS2621 - Computer Organisation · Data Representation

Character Encoding

Character encoding is a system that pairs each character from a set of characters with a specific number. This is essential for computers to process and display text. Different encoding systems have been developed to accommodate various languages and symbols.

Understanding Character Sets

A character set is a collection of characters that a computer can recognize. Common character sets include:

  • ASCII (American Standard Code for Information Interchange): This is a 7-bit character encoding that can represent 128 characters, including English letters, digits, and control characters.
  • Extended ASCII: This extends the ASCII set to 256 characters by adding additional symbols and characters from other languages.
  • Unicode: This is a universal character encoding standard that supports characters from almost all written languages. It uses various encoding forms, including UTF-8, UTF-16, and UTF-32.

ASCII Encoding

ASCII uses 7 bits to represent each character. The first 32 codes (0 to 31) are reserved for control characters (like newline and tab), while the remaining codes represent printable characters.

For example, the letter 'A' is represented by the decimal number 65 in ASCII. In binary, this is 1000001. Here is a small portion of the ASCII table:

CharacterDecimalBinary
A651000001
B661000010
C671000011

Extended ASCII

Extended ASCII uses an additional bit, allowing for 256 characters. The first 128 characters remain the same as standard ASCII, while the additional 128 characters include symbols, accented characters, and graphical characters.

For example, the character 'é' is represented by the decimal number 233 in Extended ASCII. In binary, this is 11101001.

Remember: Extended ASCII is not standardized, meaning different systems may use different characters for the additional 128 codes.

Unicode and Its Importance

Unicode was developed to address the limitations of ASCII and Extended ASCII. It can represent over 143,000 characters from various languages and scripts. Unicode is essential for global communication, as it allows text to be displayed in different languages without confusion.

Unicode can be encoded in several ways, including:

  • UTF-8: A variable-length encoding that uses 1 to 4 bytes for each character. It is backward compatible with ASCII, meaning the first 128 characters are the same as ASCII.
  • UTF-16: Uses 2 bytes for most characters and 4 bytes for additional characters. It is commonly used in modern applications.
  • UTF-32: Uses 4 bytes for every character, making it simple but less space-efficient.

UTF-8 Encoding Example

Let us look at how UTF-8 encodes characters. The character 'A' is represented in UTF-8 as a single byte: 01000001. The character '€' (Euro sign) is represented as three bytes: 11100010 10000010 10101100.

Character: A   UTF-8: 01000001   Decimal: 65
Character: €   UTF-8: 11100010 10000010 10101100   Decimal: 226 128 154

Common Mistakes in Character Encoding

Watch out: Students often confuse ASCII with Extended ASCII. Remember that while ASCII has 128 characters, Extended ASCII has 256 characters and includes additional symbols.

Applications of Character Encoding

Character encoding is vital in various applications:

  • Web Development: HTML and CSS files use character encoding to display text correctly in browsers.
  • Databases: Character encoding ensures that text data is stored and retrieved accurately.
  • Text Processing: Applications like word processors rely on character encoding for text formatting and display.

Encoding and Decoding

Encoding is the process of converting characters into their corresponding numerical values. Decoding is the reverse process, where numerical values are converted back into characters.

For example, if you have the character 'B', the encoding process converts it to the decimal value 66. The decoding process takes 66 and converts it back to 'B'.

Tip: Always ensure you know which character encoding is being used in a document or application to avoid errors.

Conclusion on Character Encoding

Character encoding is fundamental for text representation in computing. Understanding the differences between ASCII, Extended ASCII, and Unicode is crucial for working with text in various programming and application contexts.

Summary

  • Character encoding pairs characters with numerical values.
  • ASCII represents 128 characters, while Extended ASCII represents 256 characters.
  • Unicode supports a vast range of characters from different languages.
  • UTF-8 is a popular encoding format that is backward compatible with ASCII.

Check your understanding

  1. What is the main difference between ASCII and Extended ASCII?
  2. How does UTF-8 handle characters that are not in the ASCII range?
  3. Why is Unicode important in modern computing?
  4. What are the three encoding forms of Unicode mentioned in this topic?
    Character Encoding – COS2621 - Computer Organisation notes | Tyro Study