Floating Point Representation
COS2621 - Computer Organisation · Data Representation
Floating Point Representation
Floating point representation is a method used to represent real numbers in a way that can accommodate a wide range of values. This method is essential in computer systems for handling decimal numbers, especially when precision is necessary. In this topic, you will learn how floating point numbers are structured, how to convert between different representations, and how to perform arithmetic operations with them.
Structure of Floating Point Numbers
A floating point number is typically represented in a computer using three main components: the sign bit, the exponent, and the mantissa (or significand). The general form of a floating point number can be expressed as:
value = sign × mantissa × baseexponent
In most systems, the base is 2 (binary). The sign bit indicates whether the number is positive or negative. The mantissa represents the significant digits of the number, and the exponent indicates the scale of the number.
IEEE 754 Standard
The IEEE 754 standard is the most widely used standard for floating point representation. It defines several formats, but the most common are:
- Single precision (32 bits)
- Double precision (64 bits)
In single precision, the bits are allocated as follows:
- 1 bit for the sign
- 8 bits for the exponent
- 23 bits for the mantissa
In double precision, the allocation is:
- 1 bit for the sign
- 11 bits for the exponent
- 52 bits for the mantissa
Example of Single Precision Representation
Consider the number -13.25. To represent this in single precision format, follow these steps:
Step 1: Convert to Binary
First, convert the integer and fractional parts to binary:
- 13 in binary is 1101.
- 0.25 in binary is 0.01 (since 0.25 × 2 = 0.5, and 0.5 × 2 = 1.0).
Thus, -13.25 in binary is -1101.01.
Step 2: Normalize the Binary Number
Next, normalize the binary number. Normalization means adjusting the binary number so that it is in the form of 1.xxxx × 2n. For -1101.01, we shift the binary point three places to the left:
-1.10101 × 23
Step 3: Determine the Sign, Exponent, and Mantissa
The sign bit is 1 (since the number is negative). The exponent is 3. To store the exponent in IEEE 754 format, we use a bias. For single precision, the bias is 127. Thus, the stored exponent is:
exponent + bias = 3 + 127 = 130
In binary, 130 is 10000010.
The mantissa is the part after the binary point in the normalized form. Thus, the mantissa is 10101000000000000000000 (23 bits).
Step 4: Combine the Parts
Now combine the sign bit, exponent, and mantissa:
Sign: 1Exponent: 10000010Mantissa: 10101000000000000000000The final representation of -13.25 in single precision is:
1 10000010 10101000000000000000000Example of Double Precision Representation
Now consider the same number, -13.25, but represent it in double precision format. The steps are similar but with different bit allocations.
Step 1: Convert to Binary
The binary representation remains the same: -1101.01.
Step 2: Normalize the Binary Number
The normalized form is also the same: -1.10101 × 23.
Step 3: Determine the Sign, Exponent, and Mantissa
The sign bit is still 1. For double precision, the bias is 1023. Thus, the stored exponent is:
exponent + bias = 3 + 1023 = 1026
In binary, 1026 is 10000000010.
The mantissa is still 1010100000000000000000000000000000000000000000000000 (52 bits).
Step 4: Combine the Parts
The final representation of -13.25 in double precision is:
Sign: 1Exponent: 10000000010Mantissa: 1010100000000000000000000000000000000000000000000000Thus, the representation is:
1 10000000010 1010100000000000000000000000000000000000000000000000Converting Floating Point to Decimal
To convert a floating point number back to decimal, you need to reverse the steps used in the representation. For example, consider the single precision number:
1 10000010 10101000000000000000000Step 1: Extract the Sign, Exponent, and Mantissa
Sign = 1, Exponent = 10000010, Mantissa = 10101000000000000000000.
Step 2: Calculate the Exponent
Convert the exponent from binary to decimal:
10000010 = 130.
Now subtract the bias (127):
130 - 127 = 3.
Step 3: Calculate the Mantissa
The mantissa in decimal is calculated as:
1 + 1 × 2-1 + 0 × 2-2 + 1 × 2-3 + 0 × 2-4 + 1 × 2-5 = 1 + 0.5 + 0 + 0.125 + 0 + 0.03125 = 1.625.
Step 4: Calculate the Final Value
Combine the parts to get the final decimal value:
value = -1 × 1.625 × 23 = -13.25.
Arithmetic Operations with Floating Point Numbers
Arithmetic operations with floating point numbers can be complex due to the need for alignment and rounding. The following operations are commonly performed:
- Addition
- Subtraction
- Multiplication
- Division
Example: Addition of Two Floating Point Numbers
Consider adding two single precision floating point numbers: 1.5 and 2.25.
Step 1: Convert to Binary
1.5 in binary is 1.1, and 2.25 in binary is 10.01.
Step 2: Normalize
1.5: 1.1 = 1.1 × 20
2.25: 10.01 = 1.001 × 21
Step 3: Align the Numbers
Align the numbers by adjusting the exponent. We will shift 1.1 to match the exponent of 2.25:
1.1 × 20 becomes 0.11 × 21.
Step 4: Add the Mantissas
Add the mantissas:
0.11 + 1.001 = 10.111.
Step 5: Normalize the Result
Normalize 10.111:
10.111 = 1.0111 × 22.
Step 6: Determine the Resulting Exponent and Mantissa
Exponent = 2, mantissa = 0111. The final result in binary is:
0 10000010 01110000000000000000000Common Mistakes
Watch out: When converting between binary and decimal, ensure you account for both the integer and fractional parts correctly. Missing a part can lead to incorrect results.
Remember: Always check the bias value for the exponent based on whether you are using single or double precision.
Summary
- A floating point number consists of a sign bit, exponent, and mantissa.
- The IEEE 754 standard defines formats for single and double precision floating point numbers.
- Conversion between floating point and decimal requires careful extraction of the parts and proper calculations.
- Arithmetic operations require normalization and alignment of the numbers.
Check your understanding
- What are the three components of a floating point number?
- How is the exponent stored in the IEEE 754 format?
- Describe the steps to convert a floating point number to decimal.
- What is the first step in adding two floating point numbers together?