Statistics
Overview
The examiners focus on three central measures of central tendency—mean, median, and mode—along with their calculation from raw data and frequency distributions.
Mastering this topic requires understanding when to apply each measure and how to compute them quickly from grouped and ungrouped data. Questions often combine basic calculations with data interpretation, so conceptual clarity directly translates to faster solving and higher accuracy.
The key to scoring well is memorizing the direct formulas, practicing frequency table calculations, and recognizing which measure suits which data type. This topic also builds foundation for Data Interpretation questions.
Key Concepts
- Mean (Arithmetic Average) is the sum of all observations divided by the number of observations. It uses every data point and is sensitive to extreme values (outliers).
- Median is the middle value when data is arranged in ascending or descending order. It is the best measure when data has outliers or is skewed.
- Mode is the value that occurs most frequently. A dataset can have no mode, one mode (unimodal), or multiple modes (bimodal/multimodal).
- Frequency Distribution organizes data into classes with their frequencies, making calculation of mean/median/mode systematic for large datasets.
- Class Mark (Mid-value) = (Lower limit + Upper limit) / 2. This represents each class in grouped data calculations.
- Cumulative Frequency is the running total of frequencies, essential for finding median in grouped data.
- Empirical Relationship: Mode = 3 × Median − 2 × Mean. Use this shortcut when two measures are known.
- Range = Maximum value − Minimum value. A basic measure of dispersion often asked alongside central tendency.
Formulas / Key Facts
For Ungrouped Data (Raw Data):
Mean = Sum of all observations / Number of observations = Σx / n
Median:
- Arrange data in order
- If n is odd: Median = ((n+1)/2)th term
- If n is even: Median = Average of (n/2)th and (n/2 + 1)th terms
Mode = Most frequently occurring value
For Grouped Data (Frequency Distribution):
Mean (Direct Method) = Σfx / Σf where f = frequency, x = class mark
Mean (Assumed Mean Method) = A + (Σfd / Σf) where A = assumed mean, d = x − A
Mean (Step Deviation Method) = A + (Σfu / Σf) × h where u = (x − A)/h, h = class width
Median = L + ((n/2 − cf) / f) × h where:
- L = lower limit of median class
- n = total frequency (Σf)
- cf = cumulative frequency before median class
- f = frequency of median class
- h = class width
Mode = L + ((f₁ − f₀) / (2f₁ − f₀ − f₂)) × h where:
- L = lower limit of modal class
- f₁ = frequency of modal class
- f₀ = frequency of class before modal class
- f₂ = frequency of class after modal class
- h = class width
Modal Class = Class with highest frequency
Median Class = Class where (n/2)th observation lies (use cumulative frequency)
Worked Examples
Example 1: Mean from Ungrouped Data
Find the mean of: 12, 15, 18, 22, 23
Solution:
- Sum = 12 + 15 + 18 + 22 + 23 = 90
- n = 5
- Mean = 90/5 = 18
Example 2: Median from Ungrouped Data
Find the median of: 7, 3, 9, 5, 11, 8, 2
Solution:
- Arrange in order: 2, 3, 5, 7, 8, 9, 11
- n = 7 (odd)
- Median position = (7+1)/2 = 4th term
- Median = 7
Example 3: Mean from Frequency Distribution
| Class Interval | Frequency (f) |
|---|---|
| 0-10 | 5 |
| 10-20 | 8 |
| 20-30 | 12 |
| 30-40 | 10 |
| 40-50 | 5 |
Solution using Direct Method:
| Class | f | Class Mark (x) | fx |
|---|---|---|---|
| 0-10 | 5 | 5 | 25 |
| 10-20 | 8 | 15 | 120 |
| 20-30 | 12 | 25 | 300 |
| 30-40 | 10 | 35 | 350 |
| 40-50 | 5 | 45 | 225 |
| Total | 40 | 1020 |
Mean = 1020/40 = 25.5
Example 4: Median from Grouped Data
Using the same table, find median.
Solution:
- Σf = 40, so n/2 = 20
- Cumulative frequencies: 5, 13, 25, 35, 40
- 20 lies in class 20-30 (cf just before = 13)
- Median class: 20-30
- L = 20, cf = 13, f = 12, h = 10
Median = 20 + ((20 − 13)/12) × 10 = 20 + (7/12) × 10 = 20 + 5.83 = 25.83
Example 5: Mode from Grouped Data
Using the same table, find mode.
Solution:
- Modal class = 20-30 (highest frequency = 12)
- L = 20, f₁ = 12, f₀ = 8, f₂ = 10, h = 10
Mode = 20 + ((12 − 8)/(2×12 − 8 − 10)) × 10 = 20 + (4/6) × 10 = 20 + 6.67 = 26.67
Common Mistakes
- Forgetting to arrange data before finding median → Always sort data in ascending order first. Random order gives wrong middle value.
- Confusing cumulative frequency with frequency → cf is running total used to locate median class; f is the frequency of that specific class used in formula.
- Using wrong class limits in median/mode formula → Always use the lower limit (L) of the median/modal class, not upper limit or class mark.
- Applying mean formula when data has extreme outliers → Mean gets distorted by outliers. In such cases, median is the better representative measure.
- Taking wrong f₀ and f₂ in mode formula → f₀ is frequency of class immediately BEFORE modal class; f₂ is immediately AFTER. The order matters.
- Calculating class mark incorrectly for exclusive/inclusive data → For exclusive (0-10, 10-20), class mark = (0+10)/2 = 5. For inclusive (1-10, 11-20), treat as 0.5-10.5, 10.5-20.5 first.
Quick Reference
- Mean = Σfx / Σf (uses all values; affected by outliers)
- Median position: (n+1)/2 for odd; average of middle two for even
- Mode = highest frequency value; use formula for grouped data
- Median class: where cumulative frequency first crosses n/2
- Empirical relation: Mode = 3 × Median − 2 × Mean
- Class Mark = (Lower limit + Upper limit) / 2