IGCSE MATH: STATISTICS
Histograms with Unequal Class Widths
| Time (t mins) | Frequency | Class Width | Frequency Density (Frequency ÷ Width) |
|---|---|---|---|
| 0 ≤ t < 10 | 20 | 10 | 2.0 |
| 10 ≤ t < 20 | 30 | 10 | 3.0 |
| 20 ≤ t < 40 | 50 | 20 | 2.5 |
| 40 ≤ t < 60 | 30 | 20 | 1.5 |
Important: For unequal class widths, plot frequency density on y-axis, not frequency.
3. Measures of Average for Individual and Discrete Data
Mean (x̄)
Example: 12, 15, 18, 20, 25
Sum = 12+15+18+20+25 = 90
n = 5
Mean = 90 ÷ 5 = 18
✓ Uses all values
✗ Affected by outliers
✓ Can be decimal
Median
Odd n: Position = (n+1)/2
Even n: Average of two middle values
Example (even): 4, 6, 8, 10, 12, 14
n=6, positions 3 & 4: 8 & 10
Median = (8+10)÷2 = 9
✓ Not affected by outliers
✓ Always from dataset
Mode
Example: 2, 3, 3, 3, 5, 7, 7, 9
Mode = 3 (appears 3 times)
✓ Easy to find
✓ Works with categorical data
✓ Not affected by outliers
Range (Measure of Spread)
Example: 12, 15, 18, 20, 25
Range = 25 - 12 = 13
Note: Range is NOT an average - it measures spread.
Which Average to Use?
| Situation | Best Average | Reason |
|---|---|---|
| Normal distribution, no outliers | Mean | Uses all data effectively |
| Data has extreme values | Median | Not affected by outliers |
| Categorical data | Mode | Only average that works |
| Finding most popular item | Mode | Shows highest frequency |
4. Grouped Continuous Data
Frequency Tables with Midpoints
| Height (h cm) | Frequency (f) | Midpoint (x) | f × x |
|---|---|---|---|
| 140 ≤ h < 150 | 8 | 145 | 1160 |
| 150 ≤ h < 160 | 15 | 155 | 2325 |
| 160 ≤ h < 170 | 12 | 165 | 1980 |
| 170 ≤ h < 180 | 5 | 175 | 875 |
| Total | 40 | - | 6340 |
Estimated Mean for Grouped Data
Where: f = frequency, x = midpoint, Σf = total frequency
From table: Mean = 6340 ÷ 40 = 158.5 cm
Note: This is an estimate because we use midpoints.
Modal Class
Class interval with the highest frequency.
Example: In table above, modal class = 150 ≤ h < 160 (f=15)
Median from Grouped Data
Find median position: (n+1)÷2 = (40+1)÷2 = 20.5
Add cumulative frequencies to find which class contains the median.
Median class = class containing the 20.5th value.
5. Cumulative Frequency Diagrams
Cumulative Frequency Table
| Time (t mins) | Frequency | Cumulative Frequency |
|---|---|---|
| 0 ≤ t < 10 | 8 | 8 |
| 10 ≤ t < 20 | 15 | 23 |
| 20 ≤ t < 30 | 22 | 45 |
| 30 ≤ t < 40 | 12 | 57 |
| 40 ≤ t < 50 | 3 | 60 |
Median from Curve
Position = n/2 = 60/2 = 30
Draw horizontal line from 30 on y-axis to curve, then down to x-axis.
Quartiles
Lower Quartile (Q₁) = n/4 = 60/4 = 15
Upper Quartile (Q₃) = 3n/4 = 3×60/4 = 45
Read values from curve at these positions.
Interquartile Range (IQR)
Represents the middle 50% of data.
Small IQR = consistent data
Large IQR = spread out data
Using Cumulative Frequency Curves
Example Question: How many students took less than 25 minutes?
Method: Read cumulative frequency at t = 25
If reading = 35, then 35 students took less than 25 minutes.
Example Question: What percentage took more than 35 minutes?
Method: Read cumulative frequency at t = 35 (say, 54)
Number above 35 = 60 - 54 = 6
Percentage = (6/60) × 100% = 10%
6. Correlation and Scatter Diagrams
Positive Correlation
As one variable increases, the other increases.
Example: Height vs Weight
Negative Correlation
As one variable increases, the other decreases.
Example: TV hours vs Test scores
No Correlation
No clear relationship between variables.
Example: Shoe size vs IQ
Describing Correlation
Use two descriptors: strength and direction.
| Strength | Description | Example Description |
|---|---|---|
| Strong | Points close to a straight line | "Strong positive correlation" |
| Moderate | Some scatter but clear pattern | "Moderate negative correlation" |
| Weak | Points widely scattered | "Weak positive correlation" |
| Perfect | All points on a straight line (rare) | "Perfect positive correlation" |
⚠️ Important: Correlation ≠ Causation
Just because two variables are correlated does NOT mean one causes the other.
Example: Ice cream sales and drowning incidents are correlated.
This doesn't mean ice cream causes drowning!
Explanation: Both are caused by a third factor: hot weather.
7. Straight Line of Best Fit
Drawing the Line of Best Fit
Rules:
- ✓ Use a ruler
- ✓ Equal number of points above and below the line
- ✓ Line should pass through the mean point (x̄, ȳ)
- ✓ Ignore outliers when drawing the line
- ✓ Extend line slightly beyond data points
- ✓ Line should follow the trend of the data
Finding the Equation
Where: m = gradient, c = y-intercept
Finding Gradient (m)
Example: Line passes through (0, 10) and (20, 60)
m = (60 - 10) ÷ (20 - 0) = 50 ÷ 20 = 2.5
Finding y-intercept (c)
Value where line crosses y-axis (when x = 0).
Example: Line crosses at y = 10
c = 10
Equation: y = 2.5x + 10
Interpolation
Estimating values within the data range.
More reliable because you're within known data.
Example: If line gives y = 38 when x = 7 (and data has x values from 0-20).
Extrapolation
Estimating values outside the data range.
Less reliable - trend may not continue.
Example: Estimating y when x = 25 (but data only goes to x = 20).
8. Additional Important Concepts
Box-and-Whisker Plots (Box Plots)
Visual display using five-number summary:
- Minimum value
- Lower quartile (Q₁)
- Median (Q₂)
- Upper quartile (Q₃)
- Maximum value
Uses: Compare distributions, identify outliers, show spread.
Outliers using IQR Method
Any value outside these boundaries is an outlier.
Example: Q₁ = 20, Q₃ = 40, IQR = 20
Lower boundary = 20 - 1.5×20 = -10
Upper boundary = 40 + 1.5×20 = 70
Values < -10 or > 70 are outliers.
Key Formulas Summary
Mean
Mean (Grouped)
Median Position
Range
IQR
Frequency Density
Pie Chart Angle
Stratified Sampling
9. Exam Tips and Practice Questions
Top Exam Tips
✓ Always Show Working
Even for "show that" questions - markers give method marks.
✓ Use a Ruler
For graphs, lines of best fit, and straight lines.
✓ Label Everything
Axes, scales, units, keys, titles.
✓ Check Pie Charts
Angles should sum to 360°.
✓ Order Data First
Before finding median - this is a common mistake.
✓ Give Context
"The mean is 45 minutes" not just "45".
Common Mistakes to Avoid
- ✗ Using height instead of frequency density for unequal class widths in histograms
- ✗ Drawing bar charts with gaps for continuous data
- ✗ Drawing lines of best fit without a ruler
- ✗ Confusing correlation with causation
- ✗ Forgetting to include keys for pictograms and stem-and-leaf
- ✗ Not checking calculations with estimates
Practice Questions
Question 1: Averages
Find the mean, median, mode, and range of:
12, 15, 15, 18, 20, 22, 28
Mean: (12+15+15+18+20+22+28)÷7 = 130÷7 = 18.57
Median (ordered): 12, 15, 15, 18, 20, 22, 28 → 18
Mode: 15 (appears twice)
Range: 28-12 = 16
Question 2: Grouped Data
Calculate an estimate for the mean:
| Time (mins) | Frequency |
|---|---|
| 0-10 | 5 |
| 10-20 | 12 |
| 20-30 | 18 |
| 30-40 | 8 |
| 40-50 | 2 |
Midpoints: 5, 15, 25, 35, 45
f×x: 5×5=25, 12×15=180, 18×25=450, 8×35=280, 2×45=90
Σf = 45, Σ(f×x) = 1025
Mean = 1025÷45 = 22.78 minutes
Question 3: Sampling
A school has 450 Year 10s and 350 Year 11s. Calculate a stratified sample of size 40.
Total students = 450+350 = 800
Year 10: (450/800)×40 = 0.5625×40 = 22.5 → 23 students
Year 11: (350/800)×40 = 0.4375×40 = 17.5 → 17 students
Check: 23+17 = 40 ✓
Good luck with your IGCSE Statistics! 📊
Remember: Practice makes perfect. Work through lots of past paper questions!