Worked Solutions
Answer: Simple random sampling.
Advantage: It is less time-consuming and cheaper to conduct than taking a census (measuring all 800 plants).
Disadvantage: Due to natural random variation, the sample of 40 might not be perfectly representative of the entire population (e.g. it might accidentally select slightly heavier plants on average).
Answer: Continuous.
Reason: Weight is a measured quantity that can theoretically take any decimal value within a given range, depending on the accuracy of the measuring scale.
Answer: Discrete.
Reason: The number of leaves is a counted quantity. A plant can only have a whole number of fully formed leaves (e.g. 5, 6, 7), not a fractional amount.
First, order the raw data: 14, 21, 21, 21, 26, 28, 33, 34, 35, 38, 39, 42, 44, 45, 51.
Answer:
1 | 4
2 | 1 1 1 6 8
3 | 3 4 5 8 9
4 | 2 4 5
5 | 1
Key: 1 | 4 represents 14 marks.
There are \( n = 15 \) values. The median is the \( \frac{15+1}{2} = 8\text{th} \) value.
Counting through the ordered data, the 8th value is 34.
Answer: Median (\( Q_2 \)) = 34
The lower quartile is the median of the lower half (the 7 values below the median). This is the 4th value overall.
The upper quartile is the median of the upper half (the 7 values above the median). This is the 12th value overall.
Answer: Lower quartile (\( Q_1 \)) = 21, Upper quartile (\( Q_3 \)) = 42
\( \text{IQR} = Q_3 - Q_1 = 42 - 21 \).
Answer: Interquartile Range (IQR) = 21
Calculate the IQR for Class B: \( \text{IQR} = 44 - 34 = 10 \).
Calculate the outlier boundary distance: \( 1.5 \times \text{IQR} = 1.5 \times 10 = 15 \).
Lower boundary for outliers: \( Q_1 - 15 = 34 - 15 = 19 \).
Upper boundary for outliers: \( Q_3 + 15 = 44 + 15 = 59 \).
Since the lowest mark (8) is less than the lower boundary (19), it is mathematically an outlier.
Since the highest mark (58) is less than the upper boundary (59), it is not an outlier.
Answer: The calculation \( 34 - 1.5(10) = 19 \) shows that the lower boundary is 19. Since 8 < 19, it is an outlier. The upper boundary is 59, and since 58 < 59, it is not an outlier.
The box plot requires lines at the lowest non-outlier (22), \( Q_1 \) (34), Median (38), \( Q_3 \) (44), and Maximum (58). The outlier at 8 must be marked with a cross.
Sum the values: \( \sum x = 45.2 + 48.1 + 42.9 + 49.5 + 46.0 + 47.3 = 279.0 \).
Divide by \( n = 6 \): \( \bar{x} = \frac{279.0}{6} = 46.5 \).
Answer: Mean \( \bar{x} = 46.5\text{ m} \)
We can calculate the deviations from the mean for each value: \( (x - \bar{x}) \).
Deviations: -1.3, 1.6, -3.6, 3.0, -0.5, 0.8.
Square the deviations: 1.69, 2.56, 12.96, 9.00, 0.25, 0.64.
Sum of the squared deviations: \( \sum (x - \bar{x})^2 = 27.10 \).
Use the unbiased sample standard deviation formula: \( s = \sqrt{ \frac{\sum (x - \bar{x})^2}{n - 1} } \).
\( s = \sqrt{\frac{27.10}{5}} = \sqrt{5.42} \approx 2.328 \).
Answer: \( s = 2.33\text{ m} \) (to 3 s.f.)
Answer: A histogram is appropriate because the data is continuous and is grouped into classes of unequal widths.
For a histogram, the Area of the bar is proportional to the Frequency.
For the \( 10 < t \le 15 \) class: Frequency is \( 20 \). The drawn area is \( 4\text{ cm} \times 8\text{ cm} = 32\text{ cm}^2 \).
This sets our scale factor: \( \text{Area} = k \times \text{Frequency} \implies 32 = k \times 20 \implies k = 1.6 \).
Therefore, \( 1\text{ person} = 1.6\text{ cm}^2 \) of area on the paper.
Now consider the \( 20 < t \le 30 \) class: The class width is 10 (double the width of the previous class). Therefore, its drawn width will be \( 4\text{ cm} \times 2 = 8\text{ cm} \).
The frequency is \( 30 \). The required area is \( 30 \times 1.6 = 48\text{ cm}^2 \).
Since Area = width \( \times \) height, we have \( 48 = 8 \times \text{height} \).
Height = \( 6\text{ cm} \).
Answer: Width = 8cm, Height = 6cm.
(Note: This can also be calculated using Frequency Density = Freq / Class Width).
We assume the data is uniformly distributed within each class interval (linear interpolation).
From 18 to 20: This covers \( \frac{2}{5} \) of the \( 15 < t \le 20 \) class.
People = \( \frac{2}{5} \times 25 = 10 \).
From 20 to 25: This covers \( \frac{5}{10} \) (or half) of the \( 20 < t \le 30 \) class.
People = \( \frac{1}{2} \times 30 = 15 \).
Total estimated people = \( 10 + 15 = 25 \).
Answer: 25 people
Note: In an exam, reading values directly off the graph using guidelines is expected and allows for a small tolerance. The values below are derived from interpolating the exact coordinates plotted to show the precise mathematical locations.
The median is the \( \frac{80}{2} = 40\text{th} \) value on the cumulative frequency (y) axis.
Looking at the graph, y=40 falls on the line segment connecting (20,18) to (30,44).
Using linear interpolation: \( \frac{x - 20}{30 - 20} = \frac{40 - 18}{44 - 18} \implies \frac{x - 20}{10} = \frac{22}{26} \).
\( x - 20 \approx 8.46 \implies x \approx 28.5 \).
Answer: Median \( \approx 28.5\text{ mins} \)
Lower Quartile (\( Q_1 \)) is the 20th value. This falls on the same line segment (20,18) to (30,44).
\( \frac{x - 20}{10} = \frac{20 - 18}{26} = \frac{2}{26} \implies x \approx 20.8 \).
Upper Quartile (\( Q_3 \)) is the 60th value. This falls on the segment connecting (30,44) to (40,66).
\( \frac{x - 30}{10} = \frac{60 - 44}{66 - 44} = \frac{16}{22} \implies x \approx 37.3 \).
\( \text{IQR} = Q_3 - Q_1 = 37.3 - 20.8 = 16.5 \).
Answer: IQR \( \approx 16.5\text{ mins} \)
We need to find the cumulative frequency at exactly 45 minutes.
45 is the exact midpoint between x=40 and x=50. The corresponding y-values are 66 and 76.
The midpoint y-value is \( \frac{66 + 76}{2} = 71 \).
This means 71 patients waited 45 minutes or less. The number who waited longer than 45 minutes is the remainder.
\( 80 - 71 = 9 \).
Answer: 9 patients received an apology.
For Group A (\( n = 15 \)):
Mean: \( \bar{x} = \frac{\sum x}{n} = \frac{540}{15} = 36\text{ cm} \).
Sample Standard Deviation: \( s_A = \sqrt{ \frac{\sum x^2 - n\bar{x}^2}{n - 1} } \).
\( s_A = \sqrt{ \frac{20350 - 15(36^2)}{14} } = \sqrt{ \frac{20350 - 19440}{14} } = \sqrt{ \frac{910}{14} } = \sqrt{65} \).
Answer: Mean = \( 36\text{ cm} \), \( s_A \approx 8.06\text{ cm} \)
To find the combined mean, we need the total sum of all heights.
Group A sum: \( \sum x_A = 540 \).
Group B sum: \( \sum x_B = n_B \times \bar{x}_B = 25 \times 38 = 950 \).
Total sum = \( 540 + 950 = 1490 \).
Combined mean = \( \frac{1490}{15 + 25} = \frac{1490}{40} \).
Answer: Combined mean = \( 37.25\text{ cm} \)
First, work backwards to find \( \sum x^2 \) for Group B.
\( s_B^2 = \frac{\sum x_B^2 - n_B\bar{x}_B^2}{n_B - 1} \implies 5.2^2 = \frac{\sum x_B^2 - 25(38^2)}{24} \).
\( 27.04 \times 24 = \sum x_B^2 - 36100 \implies 648.96 = \sum x_B^2 - 36100 \).
\( \sum x_B^2 = 36748.96 \).
Now find the total \( \sum x^2 \) for all 40 plants:
Combined \( \sum x^2 = 20350 + 36748.96 = 57098.96 \).
Finally, calculate the combined sample standard deviation (\( n=40 \)):
\( s_{combined} = \sqrt{ \frac{57098.96 - 40(37.25^2)}{39} } \).
\( s_{combined} = \sqrt{ \frac{57098.96 - 55502.5}{39} } = \sqrt{ \frac{1596.46}{39} } = \sqrt{40.9348...} \).
Answer: Combined \( s = 6.40\text{ cm} \) (to 3 s.f.)
Mean study hours: \( \bar{x} = \frac{120}{8} = 15 \text{ hours} \).
Mean test score: \( \bar{y} = \frac{560}{8} = 70 \text{ marks} \).
Answer: \( \bar{x} = 15 \), \( \bar{y} = 70 \)
Answer: Correlation does not imply causation.
While the two variables move together, we cannot prove that studying alone caused the higher scores. There may be confounding variables at play, such as the student's natural aptitude for the subject, the quality of their teaching, or their general motivation levels.
Answer: The gradient (2.8) means that for every additional hour a student spends studying per week, the model predicts their test score will increase by 2.8 marks.
Substitute \( x = 30 \) into the regression equation: \( y = 28 + 2.8(30) = 28 + 84 = 112 \).
This estimate is highly unreliable for two reasons: