📚 Statistics
Statistics is a branch of mathematics used to collect, organize, analyze, and interpret data to draw conclusions and make decisions. In simple words, it is the science of making sense out of data.Used in: Business, Medicine, Government, Climate Research, Machine Learning, Data Science, and almost every field.
Types of Statistics
Statistics has two main branches:
Descriptive Statistics — studying and summarizing past data.
- Example: Finding average marks of students in a class.
- You are just describing what the data looks like.
Inferential Statistics — using a sample to make predictions about a larger population.
- Example: Survey 50,000 people across India to estimate average salary of all 140 crore Indians.
- You are inferring (predicting) beyond what you directly measured.
Population vs Sample
Population — the entire group you want to study. Example: All 140 crore people in India.
Sample — a small subset taken from the population. Example: 50,000 people selected randomly from all states, genders, age groups.
A good sample must be:
- Large enough (sufficient size)
- Random (no bias)
- Representative (every type of person included)
If your sample is bad, your predictions about the population will also be bad.
Data Types
Data TypesCategorical Data
- Nominal — categories with no order. Example: Gender (Male/Female), State (UP/Bihar/Maharashtra), Religion
- Ordinal — categories with some order. Example: Customer rating (Bad / Average / Good / Excellent)
Numerical Data
- Discrete — only specific values possible (no decimals). Example: How many free throws did Steph miss? (0, 1, 2, 3...)
- Continuous — any value possible, infinite decimal places. Example: Steph's height (191.3, 191.27, 191.271...)
Special case — Proportions
- Built from nominal data (made/missed), but expressed as a number. Example: Steph's 3-point % = 128 attempts, 61 made → 0.4766
Measures of Central Tendency

This tells you where the center of your data lies.
Mean (Average): Sum of all values divided by count of values. Example: Data = 3, 1, 2, 5, 4 Mean = (3+1+2+5+4) / 5 = 3
Symbols:
- Population mean → μ (mu)
- Sample mean → x̄ (x-bar)
Problem with Mean — it is heavily affected by outliers.
Example: 10 students earn ₹30,000/month. One student starts a startup and earns ₹1 crore/month. Now the class average shoots up to lakhs : which does NOT represent the class correctly.
Median: The middle value when data is sorted in order. Example: Data = 3, 1, 2, 5, 4 Sorted = 1, 2, 3, 4, 5 → Median = 3
If even number of values — average of two middle values. Advantage: not affected by outliers. Example: If Bill Gates sits in your classroom, median salary stays realistic. Mean salary becomes crores.
Always look at median salary of a college, not average salary.
Mode: The value that appears most frequently in data. Example: Data = 1, 2, 2, 3, 3, 3, 4 Mode = 3 (appears 3 times)
Most useful for Categorical data. Example: Which state do most students come from? → Most frequent state = Mode
Weighted Mean: Each value is given a different importance (weight).
Example: You have 3 ML models predicting house price:
- Linear Regression → weight 0.2 → predicted ₹10 lakh
- Random Forest → weight 0.3 → predicted ₹15 lakh
- XGBoost → weight 0.5 → predicted ₹12 lakh
Weighted Mean = (0.2 × 10) + (0.3 × 15) + (0.5 × 12) = 2 + 4.5 + 6 = ₹12.5 lakh
Trimmed Mean: Remove top and bottom X% of data (outliers), then calculate mean on remaining data. Example: Remove bottom 20% and top 20% of flat prices, then take mean of remaining 60%.
Useful when outliers exist but you still want a mean-type answer.
Measures of Dispersion
Central tendency tells you where the center is. But it does not tell you how spread out the data is.
Example:
- Column A: -1, 0, 1 → Mean = 0
- Column B: -10, 0, 10 → Mean = 0
Both have the same mean but very different spread. So you need Measures of Dispersion.
Range: Range = Maximum value − Minimum value. Example: Data = 1, 2, 3, 4, 5 → Range = 5 − 1 = 4
Problem : heavily affected by outliers. One extreme value changes everything.
Variance: Average of squared distances of each point from the mean.
Why Square Instead of |mod|?
Both squaring and taking absolute value (mod) remove the negative signs. But we prefer squaring because:
- Squaring is easier for math → no sharp corner like absolute value, so calculations (like derivatives) are smoother
- Penalizes big errors more → large differences become much larger after squaring
- Works better in formulas → simpler algebra, used in variance, regression, etc.
- Absolute value (MAD) exists → but harder to use in advanced math
Derivation of the Variance Formula
Goal: Measure how spread out data is from the mean.
Step 1 — Find the mean:
$$
\mu = \frac{1}{n} \sum_{i=1}^{n} x_i
$$
Step 2 — Find each deviation from the mean:
$$
d_i = x_i - \mu
$$
Step 3 — Why not just average the deviations?
$$
\frac{1}{n} \sum_{i=1}^{n} (x_i - \mu) = 0 \quad \text{(always)}
$$
Positive and negative deviations cancel out, so this gives nothing useful.
Step 4 — Square each deviation to remove negatives:
$$
d_i^2 = (x_i - \mu)^2
$$
Step 5 — Average the squared deviations (Variance):
$$
\sigma^2 = \frac{1}{n} \sum_{i=1}^{n} (x_i - \mu)^2
$$
Step 6 — Standard Deviation (square root of variance):
$$
\sigma = \sqrt{\frac{1}{n} \sum_{i=1}^{n} (x_i - \mu)^2}
$$
Alternate Formula for Variance
Starting from the original formula:
$$
\sigma^2 = \frac{1}{n} \sum_{i=1}^{n} (x_i - \mu)^2
$$
Step 1 — Expand the square:
$$
\sigma^2 = \frac{1}{n} \sum_{i=1}^{n} \left(x_i^2 - 2x_i\mu + \mu^2\right)
$$
Step 2 — Split the summation:
$$
\sigma^2 = \frac{1}{n} \left[ \sum x_i^2 - 2\mu \sum x_i + \sum \mu^2 \right]
$$
Step 3 — Simplify each term:
$$
\frac{1}{n} \sum x_i^2 = \overline{x^2}
$$
$$
\frac{1}{n} \cdot 2\mu \sum x_i = 2\mu \cdot \mu = 2\mu^2
$$
$$
\frac{1}{n} \sum \mu^2 = \mu^2
$$
(since $\mu$ is constant and summed $n$ times)
Step 4 — Putting it together:
$$
\sigma^2 = \overline{x^2} - 2\mu^2 + \mu^2
$$
$$
\sigma^2 = \overline{x^2} - \mu^2
$$
Variance=Mean of Squares−Square of Mean

Note — Variance is not exactly the spread. It is proportional to the spread. More variance = more spread. Less variance = more compact data.
Problem: prone to outliers (squaring makes large differences even larger).
Important: Population variance uses N in denominator. Sample variance uses N−1. This is a very common interview question.
Standard Deviation: Square root of variance.
$\sigma = \sqrt{\text{Variance}}$
Why does it exist? Variance is in squared units. Standard deviation brings it back to the original unit of data.
Example: If salary data is in ₹, variance is in ₹². Standard deviation is back in ₹ — which makes practical sense.
Coefficient of Variation (CV): Coefficient of Variation (CV) is the ratio of standard deviation to mean, used to measure relative variability.
$ CV = \frac{\text{Standard Deviation}}{\text{Mean}} \times 100$
Used when you want to compare the spread of two completely different columns.
Example: You want to compare spread of Age column vs Fare column in Titanic data. Their units and scales are totally different. CV normalizes both so you can compare them fairly.
Higher CV = more spread relative to mean. Lower CV = data is more compact around the mean.
Percentiles, Five Number Summary & Box Plots
1. Quantiles & Percentiles
What are Quantiles?
Quantiles divide your entire numerical data into equal-sized buckets, where each bucket has the same number of observations.
Types of quantiles:
| Name | Divides data into | Example |
|---|
| Quartiles | 4 equal parts | Q1 = 25%, Q2 = 50%, Q3 = 75% |
| Deciles | 10 equal parts | D1 = 10%, D2 = 20%... |
| Percentiles | 100 equal parts | P1, P2, P3... P99 |
| Quintiles | 5 equal parts | 20%, 40%, 60%, 80% |
Important rule: data must always be sorted from lowest to highest before calculating any quantile.
If you know percentiles, you can derive everything else. Q1 is just the 25th percentile. Q2 is the 50th percentile (median). Q3 is the 75th percentile.
What does a Percentile mean?
The Nth percentile is the value below which N% of observations fall.
Example: In CAT exam, if you score 99th percentile it means 99% of students scored less than you and only 1% scored more.
Percentile is NOT the same as percentage. 90% marks means you scored 90 out of 100. 90th percentile means 90% of people scored below you.
How to Calculate the Percentile Value (which score = Xth percentile?)
Formula: $L = \frac{P}{100} \times (n + 1)$
Where P = percentile you want (e.g. 75), n = total number of observations, L = location (position in sorted data).
Example: Data (sorted, 10 students) = 80, 85, 87, 88, 90, 91, 93, 95, 96, 98
Find 75th percentile value: $L = \frac{75}{100} \times (10 + 1) = 0.75 \times 11 = 8.25$.
This means the answer lies between the 8th and 9th position. 8th position value = 95, 9th position value = 96 $\text{Answer} = 95 + 0.25 \times (96 - 95) = 95.25$
How to Calculate the Percentile of a Given Value (which percentile is score $X$ at?)
1. Main Formula (More Accurate)
$\text{Percentile} = \frac{x + 0.5 \times e}{n} \times 100$
Where: $x$ = number of values below the given value
$e$ = number of values equal to the given value
$n$ = total number of values
Example: Find the percentile of score 88 in the same data.
Values below 88 = 3 (80, 85, 87). Values equal to 88 = 1. $\text{Percentile} = \frac{3 + 0.5 \times 1}{10} \times 100 = 35$
So a score of 88 is at the 35th percentile: meaning 35% of students scored below 88.
2. Simple Method
$\text{Percentile} = \frac{\text{Number of values below } x}{n} \times 100$
Example: 88 is at position 4
Values below 88 = 3. $\text{Percentile} = \frac{3}{10} \times 100 = 30$. So, 88 is at the 30th percentile
Five Number Summary
The Five Number Summary gives you 5 key values that describe the spread and distribution of your data.
| Number | What it is | Percentile |
|---|
| Minimum | Smallest value in data | 0th percentile |
| Q1 | First Quartile | 25th percentile |
| Median (Q2) | Middle value | 50th percentile |
| Q3 | Third Quartile | 75th percentile |
| Maximum | Largest value in data | 100th percentile |
These 5 numbers together tell you where the center is, how spread out the data is, and where most values lie.
IQR — Interquartile Range
IQR=Q3−Q1.
IQR represents the middle 50% of your data. It is a measure of spread that is not affected by outliers.
Why is IQR better than Range?
Because Range uses the maximum and minimum values which can be extreme outliers. IQR ignores the top 25% and bottom 25% and only looks at the middle 50%.
Example: Q1 = 72.5, Q3 = 96.5
IQR=96.5−72.5=24
Even if the smallest value was 4 instead of 40, or the largest was 1 lakh instead of 125, the IQR stays the same. That is the power of IQR. Check example below
Box Plot
A Box Plot is a visual graph built using the Five Number Summary. It gives you a lot of information about your data in one picture — center, spread, skewness, and outliers.
How to Build a Box Plot Step by Step
Step 1 — Sort your data.
Example data (sorted, 19 values): 40, 55, 60, 65, 68, 70, 72, 75, 80, 90, 92, 94, 94, 95, 95, 98, 98, 100, 125
Step 2 — Find the Median.
19 observations → median is at position 10 = 90
This splits data into left half (first 9 values) and right half (last 9 values).
Step 3 — Find Q1 (Lower Fourth).
Take the left half: 40, 55, 60, 65, 68, 70, 72, 75, 80
Median of this = 5th value = 68
Note — for even number of values in the half, average the two middle values. Left half has 9 values (odd) → Q1 = middle = 68
Step 4 — Find Q3 (Upper Fourth).
Take the right half: 92, 94, 94, 95, 95, 98, 98, 100, 125
Q3 = middle = 95
Step 5 — Calculate IQR.
$IQR = Q_3 - Q_1 = 95 - 68 = 27$
Step 6 — Find Outlier Boundaries
Upper Boundary: $\text{Upper Boundary} = Q_3 + 1.5 \times IQR = 95 + 40.5 = 135.5$
Lower Boundary: $\text{Lower Boundary} = Q_1 - 1.5 \times IQR = 68 - 40.5 = 27.5$
Any value beyond these boundaries is an outlier.
Here 125 < 135.5 so no upper outlier. 40 > 27.5 so no lower outlier.
Step 7 — Draw the Box Plot.
- Draw a horizontal axis with your data range.
- Draw a rectangle from Q1 to Q3 — this box represents the middle 50% of your data (the IQR).
- Draw a vertical line inside the box at the Median.
- Draw a line (whisker) from Q1 left to the smallest non-outlier value.
- Draw a line (whisker) from Q3 right to the largest non-outlier value.
- Plot any outliers as dots beyond the whiskers.

- The median (90) is shifted towards Q3 (95) and away from Q1 (68). This means the data is slightly left skewed — more students scored on the higher side, but a few students scored very low (like 40, 55) which pulls the left whisker far out.
- The box (IQR) is narrow (only 27 marks wide) which means the middle 50% of students scored between 68 and 95 — fairly concentrated.
- The left whisker is much longer than the right whisker — confirming the left skew.
What a Box Plot tells you?
| What you see | What it means |
|---|
| Median line in center of box | Data is symmetric |
| Median line shifted right | Data is left skewed |
| Median line shifted left | Data is right skewed |
| Long whisker on one side | More spread on that side |
| Dots beyond whiskers | Outliers exist |
| Wide box (large IQR) | High variability in middle 50% |
| Narrow box (small IQR) | Data is concentrated |
Types of Outliers
Mild Outlier — value is between $1.5 \times IQR$ and $3 \times IQR$ from the nearest quartile. Shown as a filled dot on the plot.
- Mild outlier range: $\text{Mild outlier range: } Q_3 + 1.5 \times IQR \ \text{to} \ Q_3 + 3 \times IQR$
Extreme Outlier — value is beyond $3 \times IQR$ from the nearest quartile. Shown as an open dot on the plot.
- Extreme outlier: $\text{Extreme outlier: beyond } Q_3 + 3 \times IQR$
Degrees of Freedom (df)
Degrees of freedom = the number of independent pieces of information needed to make a calculation.
It is NOT the same as sample size. It depends on:
- What calculation you are making
- What you already know before making that calculation
The more things you already know in advance, the fewer independent pieces of information you need. That is degrees of freedom.
Example 1 — Coin Toss (df = 1)
- You toss a coin. There are only 2 possible outcomes — Heads or Tails.
- If someone tells you they got Heads , you automatically know they did NOT get Tails.
- Only 1 piece of information was needed to figure out everything.
- df=2−1=1
Even if you toss the coin 100 times — you still only need to know the number of Heads to automatically know the number of Tails. So df = 1 still.
Example 2 — Traffic Light (df = 2)
- 3 possible outcomes — Red, Amber, Green.
- Someone tells you: "It is not Amber." → you still don't know. Someone tells you: "It is not Green." → now you know it must be Red.
- 2 pieces of information were needed.
- df=3−1=2
Rule for categorical data: df=Number of categories−1
Why Mean has df = n (no loss)
When you calculate the mean, you know nothing in advance. Every single value in your sample contributes independently. If any one value changes , your mean changes.So all n values are free to vary → no degree of freedom is lost.
Why Standard Deviation has df = n − 1?
To calculate standard deviation you must first calculate the mean.
$s = \sqrt{\frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n - 1}}$
Because you already know the mean, the last value is NOT free , it is fixed automatically.
Example: You have 3 numbers. You know their mean = 5. First two numbers are 4 and 5. The third number MUST be 6. It cannot be anything else.
So only n−1 values are truly independent. The last one is locked once the mean is known. df=n−1
Note: If we divided by n instead of n−1, we would include that last locked (redundant) data point which would cause us to underestimate the standard deviation. This is especially a big problem for small samples.
This is exactly why sample variance formula uses n−1 in the denominator!
Degrees of Freedom in Statistical Tests
When you run a hypothesis test (like t-test or chi-square), you first have to calculate means, totals, or variances before computing the test statistic.
Every extra thing you calculate in advance = one less degree of freedom.
| Test | What you calculate in advance | df |
|---|
| One sample t-test | Mean of 1 sample | n − 1 |
| Two sample t-test | Mean of 2 samples | n₁ + n₂ − 2 |
| Chi-square (categories) | Expected values | categories − 1 |
Why df Matters in Tests?
The p-value depends on BOTH the test statistic AND the degrees of freedom.
Example — Chi-square value of 4:
| df | p-value | Significant? |
|---|
| df = 1 | p = 0.046 | Yes (p < 0.05) |
| df = 2 | p = 0.135 | No (p > 0.05) |
Same test statistic , completely different conclusion , just because df changed!
It is like saying "my team won 9 matches":
- Out of 10 matches → very impressive
- Out of 100 matches → not so impressive
Without knowing df, the test statistic means nothing.
Graphs : Univariate Analysis (one column at a time)
1. For Categorical Column:
Frequency Distribution Table - count how many times each category appears.
| Vacation Type | Count |
|---|
| Beach | 60 |
| City | 40 |
| Adventure | 30 |
Relative Frequency - convert counts to percentages.
- Beach = 60/200 = 30%
- Use this to plot a Pie Chart.
Cumulative Frequency - running total of counts.
- Beach = 60, Beach+City = 100, +Adventure = 130...
- Use this to plot a Line Chart.
Bar Chart - plot categories on X-axis, frequency on Y-axis.
2. For Numerical Column:
Histogram - create buckets (bins) and count how many values fall in each.
- Example: Age column → Buckets: 0−10, 11−20, 21−30...
- Difference from Bar Chart — bars in histogram touch each other (continuous data). Bars in bar chart have gaps (discrete categories).
Choosing bin size matters:
- Too large bins → very few bars, lose detail
- Too small bins → too many bars, noisy
- Find a good middle ground
Histogram Shapes:
- Symmetric — most values in center, fewer on both sides (like a bell). This is Normal Distribution.
- Bimodal — two peaks. Two groups exist in the data.
- Left Skewed — tail goes left, most data is on the right. Example: marks in a very easy test.
- Right Skewed — tail goes right, most data is on the left. Example: salaries (most people earn less, very few earn a lot).
- Uniform — all values appear with roughly equal frequency.
Bivariate Analysis (two columns together)
1. Categorical + Categorical → Contingency Table (Crosstab)
- Make a table where rows = one category, columns = another category, and cells = count.
- Example: Survived (0/1) vs Passenger Class (1/2/3) in Titanic data.
- Then plot a Grouped Bar Chart or Stacked Bar Chart on top of it.
2. Numerical + Numerical → Scatter Plot
- Plot one column on X-axis and another on Y-axis. Each point = one row.
- Positive relationship — both increase together.
- Negative relationship — one increases, other decreases.
- No relationship — random scattered points.
3. Categorical + Numerical → Bar Chart with Aggregation
- Example: Male vs Female passengers — plot average age of each group.
- OR convert numerical column into buckets first, then make a Contingency Table.
Covariance & Correlation
Why do we need Covariance?
We already know:
- Mean tells us the center of data
- Variance tells us the spread of one column
But what if we have two numerical columns and we want to study the relationship between them?
Example: Flat area (sq ft) vs Price (lakhs). Does price increase as area increases?
This is where Covariance comes in.
What is Covariance?
Covariance measures the direction of the linear relationship between two numerical columns. It tells you one of three things:
| Covariance Value | Meaning |
|---|
| Positive | Both columns increase together |
| Negative | One increases, other decreases |
| Near Zero | No linear relationship |
Real life examples:
- Experience and Salary → Positive (more experience = more salary)
- Backlogs and Package → Negative (more backlogs = less package)
- Backlogs and Placement luck → Near Zero (no clear pattern)
Formula for Covariance
Population Covariance: $\sigma_{xy} = \frac{\sum (x_i - \mu_x)(y_i - \mu_y)}{N}$
Sample Covariance: $S_{xy} = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{n-1}$
Where x and y are two different numerical columns, and x̄, ȳ are their respective means.
Why n−1 for sample? Same reason as variance - because we already used one degree of freedom to calculate the mean.
Worked Example
Data: 5 employees — Experience (X) and Salary in lakhs (Y)
| Employee | Experience (X) | Salary (Y) |
|---|
| 1 | 2 | 1 |
| 2 | 5 | 2 |
| 3 | 8 | 6 |
| 4 | 12 | 8 |
| 5 | 13 | 10 |
Mean of X = 8, Mean of Y = 6 (approx for simplicity)
Step 1 — Calculate deviations and their product:
| Employee | x − x̄ | y − ȳ | (x−x̄)(y−ȳ) |
|---|
| 1 | −6 | −5 | +30 |
| 2 | −3 | −4 | +12 |
| 3 | 0 | 0 | 0 |
| 4 | +4 | +2 | +8 |
| 5 | +5 | +4 | +20 |
Sum of products = 70
Step 2 — Divide by n−1:
$S_{xy} = \frac{70}{4} = 17.5$
Positive covariance → confirms Experience and Salary are positively related.
The Quadrant Trick (Visual Shortcut)
Draw a vertical line at mean of X and horizontal line at mean of Y. This creates 4 quadrants.
|
Q2 | Q1
(−,+) | (+,+)
--------+-------- ← Mean of Y
Q3 | Q4
(−,−) | (+,−)
|
↑ Mean of X
- Points in Q1 and Q3 → product is positive → Positive covariance
- Points in Q2 and Q4 → product is negative → Negative covariance
- Points spread equally in all quadrants → cancel out → Near zero covariance
Just by seeing where most points fall, you can guess the sign of covariance!
Covariance of a Variable with Itself
What if you calculate covariance of X with X?
$Cov(X, X) = \frac{\sum(x_i - \bar{x})(x_i - \bar{x})}{n-1} = \frac{\sum(x_i - \bar{x})^2}{n-1} = \text{Variance of X}$
So covariance of a variable with itself = its own variance. This is a common interview question!
Problem with Covariance
Covariance only tells you the direction (positive or negative). It does NOT tell you the strength of the relationship.
Also covariance is affected by the scale of data. If you multiply both columns by 2, covariance becomes 4 times larger — even though the relationship between the columns has not changed at all.
Example:
- Original covariance between X and Y = 100
- Multiply both X and Y by 2 → covariance becomes 400
- But the scatter plot looks exactly the same!
This makes covariance an unreliable measure for comparing strength of relationships.
What is Correlation?
- Correlation solves the problem of covariance. It measures both the direction AND the strength of the linear relationship.
- Correlation is always between −1 and +1. No matter how you scale the data, correlation stays the same.
$r_{xy} = \frac{Cov(X, Y)}{\sigma_x \times \sigma_y}$
Where σx and σy are the standard deviations of X and Y respectively.
By dividing by the standard deviations we normalize the covariance and remove the effect of scale.
Interpreting Correlation Values
| Correlation (r) | Meaning |
|---|
| r = +1 | Perfect positive — X increases, Y increases by exact same proportion |
| r close to +1 | Strong positive relationship |
| r = 0.5 to 0.7 | Moderate positive relationship |
| r close to 0 | Weak or no relationship |
| r = −0.5 to −0.7 | Moderate negative relationship |
| r close to −1 | Strong negative relationship |
| r = −1 | Perfect negative — X increases, Y decreases by exact proportion |
Worked Example
Using the same data as above:
- Covariance (X, Y) = 17.5
- Standard deviation of X = σx (calculate from X column)
- Standard deviation of Y = σy (calculate from Y column)
- $r = \frac{17.5}{\sigma_x \times \sigma_y}$
- If this gives r = 0.82 → Strong positive correlation between experience and salary.
In Python:
df['X'].cov(df['Y']) # covariance
df['X'].corr(df['Y']) # correlation
Covariance vs Correlation
| Feature | Covariance | Correlation |
|---|
| What it tells | Direction only | Direction + Strength |
| Range | −∞ to +∞ | −1 to +1 |
| Affected by scale? | Yes | No |
| Reliable measure? | Less reliable | Very reliable |
| Used in practice? | Rarely alone | Always preferred |
Covariance exists mainly because you calculate it first, then use it to compute correlation.
Correlation Does NOT Mean Causation
This is one of the most important points in all of statistics. Just because two variables are correlated does not mean one is causing the other.
Famous Example 1 — Ice cream and murders: On hot days, both ice cream sales and murder rates go up. They are positively correlated. But eating ice cream does not cause murders! The hidden reason is hot weather — people go outside more, leading to both more ice cream sales and more crimes.
Famous Example 2 — Fire fighters and fire damage: More fire fighters are sent to bigger fires. So number of fire fighters and fire damage are positively correlated. But more fire fighters do not cause more damage — the size of the fire is the real reason.
Famous Example 3 — Experience and Salary: Experience and salary are positively correlated. But you cannot say experience causes higher salary. Other factors like talent, company, market demand also play a role.
To prove causation you need controlled experiments — correlation alone is never enough.