Recent posts in Study Materials

533
533 views

Isro cs scientist 2025 question paper

   Random text to upload file

k9vP2mQ8wX4zL7jF1bH3dN6gC5fY0sT9vK2mQ8wX4zL7jF1bH3dN6gC5fY0sT9vK2mQ8wX4zL7jF1bH3dN6gC5fY0sT9vK2mQ8wX4zL7jF1bH3dN6gC5fY0sT9vK2mQ8wX4zL7jF1bH3dN6gC5fY0sT9vK2mQ8wX4zL7jF1bH3dN6gC5fY0sT9vK2mQ8wX4zL7jF1bH3dN6gC5                                                                                                                                                      

460
460 views

I always confuse with the terms, full rank, full row rank, full column rank and relations between them. In few questions, I found out that these terms are very tricky and confusing. If you are also with me, let's just understand the things very clearly in a intuitive way.

Let's understand what actually these terms are. Let's consider a m x n matrix A.

The maximum possible rank of a matrix is min(m,n) i.e, $rank(A) \le min(m,n)$

A is said to be a

1. full rank matrix if $rank(A) = min(m,n)$

2. full column rank matrix if $rank(A) = n$

3. full row rank matrix if $rank(A) = m$


$\text{If A is \textbf{square} matrix} \rightarrow \text{m = n} \rightarrow \boxed{\text{full row rank} \equiv \text{full column rank} \equiv \text{full rank}}$


$
\begin{array}{c}
\text{If A is \textbf{rectangular} matrix} \\
\begin{array}{ccc}
\downarrow & & & & & \downarrow \\
\text{Case-1: m} \lt \text{n} & & & & & \text{Case-2: m} \gt \text{n}
\end{array}
\end{array}
$


$\textbf{Case-1: } \text{We know, m < n \& } \text{rank(A) <= min(m,n)} \xrightarrow{\text{means}} \text{rank(A) <= m} \\ \text{So, we came to a conclusion that rank(A) <= m.} \\ \text{In this case, A can be full rank only if rank(A) = m, which means A has full row rank. Full column rank is not possible in this case.}$ $$\boxed{\text{full rank} \equiv \text{full row rank}}$$


$\textbf{Case-2: } \text{We know, m > n \& } \text{rank(A) <= min(m,n)} \xrightarrow{\text{means}} \text{rank(A) <=  n} \\ \text{So, we came to a conclusion that rank(A) <= n.} \\ \text{In this case, A can be full rank only if rank(A) = n, which means A has full column rank. Full row rank is not possible in this case.}$ $$\boxed{\text{full rank} \equiv \text{full column rank}}$$

 

 

485
485 views

📚 Statistics 

Statistics is a branch of mathematics used to collect, organize, analyze, and interpret data to draw conclusions and make decisions. In simple words, it is the science of making sense out of data.Used in: Business, Medicine, Government, Climate Research, Machine Learning, Data Science, and almost every field.

Types of Statistics

Statistics has two main branches:

Descriptive Statistics — studying and summarizing past data.

  • Example: Finding average marks of students in a class.
  • You are just describing what the data looks like.

Inferential Statistics — using a sample to make predictions about a larger population.

  • Example: Survey 50,000 people across India to estimate average salary of all 140 crore Indians.
  • You are inferring (predicting) beyond what you directly measured.

Population vs Sample

Population — the entire group you want to study. Example: All 140 crore people in India.

Sample — a small subset taken from the population. Example: 50,000 people selected randomly from all states, genders, age groups.

A good sample must be:

  • Large enough (sufficient size)
  • Random (no bias)
  • Representative (every type of person included)

If your sample is bad, your predictions about the population will also be bad.

Data Types

Data Types

Categorical Data

  • Nominal — categories with no order. Example: Gender (Male/Female), State (UP/Bihar/Maharashtra), Religion
  • Ordinal — categories with some order. Example: Customer rating (Bad / Average / Good / Excellent)

Numerical Data

  • Discrete — only specific values possible (no decimals). Example: How many free throws did Steph miss? (0, 1, 2, 3...)
  • Continuous — any value possible, infinite decimal places. Example: Steph's height (191.3, 191.27, 191.271...)

Special case — Proportions

  • Built from nominal data (made/missed), but expressed as a number. Example: Steph's 3-point % = 128 attempts, 61 made → 0.4766

Measures of Central Tendency

This tells you where the center of your data lies.

Mean (Average): Sum of all values divided by count of values. Example: Data = 3, 1, 2, 5, 4 Mean = (3+1+2+5+4) / 5 = 3

Symbols:

  • Population mean → μ (mu)
  • Sample mean → x̄ (x-bar)

Problem with Mean — it is heavily affected by outliers.

Example: 10 students earn ₹30,000/month. One student starts a startup and earns ₹1 crore/month. Now the class average shoots up to lakhs : which does NOT represent the class correctly.

Median: The middle value when data is sorted in order. Example: Data = 3, 1, 2, 5, 4 Sorted = 1, 2, 3, 4, 5 → Median = 3

If even number of values — average of two middle values. Advantage: not affected by outliers. Example: If Bill Gates sits in your classroom, median salary stays realistic. Mean salary becomes crores.

Always look at median salary of a college, not average salary.

Mode: The value that appears most frequently in data. Example: Data = 1, 2, 2, 3, 3, 3, 4 Mode = 3 (appears 3 times)

Most useful for Categorical data. Example: Which state do most students come from? → Most frequent state = Mode

Weighted Mean: Each value is given a different importance (weight).

Example: You have 3 ML models predicting house price:

  • Linear Regression → weight 0.2 → predicted ₹10 lakh
  • Random Forest → weight 0.3 → predicted ₹15 lakh
  • XGBoost → weight 0.5 → predicted ₹12 lakh

Weighted Mean = (0.2 × 10) + (0.3 × 15) + (0.5 × 12) = 2 + 4.5 + 6 = ₹12.5 lakh

Trimmed Mean: Remove top and bottom X% of data (outliers), then calculate mean on remaining data. Example: Remove bottom 20% and top 20% of flat prices, then take mean of remaining 60%.

Useful when outliers exist but you still want a mean-type answer.

Measures of Dispersion

Central tendency tells you where the center is. But it does not tell you how spread out the data is.

Example:

  • Column A: -1, 0, 1 → Mean = 0
  • Column B: -10, 0, 10 → Mean = 0

Both have the same mean but very different spread. So you need Measures of Dispersion.

Range: Range = Maximum value − Minimum value. Example: Data = 1, 2, 3, 4, 5 → Range = 5 − 1 = 4

Problem : heavily affected by outliers. One extreme value changes everything.

Variance: Average of squared distances of each point from the mean.

Why Square Instead of |mod|?

Both squaring and taking absolute value (mod) remove the negative signs. But we prefer squaring because:

  • ​​​​​Squaring is easier for math → no sharp corner like absolute value, so calculations (like derivatives) are smoother
  • Penalizes big errors more → large differences become much larger after squaring
  • Works better in formulas → simpler algebra, used in variance, regression, etc.
  • Absolute value (MAD) exists → but harder to use in advanced math

Derivation of the Variance Formula

Goal: Measure how spread out data is from the mean.

Step 1 — Find the mean:

$$
\mu = \frac{1}{n} \sum_{i=1}^{n} x_i
$$

Step 2 — Find each deviation from the mean:

$$
d_i = x_i - \mu
$$

Step 3 — Why not just average the deviations?

$$
\frac{1}{n} \sum_{i=1}^{n} (x_i - \mu) = 0 \quad \text{(always)}
$$

Positive and negative deviations cancel out, so this gives nothing useful.

Step 4 — Square each deviation to remove negatives:

$$
d_i^2 = (x_i - \mu)^2
$$

Step 5 — Average the squared deviations (Variance):

$$
\sigma^2 = \frac{1}{n} \sum_{i=1}^{n} (x_i - \mu)^2
$$

Step 6 — Standard Deviation (square root of variance):

$$
\sigma = \sqrt{\frac{1}{n} \sum_{i=1}^{n} (x_i - \mu)^2}
$$

Alternate Formula for Variance

Starting from the original formula:

$$
\sigma^2 = \frac{1}{n} \sum_{i=1}^{n} (x_i - \mu)^2
$$

Step 1 — Expand the square:

$$
\sigma^2 = \frac{1}{n} \sum_{i=1}^{n} \left(x_i^2 - 2x_i\mu + \mu^2\right)
$$

Step 2 — Split the summation:

$$
\sigma^2 = \frac{1}{n} \left[ \sum x_i^2 - 2\mu \sum x_i + \sum \mu^2 \right]
$$

Step 3 — Simplify each term:

$$
\frac{1}{n} \sum x_i^2 = \overline{x^2}
$$

$$
\frac{1}{n} \cdot 2\mu \sum x_i = 2\mu \cdot \mu = 2\mu^2
$$

$$
\frac{1}{n} \sum \mu^2 = \mu^2
$$

(since $\mu$ is constant and summed $n$ times)

Step 4 — Putting it together:

$$
\sigma^2 = \overline{x^2} - 2\mu^2 + \mu^2
$$

$$
\sigma^2 = \overline{x^2} - \mu^2
$$

Variance=Mean of Squares−Square of Mean

Note — Variance is not exactly the spread. It is proportional to the spread. More variance = more spread. Less variance = more compact data.

Problem: prone to outliers (squaring makes large differences even larger).

Important: Population variance uses N in denominator. Sample variance uses N−1. This is a very common interview question.

Standard Deviation:  Square root of variance.

$\sigma = \sqrt{\text{Variance}}$

Why does it exist? Variance is in squared units. Standard deviation brings it back to the original unit of data.

Example: If salary data is in ₹, variance is in ₹². Standard deviation is back in ₹ — which makes practical sense.

Coefficient of Variation (CV): Coefficient of Variation (CV) is the ratio of standard deviation to mean, used to measure relative variability.

$ CV = \frac{\text{Standard Deviation}}{\text{Mean}} \times 100$

Used when you want to compare the spread of two completely different columns.

Example: You want to compare spread of Age column vs Fare column in Titanic data. Their units and scales are totally different. CV normalizes both so you can compare them fairly.

Higher CV = more spread relative to mean. Lower CV = data is more compact around the mean.

Percentiles, Five Number Summary & Box Plots

1. Quantiles & Percentiles

What are Quantiles?

Quantiles divide your entire numerical data into equal-sized buckets, where each bucket has the same number of observations.

Types of quantiles:

Name Divides data intoExample
Quartiles4 equal partsQ1 = 25%, Q2 = 50%, Q3 = 75%
Deciles10 equal partsD1 = 10%, D2 = 20%...
Percentiles100 equal partsP1, P2, P3... P99
Quintiles5 equal parts20%, 40%, 60%, 80%

Important rule: data must always be sorted from lowest to highest before calculating any quantile.

If you know percentiles, you can derive everything else. Q1 is just the 25th percentile. Q2 is the 50th percentile (median). Q3 is the 75th percentile.

What does a Percentile mean?

The Nth percentile is the value below which N% of observations fall.

Example: In CAT exam, if you score 99th percentile it means 99% of students scored less than you and only 1% scored more.

Percentile is NOT the same as percentage. 90% marks means you scored 90 out of 100. 90th percentile means 90% of people scored below you.

How to Calculate the Percentile Value (which score = Xth percentile?)

Formula: $L = \frac{P}{100} \times (n + 1)$

Where P = percentile you want (e.g. 75), n = total number of observations, L = location (position in sorted data).

Example: Data (sorted, 10 students) = 80, 85, 87, 88, 90, 91, 93, 95, 96, 98

Find 75th percentile value: $L = \frac{75}{100} \times (10 + 1) = 0.75 \times 11 = 8.25$. 

This means the answer lies between the 8th and 9th position. 8th position value = 95, 9th position value = 96 $\text{Answer} = 95 + 0.25 \times (96 - 95) = 95.25$

How to Calculate the Percentile of a Given Value (which percentile is score $X$ at?)

1. Main Formula (More Accurate)

$\text{Percentile} = \frac{x + 0.5 \times e}{n} \times 100$

Where:  $x$ = number of values below the given value
$e$ = number of values equal to the given value
$n$ = total number of values

Example: Find the percentile of score 88 in the same data.

Values below 88 = 3 (80, 85, 87). Values equal to 88 = 1.  $\text{Percentile} = \frac{3 + 0.5 \times 1}{10} \times 100 = 35$

So a score of 88 is at the 35th percentile: meaning 35% of students scored below 88.

2. Simple Method

$\text{Percentile} = \frac{\text{Number of values below } x}{n} \times 100$

Example: 88 is at position 4

Values below 88 = 3.  $\text{Percentile} = \frac{3}{10} \times 100 = 30$. So, 88 is at the 30th percentile

Five Number Summary

The Five Number Summary gives you 5 key values that describe the spread and distribution of your data.

NumberWhat it isPercentile
MinimumSmallest value in data0th percentile
Q1First Quartile25th percentile
Median (Q2)Middle value50th percentile
Q3Third Quartile75th percentile
MaximumLargest value in data100th percentile

These 5 numbers together tell you where the center is, how spread out the data is, and where most values lie.

IQR — Interquartile Range

IQR=Q3−Q1. 

IQR represents the middle 50% of your data. It is a measure of spread that is not affected by outliers.

Why is IQR better than Range? 

Because Range uses the maximum and minimum values which can be extreme outliers. IQR ignores the top 25% and bottom 25% and only looks at the middle 50%.

Example: Q1 = 72.5, Q3 = 96.5

IQR=96.5−72.5=24

Even if the smallest value was 4 instead of 40, or the largest was 1 lakh instead of 125, the IQR stays the same. That is the power of IQR. Check example below

Box Plot

A Box Plot is a visual graph built using the Five Number Summary. It gives you a lot of information about your data in one picture — center, spread, skewness, and outliers.

How to Build a Box Plot Step by Step

Step 1 — Sort your data.

Example data (sorted, 19 values): 40, 55, 60, 65, 68, 70, 72, 75, 80, 90, 92, 94, 94, 95, 95, 98, 98, 100, 125

Step 2 — Find the Median.

19 observations → median is at position 10 = 90

This splits data into left half (first 9 values) and right half (last 9 values).

Step 3 — Find Q1 (Lower Fourth).

Take the left half: 40, 55, 60, 65, 68, 70, 72, 75, 80

Median of this = 5th value = 68

Note — for even number of values in the half, average the two middle values. Left half has 9 values (odd) → Q1 = middle = 68

Step 4 — Find Q3 (Upper Fourth).

Take the right half: 92, 94, 94, 95, 95, 98, 98, 100, 125

Q3 = middle = 95

Step 5 — Calculate IQR.

$IQR = Q_3 - Q_1 = 95 - 68 = 27$

Step 6 — Find Outlier Boundaries

Upper Boundary: $\text{Upper Boundary} = Q_3 + 1.5 \times IQR = 95 + 40.5 = 135.5$

Lower Boundary: $\text{Lower Boundary} = Q_1 - 1.5 \times IQR = 68 - 40.5 = 27.5$

Any value beyond these boundaries is an outlier.

Here 125 < 135.5 so no upper outlier. 40 > 27.5 so no lower outlier.

Step 7 — Draw the Box Plot.

  • Draw a horizontal axis with your data range.
  • Draw a rectangle from Q1 to Q3 — this box represents the middle 50% of your data (the IQR).
  • Draw a vertical line inside the box at the Median.
  • Draw a line (whisker) from Q1 left to the smallest non-outlier value.
  • Draw a line (whisker) from Q3 right to the largest non-outlier value.
  • Plot any outliers as dots beyond the whiskers.

  • The median (90) is shifted towards Q3 (95) and away from Q1 (68). This means the data is slightly left skewed — more students scored on the higher side, but a few students scored very low (like 40, 55) which pulls the left whisker far out.
  • The box (IQR) is narrow (only 27 marks wide) which means the middle 50% of students scored between 68 and 95 — fairly concentrated.
  • The left whisker is much longer than the right whisker — confirming the left skew.

What a Box Plot tells you?

What you seeWhat it means
Median line in center of boxData is symmetric
Median line shifted rightData is left skewed
Median line shifted leftData is right skewed
Long whisker on one sideMore spread on that side
Dots beyond whiskersOutliers exist
Wide box (large IQR)High variability in middle 50%
Narrow box (small IQR)Data is concentrated

Types of Outliers

Mild Outlier — value is between $1.5 \times IQR$ and $3 \times IQR$ from the nearest quartile. Shown as a filled dot on the plot.

  • Mild outlier range: $\text{Mild outlier range: } Q_3 + 1.5 \times IQR \ \text{to} \ Q_3 + 3 \times IQR$

Extreme Outlier — value is beyond $3 \times IQR$ from the nearest quartile. Shown as an open dot on the plot.

  • Extreme outlier: $\text{Extreme outlier: beyond } Q_3 + 3 \times IQR$

Degrees of Freedom (df)

Degrees of freedom = the number of independent pieces of information needed to make a calculation.

It is NOT the same as sample size. It depends on:

  • What calculation you are making
  • What you already know before making that calculation
The more things you already know in advance, the fewer independent pieces of information you need. That is degrees of freedom.

Example 1 — Coin Toss (df = 1)

  • You toss a coin. There are only 2 possible outcomes — Heads or Tails.
  • If someone tells you they got Heads , you automatically know they did NOT get Tails.
  • Only 1 piece of information was needed to figure out everything.
  • df=2−1=1

Even if you toss the coin 100 times — you still only need to know the number of Heads to automatically know the number of Tails. So df = 1 still.

Example 2 — Traffic Light (df = 2)

  • 3 possible outcomes — Red, Amber, Green.
  • Someone tells you: "It is not Amber." → you still don't know. Someone tells you: "It is not Green." → now you know it must be Red.
  • 2 pieces of information were needed.
  • df=3−1=2

Rule for categorical data: df=Number of categories−1

Why Mean has df = n (no loss)

When you calculate the mean, you know nothing in advance. Every single value in your sample contributes independently. If any one value changes , your mean changes.So all n values are free to vary → no degree of freedom is lost.

Why Standard Deviation has df = n − 1?

To calculate standard deviation you must first calculate the mean.

$s = \sqrt{\frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n - 1}}$

Because you already know the mean, the last value is NOT free , it is fixed automatically.

Example: You have 3 numbers. You know their mean = 5. First two numbers are 4 and 5. The third number MUST be 6. It cannot be anything else.

So only n−1 values are truly independent. The last one is locked once the mean is known. df=n−1

Note: If we divided by n instead of n−1, we would include that last locked (redundant) data point which would cause us to underestimate the standard deviation. This is especially a big problem for small samples.

This is exactly why sample variance formula uses n−1 in the denominator!

Degrees of Freedom in Statistical Tests

When you run a hypothesis test (like t-test or chi-square), you first have to calculate means, totals, or variances before computing the test statistic.

Every extra thing you calculate in advance = one less degree of freedom.

TestWhat you calculate in advancedf
One sample t-testMean of 1 samplen − 1
Two sample t-testMean of 2 samplesn₁ + n₂ − 2
Chi-square (categories)Expected valuescategories − 1

Why df Matters in Tests?

The p-value depends on BOTH the test statistic AND the degrees of freedom.

Example — Chi-square value of 4:

dfp-value  Significant?
df = 1p = 0.046 Yes (p < 0.05)
df = 2p = 0.135 No (p > 0.05)

Same test statistic , completely different conclusion , just because df changed!

It is like saying "my team won 9 matches":

  • Out of 10 matches → very impressive
  • Out of 100 matches → not so impressive

Without knowing df, the test statistic means nothing.

Graphs : Univariate Analysis (one column at a time)

1. For Categorical Column:

Frequency Distribution Table - count how many times each category appears.

Vacation Type    Count
Beach60
City40
Adventure30

Relative Frequency - convert counts to percentages.

  • Beach = 60/200 = 30%
  • Use this to plot a Pie Chart.

Cumulative Frequency - running total of counts.

  • Beach = 60, Beach+City = 100, +Adventure = 130...
  • Use this to plot a Line Chart.

Bar Chart - plot categories on X-axis, frequency on Y-axis.

2. For Numerical Column:

Histogram - create buckets (bins) and count how many values fall in each.

  • Example: Age column → Buckets: 0−10, 11−20, 21−30...
  • Difference from Bar Chart — bars in histogram touch each other (continuous data). Bars in bar chart have gaps (discrete categories).

Choosing bin size matters:

  • Too large bins → very few bars, lose detail
  • Too small bins → too many bars, noisy
  • Find a good middle ground

Histogram Shapes:

  • Symmetric — most values in center, fewer on both sides (like a bell). This is Normal Distribution.
  • Bimodal — two peaks. Two groups exist in the data.
  • Left Skewed — tail goes left, most data is on the right. Example: marks in a very easy test.
  • Right Skewed — tail goes right, most data is on the left. Example: salaries (most people earn less, very few earn a lot).
  • Uniform — all values appear with roughly equal frequency.

Bivariate Analysis (two columns together)

1. Categorical + Categorical → Contingency Table (Crosstab)

  • Make a table where rows = one category, columns = another category, and cells = count.
  • Example: Survived (0/1) vs Passenger Class (1/2/3) in Titanic data.
  • Then plot a Grouped Bar Chart or Stacked Bar Chart on top of it.

2. Numerical + Numerical → Scatter Plot

  • Plot one column on X-axis and another on Y-axis. Each point = one row.
  • Positive relationship — both increase together. 
  • Negative relationship — one increases, other decreases. 
  • No relationship — random scattered points.

3. Categorical + Numerical → Bar Chart with Aggregation

  • Example: Male vs Female passengers — plot average age of each group.
  • OR convert numerical column into buckets first, then make a Contingency Table.

Covariance & Correlation 

Why do we need Covariance?

We already know:

  • Mean tells us the center of data
  • Variance tells us the spread of one column

But what if we have two numerical columns and we want to study the relationship between them?

Example: Flat area (sq ft) vs Price (lakhs). Does price increase as area increases?

This is where Covariance comes in.

What is Covariance?

Covariance measures the direction of the linear relationship between two numerical columns. It tells you one of three things:

Covariance ValueMeaning
PositiveBoth columns increase together
NegativeOne increases, other decreases
Near ZeroNo linear relationship

Real life examples:

  • Experience and Salary → Positive (more experience = more salary)
  • Backlogs and Package → Negative (more backlogs = less package)
  • Backlogs and Placement luck → Near Zero (no clear pattern)

Formula for Covariance

Population Covariance: $\sigma_{xy} = \frac{\sum (x_i - \mu_x)(y_i - \mu_y)}{N}$

Sample Covariance: $S_{xy} = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{n-1}$

Where x and y are two different numerical columns, and x̄, ȳ are their respective means.

Why n−1 for sample? Same reason as variance - because we already used one degree of freedom to calculate the mean. 

Worked Example 

Data: 5 employees — Experience (X) and Salary in lakhs (Y)

EmployeeExperience (X)Salary (Y)
121
252
386
4128
51310

Mean of X = 8, Mean of Y = 6 (approx for simplicity)

Step 1 — Calculate deviations and their product:

Employeex − x̄y − ȳ(x−x̄)(y−ȳ)
1−6−5+30
2−3−4+12
3000
4+4+2+8
5+5+4+20

Sum of products = 70

Step 2 — Divide by n−1:

$S_{xy} = \frac{70}{4} = 17.5$

Positive covariance → confirms Experience and Salary are positively related.

The Quadrant Trick (Visual Shortcut)

Draw a vertical line at mean of X and horizontal line at mean of Y. This creates 4 quadrants.

        |
   Q2   |   Q1
(−,+)   |  (+,+)
--------+--------  ← Mean of Y
   Q3   |   Q4
(−,−)   |  (+,−)
        |
        ↑ Mean of X
  • Points in Q1 and Q3 → product is positive → Positive covariance
  • Points in Q2 and Q4 → product is negative → Negative covariance
  • Points spread equally in all quadrants → cancel out → Near zero covariance

Just by seeing where most points fall, you can guess the sign of covariance!

Covariance of a Variable with Itself

What if you calculate covariance of X with X?

$Cov(X, X) = \frac{\sum(x_i - \bar{x})(x_i - \bar{x})}{n-1} = \frac{\sum(x_i - \bar{x})^2}{n-1} = \text{Variance of X}$

So covariance of a variable with itself = its own variance. This is a common interview question!

Problem with Covariance

Covariance only tells you the direction (positive or negative). It does NOT tell you the strength of the relationship.

Also covariance is affected by the scale of data. If you multiply both columns by 2, covariance becomes 4 times larger — even though the relationship between the columns has not changed at all.

Example:

  • Original covariance between X and Y = 100
  • Multiply both X and Y by 2 → covariance becomes 400
  • But the scatter plot looks exactly the same!

This makes covariance an unreliable measure for comparing strength of relationships.

What is Correlation?

  • Correlation solves the problem of covariance. It measures both the direction AND the strength of the linear relationship.
  • Correlation is always between −1 and +1. No matter how you scale the data, correlation stays the same.
$r_{xy} = \frac{Cov(X, Y)}{\sigma_x \times \sigma_y}$

Where σx and σy are the standard deviations of X and Y respectively.

By dividing by the standard deviations we normalize the covariance and remove the effect of scale.

Interpreting Correlation Values

Correlation (r)Meaning
r = +1Perfect positive — X increases, Y increases by exact same proportion
r close to +1Strong positive relationship
r = 0.5 to 0.7Moderate positive relationship
r close to 0Weak or no relationship
r = −0.5 to −0.7Moderate negative relationship
r close to −1Strong negative relationship
r = −1Perfect negative — X increases, Y decreases by exact proportion

Worked Example

Using the same data as above:

  • Covariance (X, Y) = 17.5
  • Standard deviation of X = σx (calculate from X column)
  • Standard deviation of Y = σy (calculate from Y column)
  • $r = \frac{17.5}{\sigma_x \times \sigma_y}$
  • If this gives r = 0.82 → Strong positive correlation between experience and salary.

In Python:

df['X'].cov(df['Y'])       # covariance
df['X'].corr(df['Y'])      # correlation

 Covariance vs Correlation 

FeatureCovarianceCorrelation
What it tellsDirection onlyDirection + Strength
Range−∞ to +∞−1 to +1
Affected by scale?YesNo
Reliable measure?Less reliableVery reliable
Used in practice?Rarely aloneAlways preferred

Covariance exists mainly because you calculate it first, then use it to compute correlation.

Correlation Does NOT Mean Causation

This is one of the most important points in all of statistics. Just because two variables are correlated does not mean one is causing the other.

Famous Example 1 — Ice cream and murders: On hot days, both ice cream sales and murder rates go up. They are positively correlated. But eating ice cream does not cause murders! The hidden reason is hot weather — people go outside more, leading to both more ice cream sales and more crimes.

Famous Example 2 — Fire fighters and fire damage: More fire fighters are sent to bigger fires. So number of fire fighters and fire damage are positively correlated. But more fire fighters do not cause more damage — the size of the fire is the real reason.

Famous Example 3 — Experience and Salary: Experience and salary are positively correlated. But you cannot say experience causes higher salary. Other factors like talent, company, market demand also play a role.

To prove causation you need controlled experiments — correlation alone is never enough.

 

    1,914
    1,914 views
    I am currently searching for DPPs (Daily Practice Problems) for my GATE preparation from GO Classes, but I am not sure where or how to access them. I would really appreciate some proper guidance regarding this. I am also a bit confused — are the weekly quizzes provided by GO Classes the same as DPPs, or are they completely different in terms of purpose, difficulty level, and practice strategy?
    2,543
    2,543 views

    Syllabus and Study Material of Written/Programming Test for M.Tech.(RA/RAP), M.S.(RAP), Ph.D./M.S.

    Syllabus of Programming Test for M.S.(RAP) and M.Tech.(RA/RAP) : Syllabus
    Programming Test Question Paper 2023 : Question paper 
     
    Syllabus and Study Materials for M.S./Ph.D.(including Ph.D. panels / M.S. Streams) : Syllabus
     
    The CSE department is divided into three streams. An M.S. by research candidate applies to the streams of the CSE department. (S)he chooses one interview stream in the table below.
     

     Ph.D Panels

    MS Streams

     CS1 : Software Engineering, Compilers, Programming Languages

     CS1 + CS2 = Computing Systems Stream

     CS2 :  Systems Software, Hardware and Security

     

     IS1 :  AI/ML (NLP, Speech, Text)

     IS1+IS2+IS3 = Intelligent Systems Stream

     IS2 : AI/ML (Representation, Learning, Agents)

     

     IS3 : Visual Computing

     

     TS1 : Algorithms, Complexity, Cryptography

     TS1+TS2= Theoretical Systems Stream

    TS2 : Formal methods

     


    Applicant should carefully consider the above choices of streams. They are asked to choose the stream before the test. The subject test will be divided into sections across panels. Applicants can attempt questions in any of the panels of their stream. They will not receive marks for attempting questions outside of their stream. The questions consist of MCQs and fill-in blanks.
     
    Year - MonthShortlisting Test Question Paper for Ph.D./M.S.
    2024 - DecemberQuestion paper with solutions
    2024 - MayQuestion paper with solutions
    2023 - DecemberQuestion paper
    2023 - MayQuestion paper
    2022 - DecemberQuestion paper
    2022 - MayQuestion paper
     
    IIT Bombay M.Tech.(RA)/M.S. Interview Experiences(Many such questions are shared by the applicants in their experiences) : Interview Experiences 
    7,761
    7,761 views

    Almost negligible PYQs are available for this exam (specially for CSE), so I want to share all the memory-based questions of BARC 2025

    1. HASHING (2Q+)
      Index for different keys, Chaining concepts
    2. AVL TREES
      Balance factor of a particular node after removing a node (given)
    3. ROUTING (3Q)
      IP addresses contained in a given subnet, Next hop for a IP address according to routing table
    4. PAGE REPLACEMENT ALGORITHMS (2Q+)
    5. PROCESS SCHEDULING
      RR scheduling question that asked for Average TAT & Average WT
    6. C PROGRAMMING (9Q+)
      Tell o/p of given program, Choose missing line of program, Concept of pointers, arrays, malloc, struct, union & much more
    7. FRAGMENTATION (5Q)
      Maximum data size which can be transferred for a given MTU, Tell data transferred with the help of values of given MF, Offset and some other fields (we multiply the offset by 8 and add the data size if MF=0)
    8. SYNCHRONIZATION (5 to 6Q)
      Semaphores (maximum & minimum no. of times the threads will execute), Does the thread follow bounded waiting, mutual exclusion, progress
    9. FORK() (2Q)
      No. of times print() will be executed 
    10. SORTING
      How will the list look like after 2 pass of merge sort
    11. EXPRESSION EVALUATION 
      Implementation of stack for a postfix expression
    12. LOGIC GATES 
      Identify the gate from the given state transition diagram
    13. PROPOSITIONAL LOGIC (2Q)
    14. SQL QUERY (3Q+)
    15. PNC (basic) (2Q+)
    16. STOP AND WAIT (basic formula implementation/unitary method)
    17. GRAMMAR (3Q+)
      Check if Left Recursive, Check if Ambiguous
    18. DAG (Intermediate Code Generation) (2Q)
      Identify the number of nodes & edges for the given 3 address instructions
    19. K-MAP
      No. of 1's in K-Map for given function
    20. MASTER'S THEOREM (1Q)
    21. CRC (3Q+) (a few theory-based questions regarding error detection)
    22. CYCLOMATIC COMPLEXITY (1Q)
    23. NUMERICAL RELATED TO LOC (1Q)
    24. S/W ENGINEERING (theoretical) (2Q)
    25. GRAPH THEORY (2Q+)
      No. of graphs possible with n vertices, Bipartite Graph
    26. LR PARSER
      Comparison of no. of states in different LR parsers
    27. CN (theoretical) (2Q+)
      Protocols for different layers (matching related question)
    34,164
    34,164 views

    Hi all, below are some free standard video lectures for GATE DA:

    Artificial Intelligence: 

    1.  CS60045 Artificial Intelligence (IIT KGP):

      A semester-length UG/PG level course on AI taken by two of the most renowned professors at IIT Kharagpur.    
      Instructors
       
      Prof. Pallab Dasgupta
      Prof. Partha Pratim Chakrabarti
    2. An Introduction to Artificial Intelligence (IIT Delhi):

      A semester-length UG-level course at IIT Delhi (which has also been added as an NPTEL course) taken by one of the best AI researchers in the country.  
      Instructors
        
      Prof. Mausam

     

    Machine Learning:

          1. Machine Learning Specialization

     

    The most popular course on ML. Everything is explained in simple beginner-friendly language. (Please note that this course can be audited for free in Coursera)                    
    Instructors
             
    Prof. Andrew NG

     

     

    1. Introduction to Machine Learning(Course sponsored by Aricent), IIT Madras

      Another great course                    
      Instructors
               
      Prof. Balaraman Ravindran

     

     

    Statistics and Data Analytics: 

    1.  CS61061 Data Analytics (IIT KGP):

      A semester-length PG-level course on Data Analytics taken by one of the most renowned professors at IIT Kharagpur.    
      Instructors
       
      Prof. Debasis Samanta

     

    Calculus and Optimization:

    1. Basic Calculus for Engineers, Scientists and Economists

    2. Foundations of Optimization, IIT Kanpur

                           
      Instructors
               
      Prof. Joydeep Dutta

    Programming, Data Structures and Algorithms:

    1. Programming, Data Structures and Algorithms using Python, Chennai Mathematical Institute

                           
      Instructors
               
      Prof. Madhavan Mukund

     

     

    I am not adding resources for the subjects already present in GATE CSE as there are already many blogs on gatecse.in and gateoverflow.in. Also, there are some other topics like Data Warehousing but I think some simple Google searches should be enough to learn those topics.
    All the very best and happy learning😊

     

     
    7,833
    7,833 views

    The Defence Research and Development Organisation is the premier agency under the Department of Defence Research and Development in Ministry of Defence of the Government of India, charged with the military's research and development, headquartered in Delhi, India. 

    Question Paper Answer Key Discussions Take Exam
    2022-1 Official Key GO GO
    2022-2 Official Key GO GO

     

    773
    773 views

    https://aofa.cs.princeton.edu/40asymptotic/

    When studying algorithms, this chapter covers methods for obtaining approximate solutions to problems or approaching exact answers, which allow us to generate succinct and precise approximations of quantities of interest.

    hope you find this helpful :)

    1,111
    1,111 views

    https://aofa.cs.princeton.edu/20recurrence/

    chapter focuses on the underlying mathematical aspects of various forms of recurrence relations, which are typically encountered while studying an algorithm by translating a recursive representation of a programme to a recursive representation of a function characterising its attributes.

    2,609
    2,609 views

    computer science

    23/sep/2021  BARC OCES/DGFS 2021   SET 4

     

    mostly question are previous year of gate.  i recall some of them some of them

    if packet transfer from node 1 to node b then which of the following will be change  (ttl  checksum fragment offset).

    3-4 question from  boolean simplified form

    1 question from mux

    1 question of 4 D flip flop  which have an nand gate between 2 LSB flip flop and ask after how much time i can regenerate clock.

    10-12 question of c programming

    mostly theoretical question of each subject,but easy

    5-6 question based on direct mapping associative mapping, find tag, no of line  (tag set block)

    3-4 basic of DBMS relational algebra to language

    1 question from predicate logic   may be    (all student are------ and -----) i cant remember but easy

    paper was easy to  moderate not much hard.

    2,114
    2,114 views

    Computer Networking: A Top-Down Approach

    Topic-wise video Lectures (with slide format notes) by  J.F. Kurose [Author]

    ttps://gaia.cs.umass.edu/kurose_ross/online_lectures.htm 

     

    Book → https://1lib.in/book/3422999/195b8a

    Solution (book) → https://1lib.in/book/6040932/94b2fa

    Data Communications and Network

    Book → https://1lib.in/book/2196530/3bfbaf

    Solution (book) → https://1lib.in/book/1201808/d09f03?dsource=recommend

    58,164
    58,164 views

    Here are the notes I have written during my preparation for GATE CS/IT. Future aspirants may find it helpful. These are elaborated notes, that one may go through to recapitulate stuff fast. I took help from standard books and online courses to put together these notes. As of now, only some chapters of Discrete Mathematics are uploaded. I will upload notes for all the subjects within a month. 

    https://gatecsebyjs.github.io/

    Sample note(on Combinatorics):  

    https://drive.google.com/file/d/1K5DW0rZolcWfJbAMuZvWkvvTuezs4fEP/view?usp=sharing

     

    Peace!

     

    🌟Notes of all subjects uploaded🌟

    Update 1(25/04/2021): Discrete Mathematics completely uploaded

    Update 2(26/04/2021): Computer Organization and Architecture, Compiler Design uploaded

    Update 3(27/04/2021): Theory of Computation, Databases uploaded

    Update 4(28/04/2021): C Programming, Data Structures uploaded

    Update 5(01/05/2021): Digital Logic, Algorithms uploaded

    Update 6(02/05/2021): Computer Networks uploaded

    Update 7(03/05/2021): Operating System Notes, Engg. Maths. materials uploaded

    1,643
    1,643 views
    Can anyone upload a DRDO previous year papers/ link.

    I don’t remember completely but someone had pasted the link on a blog here a year ago.

    Thanks in advance. Gateoverflow is one the most wonderful site I had ever used in my entire Journey to AIR 89.

    ---------------------------------------------------------------------------------------------------------------------------------------------------
    70,797
    70,797 views

    Books I read :
     

    SubjectName and AuthorRelevant Chapters
    Algorithm and DS
    1. Introduction to Algorithms, by CLRS  (3E)
    2. Algorithm Design, Jon Kleinberg and Éva Tardos
    1. ch 1-4, 6-9,10, 11.1-11.4, 12.1-21.3, 15, 16.1-16.3, 17, 21-25.2
    2. ch 1-6
    Discrete Mathematics
    1. Discrete mathematics and its applications by Kenneth H. Rosen (Indian 7E)
    2. Discrete mathematics with applications by Susanna S. Epp (4E)
    3. Concrete Mathematics by Donald Knuth, Oren Patashnik, and Ronald Graham (not required for GATE)
    1. ch 1,2, 4-8, 11.1-11.3
    2. ch 1, 2.1-2.3, 3, 4(optional), 5.1, 5.5-5.7, 6-10, 11-12(optional)
    3. ch 1-3, 5, 7-9
    Computer Networks
    1. Data Communications and Networking by Behrouz A. Forouzan (5E)

    1. 1.1-1.3,  2, 3.6, 8-10, 11.1-2, 12, 13.1-13.2, 17.1, 18-19.2, 20-21.2, 23-24.3, 25.1-25.2, 26
    Theory of Computation
    1. An Introduction to Formal Languages and Automata by Peter Linz (6E)

    1. ch 1.2, 1.3, 2-12, Appendix-A
    Digital Logic
    1. Digital Logic and Computer Design by M. Morris Mano

    1. 1.1-1.8, 2.1-2.7, 3-7
    Computer Organization
    1. Computer Organisation by Carl Hamacher
    2. Computer Organization and Design: the Hardware/Software Interface by David A Patterson and John L. Hennessy (5E)
    1.  ch 1.6, 2.1-2.5, 2.9, 2.10, 4.1-4.2,4.4-4.6, 5.1,5.2, 5.4-5.8, 5.9.1, 6.1-6.4, 6.7.1, 7, 8.1-8.5, 8.8,
    2. 1, 2, 4.1-4.9, 4.14, 5.1-5.10
    C
    1. The C Programming Language by Brian Kernighan and Dennis Ritchie (2E)

    1.  ch 1-8
    Operating System
    1. Operating Systems by Avi Silberschatz, Greg Gagne, and Peter Baer Galvin (International 9E)

    1.  ch 2.1-2.5, 3, 4.1-4.3, 4.6, 5.1-5.3, 6.1-6.10, 7, 8.1-8.6, 91.-9.6, 9.9, 10, 11.1-11.5, 12.1-12.6
    Databases
    1. Fundamentals of Database Systems by Ramez Elmasri and Shamkant B. Navathe (7E)
    1. ch 1.3-1.6, 2.1-2.3, 3, 5-8, 9.1, 14.1-14.5, 14.6-14.7(just overview), 15.1-15.4, 16.1-16.7, 17.1-17.6, 20.1-20.5, 21.1-21.4, 21.7

     

    1. Read every required topic and solved almost all related exercises from these books.
    2. Read every gate-related topic from underlined books, but didn’t solve any exercise problem.
    3. referred to other mentioned text for few topics and few exercise problems
    4. for all the subjects or topics I left above, didn't read any book.

    Video lectures I studied from :

    Subjectreference linkslecture numbers
    Algorithm and DS
    1. Introduction to Algorithms (SMA 5503), MIT OCW
    2. Algorithms by Shai Simonson

    3. Algorithmic Toolbox, Coursera

    4. Algorithm Specialization (Stanford)

    1.  1-7, 15-19
    2.  1-4, 6-8, 11-15     unofficial link
    3. can watch the complete course.
    4. it’s amazing, can watch it all.
    Discrete Mathematics
    1. 1-19 (all except last)  unofficial link
    Theory of Computation
    1. Theory of Computation by Shai Simonson
    1. 1-4, 6,7, 9, 11-13, 15,16, 18, 19 
    Computer Organization
    1. High Performance Computing by Prof. Matthew Jacob IISc

    2. Computer Architecture by Prof. Anshul Kumar IIT Delhi

    1. I would suggest watch all from 1-28 or GO playlist
    2. 7-12, 13-14(optional), 15,16, 17-22(recommended if you have extra time),  23-32, 34-37 

     (lots of extra things in IITD videos, you can skip as per your interest)

    Operating System
    1. High Performance Computing by Prof. Matthew Jacob IISc
    2. Operating Systems by Mythili Vutukuru IIT Bombay
    1. covered in CO section or GO playlist
    Compilers
    1. Compilers by Prof. Alex Aiken Stanford University
    1. week 2-9 (skip Cool Type Checking in week 6)
    Linear Algebra
    1. MIT 18.06 Linear Algebra by Prof. Gilbert Strang
    1. 1-10, 14,16-21
    •   for interviews I suggest you watch it all, it just amazing how intuitively he taught everything.
    Probability
    1. Probabilistic Systems Analysis and Applied Probability  (best lectures acc. to me)
    2. Statistics 110: Probability

    1.  1-8, 13-15
    2.  GO playlist
    Graph Theory
    1. Graph Theory by Dr. L. Sunil Chandran IISc
    1.  1, 2, 7, 9, 13, 15, 17,
    • Please note this is a graduate-level course, if you have less time/interest in this topic avoid watching these lectures, better go with a book.
    Group Theory
    1. Introduction to Abstract Group Theory by Krishna Hanumanthu, CMI
    1. 1-9, 16

     

     

    I also read these notes by Manu Thakur as a revision and just to check in case if I’m missing any topic, they are nicely compiled and only a few topics are not covered by these notes.

    Not everything I mentioned is important for GATE, One should be smart enough to filter out what to read/watch from these references, I haven’t added any extra reference here... I have followed every single reference mentioned, I was enjoying learning specially all these maths lectures, I watched lots of out of the syllabus stuff on these topics. I left it for you to filter the necessary topics at your convenience. One person might need a different approach, go according to what is best for you... don’t restrict yourself just to these... explore more and more.

    I solved NPTEL assignments as well, You can get those just by a Google search, If you couldn’t find them let me know, I’ll add links to those as well.

    and a special Thanks to GO, most of the resources here were recommended by GO, you can find those here: best-books and best-videos

     

    you can find all my GATE related bookmarks here, download this file and open it with any browser, you’ll get some of the additional links I followed plus links to NPTEL courses from where you can download the assignments and official page for many courses I referred, you can check their tests/assignments as well. 

    Most of the questions in these NPTEL assignments are of 1 mark level (still worth trying if you have time, you’ll find lots of interesting things), I first solved all pyqs (including TIFR problems) once before touching any extra question.

    It’s up to you how you wanna use these assignments.


    A nice website collecting most nptel courses and some additional useful links

    1,651
    1,651 views

    Hello Everyone !!!

    I am sharing a document titled "SECURITY OF CYBER SYSTEMS".

    https://docs.google.com/presentation/d/1apDOaHSVJaorwndrV35Sdz5JcjQcyLthYCdZ_9FR1VA/edit?fbclid=IwAR3l9oUnv47Nn59N9Gy4hcYMBj9YNrUMcRTnIOW1wScXJWBr76xkIdJhdKg

    This document is relevant for all aspirants appearing for DRDO Scientist-B 2020 examination as well as for cyber security enthusiasts.

    Initially, I have compiled a list of popular attacks. Based on your feedback I will extend it for rest part of DRDO syllabus.

    Feel free to report errors !!!

     

     

    69,370
    69,370 views

    Download NIELIT PDFs

    https://github.com/GATEOverflow/GO-PDFs/releases/tag/NIELIT

    Scientist  – ‘B’

    ELIGIBILITY: B.E/ B.Tech/ DOEACC B-level OR AMIE/ GIETE OR MSc OR MCA OR ME/ M.Tech OR M.Phil Electronics, Electronics, and Communication, Computer Sciences, Communication, Computer and Networking Security, Computer Application, Software System, Information Technology, Information Technology Management, Informatics, Computer Management, Cyber law, Electronics, and Instrumentation.

    Exam Dates Question Papers Answer Keys GO/AO Links Exam Links
    03 April 2022 2022(CS) : Set – D Answer Key Questions GO  
    05 Dec 2021 2021(CS) : Set – A

    Answer Key

    Changes in Answer Key

    Questions GO  
    05 Dec 2021 2021(IT) : Set – B

    Answer Key

    Changes in Answer Key

    Questions GO  
    22 Nov 2020

    2020(CS) : Set – A

    Answer Key

    Questions GO  
    01 & 02 Dec 2018 2018: Set – B Answer Key Questions GO  
     17 Dec2017 2017: Set – A Answer Key

    Section – A

     

    Section  – B

     
     22 July 2017 2017(CS): Set  – A Answer Key

    Section – A

     

    Section  – B

     
     22 July 2017 2017(IT): Set – A Answer Key

    Section – A

     

    Section  – B

     
    4 Dec 2016 2016(CS): Set – A Answer Key

    Section – A

     

    Section  – B

     
    4 Dec 2016 2016(IT): Set – A Answer Key

    Section – A

     

    Section  – B

    12-13 Mar 2016 2016: Set – A Answer Key

    Section – A

     

    Section  – B

    Section – C

     

    Scientific/Technical Assistant – ‘A’ 

    ELIGIBILITY: B.E/ B.Tech/ M.Sc./ MS/ MCA Electronics, Electronics and Communication, Electronics & Telecommunications, Computer Sciences, Computer and Networking Security, Software System, Information Technology, Informatics.

     Exam Dates Question Papers Answer Keys GO/AO Links Exam Links
    05 Dec 2021 2021(IT) : Set  – D

    Answer Key

    Changes in Answer Key

    Questions GO  
    22 Nov 2020

    2020(CS) : Set – A

    Answer Key

     

    Questions GO  
    17 Dec 2017 2017: Set –  A Answer Key

    Section – A

     

    Section – B

    Take exam
    15 Oct 2017 2017(CS): Set – A Answer Key

    Section – A

    Section – B

    Section – C

     
    15 Oct 2017 2017(IT): Set –  A Answer Key

    Section – A

    Section – B

    Section – C

     
    12-13 Mar 2016 2016(1)- Senior TA Answer Key    
    12-13 Mar 2016 2016(2) -Sample TA Answer Key    
    12-13 Mar 2016 2016(3)- Junior TA Answer Key    

     

    Scientist – ‘C’

    ELIGIBILITY

    For fulfilling the eligibility criteria, a candidate should possess one of the Essential Educational Qualifications.

    • Bachelor degree in Technology or Bachelor degree in Engineering or Associate Member of Institute of Engineers (A&B) (Computer Science or Computer Engineering or Information Technology or Electronics and Communication or Electronics and Telecommunication ) with five years (six years for Associate Member of Institute of Engineers) of relevant work experience in Ministries or Departments or Attached and Subordinate Offices of the Central Government or Statutory Bodies or Autonomous Bodies or Public Sector Undertakings or Private Sectors.
    • Master degree in Science (M.Sc.) (Physics or Electronics or Applied Electronics) with six years of relevant work experience in Ministries or Departments or Attached and Subordinate Offices of the Central Government or Statutory Bodies or Autonomous Bodies or Public Sector Undertakings or Private Sectors.
    • Department of Electronics and Accreditation of Computer Courses (DOEACC) B-Level or Graduate Institute of Electronics and Telecommunication Engineers (IETE) with six years of relevant work experience in Ministries or Departments or Attached and Subordinate Offices of the Central Government or Statutory Bodies or Autonomous Bodies or Public Sector Undertakings or Private Sectors.
    • Master in Computer Application (MCA) with six years of relevant work experience in Ministries or Departments or Attached and Subordinate Offices of the Central Government or Statutory Bodies or Autonomous Bodies or Public Sector Undertakings or Private Sectors.
    Exam Dates Question Papers Answer Keys GO/AO Links Exam Links
    26 feb 2023 2023 Answer Key Questions GO  
    07th Aug 2022 2022 Answer Key Questions AO  
    13/02/2022 2022: Set – A Answer Key Questions AO  
    10 Feb 2019 2019: Set – B Answer Key

     

    Questions AO

     
    12-13 Mar 2016 2016: Set – A Answer Key

    Section – A

    Section – B

    Section – C

     

     

    Scientist – ‘D’

    ELIGIBILITY

    For fulfilling the eligibility criteria, a candidate should possess one of the Essential Educational Qualifications.

    • Bachelor degree in Technology or Bachelor degree in Engineering or Associate Member of Institute of Engineers (A&B) (Computer Science or Computer Engineering or Information Technology or Electronics and Communication or Electronics and Telecommunication ) with eight years (nine years for Associate Member of Institute of Engineers) of relevant work experience in Ministries or Departments or Attached and Subordinate Offices of the Central Government or Statutory Bodies or Autonomous Bodies or Public Sector Undertakings or Private Sectors.
    • Master degree in Science (M.Sc.) (Physics or Electronics or Applied Electronics) with nine years of relevant work experience in Ministries or Departments or Attached and Subordinate Offices of the Central Government or Statutory Bodies or Autonomous Bodies or Public Sector Undertakings or Private Sectors;
    • Department of Electronics and Accreditation of Computer Courses (DOEACC) B-Level or Graduate Institute of Electronics and Telecommunication Engineers (IETE) with nine years of relevant work experience in Ministries or Departments or Attached and Subordinate Offices of the Central Government or Statutory Bodies or Autonomous Bodies or Public Sector Undertakings or Private Sectors.
    • Master in Computer Application (MCA) with nine years of relevant work experience in Ministries or Departments or Attached and Subordinate Offices of the Central Government or Statutory Bodies or Autonomous Bodies or Public Sector Undertakings or Private Sectors.
    Exam Dates Question Papers Answer Keys GO/AO Links Exam Links
    26 feb 2023 2023 Answer key Questions GO  
    13/02/2022 2022: Set – A Answer Key Questions AO  
    10 Feb 2019 2019: Set – B Answer Key Questions AO  
    12-13 Mar 2016 2016: Set – A Answer Key Questions AO  

     

    The official website references: https://nielit.gov.in/recruitments

    5,639
    5,639 views

    total questions recalled: 84 

     

     

    Cn:10 

    1. Router has settings as per the table: 

     

    10.196.1.0/24 

    r2 

    192.10.16.0/22 

    p1 

    192.255.240.0/20 

    r1 

    default 

    r1 

     

    A packet directed to addresses  

    (4 addresses were given and in option there were 4 destinations like: r1,r2,r1,p1) 

     

     

    1. In cidr with 192.10.16/22 how many hosts are possible: 

     

     

     

    1. Why is CSMA/CD not used in wireless networks: 

    A)a station can not differntiate between sending and receeiving signals 

    B)a station can not send and receive at the same time 

    C)signal received is very weak in power 

    D)none of the above 

     

      

      

    4. if in symmetric key encryption 20 parties want to communicate. What is the total number of keys required? 

      

    5. which is false: 

    a)in dvr the count to infinity problem is due to joining of a new node. 

    b) in link state only neighbouring topology is known 

    c) in dvr full network topology is known 

    d) in dvr only packets from neighbours are exchanged 

      

    6. in dvr of A,B,C,D,E,F,G nodes in a network. where from  

         

    7. if a 3-kHz line has a signal to noise ratio of 20dB. Then the maximum data speed is: 

    a)around 19kbits 

    b)3 kbits 

    c)6 kbits 

    d)none 

      

    8. ipv6 uses 128 bits address. So, if one million addresses are allocated every picosecond. then after how many time the addresses will exhaust: 

    a)in power of 10^18 years 

    b)in power of 10^13 years 

    c)never 

    d)very soon 

    9. which of the following digital signatures provide integrity as well as authenticity: 

    A)some sha1 based technique 

    B) sha1 

    C)rsa 

    D)none 

     10.  

    I.Tcp and UDP have same bits for port numbers. 

    II.UDP retransmits damaged packets but not lost packets. 

    Correct option: 

    A)I true and II true 

    B)I true and II false 

    C) I false and II true 

    D) I false and II false 

     

      

      

    maths:6 

    1. If laplace transformation of tf(x) is 1/(s-5). Then the laplace transformation of t3f(x) is: 

    A)2/(s-5)3 

    B)6/(s-5)6 

    C)3/(s-5)7 

    D)6/(s-5)8 

    (please recheck the coefficients in the options as the powers of (s-5), I can remember but not the constants)  

    2. if three boys and 4 girls are to be seated in a row, and all the seating arrangements are equally likely, then what is the probability of no boys together.  

      

    3. a matrix 3*3 was given, and an identity matrix I, then which of the equation is correct 

    (options like 4M-7I =0) 

    4. pnc. all letters of word  BBABCD are allowed maximum as much times as they are in the word. What is the number of all possible words of all possible lengths? 

    a)72 

    b)74 

    c)76 

    d)78 

     

    5. If X is the power-set of A and given X is subset of B. then: 

    A)2^(|A|)<=|B| 

    B)2*|A|=B 

    C)2*|A|>B 

    D) 2^(|A|)>=|B| 

     

    6.  a lattice is a complete lattice. (assertion reasoning type) 

    7.number of spanning tree in connected graph:, n nodes 

    A)2*(n*(n-1)/2) 

    B)2^(n) 

    C)n*(n-1)/2 

    D)n 

    computer graphics:2 

      

    1. find the matrix which correspons to scaling a 3-D object enlarging twice 

      

    2. vanishing point of a figure consisting of 2 cubes, one is in Ist quadrant rotated some small angle clockwise and other in 2nd quadrant rotated some small same angle anticlockwise. 

      

    artificial intelligence:4 

      

    1. 10 nodes in an cnn with 3*3 transformation function, then the number of parameters  

      

    2. sigmoid function - activation function question 

      

      

    3. for which of the following job, the clustering algo is used: 

    a)using the historic data, some conclusion drawn 

    b)e-mails as spam or not 

    c)types of customers based on their activity pattern 

    d)none of these  

      

      

    4. for the accurate prediction of digits in converting paper documents to digitl form: 

    a).....erosion 

    b).....dilution 

    c)boundary ....  

    d)gaussion …...... 

    (….. Are blank spaces, all options had 2 words, I remember only one in each) 

     

     

    c++:6 

    1. a class 

        c{} 

    and two lines: 

    I. c a1; 

    II. c a2=a1; 

    then, how many functions are activated by default. options: 

    ( included different combinations of default constructer, destructer, copy constructer) 

      

    2. c++ question one true,  

      

    3. c++ question including static variable and printing values 

      

    4.  

    #include<iostream> 

    void main() 

    { 

    static int a[]={62,61,9,0}; 

    static int b[][]={{1,2,3,4},{5,6,,78,9}}; 

    int *i=a; 

    printf("%d",i[0][2]); 

    } 

      

    output: 

      

    5. best way to reuse composite modules is: 

    A) 

    B)encapsulation inheritance 

    C)encapsulation composition 

    D)composition inheritance 

     

    6. which of the following is correct; 

    A) ->* operator in c++ can be overloaded 

    B) all operator in c++ can be overloaded 

    C) both a and b 

    D) none 

     

    7. why are abstract class are used over interfaces 

    A) to set concrete behaviour by methods that are used as it is like default by the derived classes 

    B) to instantiate a class 

    C) to have methods that are to be overridden by base classes 

    D)none of these 

    java:1 

    1. output of program simple system.out function related 

      

      

    c:5 

      

    1. #include<stdio.h> 

    int a,b,c; 

    print(void); 

    int main() 

    { 

    int a=0,b,c; 

    a+=print; 

    prinf("%d",a); 

    a++; 

    print(); 

    prinf("%d",a); 

    a+=print(); 

    prinf("%d",a); 

      

      

    } 

      

    void print(void) 

    { 

    static int a=10; 

    prinf("%d",a); 

    return a++; 

    } 

      

    output: 

    2. & unary operator is not applied to which of storage classes 

      

    3.  

    #include<stdio.h> 

    struct str 

    { 

    int a; 

    char b; 

    union {int x; char y;}t; 

      

    } 

    void main() 

    { 

    str s1; 

    s1.t.x=0; 

    si.t.y='a'; 

    printf("%d",s1.t.x); 

    printf("%c",s1.t.y);} 

      

      

    output 

      

    4.  

    #include<stdio.h> 

    void main() 

    {int x=5; 

    char *c =(char *)&x; 

    printf("%d",c); 

      

    } 

      

      

    5. the condition in (c!='a'&&z<+2=4) can be rewritten as: 

    a)!(c!='a'||z+2>4) 

    b)!(c=='a'||z+2<=4) 

    c)!(c!='a'||z+2>=4) 

    d)!(c=='a'||z+2>4) 

      

    computer architecture:4 

      

    1. 4 stage pipeline, with each stage taking 100, 90, 110,70 nanosecond,, with each stage buffer register 5nanosecond. 

    if each stage is taking one clock cycle, then after 1000 data processing, what is the time elapsed? 

      

    2. if in a microprogrammed architecture, 24 bit microinstructions, 13 bit is for micro-operation, 8 multiplexer lines, then what is  

    x: bits for address in micro-instruction 

    y: selection bits 

    z: total number of instructions 

      

    3. cao question of 4 instructions asking the time taken by instructions in number of clock cycles: 

    R2<-M[4998] 

    r1<-M[5000]+r2; 

    R3<-r1 

    M[5002]<-r3. 

    Internal diagram of cpu was shown. 

     

      

      

    4. cao question asking address saved in stack when a set of 4 opertions: 

    r1<-M[1010] 

    r2<-M[1020] 

    r1<- r1+r2; 

    M[2000]<- r1 

      

    during writing to memory an interrupt occured, then what address is at stack. if the given program is starting from 1000. 

      

      

      

      

    physics:1 

      

    1. reflection and colour of light 

      

      

    Os:14 

      

    1. file system:256gb, 200 direct blocks, each word address bits given. what is max. file size in this system? 

    2. a 2-way small cache of 4 blocks is having a refernece string:    8,0,12,6,8,12 

    how many page faults if lru is used. 

    3. if 2-level paging system. memory access time 100ns. TLB is used and tlb hit percentage0 80%. Assume tlb access time 0ns. what is effective access time of the system. 

      

    4. if in shortest seek time first the head is at 100 number cylinder. and the requests are 80,85, 90, 100, 105, 110, 134, 155. After how many request services 85 will be served.  

      

      

    5. semaphore, S=3; 7 up operations and 10 down. What is the current vlue of semaphore S? 

     

     

    6. If number of disks in raid 1 is _____  as raid 6 if(some data duplication and consistency property given): 

    A)same 

    B)2 more than 

    C)2 less than 

    D)1 more than 

    7. Bankers algorithm safe state sequence, if 24 instances of a resource is given. 

     

     

    allocated 

    Max need 

    p0 

    4 

    14 

    p1 

    8 

    10 

    p2 

    10 

    8 

    A)p0p2p1 

    B)p1p2p0 

    C)p2p0p1 

    D)p1p0p2 

     

    8. Deadlock question. Max number of resources needed if 4 processes has need of 4 instances of a resource. 

    9. in a two process scenario, deadlock can occur due to: 

    A)directly knowing of each other by a shared variable 

    B)indirectly dependent on each other 

    C)no deadlock 

    D) os primitives 

     

    10. Priority scheduling, with 0 as the lowest priority. The processes are 

     

    Process with burst time 

    Arrival time 

    priority 

    P1-9 

    0 

    2 

    P2-8 

    2 

    1 

    P3-6 

    6 

    3 

    P4-4 

    8 

    0 

    Using preemptive priority scheduling. Then what is the waiting time of p1: 

    A)10 

    B)2 

    C)12 

    D)14 

     

    11. Choose the incorrect one: 

    A)paging makes memory access slow 

    B)best fit is not really the best algorithm for continuous algorithms in all cases. 

    C)paging suffers from internal segmentation 

    D)Thrashing can be minimized by global page replacement algorithms. 

      

    12. If the reference string with all frames initially empty; (given reference string was same as one of gate questions having 9 page faults in fifo) 

    1. FIFO has 9 page faults with 3 page frames. 

    1. FIFO is better than optimal here 

    Choose the correct option: 

    A)both true 

    B)only I 

    C)both false 

    D)only II 

     

    13. In a disk system with 64 sectors per track, 8 plates with both sides used, 500 cylinders. The addressing of a sector is (a,b,c), where a is the track number from 0-63, b is the plate surface number from 0-15 and c is the cylinder number from 0-499. So, the address (20,8,31) corresponds to which sector in decimal addressing?  

     

    14. Fork system call 

    #include<stdio.h> 

    Void main() 

    { 

    Int I=1; 

    Switch(i) 

    { 
    case 1: 

     Printf(“process”); 

    Fork(); 

    Case 2: 

    Printf(“process”); 

    Fork(); 

    Case 3: 

    Printf(“process”); 

    } 

    } 

    How many times will process be printed? 

     

    dbms:6 

    1. 1. left natural outer join of R and S. 

    R(A,B)={(1,2),(3,1),(2,3),()} 

    S(C,D)={(2,3),(2,1),(1,2),(1,3)}. 

    What is the number of rows returned? 

      

    2. a database is supporting where,having,average, order by, group by, sum. Which of these will be performed first? 

      

    1.  Query output  

    1. 2 sql queries were given.and asked if both same, different, cant say, both wrong. 

    (easy queries involving some where condition and join of two tables) 

     

    1. If a table has A,B,C,D has dependency as A->B and B->D. Then the table is in : 

    A)1st normal form 

    B)2nd normal form 

    C)3rd normal form 

    D)4th normal form 

    1. Given the database table: 

     

    A 

    B 

    C 

    D 

    a1 

    b1 

    c1 

    d1 

    a2 

    b1 

    c2 

    d2 

    a1 

    b2 

    c3 

    d3 

    a2 

    b3 

    c4 

    d3 

     Then which of the following functional dependency is incorrect: 

    A)AC->B 

    B)C-.D 

    C)B->D 

    D)C-.>A 

    web-technology:1 

      

    1. which one of the following is true: 

    a) to prevent sql injection predefined queries are used..... 

    b) user validation ensures buffer not overflow 

    c) forgery of digital signatures is called cross-site..... 

    d)(something related to keys) 

      

    algorithms:9 

      

    1. if out of n elements, k- elements are inserted to form a max-heap,and out of the remaining elements each is first replacing the root and then the heap again rearranges to be a max-heap. What is the time complexity in big -oh terms? 

      

    2. by definition of asymptotic notations if f(x)=n^(1/2) and g(x)=log(n). then which of the following is true: 

    a)f=big-oh(g(x)) only 

    b)f=big-omega(g(x)) only 

    c)f=theta(g(x)) only 

    d)f=big-oh(g(x)) and f=big-omega(g(x)) 

      

    3. if f=big-oh of g(x) then: 

    a) f is increasing for a large x, never faster than g(x) 

    b) f is increasing for a large x,  faster than g(x) 

    c)cant say 

    d) f is increasing for a large x, sometimes faster or slower than g(x) 

    1. Best, average and worst Time complexity of sorting algorithms. Match 

     

    insertion 

    N*n,n*n,n*n 

    selection 

    Nlogn,nlogn,nlogn 

    merge 

    n,n*n,n*n 

    quick 

    Nlogn, nlogn,n*n 

    1.  Which of the following options contains the best sorting algorithm for small elements and large number of elements, respectively: 

    A)insertion-quick 

    B)selection-quick 

    C)merge-radix 

    D)radix-radix 

    1.  Which of the following are greedy algorithms: 

    I.bellman ford 

    II.dijkstra’s 

    III. Floyd warshall 

    IV.prims 

     A)I &II 

    B)II & IV 

    C)II & III 

    D)all of the above 

      

    data structure:6 

      

    1. if 1.2.3.4 are inserted in that order than which of the following pop sequence is wrong: 

      

      

    2. if there are two stacks with X, having 1,2,3 with 1 on top. There is on eother stack Y. If a value is popped from X, it can either be printed or be pushed to Y. But if a value is popped from Y, then it is printed. Then which of the following print sequence is not possible? 

      

    3. B+ tree 

      

      

    4. a binary search tree having 3 internal keyed nodes, then how many ways to represent it? 

    a)6 

    b)1 

    c)5 

    d)4 

     

    5. what is the inorder traversal if postorder and preorder is given 

    6.  an expression x= (a+b)*(c-d). Is given and in the syntax tree, which of the following options do not corresponds to either postorder, preorder or inorder traversal of the tree: 

    A)(inorder traversal given) 

    B) (preorder traversal given) 

    C) (postorder traversal given) 

    D)(incorrect one) 

      

      

    digital electronics:10 

      

    1. if the symbol * means 

    A        B        output 

    F        F        T 

    F        T        T 

    T        F        F 

    T        F        F 

    then AUB can be represented as: 

    a)A*~B 

    b)~A*B 

    c)~A*~B 

    d)none 

      

    2. in the given circuit diagram if s1s2s3=001 and s2s3s4=010 then the output is:(the circuit was a combination of 4 one select line MUX i.e. 2*1 MUX. ) 

      

      

    3. which of the folllowing is correct: 

    a)the range of representation in mantissa and exponent is larger than floating values 

    b)(gray code and excess-3 BCD something..) 

    c)0 has two representations in 2s complement 

    d)none 

    4. input lines A,B,C,D are applied, what is the output if: 

    I. 

    II.  

    III. 

    IV. 

    Are known. 

    (there were 4 conditions on the behaviour of A,B,C,D were given: like A is always true,if B is true A is also true ) 

     5. what is the minimized circuit if A,B,C,D are input and output f(x) is related as: 

    1. B and D always have same values 

    II.A is always 1 if C is 1 

    III. A is in always 0 if B is 1 

    Iv. A is always 1 if D is 0. 

     

    6. given f(x)=Pie(0,1,2,3,4,5,6,7,8,10,12,13,14,15) 

    A)A’B’CD+A’BCD’ 

    B)AB’C’D’+A’B’C’D 

    C) AB’CD+AB’C’D 

    D)A’B’CD+ABC’D’ 

    7. The circuit for ((~A.~B)+(~A.B))+A  is: 

    8. which circuit is same as AUB using only NAND gate, if the inputs A and B in OR gate and X and Y in NAND: 

    A)x=~A & y =~B 

    B) x=~A & y =B 

    C) x=A & y =~B 

    D)none of these 

    (though the options was in symbol of logic gates but I have given the equivalent expressions) 

     

    9. in a serial input parallel output shift register three bit, if the input is 00011011, then the output after 4 clock cycles is: 

    A)101 

    B)001 

    C)111 

    D)100 

     

    1. RAM is an example of: 

    A)combinational 

    B)sequential 

    C)both a and b 

    D)none 

    toc:4 

      

    1. the transition table is given as  

    current state        input        next state 

    A            0        B     

    A            1        C 

    B            0        A 

    B            1        D 

    C            0        D 

    C            1        A 

    D            0        C 

    D            1        B 

    If A is the initial as well as final state, then what is the above finite machine doing: 

    a)accepting languages with odd 0s and even 1s 

    b)accepting languages with odd 0s and odd 1s 

    c)accepting languages with even 0s and even 1s 

    d)accepting languages with even 0s and odd 1s 

      

    2. r= 1(1+0)*  , s=11*0  , t=1*0 

    then which of the following options is correct: 

    a)r is subset of s, r is subset of t 

    b)r is subset of s, t is subset of r 

    c)s is subset of r, s is subset of t 

    d)t is subset of r, r is subset of s 

      

    3. l= (0+1).(0+1).(0+1).(0+1).(0+1)...n times. then what is the number of states in the miniimal dfa accepting it. 

    a)n-1 

    b)n 

    c)n+1 

    d) cant say 

     

     

    4. a finite automata on E={0,1} was given accepting the multiples of 3. options: 

    A)all multiples of 3 

    B)all odd numbers 

    C)all multiples of 4 

    D)all even numbers 

    Compiler desugn:1 

    1. Given an sdt, what is the value at root, if the expression 22*10+11. 

    E->T+E    (E.val = T.val*E.val) 

    E->T  (E.val=T.val) 

    T->F*T   (T.val=F.val+T.val)) 

    T->F    (T.value=F.value) 

    F->num    (F.value=num)