$Y = X\beta + \epsilon$
$\epsilon$ (The Noise): The question tells us this is random and has a Multivariate Normal
Now,
$\hat{\beta}_n = (X^T X + \lambda I_p)^{-1} X^T Y$
$X^T X$ is always Positive Semi-Definite (like $\ge 0$).
$\lambda I_p$ (with $\lambda > 0$) is Positive Definite (like $> 0$).
Rule: (Non-negative) + (Positive) = Positive.
$\lambda > 0$ inside inverse $\Rightarrow$ ``Shrinkage/Pull'' toward 0.
Shrinkage $\ne$ Exact recovery $\Rightarrow$ Biased.
$X^T X \ge 0$. Adding $\lambda > 0$ guarantees it is strictly $> 0$ (Positive Definite).
$X^T X$ :
perfect square ($p\times p$).
perfectly symmetric.
Math: $(X^T X)^T = X^T X$.
Always ``Positive'' (Positive Semi-Definite).
eigenvalues are always $\ge 0$.
$(X^T X + \lambda I_p)$
The $-1$ outside means this is acting like a denominator (we are dividing by it).
$X^T X$: In matrix math, squaring something ($X$ times $X$) means it can never be negative. It is $\ge 0$.
$\lambda$: The question specifically states $\lambda > 0$. It is strictly positive.
$I_p$: This is the ``Identity Matrix.'' It is the matrix equivalent of the number 1.
The Equation Logic:
(Zero or higher)+(Strictly Positive Number$\times$1) = Strictly Positive = Positive Definite.
Perfect $\hat{\beta} = (X^T X)^{-1} X^T Y$
\[
\hat{\beta}_n = \underbrace{(X^T X + \lambda I_p)^{-1} X^T}_{\text{Constant Matrix } A} \cdot Y
\]
$X$ (Your Input Data): Once you collect it (like age, height), it is locked in.
$\lambda$ (Penalty): A fixed number you choose.
$I_p$ (Identity): Just fixed 1s and 0s.
Since $Y = X\beta + \epsilon$, the master equation is actually this:
$\hat{\beta}_n = \text{No $\epsilon$ here (Constant)}[(X^T X + \lambda I_p)^{-1} X^T] \times \text{The $\epsilon$ is here (Random)}(X\beta + \epsilon)$
if a chunk of math doesn't contain the random noise ($\epsilon$), we label that entire chunk a Constant.
$\hat{\beta}_n = (\text{Some Matrix}) \times Y$.
Rule:
Matrix$\times$Normal Variable=Normal Variable.
(C) Shape: Constant numbers $\times$ Normal ($Y$) = Normal. $\rightarrow$ True
Standard Regression (Perfect Match):
$[X^T X]^{-1} \times [X^T X] \times [X^T X]^{-1}$
Pattern: Outer matches Inner.
Result: It perfectly cancels down to just one $[X^T X]^{-1}$.
Rule:
\[
\operatorname{Var}(A Y) = A \operatorname{Var}(Y) A^T \quad \text{(Sandwich Rule)}
\]
To find $A^T$, use the universal matrix rule: Swap the order, flip the pieces.
$A = \text{Part 1} (X^T X + \lambda I)^{-1} \times \text{Part 2} X^T$
Step 1 (Swap): Bring Part 2 to the front. Push Part 1 to the back.
Step 2 (Flip):
Flip Part 2: $(X^T)^T \to X$
Flip Part 1: Because it's perfectly symmetric, flipping it does nothing. It stays $(X^T X + \lambda I)^{-1}$
Final Answer of $A^T$:
$A^T = X (X^T X + \lambda I)^{-1}$
This means : Outer pieces have a penalty: $(X^T X + \lambda I)$
Inner piece has no penalty: $(X^T X)$
$[X^T X + \lambda I]^{-1} \times [X^T X] \times [X^T X + \lambda I]^{-1}$
The inner piece is missing the ``$+\lambda I$''. Because they are not exact clones, they cannot cancel each other out.
Option D :
The option gives a simple inverse (no sandwich)
$\rightarrow$ Impossible.
Correct = B + C