131 views
0 0 votes

In a LASSO Regression, if the regularization parameter $\lambda$ is very high, which of the following is true?

  1.  The model can shrink the coefficients of uninformative features to exactly $O$.
  2.  The loss function is as same as the ordinary least square loss function.
  3.  The loss function is as same as the ridge regression loss function.
  4.  The bias of the model is no lower than the bias of the model with a smaller $\lambda$.

1 Answer

0 0 votes
The loss function for LASSO Regression is defined as the Ordinary Least Squares (OLS) loss plus an $l_1$ regularization penalty:
 
$$J(w) = \text{RSS}(w) + \lambda \sum_{j=1}^{p} \vert{}w_j\vert{}$$
where $\lambda$ dictates the strength of the penalty applied to the absolute values of the model's coefficients.
 
Here is the evaluation of each statement when $\lambda$ is very high:
 
  • A. The model can shrink the coefficients of uninformative features to exactly $O$: This statement is true. The defining characteristic of the $l_1$ penalty in LASSO is its ability to perform automatic feature selection. Because of the sharp vertices of the $l_1$ norm's geometric constraint space (a cross-polytope), the optimization naturally pushes the coefficients of less relevant features to exactly zero, yielding a sparse model. This shrinking effect becomes more aggressive as $\lambda$ increases. (Note: The image uses an italic 'O' which represents the number zero in this context).
  • B. The loss function is as same as the ordinary least square loss function: This statement is false. The OLS loss function consists only of the residual sum of squares (RSS) without any penalty term. The LASSO loss function is only identical to OLS when $\lambda = 0$. Since the problem states $\lambda$ is "very high", the $l_1$ penalty term fundamentally changes the objective function.
  • C. The loss function is as same as the ridge regression loss function: This statement is false. Ridge regression applies an $l_2$ penalty (the sum of squared coefficients, $\lambda \sum w_j^2$), whereas LASSO applies an $l_1$ penalty (the sum of absolute coefficients, $\lambda \sum \vert{}w_j\vert{}$).
  • D. The bias of the model is no lower than the bias of the model with a smaller $\lambda$: This statement is true. According to the bias-variance tradeoff, increasing the regularization parameter $\lambda$ strongly restricts the model's flexibility. This forces the model away from the true underlying data relationship toward zero, reducing variance but steadily increasing bias. Therefore, a model with a highly constrained complexity (very high $\lambda$) will have a bias that is strictly greater than or equal to (i.e., "no lower than") a model with a smaller $\lambda$.
The tags in "image_e5069d.png" indicate this is a #multiple-selects question. Based on the evaluation, the correct statements are A and D.
Answer:
Position:
Show:

Related questions

0 0 votes
1 1 answer
142
142 views
GO Classes asked Mar 17, 2025
142 views
Which of the following statements are true about Lasso and ridge regression?Both ridge regression and Lasso are methods used to reduce overfitting that might occur in sta...
0 0 votes
1 1 answer
177
177 views
GO Classes asked Mar 17, 2025
177 views
Which of the following statements are true about Lasso and ridge regression?Both ridge regression and Lasso are methods used to reduce overfitting that might occur in sta...
0 0 votes
1 1 answer
129
129 views
GO Classes asked Mar 13, 2025
129 views
Good practices to avoid overfitting include:  Using a two part cost function which includes a regularizer to penalize model complexity.Using a good optimizer to minimize ...
0 0 votes
0 0 answers
131
131 views
GO Classes asked Mar 13, 2025
131 views
If the model resulting from Ridge regression is currently overfitting, what are possible things to reduce overfitting ? Select all that apply. Collect new data to increas...