1,115 views

2 Answers

1 1 vote

Model 1 / w1 will be a better generalizing on test dataset. Model 2/w2 will fail because of two major problems:

  1. Grads will explode / vanish due to presence of square term
  2. Cube preserves sign, therefore there will be an absence of lower bound, meaning the optimization can even go to -inf as the lowest value
0 0 votes
option a :

1)for squared loss it will be always >=0

2)for cube loss ,it can go positive or negative so it can go towards -inf

3)so optimization may not have the finite minimum
Position:
Show:

Related questions

1 1 vote
1 1 answer
501
501 views
GO Classes asked Mar 18, 2025
501 views
Consider linear regression and logistic regression. They both use linear functions.They both can be used to solve regression prob-They both use the logistic activation fu...
1 1 vote
1 1 answer
246
246 views
GO Classes asked Mar 13, 2025
246 views
Consider a simple classification task with features $x_i \in \mathbb{R}^d$ and $k$ classes. Suppose we train a linear classifier to minimize the regularized least squared...
1 1 vote
1 1 answer
245
245 views
GO Classes asked Mar 13, 2025
245 views
Consider just the general shape of the following plots.For each of the following possible interpretations of the quantities being plotted on the $X$ and $Y$ axes, indicat...
1 1 vote
0 0 answers
184
184 views
GO Classes asked Mar 13, 2025
184 views
Consider the following training data: xy132130.5Suppose the data comes from a model $y=c x^\beta$ , for unknown constants $c$ and $\beta$. Use least squares linear regres...