Recent questions tagged machine-learning

0 0 votes
0 0 answers
118
118 views
In LASSO regression, if the regularization parameter $\lambda$ is very large and two informative featuresare highly collinear (i.e., that there exists an $\alpha$ such th...
0 0 votes
0 0 answers
206
206 views
Please assign each plot in Figure above to one (and only one) of the following regularization methods. Please answer A , B, C or D for each statements.No regularization ...
0 0 votes
0 0 answers
142
142 views
Let $x \in \mathbb{R}^n$ be the data matrix and $y \in \mathbb{R}^{n \times d}$ be the observed labels for the $n$ examples. Let $w \in \mathbb{R}^d$ be the weight parame...
0 0 votes
1 1 answer
189
189 views
Suppose you're using L2 regularization on a least squares objective. Some value $\lambda^*$ will give you the best test error among all possible $\lambda$. You train your...
0 0 votes
1 1 answer
177
177 views
Which of the following statements are true about Lasso and ridge regression?Both ridge regression and Lasso are methods used to reduce overfitting that might occur in sta...
0 0 votes
1 1 answer
132
132 views
Ridge regression can shrink all coefficients to exactly 0 if the regularization parameter $\lambda$ is large enough.(Please enter 1 for True and 0 for False).
0 0 votes
1 1 answer
167
167 views
With L1-regularization, which vector would we choose?$w_1=\left[\begin{array}{l}100 \\ 0.02\end{array}\right]$ $w_2=\left[\begin{array}{c}100 \\ 0\end{array}\right]$ $w_3...
0 0 votes
1 1 answer
183
183 views
With L2-regularization, which vector would we choose?$w_1=\left[\begin{array}{l}100 \\ 0.02\end{array}\right]$ $w_2=\left[\begin{array}{c}100 \\ 0\end{array}\right]$ $w_3...
0 0 votes
0 0 answers
124
124 views
Suppose we are minimizing $J^{\prime}(\boldsymbol{\theta})$ where$$J^{\prime}(\boldsymbol{\theta})=J(\boldsymbol{\theta})+\lambda r(\boldsymbol{\theta})$$As $\lambda$ inc...
0 0 votes
1 1 answer
132
132 views
$\text { What is the best value for lambda? }$$$\hat{\boldsymbol{\theta}}=\underset{\boldsymbol{\theta}}{\operatorname{argmin}} J(\boldsymbol{\theta})+\lambda r(\boldsymb...
0 0 votes
1 1 answer
131
131 views
Which model do you prefer, assuming both have zero training error?Model structure (for both models):$$h_{\boldsymbol{\theta}}(x)=\theta_0+\theta_1 x+\theta_2 x^2+\theta_3...
0 0 votes
1 1 answer
138
138 views
Consider a new objective function with an added regularization term:$$J_3\left(\theta, \theta_0\right)=\frac{1}{n} \sum_{i=1}^n\left(\theta^{\top} x^{(i)}+\theta_0-y^{(i)...
0 0 votes
1 1 answer
131
131 views
Suppose you are interested in predicting a one-dimensional quantitative random variable $Y$ (outcome) in terms of a two-dimensional quantitative random variable $X$ (cova...
0 0 votes
1 1 answer
180
180 views
Suppose you are interested in predicting a one-dimensional quantitative random variable $Y$ (outcome) in terms of a two-dimensional quantitative random variable $X$ (cova...
0 0 votes
0 0 answers
220
220 views
Which of the following is not a true statement about gradient descent (GD) vs. stochastic gradient descent (SGD)?Both provide unbiased estimates of the true gradient at e...
0 0 votes
0 0 answers
221
221 views
Suppose we are performing gradient descent to minimize the empirical risk of a linear regression model $y=\beta_0+\beta_1 x_1+\beta_2 x_1^2+\beta_3 x_2$ on a dataset with...
0 0 votes
1 1 answer
175
175 views
When the algorithms converge, stochastic gradient descent always finds the same solution as gradient descent.(Please enter 1 for True and 0 for False)
0 0 votes
1 1 answer
154
154 views
When $N$ is large, we typically use a small subset of the dataset to estimate the gradient - stochastic gradient descent (SGD). Explain why we use SGD instead of gradient...
0 0 votes
0 0 answers
120
120 views
Convexity is a desirable property in machine learning because it:guarantees gradient descent finds a global minimum in optimization problems for functions that have a glo...
0 0 votes
0 0 answers
126
126 views
Which of the following are true about gradient descent? (select all statements that are true.)After each iteration, we modify the weight vector in the direction of the gr...
0 0 votes
1 1 answer
156
156 views
Suppose that we are given $f(x)=x^3+x^2$ and learning rate $\alpha=1 / 4$.First of all, write down the updating rule for gradient descent in general and for this function...
0 0 votes
1 1 answer
157
157 views
Let $f: \mathbb{R} \rightarrow \mathbb{R}$ be a continuous, smooth function whose derivative $f^{\prime}(x)$ is also continuous. Suppose $f$ has a unique global minimum $...
0 0 votes
1 1 answer
175
175 views
Let $f: \mathbb{R} \rightarrow \mathbb{R}$ be a continuous, smooth function whose derivative $f^{\prime}(x)$ is also continuous. Suppose $f$ has a unique global minimum $...
0 0 votes
1 1 answer
198
198 views
Consider the following loss function based on data $x_1, \ldots, x_n$ with mean $\bar{x}$ :$$\ell(\beta)=\log \beta+\frac{\bar{x}}{\beta}+\frac{1}{n} \sum_{i=1}^n e^{-x_i...
0 0 votes
0 0 answers
147
147 views
Consider the following function of $f(\theta)$, which alternates between completely flat regions and regions of absolute slope equal to 1 . For each of the following ques...
0 0 votes
1 1 answer
163
163 views
Which of the following could affect whether or not gradient descent converges to the global minimum?Learning RateInitialization of parametersThe ordering of the data (ign...
0 0 votes
0 0 answers
113
113 views
For the dataset $\mathcal{D}=\left\{\left(x_i, y_i\right)\right\}_{i=1}^n$ and the loss function:$$L(w)=\frac{1}{n} \sum_{i=1}^n\left(y_i-\sin \left(w_0+w_1 x_i\right)\ri...
0 0 votes
0 0 answers
141
141 views
The learning rate can potentially affect which of the following? Select all that apply. Assume nothing about the function being minimized other than that its gradient exi...