0 0 votes Suppose we are performing gradient descent to minimize the empirical risk of a linear regression model $y=\beta_0+\beta_1 x_1+\beta_2 x_1^2+\beta_3 x_2$ on a dataset with 100 observations. Let $\mathcal{D}$ be the number of components in the gradient. What is $\mathcal{D}$ for the gradient used to optimize this linear regression model ?2348 Machine Learning goclasses goclasses-da-course machine-learning gradient-descent calculus + – GO Classes 222 views answer comment Share Follow Print 0 reply Please log in or register to add a comment.