321 views
2 2 votes

A bank wants to predict whether a customer will default on a loan using a classification model. The bank collects data from past customers, which includes the following features:

  • $x_1$: Credit score (scaled between $0$ and $1$)
     
  • $x_2$: Monthly income (in $\$ 1000 \mathrm{~s}$)
     
  • $x_3$: Number of late payments in the past year

The target variable $y$ represents whether the customer defaulted ( 1 for default, 0 for no default). The data is provided in the table below:

\begin{array}{|c|c|} \hline \mathbf{x} & \mathbf{y} \\ \hline [0.9,5,1] & 0 \\ \hline [0.4,3,5] & 1 \\ \hline [0.8,4,2] & 0 \\ \hline [0.2,2,7] & 1 \\ \hline [0.7,6,0] & 0 \\ \hline \end{array}

The bank uses the following linear combination to compute $z$:

$$
z=0.5 x_1+0.3 x_2-0.2 x_3
$$


The step function $u(z)$ is used to classify the outcome:

$$
u(z)= \begin{cases}1 & \text { if } z \geq 1.5 \\ 0 & \text { otherwise }\end{cases}
$$


If the threshold in the step function is changed to $z \geq 1.0$, what will happen to the misclassification rate?

  1. Increase
     
  2. Decrease
     
  3. Remain the same.
     
  4. Cannot be determined

2 Answers

1 1 vote

To determine what will happen to the misclassification rate, we need to calculate it for both the original and the new threshold.

The formula to calculate the score $z$ is: $z=0.5 x_1+0.3 x_2-0.2 x_3$
First, let's calculate the $z$ score for each customer in the dataset.

  • Customer 1: $[0.9,5,1]$

$$
z=0.5(0.9)+0.3(5)-0.2(1)=0.45+1.5-0.2=1.75
$$

  • Customer 2: $[0.4,3,5]$

$$
z=0.5(0.4)+0.3(3)-0.2(5)=0.2+0.9-1.0=0.1
$$

  • Customer 3: $[0.8,4,2]$

$$
z=0.5(0.8)+0.3(4)-0.2(2)=0.4+1.2-0.4=1.2
$$

  • Customer 4: $[0.2,2,7]$

$$
z=0.5(0.2)+0.3(2)-0.2(7)=0.1+0.6-1.4=-0.7
$$

  • Customer 5: $[0.7,6,0]$

$$
z=0.5(0.7)+0.3(6)-0.2(0)=0.35+1.8-0=2.15
$$


Step 1: Calculate the Original Misclassification Rate (Threshold $\geq 1.5$ )
Let's find the model's prediction for each customer using the original rule (Predict 1 if $z \geq 1.5$, otherwise 0 ) and compare it to the actual outcome $y$.

\[
\begin{array}{|c|c|c|c|c|}
\hline \textbf{Customer} & \textbf{z-score} & \textbf{Actual (y)} & \textbf{Prediction } (z \geq 1.5) & \textbf{Correct?} \\
\hline 1 & 1.75 & 0 & 1 & No \\
\hline 2 & 0.1 & 1 & 0 & No \\
\hline 3 & 1.2 & 0 & 0 & Yes \\
\hline 4 & -0.7 & 1 & 0 & No \\
\hline 5 & 2.15 & 0 & 1 & No \\
\hline
\end{array}
\]

With the original threshold, there are $\mathbf{4}$ misclassifications out of $\mathbf{5}$ data points.
The misclassification rate is $4 / 5 = 80\%$.
 

Step 2: Calculate the New Misclassification Rate (Threshold $\geq 1.0$ )
Now, let's use the new rule (Predict 1 if $z \geq 1.0$, otherwise 0 ) and find the new number of misclassifications.

\[
\begin{array}{|c|c|c|c|c|}
\hline \textbf{Customer} & \textbf{z-score} & \textbf{Actual (y)} & \textbf{Prediction } (z \geq 1.0) & \textbf{Correct?} \\
\hline 1 & 1.75 & 0 & 1 & No \\
\hline 2 & 0.1 & 1 & 0 & No \\
\hline 3 & 1.2 & 0 & 1 & No \\
\hline 4 & -0.7 & 1 & 0 & No \\
\hline 5 & 2.15 & 0 & 1 & No \\
\hline
\end{array}
\]

With the new threshold, there are $\mathbf{5}$ misclassifications out of 5 data points. The new misclassification rate is $5 / 5=100\%$.

The original misclassification rate was $80 \%$, and the new rate is $100\%$. Therefore, the misclassification rate will increase.

The correct option is A. Increase.

Answer:
Position:
Show:

Related questions

1 1 vote
1 1 answer
301
301 views
GO Classes asked Sep 22, 2025
301 views
A bank wants to predict whether a customer will default on a loan using a classification model. The bank collects data from past customers, which includes the following f...
1 1 vote
2 2 answers
301
301 views
GO Classes asked Sep 22, 2025
301 views
Why is data preprocessing necessary before using it for model building?The data may contain outliers or missing values due to errors in data Features may have different s...
1 1 vote
1 1 answer
290
290 views
GO Classes asked Sep 22, 2025
290 views
What is the primary risk of not performing a train-test split in a machine learning workflow?The training time will increase. The model might overfit the training data an...
3 3 votes
1 1 answer
469
469 views
GO Classes asked Sep 22, 2025
469 views
Which of the following is the best second-order polynomial that fits the dataset by the least sum of squares method?$$\min _{\theta_0, \theta_1, \theta_2} \sum_{i=1}^3\le...