83 83 votes Following table indicates the latencies of operations between the instruction producing the result and instruction using the result. $$\begin{array}{|l|l|c|} \hline \textbf {Instruction producing the result} & \textbf{Instruction using the result }& \textbf{Latency} \\\hline \text{ALU Operation} & \text{ALU Operation} & 2 \\\hline \text{ALU Operation} & \text{Store} & \text{2}\\\hline \text{Load} & \text{ALU Operation} & \text{1}\\\hline \text{Load} & \text{Store} & \text{0} \\\hline \end{array}$$ Consider the following code segment: Load R1, Loc 1; Load R1 from memory location Loc1 Load R2, Loc 2; Load R2 from memory location Loc 2 Add R1, R2, R1; Add R1 and R2 and save result in R1 Dec R2; Decrement R2 Dec R1; Decrement R1 Mpy R1, R2, R3; Multiply R1 and R2 and save result in R3 Store R3, Loc 3; Store R3 in memory location Loc 3 What is the number of cycles needed to execute the above code segment assuming each instruction takes one cycle to execute? $7$ $10$ $13$ $14$ CO & Architecture gateit-2007 co-and-architecture machine-instruction normal + – Ishrat Jahan 46.2k views answer comment Share Follow Print See all 4 Comments 4 4 Comments reply Aayushi Aggarwal commented Oct 10, 2016 reply Follow flag What is the answer ? 13 or 14? 1 1 replyShare Pravin Paikrao commented Oct 26, 2016 reply Follow flag instruction i4 is decrementing i.e arithmatic operation it is betn LOAD and ALU operation so latency is 1 you havnt included that ... 0 0 replyShare GateAnis commented Oct 27, 2016 reply Follow flag No.The explanation given above is correct and answer is 13 only. See--> I3 is ALU operation which uses result of LOAD in I2 , so latency is of 1 cycle (after I2) I4 is ALU operation which uses result of LOAD in I2 , so latency is of 1 cycle (after I2) but I3 is executing at cycle 4. Therefore. I4 will execute at cycle no. 5. I5 is ALU operation using result of ALU in I3 therefore has to wait for 2 cycles after I3 I6 is ALU and uses result of ALU in I5 ,therefore waits 2 cycles 6 6 replyShare Mr-TAGORE commented Feb 27 reply Follow flag https://youtu.be/lUcMdRwdlOw?si=HUbLUolv-kfJLV75 4 4 replyShare Please log in or register to add a comment.
Best answer 239 239 votes In the given question there are $7$ instructions each of which takes $1$ clock cycle to complete. (Pipelining may be used) If an instruction is in execution phase and any other instructions can not be in the execution phase. So, at least $7$ clock cycles will be taken. Now, it is given that between two instructions latency or delay should be there based on their operation. Ex- $1^{st}$ line of the table says that between two operations in which first is producing the result of an ALU operation and the $2^{nd}$ is using the result there should be a delay of $2$ clock cycles. $$\begin{array}{|c|c|c|c|c|c|c|c|c|c|c|c|c|c||} \hline \text {Clock cycle} & \text{T1} & \text{T2}& \text{T3 }& \text{T4} & \text{T5} & \text{T6} & \text{T7}& \text{T8 }& \text{T9} & \text{T10} & \text{T11} & \text{T12}& \text{T13 }\\\hline & \text{I1} & \text{I2} & & \text{I3}& \text{I4 } & & \text{I5} & & & \text{I6} & & & \text{I7}\\\hline \end{array}$$ Load R1, Loc 1; Load R1 from memory location Loc1 Takes 1 clock cycle, simply loading R1 on loc1. Load R2, Loc 2; Load R2 from memory location Loc2 Takes 1 clock cycle, simply loading r2 on loc2. (Add R1, R2, R1; Add R1 and R2 and save result in R1 R1=R1+R2; Hence, this instruction is using the result of R1 and R2, i.e. result of Instruction 1 and Instruction 2. As instruction 1 is load operation and instruction 3 is ALU operation. So, there should be a delay of 1 clock cycle between instruction 1 and instruction 3, which is already there due to I2. As instruction 2 is load operation and instruction 3 is ALU operation. So, there should be a delay of 1 clock cycle between instruction 2 and instruction 3. Dec R2; Decrement R2 This instruction is dependent on instruction 2 and there should be a delay of one clock cycle between Instruction 2 and Instruction 4. As instruction 2 is load and 4 is ALU, which is already there due to Instruction 3. Dec R1 Decrement R1 This instruction is dependent on Instruction 3 As Instruction I3 is ALU and I5 is also ALU so a delay of 2 clock cycles should be there between them of which 1 clock cycle delay is already there due to I4 so one clock cycle delay between I4 and I5. MPY R1, R2, R3; Multiply R1 and R2 and save result in R3 R3=R1*R2; This instruction uses the result of Instruction 5, as both instruction 5 and 6 are ALU so there should be a delay of 2 clock cycles. Store R3, Loc 3 Store R3 in memory location Loc3 This instruction is dependent on instruction 6 which is ALU and instruction 7 is store so there should be a delay of 2 clock cycles between them. Hence, a total of 13 clock cycles will be there. Correct Answer: $C$ Madhab answered Dec 31, 2016 • edited May 16, 2019 by Naveen Kumar 3 Madhab comment Share Follow See all 31 Comments 31 31 Comments reply prasitamukherjee commented Jan 15, 2017 reply Follow flag well explained!! 1 1 replyShare vamp_vaibhav commented May 26, 2017 reply Follow flag best explaination !! 0 0 replyShare VS commented Jul 2, 2017 reply Follow flag I just have one doubt why the Decrement instructions are considered as ALU instructions? Increment and Decrement can also be done using counter circuits without the need of using the ALU for the same. 18 18 replyShare atul_21 commented Jul 18, 2017 reply Follow flag Wonderful explanation 0 0 replyShare bharti commented Aug 17, 2017 reply Follow flag Greatly explained! Thank you 0 0 replyShare krish__ commented Dec 7, 2017 i edited by krish__ Dec 7, 2017 reply Follow flag If DEC was not considered to be an ALU operation then the answer would be 10 because there would be no time dependency between DEC and MPY which would reduce the answer by 2 and the additional delay that was added between DEC and DEC can be reduced. 2 2 replyShare talha hashim commented Apr 23, 2018 reply Follow flag what a lovely explanation 0 0 replyShare Shivam Kasat commented Aug 29, 2018 reply Follow flag why aren't we considering dependencies here!? as there is WAR dependency between I3 and I4! @madhab 1 1 replyShare sudharshan commented Sep 6, 2018 reply Follow flag the first two instructions are using store and load which takes a latency of 0 according to given table in the question.but in your explanation you have taken it as latency 1.why? please reply. 0 0 replyShare suraj777 commented Oct 6, 2018 reply Follow flag plz help ,The difference between Ints4 and Inst2 is 1 right I got it but if you look in the diagram there is a difference of 2 cycles that is T3 and T4 between I4 and I2.plz tell where i get it wrong 1 1 replyShare Mostafize Mondal commented Oct 6, 2018 i edited by Mostafize Mondal Oct 7, 2018 reply Follow flag @sudharshan Already mentioned in question that each instruction takes one cycle to execute. So,for those instructions, you must add one more clock cycle with whatever you got from instruction operation. 3 3 replyShare suraj777 commented Oct 6, 2018 i edited by suraj777 Oct 6, 2018 reply Follow flag @Mostafize Mondal so what are you saying is that T4 is not counted as latency for Instrn.4 if that is the case then how would you define the the number of cycles between Inst5 and Inst3 is it 1 or 2 . Pictorially the difference in no. of cycles between instr2 and Inst4 is similiar to Instr.3 and Instr5 but it is not it is 1 and 2 respectively . plz help thanks 0 0 replyShare Mostafize Mondal commented Oct 7, 2018 reply Follow flag @suraj777 Instruction 4 is dependent on instruction 2.As instruction 2 is load and instruction 4 is ALU.So,there should be delay of one clock cycle between those instructions,which is already there due to instruction 3. We know that CPU executes instructions one by one.Becoz of pipeline,we can do parallel works at a time.So, instruction 4 will complete after completing of instruction 3. 0 0 replyShare rohan30 commented Oct 12, 2018 reply Follow flag Amazing Explanation :) 0 0 replyShare mehul vaidya commented Oct 27, 2018 reply Follow flag Very good & Detailed answer 0 0 replyShare shankargadri commented Dec 18, 2018 reply Follow flag beautiful explanation loved it. 0 0 replyShare gmrishikumar commented Dec 20, 2018 reply Follow flag Why do we consider DEC as an ALU operation? 0 0 replyShare Punit Sharma commented Dec 7, 2019 reply Follow flag @gmrishikumar I think unless specifically mentioned we should consider it as an ALU Operation only... http://www.learnabout-electronics.org/Digital/dig58.php#:~:targetText=The%20ALU%20can%20also%20perform,decrement%2C%20subtract%201%20from%20it.&targetText=Putting%20the%20correct%20pattern%20of,input%20at%20A%20and%20B. 1 1 replyShare Harshii97 commented Dec 30, 2019 reply Follow flag For I3 why we are not considering 2 more cycle delay as we are performing Alu opn and then storing it... According to table we need 2 cycle delay in between them 0 0 replyShare pvijay commented Apr 12, 2020 reply Follow flag 6. MPY R1, R2, R3; Multiply R1 and R2 and save result in R3 R3=R1*R2; This instruction uses the result of Instruction 5, as both instruction 5 and 6 are ALU so there should be a delay of 2 clock cycles. i have a doubt that for MPY instruction uses the result of both DEC R2 and DEC R1, so why did we consider only instruction 5 for delay. shouldn't it be 3 cycles. 2 for instr 4 and 6 and 2 cycle for instr 5 and 6. as instruction 5 is in between so it will take 1 clock so (2-1)+2 = 3. ? 0 0 replyShare o commented Nov 29, 2021 i edited by o Nov 29, 2021 reply Follow flag @Shivam Kasat Because in title/heading of the table it is written that latency is to be considered when an “instruction is using the result” from “instruction producing the result” In I4, R2 is being used by I4, which is also used by I3, but it is not a result produced by I3. 0 0 replyShare ankit3009 commented Dec 5, 2021 i edited by ankit3009 Dec 5, 2021 reply Follow flag I4 MUST be taking 2 cycles right (1 execution cycle + 1 cycle for operation between load and ALU instruction), then why you have allotted only 1 cycle? I am totally fine with other operation. But confused about I4 taking only 1 cycle. NOTE : It’s mentioned that each instruction takes 1 execution cycle. For me it’s : (execution time + time consumed between 2 instructions) : Total time : ((1+[0])[load(I1) + store(I1)] + (1+[0])[load(I1) + store(I1)] + (1+[1])[load(1) + ALU(I3)] + (1+[0])[load(I2) + store(I4)] + (1+[2])[store(I5) + ALU(I3)] + (1+[2])[store(I4, I5) + ALU(I6)] +(1+[2])[store(I7) + ALU(I6)]) = 14 cycles. Decrement operation can be done using hardware of which executed value can be then stored directly. Option (D) Please help where I am going wrong. 0 0 replyShare Manisha Jaishwal commented May 6, 2022 reply Follow flag I am having the same doubt ....we don't consider it because of parallel execution ?? Means both the operands are fetched parallely so the maximum is 2 cc ?? @Arjun Sir please clarify ... 0 0 replyShare svas7246 commented Aug 30, 2022 reply Follow flag Simply superb 0 0 replyShare Pranavpurkar commented Dec 6, 2022 reply Follow flag Very well explained ! Thanks @Madhab. 0 0 replyShare Kshitij Sharma commented Jan 1, 2025 reply Follow flag Bohot pyara likha hai Bhai. 0 0 replyShare Tejaswee_Bommaluleni commented Jun 16, 2025 reply Follow flag This is the best answer ! I wasted my 1 hour time by ignoring this answer and searching in some other sources😵 3 3 replyShare Binit_Gudhka commented Jul 1, 2025 reply Follow flag great explanation bro 1 1 replyShare Yedilmaagemore commented Jul 24, 2025 reply Follow flag bhai literally you turned that complex logic into simple logic 0 0 replyShare N210479_GEDELA_LEELA commented Dec 20, 2025 reply Follow flag As I6 is using result of I5 so there should be delay of 2 cycles , at the same time I6 is also using result of I4 , as both are ALU operations there should be 2 delay , out of 1 delay is already ther due to I5 so there should be extra one delay i.e at I6 there should be 2+1 = 3 extra dely ??? 0 0 replyShare EagerLearner commented Jul 22 reply Follow flag N210479_GEDELA_LEELANo, it's the overall delay that needs to be counted...Just think R2 needs to wait only 1 cycle and R2 needs to wait for 2 cycles....Since 1 delay is already there due to I5, after another delay(clock cycle) R2 is available for us to use independently...So we only have to wait for one more clock cycle so that R1 is also available 0 0 replyShare Please log in or register to add a comment.
21 21 votes Answer is (C)Here each instruction takes $1$ cycle but apart from that we have to consider latencies b/winstruction: If there are two ALU operations by $I1$ and $I2$ such that $I2$ uses the valueproduced by $I1$ in some register then $I2$ will be executed ONLY after waiting TWO morecycles after $I1$ has executed because latency b/w two ALU operations is $2$See here:Clock12345678910111213InstructionI1I2-I3I4-I5--I6--I7 $I3$ is ALU operation which uses the result of LOAD in $I2,$ so latency is of $1$ cycle.$I5$ is ALU operation using result of ALU in I3, therefore, has to wait for $2$ cycles after $I3$$I6$ is ALU and uses result of ALU in I5, therefore waits for $2$ cycles Sandeep_Uniyal answered Jan 20, 2015 • edited Jan 26 by Shubham Sharma 2 Sandeep_Uniyal comment Share Follow See all 5 Comments 5 5 Comments reply Show 2 previous comments GateAnis commented Oct 21, 2016 reply Follow flag I3 is ALU operation which uses result of load in I2 hence latency should be two cycles I4 is also ALU opn which uses result of load in I2 hence latency 2 I5 is ALU - ALU hence 2 I6 latency 2 I7 latency 1 In this way answer comes to be 16. Can you please clear my doubt sir? 0 0 replyShare Pravin Paikrao commented Oct 26, 2016 reply Follow flag instruction i4 is decrementing i.e arithmatic operation it is betn LOAD and ALU operation so latency is 1 you havnt included that . 0 0 replyShare Thanneeru_Venkateswa commented Dec 26, 2025 reply Follow flag Short and Crisp answer 0 0 replyShare Please log in or register to add a comment.
14 14 votes cool answer...>> blackcloud answered Feb 3, 2020 blackcloud comment Share Follow See all 2 Comments 2 2 Comments reply Vandit Shah commented Jun 18, 2021 reply Follow flag totally not cool. For such a good answer to be at bottom! 2 2 replyShare ankita _Mishra commented Jan 5, 2023 reply Follow flag beautifully explained :) 0 0 replyShare Please log in or register to add a comment.
2 2 votes As per the given table and the assumption that The instruction of Type R1 <- R1 + R2 requires 2 x (load - alu op) type1 x (alu op - store) type Then the answer should be 14 Analysis: Instr | Cost ==== ==== 1. 0 2. 0 3. 1+1+2 = 4 4. 1+2 = 3 5. 1+2 = 3 6. 1+1+2 = 4 7. 0 AND if, R1 <- R1+R2 is only an (alu-store) type.. then intrs. 3 and 6 take 2 time units each resulting in the answer as 10.. Unless, somebody presents a different interpretation.. Ravi Thakur answered Jan 8, 2015 Ravi Thakur comment Share Follow See all 2 Comments 2 2 Comments reply Vicky Bajoria commented Jan 9, 2015 reply Follow flag I m not sure between (A) and (B).. If go with table.. it will come out to be 10 clocks, assuming 1 latency is 1 clock. Next it is clearly given that "Assuming each instruction takes one cycle to execute".. here there are 7 instructions.. So 7 cycle.. isn't it?.. So which should we follow the table or the statement that given..? 0 0 replyShare sushant suman commented Aug 2, 2018 reply Follow flag made easy soluion 1 1 replyShare Please log in or register to add a comment.
2 2 votes explanation using clock cycle graph Mohitdas answered Oct 11, 2021 Mohitdas comment Share Follow 0 reply Please log in or register to add a comment.
0 0 votes Video Analysis (Good Analysis)https://youtu.be/lUcMdRwdlOw?si=jZsmWFuXRFiB_HTc Rishabh_Chaudhari answered Jul 29 Rishabh_Chaudhari comment Share Follow 0 reply Please log in or register to add a comment.