A team prunes 80 percent of individual weights from a recommender model, but GPU inference latency barely changes. Profiling shows dense kernels still execute full matrix operations. What should the team conclude before claiming a serving-speed improvement?
Your Answer
Explanation
FREE AI FEEDBACK
Guest checks show correctness and the answer key. Log in to save history and unlock evaluator notes.
i
Review the explanation and try similar questions to strengthen your understanding.
Issue reporting
Report an issue
Clara
How can I help?
Search ARTiBA AI Certs Reviewer
Search across lessons, syllabus topics, provider capabilities, and certification questions.