Google Cloud Professional Machine Learning Engineer PMLE — Question 314
Topic 1 · Question 314 of 339
Topic 1 · Question 314
You are training a large-scale deep learning model on a Cloud TPU. While monitoring the training progress through Tensorboard, you observe that the TPU utilization is consistently low and there are delays between the completion of one training step and the start of the next step. You want to improve TPU utilization and overall training performance. How should you address this issue?