A company uses Amazon SageMaker AI to generate article summaries in multiple languages. The company needs a metric to evaluate the quality of the summary translations in multiple languages. Which evaluation metric will meet these requirements?
- ARecall-Oriented Understudy for Gisting Evaluation (ROUGE)
- BBilingual evaluation understudy (BLEU) (correct answer)
- CArea Under the ROC Curve (AUC)
- DPrecision
Reveal answer & explanationHide answer
The correct answer is B. Option B: Bilingual evaluation understudy (BLEU).