To meet these performance targets, organizations need a systematic approach to AI evaluation and real-time bottleneck identification. That is why we are sharing the automated evaluation framework and performance optimization techniques that helped KDDI successfully launch their application.
The results were inspiring: KDDI reduced total application response latency by 38%, successfully hitting their target response performance. They also achieved a nearly 18% improvement in TTFT.
“Our vision hinged on a platform where content, once ingested, would instantly function as a working RAG system. Google’s careful, hands-on guidance made that a reality — we’re sincerely grateful for their support.” — Shunya Onoda, AI Product Department, KDDI.
With these performance and accuracy improvements, Buffmee now empowers users to safely explore their favorite media through interactive Q&A and deep-dive analysis, delivering a highly personalized experience while maintaining strict trust and compliance for content providers.
Let’s deep dive into how they achieved these results.
Establish automated evaluation for diverse content
Traditional manual testing requires immense effort and cannot scale to accommodate a large content library. To solve this, the development team designed a systematic AI evaluation process using Gemini Enterprise Agent Platform Evaluation Service.
By implementing automated evaluation frameworks like LLM-as-a-Judge and the Rule of Hundreds, the team replaced labor-intensive manual testing with a data-driven process. They ingested their extensive document corpus, constructed hundreds of automated evaluation tests, and built a comprehensive benchmark dataset to measure the reliability of answers for each use case. As a result, the team improved their groundedness scores by 25%, helping deliver highly accurate and reliable outputs.
Source Credit: https://cloud.google.com/blog/topics/customers/how-kddi-optimized-rag-performance-with-agent-development-kit/
