The evaluation primarily consists of two core components: human evaluation on multi-turn conversation scenarios, and blind ranking on public arenas.
In the human evaluation, the research team used a held-out test set of 500 diverse scenarios, each describing a series of multi-turn requests for the model (e.g., brainstorming, planning, or learning new knowledge), with an average of 8.4 interaction turns between the user and the model. The evaluation results show that, compared to Gemma 1.1 7B, human evaluators rated Gemma 2 9B and 27B models significantly higher on both conversational satisfaction and goal achievement rate. Furthermore, Gemma 2 models better maintain high-quality responses from the beginning of a conversation through subsequent turns. These evaluation results are directly linked to improvements in the post-training stage: Gemma 2 used a much larger reward model in the RLHF stage, specifically optimized for conversational capability (particularly multi-turn interactions), to better align with human preferences.
In automated evaluation and public rankings, Gemma 2's instruction-tuned models achieved strong results on the authoritative blind-testing platform LMSYS Chatbot Arena. The 27B Gemma 2 instruction model outperformed much larger models such as Llama 3 70B, ranking first among all open-weight models; the 9B model was the best-performing model under 15B parameters at the time. The subsequently released 2B model also performed remarkably well, surpassing GPT-3.5 as well as models with tens of times more parameters, such as Mixtral 8x7B and Llama 2 70B.