To ensure fair comparison, the paper's authors re-ran all model benchmarks using their own evaluation pipeline, covering a broad range of widely recognized standard tests including commonsense reasoning (0-shot), world knowledge (5-shot), reading comprehension (0-shot), mathematics (using majority voting), code generation, and aggregate metrics like MMLU. Detailed evaluation data shows that Mistral 7B outperforms the larger Llama 2 13B model on all evaluation metrics -- for example, achieving 60.1% on the comprehensive MMLU benchmark. It particularly excels in code, mathematics, and reasoning, even surpassing the 340B-parameter Llama 1 34B model in these areas, and approaching the performance of the code-specialized Code-Llama 7B on HumanEval. Further efficiency analysis introduces the concept of "equivalent model size," noting that Mistral 7B performs on reasoning and understanding tasks at a level comparable to Llama 2 models more than 3 times its size, convincingly demonstrating the efficiency of its architecture.
Note: No module data