The SFT phase of Mistral 7B is designed to evaluate the generalization capability of its base model, with the entire process emphasizing simplicity and transparency. The researchers only performed supervised fine-tuning on publicly available instruction datasets from the Hugging Face repository, using no proprietary data or complex training tricks. The goal was to demonstrate that a high-quality base model can achieve excellent performance through a simple pipeline. Evaluation results show that the resulting Mistral 7B - Instruct model outperforms all comparably sized 7B models on the MT-Bench automated benchmark and performs on par with 13B-scale chat models. More importantly, on an independent anonymous human preference evaluation platform (LLM Boxing), as of October 6, 2023, Mistral 7B's generated responses received 25,281 preferences, significantly higher than Llama 2 13B Chat's 21,257, strongly demonstrating that its SFT-tuned responses are of higher quality in human judgment.