ENHANCING PROMPT QUALITY IN LANGUAGE MODELS THROUGH SUPERVISED FINE-TUNING: AN EXPLORATION WITH MISTRAL-7B
Keywords:
Prompt Engineering, Large Language Models, Supervised Fine-Tuning, Mistral-7B, G-Eval, Natural Language Processing, Direct Preference Optimization, Automated Prompt Generation.Abstract
When it comes to the field of Natural Language Processing (NLP), the input prompts are a critical factor that influences how the human user interacts with the Large Language Model (LLM). The formulation of prompts has changed from basic, rule-based to sophisticated, context-sensitive prompts. But existing techniques have one drawback: the need for a lot of domain knowledge and manual prompting in creating the prompts. This gap between the theoretical capabilities of LLMs and their actual use by everyday users without AI expertise is a significant hurdle that needs to be addressed.
This paper proposes an automated prompt generation and refinement system that uses Supervised Fine-Tuning (SFT) of the Mistral-7B large language model to close this gap. The proposed system acts as a smart mediator between the user and the target LLM, taking input from the user, converting it into a well-constructed, high-quality prompt following best practices of prompt engineering, and providing the refined prompt to the target LLM. The system can be used as a standalone application or as a middleware element that can be placed into any human-AI interaction pipeline, thereby making the job of prompt optimization easier for the end user and restoring accessibility and usefulness to the system.
To fine-tune, a custom 2500 paired raw and refined prompts were created. The proposed system's effectiveness is quantitatively assessed with the specialized evaluation framework G-Eval, which consists of GPT-4 and the chain of thought evaluation protocol, with scores ranging from 1 to 100 for three dimensions: completeness, clarity, and overall quality. Experimental results show that the fine-tuned Mistral-7B SFT model outperforms the unmodified Mistral-7B base model and a Direct Preference Optimization (DPO) fine-tuned model for all elements of the evaluation in terms of completeness, clarity, and overall.
This work proposes an architecture to refine prompts for deployment, a curated prompt pair training set, and empirical data showing that supervised fine-tuning of prompt pair data can be an effective approach to automating prompt optimization. The knowledge acquired suggests valuable future research opportunities in NLP related to the creation of models that would better comprehend and react to the complete meaning in natural human speech.


