Optimierung von Anforderungen im Requirements Engineering durch den Einsatz von generativer KI und Prompt Engineering
Masterarbeit, Jonah Heisterkamp, 2025
Kurzfassung
Die Qualität von Anforderungsspezifikationen ist entscheidend für den Erfolg von Softwareentwicklungsprojekten. Unklare, widersprüchliche oder unvollständige Anforderungen führen häufig zu Fehlinterpretationen, Projektverzögerungen und erhöhten Entwicklungskosten. In den letzten Jahren hat sich gezeigt, dass Methoden der künstlichen Intelligenz (KI), insbesondere große Sprachmodelle (LLMs), vielversprechende Ansätze zur Unterstützung des Requirements Engineering bieten.
Diese Arbeit untersucht, inwiefern LLMs das Requirements Engineering bei der Optimierung von Anforderungen unterstützen können und welche Prompt-Engineering-Methoden dabei wirksam sind. Dafür wird zunächst mittels einer Literaturanalyse der Stand der Forschung des Prompt Engineering untersucht und es wird ermittelt, welche Best-Practices es dabei im Requirements Engineering gibt. In eigenen Tests wird im zweiten Teil der Arbeit exemplarisch die Fähigkeit von ChatGPT (GPT-5) getestet, Widersprüche in Anforderungsspezifikationen zu identifizieren. Dabei werden in zwei Testszenarien die Wirksamkeit eines einfachen Zero-Shot-Prompts mit einem umfangreichen Few-Shot-Prompt miteinander verglichen.
Die Ergebnisse der Literaturanalyse zeigen eine breite Übereinstimmung hinsichtlich einem allgemein empfohlenen Promptaufbau, der prägnant und klar strukturiert ist und bei dem das Wichtigste am Anfang oder Ende steht. Außerdem können Methoden wie das Chain-of-Thought- und Self-Consistency-Prompting für bessere Ergebnisse bei besonders logisch anspruchsvollen Aufgaben genutzt werden. Für das Requirements Engineering und die Analyse von Anforderungen hinsichtlich ihrer Qualitätskriterien gibt es klare Empfehlungen, das Vorgehen je nach Qualitätskriterium konkret zu spezifizieren sowie ein einheitliches Ausgabeformat im Checklistenformat anzugeben, um die Auswertbarkeit zu verbessern.
In dem praktischen Vergleichstest identifiziert ChatGPT den Großteil der Widersprüche in einer Anforderungsspezifikation und zeigt damit seinen Mehrwert für das Prompt Engineering. Überraschend schnitten die Tests mit dem einfacheren Zero-Shot-Prompt durchweg besser ab als die Tests mit dem aufwändigeren Few-Shot-Prompt. Als mögliche Ursachen werden Bias bzw. Verhaltenseingrenzungen durch konkretere Vorgaben sowie Modellfortschritte diskutiert. Die Ergebnisse sind daher nicht ohne Weiteres verallgemeinerbar, zeigen jedoch das Potenzial von LLMs für die Widerspruchsprüfung, auch ohne aufwändige Prompts.
Insgesamt zeigen sowohl die Fachliteratur als auch die Tests, dass LLMs derzeit leistungsfähige Assistenzwerkzeuge sind, die jedoch weiterhin eine menschliche Kontrollinstanz erfordern.
Abstract
The quality of requirements specifications is critical to the success of software development projects. Ambiguous, contradictory, or incomplete requirements often lead to misinterpretations, project delays, and increased development costs. In recent years, methods from artificial intelligence (AI), in particular large language models (LLMs), have emerged as promising approaches to supporting requirements engineering.
This thesis investigates the extent to which LLMs can support requirements engineering in optimizing requirements and which prompt-engineering methods are effective. To this end, a literature review first examines the state of the art in prompt engineering and identifies best practices for its application in requirements engineering. In the second part, proprietary tests exemplarily assess ChatGPT’s (GPT 5) ability to identify contradictions in requirements specifications. Two
test scenarios compare the effectiveness of a simple zero-shot prompt with that of a more extensive few-shot prompt.
The literature review reveals broad agreement on a generally recommended prompt structure that is concise and clearly organized, with the most important information placed at the beginning or the end. In addition, methods such as chain-of-thought and self-consistency prompting can be used to achieve better results on particularly demanding logical tasks. For requirements engineering and the analysis of requirements with respect to quality criteria, there are clear recommendations to specify the procedure concretely for each criterion and to prescribe a standardized checklist-style output format to facilitate evaluation.
In the practical comparison test, ChatGPT identified most of the contradictions in a requirements specification, demonstrating its added value for prompt engineering. Surprisingly, the tests using the simpler zero-shot prompt consistently outperformed those using the more elaborate few-shot prompt. Potential explanations discussed include bias or behavioral constraints induced by more specific instructions, as well as advances in the model. The results are therefore not readily generalizable, but they do highlight the potential of LLMs for contradiction detection, even without elaborate prompts.
Overall, both the literature and the experiments indicate that LLMs are currently powerful assistant tools, but still require human oversight.