Bachelor's Thesis: LLMs × BPMN
Systematic comparison of eight LLMs generating BPMN 2.0 models from natural language – with a custom evaluation framework and real-world datasets.
Details
- Role
- Author, B.Sc. Business Informatics (IWi at DFKI, Saarland University)
- Period
- 01/2025
- Technologies
- LLM evaluation, BPMN 2.0, Prompt engineering, F1 score, Camunda, hdBPMN
Research
Can freely available language models generate valid, usable BPMN 2.0 models from natural-language descriptions?
20+ tasks in two task classes, including seven pseudo-BPMN tasks targeting specific specifications; identical prompting strategy for all models.
- Models evaluated
- o1, GPT-4o, GPT-4o-mini, GPT-4, Claude, Gemini, Gemini 2.0, Mistral 7B
- Datasets
- Camunda, hdBPMN
- Metrics
- Valid BPMN 2.0 XML, Completeness, Precision, F1 score, Logical correctness (soundness)
Findings
- Clear instructions produced near error-free models.
- Unstructured process descriptions clearly degraded the results of every model.
- Producing valid XML proved harder than understanding the process.
- Human review remains necessary for production use.