Bachelor's Thesis: LLMs × BPMN

Research · 01/2025

Systematic comparison of eight LLMs generating BPMN 2.0 models from natural language – with a custom evaluation framework and real-world datasets.

Details

Role
Author, B.Sc. Business Informatics (IWi at DFKI, Saarland University)
Period
01/2025
Technologies
LLM evaluation, BPMN 2.0, Prompt engineering, F1 score, Camunda, hdBPMN

Research

Can freely available language models generate valid, usable BPMN 2.0 models from natural-language descriptions?

20+ tasks in two task classes, including seven pseudo-BPMN tasks targeting specific specifications; identical prompting strategy for all models.

Models evaluated
o1, GPT-4o, GPT-4o-mini, GPT-4, Claude, Gemini, Gemini 2.0, Mistral 7B
Datasets
Camunda, hdBPMN
Metrics
Valid BPMN 2.0 XML, Completeness, Precision, F1 score, Logical correctness (soundness)

Findings

  • Clear instructions produced near error-free models.
  • Unstructured process descriptions clearly degraded the results of every model.
  • Producing valid XML proved harder than understanding the process.
  • Human review remains necessary for production use.

← Back to overview