Can LLMs Write Correct TLA + Specifications? Evaluating Natural-Language-to-TLA + Generation
Published in Proceedings of the 21st International Conference on Software Technologies, 2026
This paper presents the first systematic evaluation of LLM-based TLA+ specification synthesis from natural language. Our study evaluates 30 LLMs across eight families on a curated dataset of 205 TLA+ specifications: 25 open-weight models across four prompting strategies (2,600 runs) and 5 proprietary models under few-shot prompting (130 runs), all validated by the SANY parser and TLC model checker. LLMs achieve up to 26.6% syntactic correctness but only 8.6% semantic correctness, with successes exclusive to progressive prompting. Results show that model size does not predict quality, e.g., DeepSeek r1:8b outperforms its 70B variant across all strategies, which suggests the importance of reasoning alignment for formal languages.
Recommended citation: Bisharat, A., Ortiz, B., Spencer, E., Bhadauria, K., Wang, T., Thiruvathukal, G. K., Läufer, K. and Abuhamad, M. (2026). Can LLMs Write Correct TLA + Specifications? Evaluating Natural-Language-to-TLA + Generation. In Proceedings of the 21st International Conference on Software Technologies - ICSOFT; ISBN 978-989-758-855-6; ISSN 2184-2833, SciTePress, pages 39-50. DOI: 10.5220/0015070400004088
Download Paper
