TLA+-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA+ Specification Generation

Published in arXiv, 2026

Download paper here

Large language models increasingly write TLA+ formal specifications from natural-language descriptions, but progress is hard to measure: existing resources grade by resemblance to a reference or by whether the output parses, neither of which shows correctness. We present TLA+-Bench, a dataset and benchmark that grades by execution. Every gold specification ships a configuration the TLA+ model checker runs over the full reachable state space, deciding exactly whether the specification holds the properties that configuration names. The dataset holds 403 model-checked gold and 897 parse-only silver specifications from 13 public repositories, subsumes prior TLA+ generation data, and carries four model-written descriptions in two styles from two providers, with difficulty and category labels.

Recommended citation: Bisharat, A., Spencer, E., Ortiz, B., Bhadauria, K., Nazari, M., Santos, B., Ramos, A., Wang, T., Thiruvathukal, G.K., Läufer, K. and Abuhamad, M. (2026) TLA+-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA+ Specification Generation. arXiv:2607.23425. Available at: https://arxiv.org/abs/2607.23425.

Recommended citation: Bisharat, A., Spencer, E., Ortiz, B., Bhadauria, K., Nazari, M., Santos, B., Ramos, A., Wang, T., Thiruvathukal, G.K., Läufer, K. and Abuhamad, M. (2026) TLA+-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA+ Specification Generation. arXiv:2607.23425. Available at: https://arxiv.org/abs/2607.23425.
Download Paper