A task with its own rules.
A benchmark for agentic software engineering tasks executed in real terminal environments. DeepSeek reports it in the agentic section, while BenchLM also mirrors it in coding for models that publish it as a developer-task signal.
- Organisation
- Terminal-Bench contributors
- Version
- 2026
