source
Terminal-Bench
End-to-end task success on the Terminal-Bench 4.0 task set with the published agent scaffold and reasoning configuration.
Scores from earlier Terminal-Bench versions use different tasks and cannot be compared directly. IDE-native workflows, code review quality, and non-terminal product engineering work.