HarnessSecurity-Bench: Do Security Mechanisms Really Protect Coding Agent Harnesses?

Published in arXiv preprint arXiv:2610.07639, 2026

We study how coding agent harnesses protect tool execution and balance security with task utility. HarnessSecurity examines 40 harnesses and organizes their controls into ten security mechanisms.

HarnessSecurity-Bench provides 23 tasks across five attack surfaces, with separate checks for legitimate task completion and attack effects. We evaluate nine mechanisms across six harnesses in 2,500 trials.

The results show that security controls have different effects on protection and utility, and that alternative execution paths can bypass restrictions. The benchmark supports evaluating these tradeoffs under controlled security settings.

Recommended citation: Zhengyang Zhu, Liming Huang, Runmin Ji, Mingxi Ye, Zihan Zhou, Hanyang Guo, Jingwen Wu, Yuhan Ye, Yuming Feng, Hong-Ning Dai, Zibin Zheng. "HarnessSecurity-Bench: Do Security Mechanisms Really Protect Coding Agent Harnesses?" arXiv preprint arXiv:2610.07639, 2026.
PDF

Project website

Direct Link