Towards CSI: What's the best harness? (arXiv 2026)
We studied a question that receives surprisingly little attention:
Does the agent harness matter as much as the underlying LLM?
We benchmarked five different cybersecurity scaffolds while keeping the model fixed (alias2-mini) across all 33 CyBench challenges.
Key findings:
No single scaffold performs best across every challenge. Combining heterogeneous scaffolds consistently improves coverage. A shared blackboard architecture solves 19/33 challenges (57.6%), outperforming every individual harness while reducing execution time.
Paper: https://arxiv.org/pdf/2605.28334
Happy to answer technical questions or discuss the benchmarking methodology. submitted by /u/Obvious-Language4462