Andreas Belitz
Blog Notes Projects About

Tagged: swe-bench

All AI Engineering Tools & Workflow 3D Printing & Making LLM Patterns Career & Thinking Infrastructure
Measuring my agent harness against plain Claude Code
Sep 29, 2026 · 8 min read · AI Engineering

Measuring my agent harness against plain Claude Code

I built a multi-agent harness and ran it head to head against a plain control on a real benchmark. On these tasks it cost about five times more and did not score better. What it bu...

agents evaluation benchmarking swe-bench harness

© 2026 Andreas Belitz v:1ed2117

RSS GitHub LinkedIn