The State Of AI Harness Engineering 2026

The State Of AI Harness Engineering 2026

François Zaninotto
• 28 min read

There is a lot of talk about harness engineering (building the tools, context, memory, and control loops to make an agent reliable) but what are people actually doing, and how is that working out for them?

One number summarizes the importance of this question. Someone ran the same model through 8 different harnesses on the same 25 tasks, with the same provider and the same tools. Success rate went from 68% to 88%. As we’ll see, it’s about the only thing anyone has properly measured.

We’ve audited 246 repositories and read 57 publications to find out. We’ve limited the research to harnesses used for software development and coding agents, as it’s what we know the best. This post summarizes our findings about the main tendencies, best practices and actionable learnings of this young discipline.

Our sources

We’ve selected open-source repositories that describe themselves as harnesses, as well as large and popular repositories that contain a harness. This boils down to 246 open-source repositories. We’ve cloned each of them to get its real first commit date, its last activity, and the number of files it versions under .claude/.

RepositoryKindLanguageCreatedActivityCommitsHarness
obra/superpowersCollection288kJavaScript2025-10-09 (11 mo)1 mo681
affaan-m/ECCCollection260kJavaScript2026-01-17 (8 mo)active2.7k13
NousResearch/hermes-agentCollection246kPython2025-07-22 (1.2 yr)active36k
anomalyco/opencodeStandalone agent208kTypeScript2025-03-21 (1.5 yr)active16k
Significant-Gravitas/AutoGPTStandalone agent187kPython2023-03-16 (3.5 yr)active9.0k67
langgenius/difyStandalone agent156kTypeScript2023-05-15 (3.3 yr)active13k6
anthropics/claude-codeStandalone agent142kTypeScript2025-02-22 (1.6 yr)active8503
DietrichGebert/ponytailCollection141kJavaScript2026-06-12 (3 mo)active2241
github/spec-kitOrchestrator137kPython2025-08-21 (1.1 yr)active2.0k
farion1231/cc-switchBuilding block133kRust2025-08-04 (1.1 yr)active2.4k
openai/codexStandalone agent119kRust2025-04-16 (1.4 yr)active11k
Graphify-Labs/graphifyBuilding block119kPython2026-04-03 (5 mo)active1.8k
VoltAgent/awesome-design-mdCollection116kMarkdown2026-03-31 (6 mo)2 mo61
earendil-works/piStandalone agent106kTypeScript2025-08-09 (1.1 yr)active6.4k
JuliusBrussee/cavemanBuilding block106kGo2026-04-04 (5 mo)active686
nexu-io/open-designBuilding block97kTypeScript2026-04-28 (5 mo)active3.6k23
addyosmani/agent-skillsCollection96kJavaScript2026-02-15 (7 mo)active51210
thedotmack/claude-memBuilding block94kTypeScript2025-09-06 (1.0 yr)active2.7k5
infiniflow/ragflowStandalone agent91kGo2023-12-12 (2.8 yr)active9.4k
OpenHands/OpenHandsStandalone agent88kTypeScript2024-03-13 (2.5 yr)active8.3k
Leonxlnx/taste-skillCollection88kJavaScript2026-02-19 (7 mo)active157
Egonex-AI/Understand-AnythingBuilding block83kTypeScript2026-03-14 (6 mo)active846
bytedance/deer-flowStandalone agent83kPython2025-04-07 (1.4 yr)active3.3k18
Panniantong/Agent-ReachBuilding block83kPython2026-02-24 (7 mo)active376
rtk-ai/rtkBuilding block81kRust2026-01-22 (8 mo)active2.0k44
shareAI-lab/learn-claude-codeStandalone agent77kPython2026-02-21 (7 mo)active233
ComposioHQ/awesome-claude-skillsCollection75kPython2025-10-17 (11 mo)2 mo77
microsoft/ai-agents-for-beginnersCollection75kJupyter2024-11-28 (1.8 yr)active2.0k
ruvnet/rufloOrchestrator73kTypeScript2025-06-02 (1.3 yr)active7.5k371
headroomlabs-ai/headroomBuilding block73kPython2026-01-06 (8 mo)active2.8k1
colbymchenry/codegraphBuilding block71kC2026-01-18 (8 mo)active1.0k4
stablyai/orcaOrchestrator71kTypeScript2026-03-16 (6 mo)active11k
FoundationAgents/MetaGPTStandalone agent70kPython2023-06-30 (3.2 yr)dormant6.4k
code-yeongyu/oh-my-openagentBuilding block69kTypeScript2025-12-03 (9 mo)active17k4
cline/clineStandalone agent68kTypeScript2024-07-05 (2.2 yr)active7.3k10
openinterpreter/openinterpreterStandalone agent68kRust2025-04-16 (1.4 yr)active11k
asgeirtj/system_prompts_leaksCollection67kJavaScript2025-05-03 (1.4 yr)active778
diegosouzapw/OmniRouteBuilding block67kTypeScript2026-02-18 (7 mo)active8.7k
Mintplex-Labs/anything-llmStandalone agent66kJavaScript2023-06-03 (3.3 yr)active2.4k
shanraisshan/claude-code-best-practiceCollection66kHTML2025-10-31 (11 mo)active2.2k113
gsd-build/get-shit-doneOrchestrator65kJavaScript2025-12-14 (9 mo)4 mo2.9k
microsoft/autogenStandalone agent61kPython2020-12-04 (5.8 yr)5 mo3.8k
FlowiseAI/FlowiseStandalone agent56kTypeScript2023-04-06 (3.4 yr)1 mo3.6k
hesreallyhim/awesome-claude-codeCollection54kPython2025-04-19 (1.4 yr)active1.8k
block/gooseStandalone agent54kRust2024-08-23 (2.1 yr)active5.7k
bmad-code-org/BMAD-METHODOrchestrator53kPython2025-04-13 (1.4 yr)active2.2k
router-for-me/CLIProxyAPIBuilding block52kGo2025-07-02 (1.2 yr)active3.9k
Aider-AI/aiderStandalone agent49kPython2023-04-03 (3.5 yr)4 mo13k
HKUDS/nanobotStandalone agent48kPython2026-02-01 (7 mo)active4.4k3
bojieli/ai-agent-bookCollection48kPython2025-09-09 (1.0 yr)active1.8k
zhayujie/CowAgentStandalone agent47kPython2022-08-10 (4.1 yr)active2.8k
DeusData/codebase-memory-mcpBuilding block44kC2026-02-25 (7 mo)active3.0k
Hmbown/CodewhaleStandalone agent41kRust2026-01-20 (8 mo)active9.8k
tinyhumansai/openhumanStandalone agent40kRust2026-01-27 (8 mo)active21k16
wshobson/agentsCollection40kPython2025-07-24 (1.1 yr)active5763
herdrdev/herdrStandalone agent39kRust2026-03-23 (6 mo)active1.7k
continuedev/continueStandalone agent36kTypeScript2023-05-23 (3.3 yr)2 mo22k1
esengine/DeepSeek-ReasonixStandalone agent36kGo2026-05-29 (4 mo)active7.1k
VoltAgent/awesome-agent-skillsCollection35kMarkdown2025-10-28 (11 mo)active649
iOfficeAI/AionUiStandalone agent33kTypeScript2025-08-07 (1.1 yr)active6.0k8
alibaba/open-code-reviewStandalone agent33kGo2026-05-18 (4 mo)active7112
agentscope-ai/agentscopeStandalone agent32kPython2025-08-15 (1.1 yr)active597
openai/openai-agents-pythonStandalone agent30kPython2025-03-11 (1.5 yr)active2.3k
rohitg00/agentmemoryBuilding block29kTypeScript2026-02-25 (7 mo)active482
google-labs-code/design.mdCollection28kTypeScript2026-04-10 (5 mo)2 mo62
QwenLM/qwen-codeStandalone agent28kTypeScript2025-04-15 (1.4 yr)active9.8k
openai/symphonyOrchestrator27kElixir2026-03-04 (6 mo)active47
Fosowl/agenticSeekStandalone agent27kPython2025-02-19 (1.6 yr)active992
OthmanAdi/planning-with-filesOrchestrator27kShell2026-01-03 (8 mo)active42636
xai-org/grok-buildStandalone agent27kRust2026-07-16 (2 mo)active45
deepset-ai/haystackStandalone agent27kPython2019-11-14 (6.8 yr)active6.3k
VoltAgent/awesome-claude-code-subagentsCollection25kShell2025-07-30 (1.1 yr)active5171
letta-ai/lettaStandalone agent25kPython2023-10-11 (2.9 yr)active7.5k
agentsmd/agents.mdCollection24kTypeScript2025-08-19 (1.1 yr)active38
browserbase/stagehandStandalone agent24kTypeScript2024-03-20 (2.5 yr)active1.5k
RooCodeInc/Roo-CodeStandalone agent24kTypeScript2024-07-05 (2.2 yr)4 mo7.1k
mksglu/context-modeBuilding block23kTypeScript2026-02-23 (7 mo)active2.2k10
PrimeIntellect-ai/prime-agentStandalone agent21kTypeScript2025-08-09 (1.1 yr)active4.8k1
jnMetaCode/agency-agents-zhOrchestrator21kShell2026-03-06 (6 mo)active273
pydantic/pydantic-aiStandalone agent20kPython2024-06-14 (2.3 yr)active3.0k9
Kilo-Org/kilocodeStandalone agent20kTypeScript2025-03-21 (1.5 yr)active31k
1jehuang/jcodeStandalone agent20kRust2026-01-05 (8 mo)active7.5k1
emcie-co/parlantStandalone agent18kPython2024-02-15 (2.6 yr)3 mo5.5k
microsoft/SkillOptCollection17kPython2026-05-21 (4 mo)active531
cft0808/edictOrchestrator17kPython2026-02-23 (7 mo)4 mo149
HKUDS/DeepCodeOrchestrator17kPython2025-07-20 (1.2 yr)active513
HKUDS/OpenHarnessOrchestrator16kPython2026-04-01 (6 mo)3 mo4295
wasp-lang/open-saasProduct-embedded16kMDX2023-03-29 (3.5 yr)1 mo637
walkinglabs/learn-harness-engineeringOrchestrator15kTypeScript2026-03-30 (6 mo)active252
yc-software/qmOrchestrator15kTypeScript2026-07-29 (2 mo)active4985
travisvn/awesome-claude-skillsCollection15kMarkdown2025-10-16 (11 mo)5 mo42
mindfold-ai/TrellisBuilding block15kTypeScript2026-01-26 (8 mo)active1.4k96
NanmiCoder/cc-hahaBuilding block15kTypeScript2026-03-31 (6 mo)active2.0k
lsdefine/GenericAgentStandalone agent14kPython2026-01-16 (8 mo)active1.4k
The-Pocket/PocketFlowStandalone agent11kPython2024-12-24 (1.7 yr)2 mo64120
omnigent-ai/omnigentOrchestrator10kPython2026-06-13 (3 mo)active3.8k11
diet103/claude-code-infrastructure-showcaseProduct-embedded10kTypeScript2025-10-29 (11 mo)2 mo1488
MervinPraison/PraisonAIStandalone agent9.1kPython2024-03-19 (2.5 yr)active9.4k
revfactory/harnessOrchestrator9.0kMarkdown2026-03-27 (6 mo)3 mo45
iflytek/astron-agentStandalone agent9.0kJava2025-09-22 (12 mo)active3.2k
automazeio/ccpmOrchestrator8.4kShell2025-08-19 (1.1 yr)dormant87
YaoApp/yaoStandalone agent8.0kGo2021-09-06 (5.0 yr)active4.3k
max-sixty/worktrunkBuilding block7.9kRust2025-10-16 (11 mo)active5.1k6
SWE-agent/mini-swe-agentStandalone agent7.7kPython2025-06-28 (1.2 yr)active1.0k
mnfst/llm-gatewayStandalone agent7.5kTypeScript2022-09-27 (4.0 yr)active6.2k2
Gentleman-Programming/gentle-aiOrchestrator6.9kGo2026-02-28 (7 mo)active3.5k1
builderz-labs/mission-controlOrchestrator6.2kTypeScript2026-02-23 (7 mo)active536
dontriskit/awesome-ai-system-promptsCollection6.2kTypeScript2025-03-05 (1.5 yr)dormant67
ChrisWiles/claude-code-showcaseProduct-embedded6.1kJavaScript2026-01-06 (8 mo)dormant422
FlorianBruniaux/claude-code-ultimate-guideCollection6.0kPython2026-01-09 (8 mo)active1.0k48
ModelEngine-Group/nexentStandalone agent5.9kPython2025-04-28 (1.4 yr)active5.8k30
SWE-bench/SWE-benchLab / bench5.9kPython2023-10-10 (2.9 yr)active746
the-open-agent/openagentStandalone agent5.6kGo2022-03-31 (4.5 yr)active2.4k
kodu-ai/claude-coderStandalone agent5.2kTypeScript2024-07-05 (2.2 yr)dormant1.0k
entireio/cliBuilding block5.1kGo2026-01-03 (8 mo)active8.8k37
campfirein/byterover-cliBuilding block5.0kTypeScript2025-10-07 (11 mo)3 mo3.1k1
darrenhinde/OpenAgentsControlOrchestrator4.9kTypeScript2025-08-14 (1.1 yr)2 mo224
Kodezi/ChronosStandalone agent4.9kJava2025-07-26 (1.1 yr)dormant5
Agenta-AI/agentaStandalone agent4.8kTypeScript2023-04-27 (3.4 yr)active28k25
vijaythecoder/awesome-claude-agentsOrchestrator4.4kMarkdown2025-07-26 (1.1 yr)dormant43
gptme/gptmeStandalone agent4.4kPython2023-03-24 (3.5 yr)active4.5k
ai-boost/awesome-harness-engineeringCollection4.3kMarkdown2026-03-29 (6 mo)active259
anthropics/claude-plugins-communityCollection4.2kPython2026-03-20 (6 mo)active2.3k
parcadei/Continuous-Claude-v3Orchestrator3.9kPython2026-01-10 (8 mo)dormant117718
matt1398/claude-devtoolsBuilding block3.9kTypeScript2026-02-11 (7 mo)4 mo34911
gadievron/raptorCollection3.8kPython2025-11-21 (10 mo)active7.4k151
gotalab/cc-sddOrchestrator3.7kTypeScript2025-07-17 (1.2 yr)5 mo429
gemini-cli-extensions/conductorOrchestrator3.7kPython2025-12-17 (9 mo)active133
smallcloudai/refactStandalone agent3.5kRust2022-07-24 (4.2 yr)4 mo11k
davepoon/buildwithclaudeCollection3.5kPython2025-07-25 (1.1 yr)active566
foryourhealth111-pixel/Vibe-SkillsCollection3.3kPython2026-02-22 (7 mo)active1.4k
AutoCodeRoverSG/auto-code-roverStandalone agent3.1kPython2024-04-09 (2.4 yr)dormant219
jsynowiec/node-typescript-boilerplateProduct-embedded3.0kTypeScript2016-08-20 (10.1 yr)3 mo347
WorldFlowAI/everything-claude-codeCollection3.0kJavaScript2026-01-17 (8 mo)dormant271
wesammustafa/Claude-Code-Everything-You-Need-to-KnowCollection3.0kPython2025-08-17 (1.1 yr)2 mo6824
microsoft/skillsCollection3.0kTypeScript2026-01-16 (8 mo)active7501
rohitg00/pro-workflowOrchestrator2.9kJavaScript2026-02-01 (7 mo)2 mo86
intellectronica/rulerBuilding block2.9kTypeScript2025-05-20 (1.3 yr)active1.1k
ciembor/agent-rules-booksCollection2.8kMarkdown2026-04-16 (5 mo)active62
lopopolo/harness-engineeringProduct-embedded2.7kMarkdown2026-07-18 (2 mo)2 mo2
zubair-trabzada/ai-marketing-claudeCollection2.7kPython2026-03-01 (7 mo)dormant1
centminmod/my-claude-code-setupProduct-embedded2.6kMarkdown2025-07-08 (1.2 yr)active301111
rohitg00/awesome-claude-code-toolkitCollection2.6kJavaScript2026-02-04 (7 mo)4 mo700
humanlayer/advanced-context-engineering-for-coding-agentsLab / bench2.6kMarkdown2025-08-29 (1.1 yr)1 mo45
AgentsMesh/AgentsMeshStandalone agent2.4kGo2026-01-08 (8 mo)1 mo1.1k18
supabitapp/supacodeBuilding block2.4kSwift2026-01-20 (8 mo)active2.0k
0xSteph/pentest-ai-agentsCollection2.2kShell2026-03-28 (6 mo)1 mo56
iannuttall/claude-agentsCollection2.0kMarkdown2025-07-25 (1.1 yr)dormant1
standardagents/dmuxBuilding block1.8kHTML2025-08-19 (1.1 yr)1 mo7472
kunchenguid/treehouseBuilding block1.7kGo2026-03-14 (6 mo)active1101
google/mantisBuilding block1.6kPython2026-06-15 (3 mo)active72
openedclaude/claude-reviews-claudeBuilding block1.6kTypeScript2026-03-31 (6 mo)6 mo42
Paritok-official/paritok-4b-v1Standalone agent1.5kPython2026-07-15 (2 mo)active147
CloudAI-X/claude-workflow-v2Orchestrator1.4kPython2026-01-01 (9 mo)4 mo32
yohey-w/multi-agent-shogunOrchestrator1.4kShell2026-01-25 (8 mo)1 mo3722
openai/SWELancer-BenchmarkLab / bench1.4kPython2025-02-18 (1.6 yr)dormant55
first-fluke/oh-my-agentOrchestrator1.3kTypeScript2026-01-27 (8 mo)active3.6k
obra/superpowers-marketplaceCollection1.3kMarkdown2025-10-09 (11 mo)active1271
coollabsio/jeanBuilding block1.3kTypeScript2026-01-23 (8 mo)active1.7k8
Shopify/roastOrchestrator1.2kRuby2025-04-22 (1.4 yr)1 mo8892
Gentleman-Programming/agent-teams-liteOrchestrator1.2kShell2026-02-16 (7 mo)6 mo73
hoangnb24/repository-harnessProduct-embedded1.2kRust2026-05-05 (4 mo)1 mo312
OpenHands/software-agent-sdkStandalone agent1.1kPython2025-08-23 (1.1 yr)active2.4k
fivetaku/gptaku_pluginsCollection1.1kPython2026-03-02 (7 mo)active307
numman-ali/n-skillsCollection1.0kTypeScript2026-01-02 (8 mo)active46
zhu1090093659/spec_driven_developOrchestrator978Shell2026-03-21 (6 mo)2 mo70
data-goblin/power-bi-agentic-developmentCollection918C#2026-01-13 (8 mo)1 mo455
milisp/codexiaBuilding block918TypeScript2025-08-13 (1.1 yr)active1.6k
michaelshimeles/skillsProduct-embedded903Python2026-07-13 (2 mo)active33
china-qijizhifeng/agentic-harness-engineeringOrchestrator892Python2026-04-26 (5 mo)2 mo47
augmentcode/augment-swebench-agentStandalone agent884Python2025-03-28 (1.5 yr)dormant12
microsoft/power-platform-skillsCollection878JavaScript2026-01-21 (8 mo)active2491
sangrokjung/claude-forgeOrchestrator837Shell2026-02-23 (7 mo)active134
marckohlbrugge/37signals-skillsCollection715Markdown2025-12-17 (9 mo)3 mo35
gmickel/flow-nextOrchestrator699Python2025-12-26 (9 mo)active1.9k
shotgun-sh/shotgunOrchestrator684Python2025-09-10 (1.0 yr)5 mo4883
ThibautBaissac/rails_ai_agentsOrchestrator663Shell2025-12-09 (9 mo)4 mo90148
QuantaAlpha/RepoMasterStandalone agent553Python2025-08-28 (1.1 yr)dormant24
scaleapi/SWE-bench_Pro-osLab / bench526Python2025-09-05 (1.0 yr)4 mo75
multi-swe-bench/multi-swe-benchLab / bench362Python2025-04-01 (1.5 yr)dormant312
CronusL-1141/AI-companyOrchestrator361Python2026-03-12 (6 mo)active728
kaanozhan/FrameOrchestrator355JavaScript2026-01-21 (8 mo)active7552
athola/claude-night-marketOrchestrator337Python2025-11-23 (10 mo)active1.6k39
fstandhartinger/ralph-wiggumOrchestrator296Shell2026-01-14 (8 mo)4 mo463
huangjia2019/agent-design-patternsCollection264HTML2026-05-18 (4 mo)active113
QuantaAlpha/GitTaskBenchLab / bench258Python2025-05-15 (1.3 yr)dormant84
NeuZhou/awesome-ai-anatomyCollection240D22026-04-04 (5 mo)5 mo212
RUCAIBox/awesome-agent-harnessCollection198Markdown2026-03-15 (6 mo)4 mo44
sudokar/openspec-plusOrchestrator196Markdown2026-06-22 (3 mo)active17
mattgierhart/PRD-driven-context-engineeringProduct-embedded182HTML2025-07-09 (1.2 yr)2 mo219209
nwiizo/ccswarmOrchestrator152Rust2025-09-21 (12 mo)active15637
AnastasiyaW/codex-claude-code-configProduct-embedded150Python2026-03-31 (6 mo)active3673
aaddrick/claude-pipelineOrchestrator128Markdown2026-02-06 (7 mo)dormant5122
closedloop-ai/claude-pluginsOrchestrator103Python2026-03-04 (6 mo)active51915
GarrickZ2/groveOrchestrator47TypeScript2026-01-13 (8 mo)active835
TaylorHuston/ai-toolkitOrchestrator13Markdown2025-08-21 (1.1 yr)2 mo283
cisco-open/ai-harness-toolkitOrchestrator12Python2026-04-30 (5 mo)active4
tylerburleigh/claude-sdd-toolkitOrchestrator9Python2025-10-24 (11 mo)dormant125
nesquikm/dev-process-toolkitOrchestrator7TypeScript2026-03-19 (6 mo)active9655
QuasarByte/codex-feature-driven-flowOrchestrator7PowerShell2026-02-21 (7 mo)dormant7
moruno21/agents-concertoOrchestrator2Markdown2026-07-23 (2 mo)2 mo311
microsoft/TaskWeaverOrchestratorPython2023-09-11 (3.0 yr)6 mo657
ScaleML/AgentSPEXOrchestratorPython2026-04-19 (5 mo)3 mo6
statewright/statewrightOrchestratorTypeScript2026-05-03 (5 mo)active4431
cobusgreyling/loop-engineeringOrchestratorMarkdown2026-06-09 (3 mo)active514
Tianshi-Xu/Life-HarnessOrchestratorPython2026-05-21 (4 mo)active24
reshashi/claude-orchestratorOrchestratorMarkdown2026-01-11 (8 mo)dormant52
pedrohcgs/claude-code-my-workflowOrchestratorTeX2026-02-06 (7 mo)active252159
marmelab/atomic-crm Product-embeddedTypeScript2024-02-07 (2.6 yr)active1.7k146
codejunkie99/agentic-stackProduct-embeddedMarkdown2026-04-15 (5 mo)active189
aattaran/deepclaudeStandalone agentTypeScript2026-05-03 (5 mo)4 mo15
ithiria894/awesome-claude-code-workflowsCollectionMarkdown2026-03-23 (6 mo)6 mo12
sickn33/antigravity-awesome-skillsCollectionTypeScript2026-01-14 (8 mo)active2.8k
mgechev/skillgradeCollectionTypeScript2026-02-26 (7 mo)active75
QwenLM/Qwen-MM-PluginsCollectionPython2026-08-03 (1 mo)active209
rpamis/cometCollectionJava2026-05-14 (4 mo)active35615
upstash/context7Building blockTypeScript2025-03-29 (1.5 yr)active976
mem0ai/mem0Building blockPython2023-06-20 (3.2 yr)active2.6k
getzep/zepBuilding blockGo2023-04-29 (3.4 yr)active3961
topoteretes/cogneeBuilding blockPython2023-08-16 (3.1 yr)active10k7
Gentleman-Programming/engramBuilding blockGo2026-02-16 (7 mo)active845
Infisical/agent-vaultBuilding blockGo2026-03-27 (6 mo)active3283
modelcontextprotocol/serversBuilding blockTypeScript2024-11-19 (1.8 yr)active4.2k
microsoft/playwright-mcpBuilding blockTypeScript2025-03-21 (1.5 yr)active5821
ChromeDevTools/chrome-devtools-mcpBuilding blockTypeScript2025-09-11 (1.0 yr)active1.2k
ComposioHQ/composioBuilding blockPython2025-04-11 (1.4 yr)active5.3k4
lastmile-ai/mcp-agentBuilding blockPython2024-12-17 (1.7 yr)dormant767
agentgateway/agentgatewayBuilding blockRust2025-03-18 (1.5 yr)active2.7k
microsoft/LLMLinguaBuilding blockPython2023-07-07 (3.2 yr)active87
dottxt-ai/outlinesBuilding blockPython2023-03-17 (3.5 yr)1 mo1.3k
volcengine/OpenVikingBuilding blockPython2026-01-29 (8 mo)active2.4k
langchain-ai/openwikiBuilding blockTypeScript2026-06-22 (3 mo)active4051
MinishLab/sembleBuilding blockPython2026-04-06 (5 mo)active144
dirac-run/diracBuilding blockRust2026-04-09 (5 mo)active954
Mibayy/token-saviorBuilding blockTypeScript2026-03-25 (6 mo)1 mo535
strukto-ai/mirageBuilding blockGo2026-05-06 (4 mo)active2.7k
NanoNets/GraftBuilding blockPython2026-07-03 (2 mo)active4805
HKUDS/CLI-AnythingBuilding blockPython2026-03-08 (6 mo)active885
vercel-labs/zerolangBuilding blockTypeScript2026-05-15 (4 mo)active1.2k
onesuper/tui-useBuilding blockPython2026-04-06 (5 mo)5 mo11313
BloopAI/vibe-kanbanBuilding blockTypeScript2025-06-14 (1.3 yr)active2.1k
Justin0504/AegisBuilding blockPython2026-03-03 (7 mo)active401
manuelschipper/nahBuilding blockTypeScript2026-08-01 (2 mo)active449
facebookresearch/cca-swebenchLab / benchPython2025-12-26 (9 mo)4 mo2
VILA-Lab/Dive-into-Claude-CodeLab / benchMarkdown2026-04-11 (5 mo)active80
aws/agent-toolkit-for-awsLab / benchPython2026-04-23 (5 mo)active240

Repositories that call themselves harnesses are a biased sample, so we’ve run a counter-test: 145 large open-source projects that have nothing to do with AI (Symfony, Rails, Django, React, Kubernetes, Rust, Spring, Odoo, the competing CRMs), cloned and inspected the same way. 26% of them have a non-empty .claude/ directory (38).

Open source has adopted the documentation, but doesn’t constrain agents yet. And 54 of the 145 have nothing at all: symfony, laravel, vuejs/core, golang/go, ruby/ruby, spring-boot, flask, fastapi, tailwindcss, express, nestjs, tokio, redis, postgres, curl, terraform, odoo.

Then we looked for the harnesses that check themselves. 97 repositories remained, taking everything from the initial set of 246 repositories or from the 145 counter-test that ships at least 3 harness files, plus the well-known harnesses distributed as plugins.

We then eliminated the ones that don’t test their harness, the ones with less than 0.5 commit per harness file, the ones with a single author, and the ones whose harness lived less than 60 days or has been dormant for 90. The iteration filter matters more than it sounds: parcadei/Continuous-Claude-v3 ships 718 harness files, the most of the whole inventory, and fails three criteria at once (0.09 commit per file, built in 16 days, nothing since). A ranking on volume would have put it first.

We ended up with 11 repositories:

Repositoryfilestests% of code testedevalsexec. hooksdenyCIiterationauthorslife
gmickel/flow-next102030676 %21.0811262 d
gadievron/raptor1562747 %222.9814298 d
statewright/statewright2167738 %51.43136 d
marmelab/atomic-crm 1483336 %77yes0.886259 d
AnastasiyaW/codex-claude-code-config4824226 %299210.535161 d
bmad-code-org/BMAD-METHOD2621718 %0.595466 d
rtk-ai/rtk72215 %1221.9637229 d
affaan-m/ECC7741912 %141.06212238 d
ruvnet/ruflo1049308 %22yes212.7213463 d
ClickHouse/ClickHouse9013 %125.7933384 d
closedloop-ai/claude-plugins53482 %210.7811194 d

“Iteration” is commits on the harness divided by harness files, ”% of code tested” relates test files to the harness’s executable files. The file counts here are a few units off the table above because this pass also counts harness files kept outside of .claude/, which matters for anything shipped as a plugin.

As for the articles, we’ve found 57 publications dealing with the matter.

Among them, 8 founding texts:

9 field reports from products:

12 articles about methods & patterns:

8 articles about the loop and its variants:

12 research articles:

8 surveys & lists:


Disclaimer: We write AI harnesses ourselves, for our customer and open-source projects. Most notably, we’re the author of Atomic CRM, an open-source CRM with a sophisticated harness that lets non-developers modify the software through a supervised agent pipeline. This is the reason why we’ve made this study: we wanted to compare our harness engineering practices with the best ones in the field. Atomic CRM is one of the 11 finalists above, so read that for what it is.

A Few Words Of Caution

Harness Engineering Is A New Field: The term “Harness Engineering” appeared in a February 2026 article by OpenAI. This ignited an exploration in many different directions, and it’s safe to say that the only consensus is about what Harness Engineering is (Agent = Model + Harness, a formula we owe to Birgitta Böckeler). As for the how to reliably build an AI harness, this is still an open subject. The median repository in our inventory is 8.7 months old, so nobody has a structural lead.

So if you don’t already have your own AI harness, don’t panic! Most teams don’t have one either: out of 145 large open-source projects, only 4 declare a subagent.

Harness Engineering Is A Moving Target: An AI harness constrains, guides, and elevates an AI model. As these models evolve, the shape of harnesses evolves, too. What used to be a good harness in early 2026 is no longer relevant for the last generation of frontier models. Anthropic published the cleanest example of this: “We added resets to clear the context window in order to address this ‘context anxiety.’ With Opus 4.5, the behavior was gone. The context resets we built to compensate had become dead weight in the agent harness.” Their own conclusion is the general rule: a harness encodes assumptions about what the model can’t do on its own, and those assumptions rot as the model improves.

So any definition of a good harness (such as this article) is doomed to age poorly. If your harness doesn’t follow the practices below, no problem! You may be right and the sources we selected may be wrong. But we’d love to hear from you, as we’re trying to understand what works and what doesn’t.

We Used AI To Explore the Sources: We used AI to gather, rank, and aggregate trends from the numerous sources we’ve identified. Without AI, this work would have taken weeks.

Harness Engineering In Numbers

A fast-growing field

  • 73,400 GitHub repositories carry the claude-code topic, and 21,500 “agent harness” repositories were created in 2026 alone.
  • The median harness repository is 8.7 months old. 83% were created in 2025 or 2026.

Everybody writes instructions, almost nobody enforces them

  • 63% of the 145 large open-source projects (Rails, React, Kubernetes, Django…) ship an agent instruction file. Only 4 declare a subagent, and there are 14 hooks in total across all 145.
  • 57% of the root instruction files we read are longer than 100 lines (the size OpenAI settled on for theirs). The median is 123 lines.
  • Only 12 of the 391 repositories commit a single deny rule in their Claude Code settings. 22 declare a PreToolUse hook, the kind that can veto an action before it runs.
  • In the median repository that uses subagents, only 1 in 4 subagents has a restricted list of tools. The others see everything.

AGENTS.md is winning

  • 189 repositories have an AGENTS.md at the root, 155 a CLAUDE.md. In the large open-source projects, 36 of the 53 CLAUDE.md files are just a symlink or a short redirect to AGENTS.md.
  • 46% of the repositories write instructions for at least two different agents (Claude Code, Codex, Copilot, Cursor, Gemini…).

Few harnesses check themselves

  • Out of the 97 harnesses big enough to test, 60% have neither a test nor an eval. 21% test themselves, 27% version evals, and only 7 do both.

The agent writes its own harness

  • 25% of the commits that touched a harness since January 2025 are explicitly signed by an AI. That’s a floor: many tools and people strip the signature. In 52 of 79 repositories, the harness is more AI-written than the code it governs (median 16% against 9%).

Now, on to the key learnings.

Test / Eval Your Harness

Just like code, a harness may contain bugs - or overly verbose or ambiguous instructions that don’t work all the time. Instructions in an AGENTS.md may never be followed. To make sure the harness works, it must be tested.

This is a rare practice: 60% of the harnesses we looked at have neither a test nor an eval. They’re markdown instructions, and nobody knows whether they hold.

Tests and evals are two different jobs, and an AI harness needs both.

A test tells you a script still works. A rule or a hook (a script that inspects what the agent is about to do and refuses it) may fail. A test is a reproducible way to prove it works. It uses code rather than an LLM. A good example is hooks tests in Atomic CRM.

Tests may fail in two ways: letting through what it should stop, or stopping what it should allow. So each rule should be tested twice, once to make sure it fires when required, and once to make sure it doesn’t fire when it shouldn’t. If your tests only contain things that must be blocked, tightening the guard is always safe and loosening it is invisible, so the harness slowly drifts towards blocking everything and strangling the agent. Anthropic puts it in one line: “Test both the cases where a behavior should occur and where it shouldn’t. One-sided evals create one-sided optimization.”

The clearest example in the corpus is codex-claude-code-config, which files 63 cases against 12 guards: 29 that must be blocked, 19 that must be allowed, and 15 that check the exact text a router emits. rm -rf /tmp/scratch-dir is allowed but rm -rf / is blocked. A runner replays every case against the real script.

A check can also pass because its trigger never fired, or because it quietly read something from your own machine (a personal settings file, an override variable, a warm cache). flow-next planted a file its tests shouldn’t have been able to see, and the harness “reported the global owner block”. The leak was real, which is why they changed how settings are loaded.

Automated harness tests also detect regressions.

An eval tells you the harness helps. Freeze a set of tasks, run them end to end, score the outcome. For a coding agent, this can be a series of bug fixes (faulty code + bug report) that an agent without harness fails to fix properly. Pick 20 to 50 tasks from failures you saw, each starting from a clean state with a known good answer. Keep the model fixed, change one part of the harness at a time, then remove each part in turn to see which ones were pulling their weight.

Out of the 26 repositories containing evaluation cases, only one (flow-next) publishes a before/after comparison with a replication and an owned null result. If you can’t prove a rule improves the agent’s behaviour, you can’t prove the agent needs it.

Reduce The Harness To The Strict Minimum

It’s tempting to download skills from popular repositories, or specialized subagents for security research. But out of the box, the last generation of coding agents is already very capable. So a harness should be built in reaction to an agent failure, not based on an assumption that the agent can’t do something properly.

That’s why the best source for harness instructions are past sessions. If you had to correct the agent, it may be because it didn’t have enough guidance or control, so this deserves an addition in the harness. flow-next does this in the open: 98 bug write-ups committed under .flow/memory/bug/, each carrying a root_cause field and a Prevention paragraph naming the rule that should have caught it.

That’s also why we think a harness should be built by a human, instead of an agent. If an agent needs supervision, how can it decide the supervision it needs? This is our opinion rather than a finding: the only ablation study we found (NLAH, arXiv 2603.25723) measures the opposite, with a self-improving harness gaining 4.8 and 2.7 points on two benchmarks.

Change one thing at a time, then check that the harness is actually better after the change (that’s what the eval is for). The risk is to add too many useless rules, that would increase the cost and the latency, and make the agent response less relevant due to context rot. And instructions are not free. ETH Zurich tested context files across 138 real-world tasks: machine-generated ones reduce task success compared to giving the agent no context at all, while increasing inference cost by over 20%. Human-written ones helped by about 4%. Either way, the agent spends 14 to 22% more reasoning tokens to get there.

Keep agent rules small and don’t put too many of them in your harness. OpenAI tried the single big instruction file and named four ways it failed: it crowds out the actual task, “when everything is ‘important,’ nothing is”, “it rots instantly… a graveyard of stale rules”, and “a single blob doesn’t lend itself to mechanical checks.” They replaced it with roughly 100 lines. The main AGENTS.md should act as a table of contents of the rule tree.

The same goes for tools. How many tools and skills the agent can see is a performance setting, and one of the very few changes in this survey that moves success rate, token cost and latency in the same direction. Vercel “stripped 80% of the tools out of an agent and watched its success rate go from 80% to 100% on the same model, with tokens more than halved and latency down from 724 seconds to 141.” Microsoft passed 100 tools in two weeks on its Azure operations agent and had to collapse them into two broad ones. Treat the direction as solid and the magnitude as unverified - Vercel’s page marks that figure as second-hand and nobody reproduces it. It’s also why no repository showed us this practice: restraint leaves no trace in git.

Double Or Replace Rules With Scripts

Agents may or may not follow the instructions written in an AGENTS.md, Skill or Rule files. This is especially true if a harness contains many of them, as the agent will start each session with a large context full of (sometimes contradicting) rules.

Across 481 public CLAUDE.md files, researchers checked how many written security rules had a real effect (arXiv 2608.23550). The results oscillate between 4% and 16%, which means the corresponding harness is as good as blind. Rules written in prose are almost never actually enforced.

The most effective harnesses enforce their rules via executable guards (e.g., git hooks, linter rules, static code analysis, custom scripts). These checks execute in the computer running the harness AND in the CI, as not every developer uses the same harness. Some developers manage to integrate these checks in the agentic loop (usually using Hooks), in which case they can limit the instructions to the minimum.

The gap is not cosmetic. codex-claude-code-config’s destructive-command-guard.py refuses rm --recursive --force /, rm / --recursive --force and rm -rf /; echo done using regular expressions. An LLM looking for a text description of this pattern cannot catch them all.

And to avoid hallucinations, make the agent cite something a script can check. Have deterministic code produce the raw facts first (the files and functions, the lines the change touched), with no judgement attached, let the agent reason over that, then have a script match every claim back against it. Scripts measure and never judge; the agent judges and never measures.

Let The Agent Run The App

Agents produce code faster than humans can review it, so the bottleneck is no longer the production of code but its verification. You can delegate part of this verification to the agent by letting it start its own instances of the app, open it in a real browser, click through the interface and act like a real user would do. Let the agent read the logs to locate bugs and regressions.

OpenAI made their app startable per working copy, wired browser control into the agent’s runtime, and gave each working copy a throwaway logging and metrics stack the agent can query. The agent reproduces the bug, records the failure, fixes it, then drives the app again to demonstrate the fix. It also turns instructions into actual checks: “ensure service startup completes in under 800ms”, or “no span in these four critical user journeys exceeds two seconds”. The same article describes a smart trick: write your custom checks so that the error message contains the fix. “Because the lints are custom, we write the error messages to inject remediation instructions into agent context.”

Once the agent has successfully tested a feature or a bug fix, it should convert the test to a reproducible script that doesn’t need an LLM to run - a smoke test. This reduces token costs and verification time.

The faster a test harness starts, the shorter the agent feedback loop is. So make sure your test app starts in less than a second, mock unstable external dependencies. Also, if you ever want to have more than one agent work on the codebase, modify it so that it can exist in several different instances (with varying port and database schema, or better, different containers).

Measure How The Harness Is Actually Used

Evals tell you whether a change helps on a set of tasks you picked. They don’t tell you what happens in real sessions: which skills actually fire, how often a hook blocks something, how many iterations a session needs depending on the tools it gets. For that, you need observability: measure what actually happens when the harness runs, like any other piece of software.

The VS Code team runs offline benchmarks before shipping a change, then keeps measuring once it’s live, with A/B tests, aggregate usage signals and weekly reporting. Anthropic’s guide to evals says the same: production monitoring catches the drift that evals miss, and they don’t trust an eval score until someone has read some of the transcripts behind it.

It doesn’t take much. Claude Code can log to any OpenTelemetry backend. n8n adds a hook that fires after every skill call and sends a “Claude Code skill activated” event to its own telemetry server. claude-code-infrastructure-showcase writes every skill suggestion, activation and block to a local metrics.jsonl file, with a script to report on it (except for benchmark sessions, to avoid polluting the numbers).

Yet it’s currently a rare practice: only 5 of the 391 repositories record anything about how their harness is used. The Atomic CRM Builder logs every step of every session, and lets the user parse these logs to understand skill and tool usage.

A Harness Rots Quickly

A codebase evolves quickly, especially when agents take care of it. Some architectural changes may invalidate a rule in the harness. Or you may upgrade to a more recent model that already does the right thing without the need for additional rules. As a consequence, even though a harness brings value at one moment, this value decreases over time.

Most controls exist to compensate for something the model used to get wrong. When it stops getting it wrong, the control is pure cost. So give every control an expiry condition, and write down the incident it exists for. One line per hook naming the failure behind it is the cheap version. We don’t do it: Atomic CRM ships 148 harness files and zero recorded decisions, so nothing says why any control exists, and none of them can be safely removed. flow-next does: 9 decision records, each dated and named after what it settled, plus a declined/ folder for what the project deliberately chose not to build.

This means you should treat the harness like a garden: regularly scan for stale or obsolete documentation that does not reflect the real code behavior, remove old rules before planting new ones, and make sure the harness works fine after every major change. OpenAI automates the first half: “a recurring ‘doc-gardening’ agent scans for stale or obsolete documentation that does not reflect the real code behavior and opens fix-up pull requests.”

A good practice is to scan past sessions (JSONL files in Claude Code) for erratic behavior that should have been caught by the harness. Whenever such rot appears, it’s time to take care of the harness. This task can be (partially) delegated to a gardener agent.

What Doesn’t Work

Another valuable thing we got out of the literature is the feedback of failed attempts.

Adding a reviewer agent makes the code worse. One paper (NLAH, arXiv 2603.25723) removed each harness component in turn and measured the difference. On a benchmark where the base agent solved 41% of tasks, adding a second reviewer agent lowered the success rate by 8%. An agent team isn’t necessarily better than a single agent either: a four-role team resolves 72.2% of SWE-bench 500 where the best single agent resolves 71.8%. This came as a surprise, as Atomic CRM does have a reviewer role, and it does often improve the outcome.

Past 4 agent-to-agent handovers, it stops working. Microsoft grew its Azure operations agent team to more than 50 specialists. It wasn’t a good idea: “problems requiring more than four handoffs almost always failed.”. They eventually rolled back to a few generalist agents. Atomic CRM chains 7 roles.

Writing permission rules in advance works worse than approving each action. A study with 113 participants compared two ways of supervising an agent: writing the rules up front, or approving each action as it came. The rule writers blocked 20.1 percentage points fewer bad actions. They had set most of their rules to “ask me”, then approved the prompts anyway, as most people do (93% of permission prompts get approved). A rule that ends in a prompt isn’t a rule.

When you ship a lot of code, waiting for every check to pass costs more than it saves. Agents open far more pull requests than humans do, and the reflex is to add more checks before a merge. OpenAI does the opposite: few mandatory checks, short-lived PRs, and a re-run for a flaky test instead of a blocked merge. In their words, “corrections are cheap, and waiting is expensive”. A bad merge costs one extra PR to fix, while a blocked queue costs everyone, every time. But they warn that the same choice “would be irresponsible in a low-throughput environment”: the right answer depends on how much you ship. Atomic CRM still requires a human review for every PR.

Looping works. Throwing away the context each time is unproven. OpenAI runs the loop in production (they call it a Ralph Wiggum Loop) at the scale of 1,500 merged PRs. The fresh-context variant rests on one uncontrolled run of a toy app, with no baseline. What actually helps isn’t the wipe, it’s that the work is written down in files.

Conclusion

The best practices outlined in this article aren’t necessarily the ones we use in Atomic CRM. Working on this study helped us clarify the trade-offs we made, and traced the path to future improvements in our harness engineering approach.

Every strong measurement we quoted scores an autonomous coding agent on a public benchmark. A harness like Atomic CRM does something else: it governs a process (separate working copies, review gates, output contracts, merge discipline) on a task it doesn’t control. We think it’s the same mechanism in both cases: deterministic scaffolding around a probabilistic step. But it’s still an analogy, and only your own task set can turn it into a measurement. Which happens to be the first item on the list.

If your harness contradicts any of this, please tell us! That’s the part we can’t get from a repository. And the developer community very much needs well-sourced insights from real-world experiences to progress further in AI harness engineering.

Authors

François Zaninotto

Marmelab founder and CEO, passionate about web technologies, agile, sustainability, leadership, and open-source. Lead developer of react-admin, founder of GreenFrame.io, and regular speaker at tech conferences.

Ready to build something extraordinary?
Our team of talented full-stack developers is ready to tackle your next web or mobile project. Let's build it together!