Harness Engineering: How to Build AI Agents That Don’t Fall Apart
- URL: https://x.com/0xwhrrari/status/2093685107534000560
- Date Saved: 2026-09-03
- Source: X (Twitter Article)
- Tags: ai-engineering
- Author: rari (@0xwhrrari) — building @kollectivexyz, AI + prediction markets
Summary
1.1M views, 2K+ bookmarks 的长文。核心论点:Agent 失败时不要改 prompt/换模型/加 context window,问题往往出在 harness(围绕模型的运行环境)。
引用 Dario Amodei: “You need an interface, you need a harness to use them” 引用 OpenAI Codex 团队: “The environment was underspecified”
三层区分
- Prompt engineering → 改指令
- Context engineering → 改模型看到什么
- Harness engineering → 构建模型行动的世界
生产级 Harness 的 7 个职责
- Turn request into contract — 把请求转为有边界的对象(goal/inputs/output/constraints/done_when),防止 agent 做完不同的事还宣称成功
- Give agent a map — 用 AGENTS.md 做根导航文件指向架构/测试/产品/安全规则,按需加载而非全塞 context
- Scope tools with permissions — 读文件默认允许,部署/删数据需 approval。MODEL SUGGESTS → POLICY CHECKS → TOOL EXECUTES
- Store state outside conversation — 决策/artifacts/failures 存在对话外的持久状态,避免跨 session 丢失
- Verify before accepting — CODE→tests+lint, UI→screenshot+visual, RESEARCH→source check, DATA→schema+range
- Bound the loop — retry cap + budget 上限 + 失败后 escalate to human
- Record traces — 完整记录 context/tools/state changes/cost/rollback point,让 model 升级可比较、regression 可追溯
Harness 成熟度分级
- Level 0: prompt + model
- Level 1: project guide + tools
- Level 2: structured state + tests + bounded loop
- Level 3: permissions + traces + recovery + human gates
故障排查速查表
- Missing context → add map/retrieval rule
- Wrong tool → improve tool description/routing
- Bad output → add validator/stronger contract
- Repeated loop → add retry cap + escalation
- Unsafe action → add permission gate
- Lost decision → store in durable state
- Unknown failure → add tracing
12 项 Harness Checklist
涵盖:success 是否预先定义、知识是否按需加载、工具是否有 failure state、执行是否隔离、决策是否持久化、循环是否有 cap、能否 rollback 等。