X (Twitter Article)
ai-engineering
Original source

Harness Engineering: How to Build AI Agents That Don’t Fall Apart

Summary

1.1M views, 2K+ bookmarks 的长文。核心论点:Agent 失败时不要改 prompt/换模型/加 context window,问题往往出在 harness(围绕模型的运行环境)。

引用 Dario Amodei: “You need an interface, you need a harness to use them” 引用 OpenAI Codex 团队: “The environment was underspecified”

三层区分

  • Prompt engineering → 改指令
  • Context engineering → 改模型看到什么
  • Harness engineering → 构建模型行动的世界

生产级 Harness 的 7 个职责

  1. Turn request into contract — 把请求转为有边界的对象(goal/inputs/output/constraints/done_when),防止 agent 做完不同的事还宣称成功
  2. Give agent a map — 用 AGENTS.md 做根导航文件指向架构/测试/产品/安全规则,按需加载而非全塞 context
  3. Scope tools with permissions — 读文件默认允许,部署/删数据需 approval。MODEL SUGGESTS → POLICY CHECKS → TOOL EXECUTES
  4. Store state outside conversation — 决策/artifacts/failures 存在对话外的持久状态,避免跨 session 丢失
  5. Verify before accepting — CODE→tests+lint, UI→screenshot+visual, RESEARCH→source check, DATA→schema+range
  6. Bound the loop — retry cap + budget 上限 + 失败后 escalate to human
  7. Record traces — 完整记录 context/tools/state changes/cost/rollback point,让 model 升级可比较、regression 可追溯

Harness 成熟度分级

  • Level 0: prompt + model
  • Level 1: project guide + tools
  • Level 2: structured state + tests + bounded loop
  • Level 3: permissions + traces + recovery + human gates

故障排查速查表

  • Missing context → add map/retrieval rule
  • Wrong tool → improve tool description/routing
  • Bad output → add validator/stronger contract
  • Repeated loop → add retry cap + escalation
  • Unsafe action → add permission gate
  • Lost decision → store in durable state
  • Unknown failure → add tracing

12 项 Harness Checklist

涵盖:success 是否预先定义、知识是否按需加载、工具是否有 failure state、执行是否隔离、决策是否持久化、循环是否有 cap、能否 rollback 等。