研究報告

Research Report · 2026-08-23

Boris Cherny 的 Agent Harness 與 Loop 設計方法

研究 Boris Cherny 如何使用 Claude Code 的 loop、Sub-agent、Dynamic Workflows、worktree 與驗證迴路,並區分可證實資訊、合理重建與社群推論。

從「同時開 5–10 個 Claude」到「讓 loops 自己提示 Claude」:拆解排程、Sub-agent fan-out、Dynamic Workflows、worktree 隔離、驗證與停止條件。

閱讀時間
14 分
更新日期
開啟原始報告
語言
繁體中文(台灣)
資料截點
2026-08-23
來源
原始訪談、Boris 貼文、Anthropic 文件、GitHub、Reddit、Hacker News

01章節

結論先行

Boris 的核心不是「寫更厲害的 prompt」,而是把自己從逐輪操作員,提升成定義目標、邊界、驗證、停止條件與資源配置的人。

1把計畫移出人的腦

人工只定義 outcome、guardrails、exit criteria;執行路徑交給模型與 workflow script。

2把中間狀態移出主 context

工作清單、分支與中間結果放進 script variables、Git、PR、進度檔與各自的 Sub-agent context。

3把「完成」交給證據

CI、test exit code、browser、benchmark、獨立 verifier 或 Stop hook,而不是 worker 自稱完成。

這套方法最適合「驗證便宜」的任務:CI 修復、auto-rebase、flaky test、migration、批次 refactor、效能 hill-climbing、回饋聚類與 repo audit。若成功無法被機器可靠判定,多跑 Agent 只會放大錯誤與審查債務。

02章節

身分與證據邊界

結論證據強度說明
他同時維持 5–10 個 session,session 內再有多個 agents第一手自述2026 Sequoia 訪談片段;他展示手機中的 Claude Code sessions。
白天有數百、夜間有數千 agents第一手自述,未獨立稽核是 Boris 公開說法,但沒有公開完整 run logs、成本、成功率或每個 agent 的任務明細。
/loop 以 cron/排程重複執行 prompt官方文件確認目前是 bundled skill;session 開啟時在本機運行,最小間隔 1 分鐘。
Dynamic Workflows 可在單次 run 使用數十至數百 agents官方文件確認目前上限為單次 run 累計 1,000 agents、同時最多 16 agents;不是 1,000 個真正同時並行。
Boris 有公開完整的私人 harness 原始碼無證據沒有找到可驗證的完整 repo。能重建的是原則、公開設定片段、產品 primitives 與工作範例。

03章節

方法演進:從人工並行到程式化編排

  1. 2026-01-02 · 「Surprisingly vanilla」個人設定

    5 個 terminal sessions,加上 5–10 個 web/mobile sessions;共享 CLAUDE.md、Plan mode、slash commands、固定 Sub-agents、hooks、permissions 與 verification。

  2. 2026-01-31 · Claude Code 團隊的十條做法

    把 parallel worktree 視為最大生產力槓桿;複雜任務先 plan;重複流程變 Skills;明確要求使用 Sub-agents。

  3. 2026-04 · 從等待 permission 轉向 Auto mode

    長任務不再需要人在旁逐次批准;強調 E2E verification,並把 verify → simplify → PR 鏈成可重複的 /go 類 Skill。

  4. 2026-05 · Sequoia:數百/數千 agents 與數十條 loops

    公開表示多數工作改由手機管理;loop 用於 PR babysitting、CI health 與每 30 分鐘整理 X feedback;Routines 則把同樣概念搬到 server。

  5. 2026-05-28 · Dynamic Workflows 公開

    Claude 即時寫 JavaScript orchestration script,以數十至數百 Sub-agents 執行 audit、migration、research 與交叉驗證。

  6. 2026-06 · 「My job is to write loops」

    他的抽象層級從手動提示多個 session,轉為撰寫會自動發現、分派、驗證與繼續工作的 loops。

  7. 2026-07 · 簡化 harness:刪除過時 scaffolding

    YC 訪談主張定期對 system prompt、CLAUDE.md、Skills 與 hooks 做 ablation;模型變強後,舊限制可能反而妨礙能力。

04章節

Boris 的七個設計思維

1. Outcome over choreography

描述要達成的結果、不可跨越的邊界與完成證據;不要把 1→2→3→4 的人類路徑硬編進 prompt。

2. Plan deeply, then delegate broadly

複雜 PR 先用 Plan mode 反覆修正計畫;計畫穩定後切入 Auto mode,減少執行途中反覆 steering。

3. Repetition becomes infrastructure

一天做多次的動作變 slash command/Skill;常見 PR 工作變 Sub-agent;生命週期規則變 hook。

4. Verification is the real multiplier

他公開估計可靠 feedback loop 可讓結果品質提高 2–3 倍。測試與 verifier 不是收尾,而是 harness 的控制器。

5. Parallelism requires isolation

agent 數量不是第一問題;共享檔案狀態才是。以 checkout/worktree 把每個寫入者隔離,再用 PR 合併。

6. Context must compound, not bloat

CLAUDE.md 記錄反覆犯錯與專案慣例;Skills 延遲載入程序知識;Sub-agent 把探索結果留在獨立 context,只回傳摘要。

7. Harness assumptions expire

模型每代變強,舊 prompt、hooks 與流程可能變成負擔。定期清空、做 ablation,只把能證明有用的規則加回來。

設計上的隱含原則

讓模型負責「如何做」,讓程式負責「何時做、能做什麼、花多少、何時停」,讓外部證據負責「是否完成」。

“My job is to write loops.”

— Boris Cherny,2026 公開訪談;「loop engineering」這個名詞則主要由後續社群系統化。

05章節

可重建的 Agent Harness 架構

下面不是 Boris 公開的原始碼,而是把他的第一手描述、官方 Claude Code primitives 與現行文件組合後,得到的最小一致架構。

可重建的 Agent Harness 架構
  1. 1 Work Discovery

    GitHub PR/CI、Slack、Sentry、X feedback、BigQuery、repo state。由 MCP、CLI、webhook 或排程取得。

  2. 2 Trigger Layer

    /loop:session 內定時輪詢。/goal:每輪結束後判斷條件。Routines/Desktop Tasks:持久排程。GitHub Actions:事件驅動。

  3. 3 Orchestrator

    Plan mode 建立方案;Dynamic Workflow 把方案寫成 JavaScript;agent()、pipeline()、分支、重試與聚合持有 orchestration state。

  4. 4 Worker Fleet

    專門化 Sub-agents:explorer、implementer、code-simplifier、verify-app、reviewer。寫入任務使用 checkout/worktree 隔離。

  5. 5 Shared Memory

    CLAUDE.md 放穩定專案規則;Skills/commands 放程序;Git/PR/issue/progress file 放 durable state;避免把所有中間資料灌回主 context。

  6. 6 Verification

    CI、unit/integration/E2E tests、browser、simulator、benchmark、independent verifier、adversarial review。輸出 pass/fail 與可追溯證據。

  7. 7 Control Plane

    permissions/sandbox、agent cap、token/budget cap、timeout、no-progress rule、audit log、workflow progress、pause/resume/kill。

最關鍵的控制關係

text
人類:定義 Outcome + Guardrails + Exit Criteria + Budget
  ↓
Trigger:時間到、事件到、或上一輪結束
  ↓
Orchestrator:發現工作 → 分解 → fan-out → 聚合 → 交叉驗證
  ↓
Worker:在隔離環境執行,產生 commit / PR / report
  ↓
Verifier:依外部可觀測證據判定 PASS / RETRY / STOP
  ↺ 若 RETRY,帶著「失敗證據」進入下一輪

06章節

他使用 /loop 的方式

Boris 描述的是「讓 Claude 用 cron 在未來反覆執行一個 job」。他的實際案例都有低成本、可輪詢、可修復或可聚合的輸出。

用途觸發Agent 行為可驗證訊號
PR babysitting每幾分鐘讀 CI/review/merge state;修 CI、auto-rebase、處理 commentCI 綠燈、無 conflict、review queue 清空
維持 CI health持續/定時發現 flaky test、重現、修復、重新跑測試重跑穩定、flaky 頻率下降、test pass
產品回饋聚類每 30 分鐘抓取 X feedback、分群、摘要與排序新回饋批次、cluster coverage、重複項去除
夜間 deeper work夜間排程/Routinessession 內再以 Sub-agent 或 workflow 分解大型任務測試、PR、report、benchmark;公開資料未揭露完整任務清單

現行 Claude Code 語法

text
# 固定每 5 分鐘執行
/loop 5m check whether CI passed and address any review comments

# 不指定間隔:Claude 每輪依狀態選擇 1 分鐘至 1 小時
/loop check the deploy and fix any issue you can verify

# 定時重跑一個可由模型自行呼叫的 Skill
/loop 20m /review-pr 1234

# 使用內建 maintenance prompt,或專案的 .claude/loop.md
/loop

# 查看/取消排程,可直接用自然語言
what scheduled tasks do I have?
cancel the deploy check job

/loop、/goal、Routines、Dynamic Workflows 的差異

Primitive下一輪如何開始停止方式適用範圍
/loop時間間隔到Esc、Claude 自停、或 7 天到期開啟中的本機 session;quick polling
/goal上一輪結束,獨立 evaluator 判定未完成條件達成、不可能、錯誤或手動 clear有清楚、可證明終態的長任務
Routines雲端 schedule、GitHub event 或 API持久設定/停用關機後仍要運行的正式 automation
Dynamic Workflowsworkflow script 的 loop/branch/pipelinescript 條件、agent cap、手動 stop數十至數百 workers 的大型 fan-out 與交叉驗證

07章節

「數百個 Sub-agent」到底如何成立

不是把數百個 agents 全部塞進同一個對話。現行 Dynamic Workflow 把 orchestration 寫成 plain JavaScript,讓中間結果存在 script variables;主對話只收到最後聚合結果。官方文件把這件事描述為「plan moved into code」。

javascript
export const meta = {
  name: 'audit-routes',
  description: 'Audit route handlers, verify findings, return one report'
}

const inventory = await agent('List every route handler.', {
  schema: routeListSchema
})

const findings = await pipeline(inventory.files, file =>
  agent(`Audit ${file}. Return structured evidence.`, { label: file })
)

const verified = await pipeline(findings.filter(Boolean), finding =>
  agent(`Independently challenge this finding: ${JSON.stringify(finding)}`)
)

return await agent(`Rank and deduplicate verified findings: ${JSON.stringify(verified)}`)

廣度

單次 run 最多累計 1,000 agents;預設 size guideline 為 medium,目標少於 15 agents。

同時性

同時最多 16 agents,並受 CPU 資源影響。因此「數千夜間 agents」更可能是跨多個 sessions/loops/runs 的累計。

成本訊號

超過 25 agents 或預估超過 150 萬 tokens 時會標示 Large workflow;警告不會自動停止。

為什麼這種結構比「主 Agent 一直開 Sub-agent」穩定

  • **決策狀態在程式裡:**loop、branch、重試、schema validation 與聚合不必佔主模型 context。

  • **每個 worker 只拿局部問題:**減少 context 污染,也能使用結構化輸出。

  • **品質模式可重複:**maker → independent reviewer → synthesizer,不靠每次臨場想起。

  • 可觀測:/workflows 顯示 phase、agent count、token、elapsed time,能 pause、resume、restart、kill。

  • **可保存:**成功的動態 script 可存進 .claude/workflows/,變成團隊共用 slash command。

08章節

Boris 方法的最小可重建藍圖

以下是從其公開方法推導出的 production-oriented 規格;不是宣稱為 Boris 私人設定的逐字複製。

A. Loop Contract

text
Objective: 保持指定 PR 可合併
Scope: 僅能修改該 PR branch;不得碰 production secrets
Discovery: 讀取 CI、review comments、merge conflict、branch freshness
Action: 一次只處理一個可重現 failure;建立最小修正
Verification: 對應 test + full required CI;列出 command 與 exit code
Stop: CI 全綠、無未處理 review、無 conflict;或連續兩輪無進展
Budget: 最多 12 輪/指定 token 或美元上限
Escalation: 權限不足、需求矛盾、production 風險時停止並留下報告
Artifact: commit、PR comment、執行證據、下一步

B. 三層工作分配

Tier 1Session loop

/loop 負責短期輪詢與維護;不需要大型 fan-out。

Tier 2Durable routine

雲端 schedule/event 觸發;每次建立獨立 session,適合夜間與跨裝置。

Tier 3Workflow fan-out

inventory → parallel workers → verifier → synthesis;只用於範圍足以抵銷成本的任務。

C. 控制面必要欄位

  • 未完成:

    每個 run 有唯一 ID、來源 trigger、owner、repo、branch/worktree。

  • 未完成:

    worker 的 tool allowlist、denylist、sandbox 與 secrets scope 明確。

  • 未完成:

    有 max turns、max budget、timeout、agent concurrency、total agent cap。

  • 未完成:

    每輪記錄 input、tool calls、diff、tests、verdict、token、cost、terminal state。

  • 未完成:

    成功、失敗、無進展、需人工、budget exhausted 分成不同終態。

  • 未完成:

    寫入工作以 PR/commit 為輸出,不直接合併高風險變更。

  • 未完成:

    maker 與 verifier 的角色/context 分離;高風險驗證使用 deterministic checks。

  • 未完成:

    每次模型升級重跑 ablation;移除已不再產生可測量增益的 instructions。

09章節

網路討論:支持、質疑與裁決

討論主題支持觀點質疑觀點本報告裁決
「Loops 取代 prompts」人不必逐輪觸發;工作可依 schedule/condition 自動推進。每個 loop 仍由 prompt、context 與 tool policy 組成;只是抽象層上移。不是 prompt 消失,而是 prompt 被封裝進可重複、可驗證的控制系統。
數百/數千 agents官方 Dynamic Workflows 確實支援數十到數百、單 run 累計上限 1,000。Boris 的個人數字缺乏公開 cost、success rate、task granularity 與 logs。技術上可行;個人規模只能標示為第一手自述,不能視為已稽核 benchmark。
生產力worktree、fan-out 與機器驗證能有效處理大型 migration、audit、CI 維護。Agent 數增加不等於有效吞吐;coordination、merge、review 與錯誤清理會非線性增加。衡量指標應是 verified outcomes/成本,不是 PR 數或 agent 數。
成本prompt cache、局部 worker、model routing 與先小規模測試可控制支出。Reddit 將 Dynamic Workflows稱為 token black hole;單 prompt 可消耗數十萬到百萬 tokens。成本是此模式的主要限制。沒有硬 budget 與 no-progress stop,不具 production 資格。
可靠性獨立 verifier、CI、E2E、browser 與 adversarial review 可提高可信度。LLM-as-judge 可能與 worker 共享盲點;自我驗證不等於獨立真實性。可機器判定的 check 最可靠;模型 verifier 只能補充,不能取代 critical deterministic gate。
安全permissions、sandbox、worktree、tool allowlist 與 Auto mode classifier 降低 unattended risk。connectors、外部內容與寫入工具擴大 prompt injection、secret exposure 與 supply-chain 面積。讀取廣、寫入窄;credentials 最小權限;不可逆動作與 production deploy 保留人工 gate。
利益衝突Boris 是最接近產品內部能力與 dogfooding 的第一手來源。他同時是產品負責人,有明顯宣傳誘因;Anthropic 內部 token 條件也不代表一般用戶。方法論可採用;所有產能與規模數字必須以自己的 eval、成本與錯誤率重測。

10章節

限制、風險與適用邊界

適合

有可列舉工作集合、局部寫入範圍、快速測試、可重試、可回滾、結果能由 exit code/diff/benchmark 判定。

不適合

需求仍在探索、成功高度主觀、涉及不可逆 production 操作、法規/安全責任重大、資料來源可能含敵意內容。

六個常見失敗模式

  1. 沒有真正的 exit condition:「讓 codebase 更好」會形成無限工作面。

  2. **同一模型自產自評:**worker 與 judge 共享盲點,容易把合理敘述誤當證據。

  3. **共享 checkout:**平行 agents 互相覆寫,最後成本轉移到 merge conflict。

  4. **重試沒有新資訊:**同 prompt、同 context、同工具反覆執行,只會穩定重現失敗。

  5. **token 沒有上限:**fan-out × retry × evaluator 會快速乘法放大。

  6. **人類喪失理解:**大量未閱讀 code 形成 comprehension debt;出現事故時沒有人能快速判斷變更意圖。

11章節

來源與可信度

搜尋排除中文網站。內容優先使用 Boris 第一手貼文/訪談、Anthropic/Claude 官方文件、GitHub,再以 Reddit、Hacker News 與獨立技術文章補充批判觀點。

級別來源本報告使用內容
第一手Boris Cherny:How I use Claude Code(2026-01-02)5+5–10 sessions、CLAUDE.md、Plan mode、commands、Sub-agents、hooks、permissions、verification。
第一手Boris Cherny:Claude Code team tips(2026-01-31)worktrees、plan review、Skills、Sub-agent 與團隊內部使用模式。
第一手影片Sequoia:Why Coding Is Solved, and What Comes Next手機工作、5–10 sessions、數百/數千 agents、loop 案例、Routines。
第一手影片Acquired Unplugged:Claude Code & the Future of Engineering從 manual prompting 轉為 writing loops 的抽象層轉變。
第一手影片Y Combinator:Building Claude Code(2026-07)outcome/guardrails/exit criteria、反對 micromanagement、定期 ablation 舊 harness。
官方文件Run prompts on a schedule/loop 語法、dynamic interval、loop.md、task limits、expiry、jitter。
官方文件Keep Claude working toward a goal/goal 的獨立 evaluator、停止條件、與 /loop 的差異。
官方文件Dynamic WorkflowsJavaScript orchestration、agent/pipeline、16 concurrency、1,000 total、cost warnings、保存與監控。
官方文件Run agents in parallel;WorktreesSub-agent、Agent Teams、worktree、/batch 與 workflow 的邊界。
官方文章Getting started with loopsturn-based、goal-based、time-based、proactive 四種 loop 分類。
官方文章Effective harnesses for long-running agentsinitializer/coding agent、progress artifact、跨 context 交接與增量工作。
官方文章How we built our multi-agent research systemlead agent、parallel workers、coordination、observability、eval 與 production lessons。
獨立分析Addy Osmani:Loop Engineering;Own the Outer Loop五個構件、inner/outer loop、以 evidence 劃分 Agent 與人類責任。
獨立分析Loops Win Where Verification Is Cheapcheck-shaped loops、驗證成本與 unattended error amplification。
社群Reddit:Dynamic Workflows 討論token black hole、品質、預設啟用與成本焦慮。
社群Hacker News:Boris setup 討論可靠性、利益衝突、一般用戶使用上限、並行與 worktree 的實務質疑。
GitHubanthropics/claude-code Issues用作社群可靠性與可觀測性爭議的問題庫;不以 issue 數直接推導產品品質。

本報告對「數百/數千 agents」保留自述標記;未將無公開 logs 的規模宣稱轉寫成已驗證事實。

參考資料