STEP-BY-STEP TUTORIAL

Claude Code 省 Token 多模型工作流
從零設定到日常使用

2026-07-10 | 適用:Claude Code CLI | 概念來源:Diego (@diegocabezas01)
目錄
  • 觀念:為什麼這樣做能省 Token
  • Step 1:把 Fable 5 設為主模型、推理拉滿
  • Step 2:建立 deep-reasoner 子代理(Opus)
  • Step 3:建立 fast-worker 子代理(Sonnet)
  • Step 4:安裝 OpenAI Codex plugin
  • Step 5:在 CLAUDE.md 加入指揮規則
  • Step 6:驗證整套設定
  • 日常使用:怎麼下 Prompt
  • 四個會讓省 Token 效果歸零的錯誤
  • 觀念:為什麼這樣做能省 Token

    核心只有一句話:最貴的模型只當指揮,不做粗活。就像 tech lead 不該自己刻 boilerplate 一樣,Fable 5 只負責規劃、拆解、整合,實際的推理與執行分派給更便宜的模型。

    🎯 Orchestrator 總指揮
    FABLE 5(MAX REASONING)

    規劃、拆解任務、整合結果。context 必須保持精簡——它是最貴的模型。

    🧠 deep-reasoner 深度推理
    OPUS

    架構設計、複雜 debug、演算法設計。想得深,但只回傳精簡結論。

    ⚡ fast-worker 機械執行
    SONNET

    boilerplate、測試、格式化、簡單修改。快速執行,不多想。

    👥 Codex 同儕工程師
    OPENAI CODEX CLI

    不同視角的資深工程師「同儕」(不是審查者)。走 OpenAI 額度,不吃 Claude token。

    省 Token 的三層原理:

    ① 子代理各自燒自己的 context window——讀大檔案、跑長推理都發生在子代理裡,只把精簡結論回傳,總指揮的 context 始終乾淨。

    ② 貴的模型(Fable 5)的 token 只花在規劃與整合;推理消耗轉嫁給 Opus、機械消耗轉嫁給 Sonnet。

    ③ Codex 的第二意見完全不吃 Claude 用量。

    1把 Fable 5 設為主模型、推理拉滿

    打開終端機、進入你的專案資料夾、啟動 Claude Code 之後,輸入兩個指令:

    /model
    → 在選單中選 Fable 5
    
    /effort
    → 選 max

    ✓ 完成判斷:Claude Code 介面下方的狀態列顯示目前模型為 Fable 5。

    2建立 deep-reasoner 子代理(Opus)

    方式 A——用互動選單建立:

    /agents
    → 選 Create new agent
    → 名稱輸入:deep-reasoner
    → 模型選:opus
    → 描述貼上:
    Use for reasoning-heavy phases: architecture, debugging complex
    issues, algorithm design. Think thoroughly, return a concise
    conclusion the orchestrator can act on.

    方式 B——直接建檔(效果相同,方便進版控)。在專案根目錄建立 .claude/agents/deep-reasoner.md:

    ---
    name: deep-reasoner
    description: Use for reasoning-heavy phases: architecture, debugging complex issues, algorithm design. Think thoroughly, return a concise conclusion the orchestrator can act on.
    model: opus
    ---
    
    You are the deep-reasoning subagent. The orchestrator delegates you
    the hardest thinking: architecture decisions, complex debugging,
    algorithm design.
    
    - Read the relevant files yourself; your context is disposable,
      the orchestrator's is not.
    - Your FINAL answer must be a concise, actionable conclusion,
      plus a short "why". No reasoning dumps, no long file pastes —
      cite file:line instead.

    ✓ 完成判斷:再次輸入 /agents,在 Library 分頁看到 deep-reasoner · opus。

    3建立 fast-worker 子代理(Sonnet)

    同樣方式,建立第二個子代理。檔案版 .claude/agents/fast-worker.md:

    ---
    name: fast-worker
    description: Use for mechanical tasks: boilerplate, tests, formatting, simple edits. Execute efficiently.
    model: sonnet
    ---
    
    You are the fast-worker subagent. The orchestrator delegates you
    mechanical, well-specified work: boilerplate, test scaffolding,
    formatting, renames, simple multi-file edits.
    
    - Execute efficiently. Do not over-think or expand scope — if the
      task needs a design decision, stop and report back.
    - Follow existing project conventions exactly.
    - FINAL answer = terse completion report: files changed + anything
      that failed. No prose recap.

    ✓ 完成判斷:/agents Library 裡同時看到 deep-reasoner · opus 和 fast-worker · sonnet。

    4安裝 OpenAI Codex plugin

    前提:先在電腦上安裝並登入 Codex CLI(需要 OpenAI 帳號):

    npm install -g @openai/codex
    codex        # 首次執行會引導登入

    然後回到 Claude Code,依序輸入三個指令:

    /plugin marketplace add openai/codex-plugin-cc
    /plugin install codex@openai-codex
    /codex:setup

    ✓ 完成判斷:/agents 的 Plugin agents 區出現 codex:codex-rescue。

    注意:Codex 是選配。如果你沒有 OpenAI 帳號,跳過這一步也能運作——只是少了「不同視角的第二意見」,高風險決策時只剩 deep-reasoner 單一意見。

    5在 CLAUDE.md 加入指揮規則

    把下面這段原封不動加到專案根目錄的 CLAUDE.md(沒有這個檔就新建一個)。這段文字讓 Fable 5 每次開對話都自動知道自己是總指揮:

    ## Orchestration workflow
    
    You (Fable) are the orchestrator. Plan, decompose, synthesize.
    Reasoning-heavy phases → deep-reasoner
    Mechanical work → fast-worker
    Codex (/codex:rescue --background) is a cracked engineer on par with
    deep-reasoner, from a different perspective. Treat as a peer, not a
    reviewer.
    High-stakes decisions: task Opus + Codex on the same problem in
    parallel, synthesize the best of both, without showing either the
    other's answer. Keep your own context lean.

    ✓ 完成判斷:重新啟動 Claude Code 後隨便問一句「你的分工規則是什麼」,它能複述出上面的分派邏輯。

    6驗證整套設定

    用一個小任務跑一次完整流程,確認分派真的有發生:

    Goal: Add input validation to the signup form and unit tests for it.
    Context: src/components/SignupForm.tsx, we use zod + vitest.
    You're the lead. Delegate reasoning to deep-reasoner, grunt work to
    fast-worker. Show me your plan first, then execute.

    觀察重點:

    應該看到代表
    Fable 先輸出一份計畫,等你確認plan-first gate 生效
    執行時出現 Task(deep-reasoner) / Task(fast-worker) 呼叫分派真的發生,不是總指揮自己做
    子代理回報都很精簡concise-return 規則生效

    ✓ 全部符合,設定完成。

    日常使用:怎麼下 Prompt

    每次開任務,像對 tech lead 交辦一樣使用這個模板:

    Goal: [你要達成的目標+成功標準]
    Context: [相關檔案、限制條件]
    You're the lead. Delegate reasoning to deep-reasoner, grunt work to
    fast-worker, fresh-perspective problems to Codex. Show me your plan
    first, then execute.

    什麼任務派給誰(口訣)

    情境派給
    「這個 race condition 怎麼發生的」「重新設計這個模組」deep-reasoner
    「12 個檔案的 import 改成新路徑」「補單元測試」fast-worker
    「卡住了,換個腦袋」「這設計有沒有我沒想到的坑」Codex(/codex:rescue --background)
    資料庫 schema、對外 API 這種難回頭的決策deep-reasoner + Codex 平行雙盲,Fable 整合

    高風險決策的「雙盲模式」怎麼下

    High-stakes decision: [描述決策,例如 choose order-system schema:
    single-table JSONB vs normalized tables]
    Task deep-reasoner and Codex on this in parallel. Do NOT show either
    one the other's answer. Then synthesize the best of both and give me
    your recommendation with the single most important deciding factor.

    四個會讓省 Token 效果歸零的錯誤

    ❌ 錯誤 1:讓總指揮自己讀大檔案。

    正確做法:在 Prompt 裡明寫「Never read large files yourself; delegate reading + summarizing.」由子代理讀完回報摘要與 file:line。

    ❌ 錯誤 2:子代理回傳整份思考過程。

    正確做法:每個子任務都限定回傳格式與行數上限,例如「return: root cause + fix location, ≤10 lines」。

    ❌ 錯誤 3:跳過「Show me your plan first」。

    正確做法:永遠先看計畫。方向錯誤在計畫階段攔下來,成本是幾百 token;執行到一半才發現,成本是幾萬 token。

    ❌ 錯誤 4:把 Codex 當 code reviewer。

    正確做法:Codex 的價值是「不同視角的獨立解法」,把它當同儕、丟同一個問題給它,而不是事後叫它挑錯。

    完成後的專案結構

    your-project/
    ├── CLAUDE.md                      ← 含 Orchestration workflow 段落
    └── .claude/
        └── agents/
            ├── deep-reasoner.md       ← pinned to opus
            └── fast-worker.md         ← pinned to sonnet

    就這樣。(That's it.)