CONTENTS · OWNER-CREATED WRITING
ออกแบบ Cost-Aware Distributed Agent Control Plane
Execution Budget, Model Routing, Worker Isolation, State, Observability และ Human Approval
CORE ARGUMENT
Agent ที่ทำงานขนานเพิ่ม Throughput ได้ก็ต่อเมื่อ Budget, State Ownership, Verification และ Stop Condition ชัดเจน
Core thesis: การเพิ่มจำนวนเอเจนต์หรือเพิ่มจำนวนเครื่องไม่ได้แก้ปัญหาโดยอัตโนมัติ สิ่งที่ขาดจริงคือระบบควบคุมที่รู้ว่า งานใดควรคิดลึก งานใดควรตอบทันที งานใดควรกระจาย และงานใดไม่ควรสร้างซับเอเจนต์ตั้งแต่แรก
จุดเริ่มต้น: เมื่อ “งานง่าย” เริ่มใช้เวลาราวกับงานสถาปัตยกรรม
AI coding agents รุ่นใหม่สามารถสร้างซับเอเจนต์ แบ่งงานขนาน อ่านโค้ด รันคำสั่ง แก้ไฟล์ และรวมผลลัพธ์กลับมาได้แล้ว แต่ประสบการณ์ใช้งานจริงกลับเผยปัญหาอีกด้านหนึ่ง: คำถามง่ายที่เคยตอบได้แทบจะทันที เริ่มใช้เวลาตั้งแต่หลายสิบวินาทีไปจนถึงหลายนาที
ปัญหาไม่ได้แปลว่าโมเดลฉลาดน้อยลง ตรงกันข้าม ความสามารถของโมเดลเพิ่มขึ้นจน harness ให้อิสระมากขึ้น ทั้งการวางแผน การใช้เครื่องมือ การตรวจสอบซ้ำ และการสร้างซับเอเจนต์ ผลลัพธ์คือระบบมีศักยภาพสูงขึ้น แต่ ค่า coordination, context และ reasoning โตตามไปด้วย
ผมเรียกต้นทุนส่วนนี้ว่า Harness Tax — เวลาที่เสียไปกับการเตรียมบริบท เลือกเครื่องมือ วางแผน แตกงาน รอผล ตรวจสอบ และสรุป มากกว่าตัวงานจริง
ตัวอย่างที่เห็นได้ชัดคือคำถามประเภท:
- ชื่อ method นี้เหมาะไหม
- จุดนี้มีโอกาส header ซ้ำหรือไม่
- ควรใช้ alias
plan,reviewหรือanalyze - ค้นหา implementation ของ interface หนึ่งตัว
งานเหล่านี้อาจต้องการเพียงการอ่านหนึ่งไฟล์หรือใช้ grep หนึ่งถึงสองครั้ง แต่ harness ที่ไม่มีงบประมาณชัดเจนอาจเลือกใช้โมเดลระดับสูง เปิด reasoning ยาว สร้าง explorer และ reviewer แล้วจึงรวมคำตอบกลับมา
คำถามที่ถูกต้องจึงไม่ใช่เพียง “เอเจนต์ทำงานนี้ได้ไหม” แต่คือ:
“ควรใช้ intelligence, เวลา, token, tool calls และจำนวน worker เท่าใดจึงจะคุ้มกับความซับซ้อนของงานนี้?”
ความจริงเรื่องหลายเครื่อง: ช่วยได้ แต่ช่วยคนละคอขวด
เมื่อใช้ Codex หรือ Claude Code ร่วมกับโมเดลคลาวด์ การคิดของโมเดลเกิดบนโครงสร้างพื้นฐานของผู้ให้บริการ เครื่องของเรารับผิดชอบส่วน local execution เช่น:
- อ่านและเขียนไฟล์
- Git, worktree และ branch
- build, test และ lint
- Docker และ integration environment
- browser automation
- MCP servers และ local tools
ดังนั้น การมีเครื่องที่สองหรือเครื่องที่สามช่วยได้มากกับงาน build/test ที่หนัก งาน browser test หลายชุด งานที่ต้องใช้ GPU ภายใน หรือการแยก environment ของ worker แต่ ไม่ได้ทำให้ request ที่กำลังรอ model reasoning บน cloud เร็วขึ้นโดยตรง
นี่คือเหตุผลที่ distributed agents ต้องออกแบบให้แยกสอง plane ออกจากกัน:
- Reasoning plane — เลือกโมเดล ระดับ reasoning บริบท และจำนวน agent
- Execution plane — เลือกเครื่อง environment capability และ resource quota
การโยนงานทุกอย่างไปหลายเครื่องโดยไม่มี policy ไม่ใช่ scale-out ที่ดี มันเป็นเพียงการขยายความไม่มีประสิทธิภาพจากหนึ่งเครื่องไปสามเครื่อง
สถาปัตยกรรมที่ต้องการ: Agent Control Plane
ระบบที่ตอบโจทย์นี้ประกอบด้วยส่วนสำคัญดังต่อไปนี้
1. Agent Gateway
รับ intent จากผู้ใช้ IDE หรือ CI แล้วแปลงเป็น task contract ที่มีขอบเขตชัดเจน เช่น repository, base commit, allowed files, expected output และ acceptance criteria
Gateway ไม่ควรส่ง prompt ตรงไปยังโมเดลทันที แต่ต้องจำแนกก่อนว่าเป็น:
- fast answer
- repository lookup
- bounded code change
- diagnosis
- review
- architecture/design
- parallelizable workload
2. Cost-Aware Model Router
Router เลือกโมเดลจากความต้องการจริง ไม่ใช่ใช้ frontier model เป็นค่าเริ่มต้นทุกงาน
ตัวอย่าง policy:
| Task class | Model profile | Reasoning | Tool budget | Subagents |
|---|---|---|---|---|
| Rename / simple question | Fast | Low | 0–2 | 0 |
| Bounded repository lookup | Fast explorer | Low | 2–6 | 0–1 |
| CRUD / isolated change | Coding executor | Medium | 6–12 | 0–1 |
| Cross-layer diagnosis | Strong model | High | 10–20 | 1–2 |
| Architecture / security | Frontier reviewer | High | Bounded | 1–3 |
แนวคิดสำคัญคือ งานต้องพิสูจน์ความซับซ้อนก่อน จึงจะได้รับงบ reasoning เพิ่ม ไม่ใช่ได้รับงบสูงสุดตั้งแต่ต้น
3. Scheduler และ Capability Registry
Scheduler เลือก worker จาก capability เช่น:
workers:
machine-1:
capabilities: [dotnet, sql-server, docker]
max_jobs: 2
machine-2:
capabilities: [node, angular, playwright]
max_jobs: 3
machine-3:
capabilities: [gpu, local-llm, embeddings]
max_jobs: 1งาน Playwright ไม่ควรถูกส่งไปเครื่องที่ไม่มี browser dependencies และงาน .NET integration test ไม่ควรถูกส่งไป worker ที่ไม่มี SQL Server test environment การเลือก worker จึงต้องพิจารณา capability, queue depth, lease, timeout และ resource pressure
4. Isolated Worker Runtime
Worker แต่ละตัวควรทำงานใน Git worktree, branch หรือ container ที่แยกจากกัน ไม่ควรให้หลายเอเจนต์เขียนลง working directory เดียว เพราะจะเกิด race condition, context drift และ patch conflict
หลักปฏิบัติสำคัญ:
- ทุกงานอ้างอิง base commit SHA
- จำกัด write scope ตาม module หรือไฟล์
- ห้าม worker merge เอง
- ส่งผลเป็น patch พร้อม test evidence
- ยกเลิกงานเมื่อ lease หรือ heartbeat หมดอายุ
5. Result Store และ Aggregator
ผลลัพธ์ของ worker ต้องเป็น structured contract ไม่ใช่ข้อความว่า “เสร็จแล้ว”
{
"taskId": "AGM-20260718-001",
"workerId": "machine-2",
"baseCommit": "a92f30c",
"status": "completed",
"summary": "Updated request resolver",
"evidence": [
"src/Extensions/HttpRequestValueExtensions.cs:42"
],
"branch": "agent/AGM-20260718-001/machine-2",
"tests": {
"passed": 47,
"failed": 0
},
"metrics": {
"durationMs": 84213,
"toolCalls": 11,
"inputTokens": 18342,
"outputTokens": 2150
}
}Aggregator มีหน้าที่ตรวจหลักฐาน รัน verification ที่จำเป็น แก้ conflict และสังเคราะห์คำตอบสุดท้าย ไม่ควรเชื่อคำสรุปของ worker โดยไม่มี patch, test หรือ file reference รองรับ
ปัญหาหลักของ harness ปัจจุบัน
Context inflation
เมื่อ session ยาวขึ้น ประวัติสนทนา tool schema และผลลัพธ์ก่อนหน้าถูกส่งกลับเข้าโมเดลมากขึ้น คำถามสั้นจึงอาจแบกบริบทจากงานก่อนหน้าหลายหมื่น token
Model inheritance
หากซับเอเจนต์ inherit โมเดลหลัก งาน exploration ง่ายอาจถูกส่งไปโมเดลราคาแพงและ reasoning สูงโดยไม่ตั้งใจ นอกจากนี้ thinking configuration อาจถูก inherit ตาม session หลักด้วย
Unbounded fan-out
การแตกงานขนานช่วยเมื่อ work units เป็นอิสระ แต่ถ้างานเล็กหรือมี dependency สูง coordination cost จะมากกว่าประโยชน์ การให้ซับเอเจนต์สร้างซับเอเจนต์ต่อโดยไม่มี depth limit ยิ่งเพิ่ม token, latency และความยากในการติดตาม
Verification without marginal value
เอเจนต์มักถามว่า “ตรวจเพิ่มได้ไหม” แทนที่จะถามว่า “การตรวจเพิ่มนี้มีโอกาสเปลี่ยนคำตอบหรือไม่” จึงเกิดการอ่านซ้ำ วางแผนซ้ำ และ review ซ้ำแม้ evidence เพียงพอแล้ว
Policy ที่ควรฝังในระบบ
Do not spawn a subagent when the request can be answered
without repository inspection or with at most two tool calls.
Spawn subagents only when:
1. There are at least two independent work units.
2. Each unit has a bounded output contract.
3. Parallel execution saves more time than coordination costs.
4. Workers do not write to the same files or working tree.
Do not allow recursive delegation unless the operation
explicitly permits it.นอกจากนี้ operation แต่ละประเภทควรมี execution budget ของตัวเอง เช่น:
| Operation | Agent policy | Reasoning | Parallelism |
|---|---|---|---|
agm-fast | Main agent only | Low | 0 |
agm-analyze | Main + explorer เมื่อจำเป็น | Medium | 0–1 |
agm-diagnose | Explorer + verifier | High เฉพาะ final | 1–2 |
agm-review | Reviewer แยก concern | Medium/High | 2–3 |
agm-plan | Architect + feasibility check | High | 1–2 |
agm-design | Architect + critic | High | 1–2 |
agm-execute | Bounded executor | Low/Medium | ตาม module |
Observability: ต้องเห็นว่าเอเจนต์กำลังทำอะไร โดยไม่ต้องเปิดเผย chain-of-thought
สิ่งที่ระบบควรแสดงไม่ใช่ private reasoning แต่เป็น operational state:
CREATED → QUEUED → ASSIGNED → RUNNING → VERIFYING → COMPLETED
├─ WAITING_FOR_TOOL
├─ BLOCKED
├─ RETRYING
└─ TIMED_OUT / CANCELLED / FAILEDแต่ละ state transition ควรปล่อย event ที่ประกอบด้วย:
- task ID, agent ID และ parent agent ID
- model และ reasoning effort
- start time, duration และ time-to-first-token
- tool name, command duration และ exit code
- input, output และ cache tokens
- retry count และ timeout
- files touched และ patch reference
- test result และ final exit reason
OpenTelemetry เหมาะกับโจทย์นี้เพราะสามารถรวม metrics, logs และ traces เข้ากับ Grafana, Loki และ Tempo ได้ ทำให้ตอบคำถามที่ harness ปัจจุบันตอบได้ไม่ชัด เช่น:
- เอเจนต์ยังทำงานอยู่หรือค้าง
- เวลาหายไปที่ model หรือ tool
- agent ตัวใดสร้าง token มากผิดปกติ
- งานใด retry ซ้ำ
- subagent ตัวใดไม่ส่ง heartbeat
- fan-out รอบใดไม่คุ้มกับเวลาที่ประหยัดได้
Metrics ที่ควรวัด
อย่าวัดเพียงจำนวนงานที่สำเร็จ แต่ควรวัด economics ของการตัดสินใจด้วย:
- Time to First Useful Result — เวลาจนมี evidence ชิ้นแรกที่ใช้ได้
- Total Wall-Clock Time — เวลาตั้งแต่รับงานจน verify เสร็จ
- Coordination Ratio — เวลา orchestration เทียบกับเวลา execution จริง
- Token per Accepted Change — token ต่อ patch ที่ถูกยอมรับ
- Rework Rate — สัดส่วนงานที่ verifier ส่งกลับไปแก้
- Parallel Efficiency — เวลาที่ลดได้หารด้วยต้นทุน agent เพิ่ม
- Wrong-Model Rate — งานที่ถูก escalate หรือ downgrade หลังเริ่มทำ
- Idle Worker Time — worker ว่างเพราะ scheduler เลือก capability ผิดหรือ dependency ยังไม่พร้อม
Roadmap ที่สมเหตุผล
Phase 1 — Policy ก่อน Distribution
- เพิ่ม task classification
- กำหนด reasoning/tool/agent budget
- จำกัด subagent depth
- เพิ่ม explicit stop conditions
Phase 2 — Structured Telemetry
- นิยาม task state machine
- ส่ง heartbeat และ tool events
- แสดง token, latency และ retry ต่อ agent
- ทำ dashboard เปรียบเทียบ fast path กับ deep path
Phase 3 — Isolated Local Workers
- ใช้ Git worktree แยก agent
- บังคับ output contract
- เพิ่ม verifier ก่อน merge
- ทดลองหลาย worker บนเครื่องเดียวก่อน
Phase 4 — Remote Worker Pool
- เพิ่ม queue และ capability registry
- ส่งงานไปเครื่องสองหรือเครื่องสาม
- เพิ่ม lease, cancellation และ artifact store
- scale ตาม bottleneck ที่วัดได้จริง
Phase 5 — Learning Router
- ใช้ historical telemetry ปรับ model routing
- คาดการณ์เวลางานและโอกาส rework
- ทำ graduated autonomy ตามสถิติของ task category
บทสรุป
ปัญหาของยุค multi-agent ไม่ใช่การขาดโมเดลที่ฉลาดหรือการขาดเครื่องให้รัน แต่คือการขาดระบบตัดสินใจว่า ควรใช้ความฉลาดและทรัพยากรมากเท่าใดในแต่ละงาน
การเพิ่มเครื่องโดยไม่มี governance อาจเปลี่ยนจากเอเจนต์หกตัวที่คิดเกินเหตุบนเครื่องเดียว เป็นเอเจนต์สิบแปดตัวที่คิดเกินเหตุบนสามเครื่อง
ดังนั้นหลักการสำคัญที่สุดคือ:
Scale the control plane before scaling the worker pool.
นี่คือจุดเปลี่ยนจากการ “ใช้ AI เขียนโค้ด” ไปสู่การออกแบบ AI Platform และ Software Delivery System ที่วัดต้นทุน ตรวจสอบได้ และควบคุม autonomy อย่างมีวินัย
EVIDENCE NOTE
Claim boundary
เนื้อหาหน้านี้นำมาจากไฟล์บทความต้นฉบับใน Information โดยคงลำดับเหตุผลและข้อจำกัดเดิม ไม่ได้อ้างว่าทุก Pattern ถูกใช้ใน Production และไม่สร้าง Metric เพิ่มจาก Source