/investigate — Debug với iron law

Skill mới trong bộ claudekit — /investigate tuân theo Iron Law: “Không có fix mà không có root cause”. 4 pha systematic: investigate (thu thập evidence) → analyze (tìm pattern) → hypothesize (dự đoán gốc) → implement (fix + verify).

Tổng quan

/investigate là phương pháp debugging hệ thống thay vì trial-and-error. Thay vì thử nhiều fix, /investigate buộc dev tìm hiểu tại sao bug xảy ra trước khi code. Iron Law: Mỗi fix phải có root cause đi kèm. Nếu không biết gốc rễ → không dùng /fix, dùng /investigate trước. Quy trình:
  1. Investigate — Tập hợp dữ liệu: logs, stack trace, behavior pattern
  2. Analyze — Suy luận: đâu là điểm khác biệt giữa success & fail?
  3. Hypothesize — Dự đoán: root cause là gì? Cách test hypothesis?
  4. Implement — Viết fix + test tạo scenario fail trước, pass sau

Prefix & flags

Chỉ có một cú pháp duy nhất:

Workflow

1

Pha 1: Investigate

Thu thập dữ liệu:
  • Đọc error message, stack trace
  • Replay bug: exact steps
  • Check logs, network tab, console
  • Kiểm tra recent changes
Output: Issue timeline + evidence
2

Pha 2: Analyze

Tìm pattern:
  • Khi nào bug xảy ra? (always / sometimes / specific condition)
  • Cái gì thay đổi? (new code / config / data)
  • Liên quan đến gì? (DB / API / frontend / environment)
  • Có ai khác bị không?
Output: Pattern analysis + suspect list
3

Pha 3: Hypothesize

Dự đoán root cause:
  • Nếu pha 2 → suspect list = [A, B, C]
  • Dự đoán: root cause là A (vì X)
  • Plan test: “Nếu là A, thì Y sẽ happen”
  • Test hypothesis → confirm / refute
Output: Confirmed root cause + rationale
4

Pha 4: Implement

Code fix:
  • Write test: “Nếu root cause fixed, test pass”
  • Implement fix
  • Verify test pass + regression test
  • Link PR: “Fixes #XXX, root cause: Y”
Output: PR với root cause documented

Ví dụ thực tế

Case 1: 500 error trên staging

Workflow:
  • Investigate: Check logs → “TypeError: cannot read property ‘title’ of undefined”
  • Analyze: Xảy ra khi create page với empty title. Mới cập nhật validation schema hôm qua
  • Hypothesize: Schema validation không match database trigger → null title bypass validation
  • Implement: Add NOT NULL constraint, update validation, test pass

Case 2: Flaky test

Workflow:
  • Investigate: Run test 10 lần → fail 2 lần. Stack trace: “Timeout waiting for API”
  • Analyze: Flaky = timing issue. API mock occasionally slow, test timeout 5s
  • Hypothesize: Test không handle async properly. Promise.all vs sequential
  • Implement: Use waitFor() instead of hardcoded timeout

Case 3: iOS Safari checkout issue

Workflow:
  • Investigate: iOS Safari console → “Form.submit() not supported”. Button click handler missing
  • Analyze: Desktop: <form onSubmit>. iOS Safari: form submit event not firing
  • Hypothesize: Missing event handler on submit button (refactored last week, forgot iOS compat)
  • Implement: Add fallback onClick handler for iOS

So sánh với skill khác

Sự khác nhau:
  • /investigate = structured, 4-pha, documented root cause
  • /ck-debug = interactive, real-time, explores code live
  • /fix = direct implementation, cần root cause đã biết sẵn

Common pitfalls

Sai lầm phổ biến:
  • Skip investigate, jump to fix: Thử fix ngẫu nhiên = waste time. Root cause phải clear
  • Không document root cause: Fix merged nhưng nhóm không hiểu lý do → bug lặp lại sau
  • Investigate mà không test hypothesis: “Tôi nghĩ là…” ≠ “Tôi chứng minh được”. Test trước fix
  • Dùng /investigate cho non-bugs: “Why does feature X work this way?” → dùng /ck-debug / docs, không phải /investigate
  • Chạy investigate mà code không commited: Nếu code dirty, evidence không accurate. Commit trước hoặc stash

FAQ

A: /investigate loop lại:
  1. Pha 3 hypothesis sai → tìm suspect khác
  2. Quay lại pha 1 gather thêm evidence
  3. Repeat cho tới khi root cause rõ ràng
Nếu stuck quá lâu (>2 hours): escalate, dùng /ck-debug interactive mode
A: Có. Một quy trình thường:
Nhưng /investigate output phải rõ ràng (root cause + test plan)
A: /investigate vẫn chạy:
  • Investigate: Check logs, API response, timeout
  • Analyze: External service behavior vs expected
  • Hypothesize: Rate limit? Schema change? Migration issue?
  • Implement: Fallback logic / retry / alerting
(Fix có thể là workaround nếu không control external service)
A: Yes. PR message phải có:
Nó giúp code review hiểu tại sao fix cần thiết.
A: Dừng khi:
  • Root cause clear & documented ← 90% confidence
  • Có test case confirm root cause
  • Có fix plan (cách implement)
Nếu còn < 80% confident → continue investigating

Xem thêm