Incident Response — Xử lý sự cố production

Quy trình xử lý sự cố production có hệ thống: emergency triage, coordination, root cause analysis, postmortem.
Cập nhật bộ mới: /investigate (phân tích root cause — iron law: không fix mà không biết nguyên nhân), /hotfix (fix khẩn cấp), /day-one-patch (patch trong 24h), /guard (debug production an toàn).

Cú pháp


Luồng hoạt động

1

Bước 1: Triage

Đánh giá severity (P1-P4), impact scope, affected services.
  • P1: Service down, data loss — tất cả hands on deck
  • P2: Major feature broken — team lead + on-call
  • P3: Minor degradation — on-call engineer
  • P4: Cosmetic issue — normal sprint
2

Bước 2: Coordinate

Phân công roles: Incident Commander, Tech Lead, Communicator. Tạo war room (Lark group/thread).
3

Bước 3: Mitigate

Ưu tiên giảm impact trước, fix root cause sau. Options: rollback, feature flag off, scale up, hotfix.
4

Bước 4: Root Cause Analysis

Sau khi mitigate xong → tìm root cause. Dùng 5-Whys hoặc fault tree analysis.
5

Bước 5: Postmortem

Viết postmortem report: timeline, root cause, action items. Blameless — focus vào process, không blame người.

Severity Levels


Ví dụ thực tế


Khi nào dùng / không dùng


Skill bổ trợ

Prefix: /investigate, /investigate --deep, /investigate --5whyPhân tích nguyên nhân sâu. Iron law: Không fix mà không biết root cause.
Khi dùng: Sau khi mitigate incident, trước khi fix.
Prefix: /hotfix, /hotfix --verify, /hotfix --rollbackFix nhanh cho production incident. Tối giản scope, test từng bước.
Khi dùng: P1/P2 incident cần fix ngay trong vài phút.
Prefix: /day-one-patch, /day-one-patch --communicateFix được phát triển + deploy trong 24 tiếng, kèm communication.
Khi dùng: P2 incident, cần fix kỹ lưỡng trong vòng 1 ngày.
Prefix: /guard, /guard --logs, /guard --breakpointDebug trực tiếp production mà không gây downtime: inspect state, logs, breakpoint.
Khi dùng: Cần inspect production state để xác định root cause.