Production Readiness
by @wpank
Meta-skill that orchestrates logging, monitoring, error handling, performance, security, deployment, and testing skills to ensure a service is fully production-ready before launch. Use before first deploy, major releases, quarterly reviews, or after incidents.
clawhub install production-readinessπ About This Skill
name: production-readiness model: reasoning description: Meta-skill that orchestrates logging, monitoring, error handling, performance, security, deployment, and testing skills to ensure a service is fully production-ready before launch. Use before first deploy, major releases, quarterly reviews, or after incidents.
Production Readiness (Meta-Skill)
Coordinates all operational concerns into a single readiness review. Instead of duplicating domain expertise, this skill routes to specialized skills and agents for each area, then synthesizes results into a unified go/no-go assessment.
Installation
OpenClaw / Moltbot / Clawbot
npx clawhub@latest install production-readiness
Purpose
Ensure a service is production-ready by systematically checking every operational concern β logging, error handling, performance, security, deployment, testing, and documentation β before traffic hits it.
A production-ready service:
When to Use
| Trigger | Context | |---------|---------| | Before first deploy | New service going to production for the first time | | Before major release | Significant feature or architectural change shipping | | Quarterly production review | Scheduled audit of existing services | | After incident | Post-incident hardening to prevent recurrence | | Dependency upgrade | Major framework, runtime, or infrastructure change | | Team handoff | Transferring ownership of a service to another team |
Orchestration Flow
Run each area sequentially or in parallel. Each step delegates to a specialized skill or agent β this skill does not re-implement their logic.
βββββββββββββββββββββββββββββββββββββββββββββββββββ
β Production Readiness Review β
βββββββββββββββββββββββββββββββββββββββββββββββββββ€
β β
β 1. Logging & Observability βββΊ logging-observability skill
β 2. Error Handling ββββββββββββΊ error-handling-patterns skill
β 3. Performance βββββββββββββββΊ performance-agent
β 4. Security ββββββββββββββββββΊ security-review meta-skill
β 5. Deployment ββββββββββββββββΊ deployment-agent + docker-expert skill
β 6. Testing βββββββββββββββββββΊ testing-workflow meta-skill
β 7. Documentation βββββββββββββΊ /generate-docs command
β β
β βββΊ Synthesize results into go/no-go report β
βββββββββββββββββββββββββββββββββββββββββββββββββββ
Step Details
1. Logging & Observability β Structured logging, log levels, correlation IDs, metrics endpoints, distributed tracing, alerting rules 2. Error Handling β Global error boundaries, retry policies, dead-letter queues, error classification, user-facing error messages 3. Performance β Load testing results, P95/P99 latency baselines, memory/CPU profiling, database query analysis, caching strategy 4. Security β Auth/authz verification, input validation, dependency audit, secrets management, OWASP top-10 review 5. Deployment β Container hardening, rollback strategy, blue-green/canary configuration, infrastructure-as-code review 6. Testing β Unit/integration/e2e coverage, contract tests, chaos/failure injection, smoke tests in staging 7. Documentation β API docs, runbooks, architecture diagrams, on-call playbooks, ADRs for key decisions
Skill Routing Table
| Concern | Skill / Agent | Path |
|---------|--------------|------|
| Logging & Observability | logging-observability | ai/skills/tools/logging-observability/SKILL.md |
| Error Handling | error-handling-patterns | ai/skills/backend/error-handling-patterns/SKILL.md |
| Performance | performance-agent | ai/agents/performance/ |
| Security | security-review | ai/skills/meta/security-review/SKILL.md |
| Deployment (containers) | docker-expert | ai/skills/devops/docker/SKILL.md |
| Deployment (pipelines) | deployment-agent | ai/agents/deployment/ |
| Testing | testing-workflow | ai/skills/testing/testing-workflow/SKILL.md |
| Rate Limiting | rate-limiting-patterns | ai/skills/backend/rate-limiting-patterns/SKILL.md |
| Documentation | /generate-docs | ai/commands/documentation/ |
> Routing rule: Read the target skill first, follow its instructions, then return results here for synthesis.
Production Readiness Checklist
Health & Lifecycle
/healthz or /health) returns dependency statusResilience
Configuration & Secrets
Data Safety
Operational Readiness
Maturity Levels
| Level | Name | Requirements | |-------|------|-------------| | L1 | MVP | Health check, basic logging, error handling, manual deploy, unit tests, README | | L2 | Stable | Structured logging, metrics, graceful shutdown, CI/CD pipeline, integration tests, runbooks | | L3 | Resilient | Distributed tracing, circuit breakers, auto-scaling, chaos testing, SLOs, on-call rotation | | L4 | Optimized | Adaptive rate limiting, predictive alerting, canary deploys, full observability, error budgets, postmortem culture |
Progression Guidance
Incident Response
On-Call Rotation
Escalation Matrix
| Severity | Response Time | Escalation After | Stakeholder Notification | |----------|--------------|-------------------|--------------------------| | SEV-1 (outage) | 15 min | 30 min | Immediate β exec + customers | | SEV-2 (degraded) | 30 min | 1 hour | Within 1 hour β eng lead | | SEV-3 (minor) | 4 hours | Next business day | Daily standup | | SEV-4 (cosmetic) | Next sprint | N/A | Backlog |
Postmortem Template
## Incident: [Title]
Date: YYYY-MM-DD | Duration: X hours | Severity: SEV-NSummary
One-paragraph description of what happened and impact.Timeline
HH:MM β First alert fired
HH:MM β Engineer paged, investigation started
HH:MM β Root cause identified
HH:MM β Mitigation applied
HH:MM β Full resolution confirmed Root Cause
What broke and why. Link to code/config change if applicable.Impact
Users affected: N
Revenue impact: $X (if applicable)
SLO budget consumed: X% Action Items
| Action | Owner | Due Date | Status |
|--------|-------|----------|--------|
| Fix X | @eng | YYYY-MM-DD | Open |Lessons Learned
What went well
What went poorly
Where we got lucky
NEVER Do
1. NEVER skip health checks β every service must expose health endpoints; no exceptions for "simple" services 2. NEVER store secrets in code or container images β use a secrets manager; never default env vars with real values 3. NEVER deploy without a rollback plan β if you cannot roll back in under 5 minutes, you are not ready to deploy 4. NEVER ignore error budget violations β when the error budget is exhausted, freeze feature work and fix reliability 5. NEVER treat logging as optional β a service without structured logging is a service you cannot debug at 3 AM 6. NEVER go to production without runbooks β if on-call cannot resolve the top 5 failure modes without the original author, the service is not production-ready