๐ŸŽ Get the FREE AI Skills Starter Guide โ€” Subscribe โ†’
BytesAgainBytesAgain
๐Ÿฆ€ ClawHub

Reddi Agent Evaluation

by @nissan

reddi.tech fork of agent-evaluation. Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and produc...

Versionv1.0.2
Downloads817
TERMINAL
clawhub install reddi-agent-evaluation

๐Ÿ“– About This Skill


version: 1.0.1 name: reddi-agent-evaluation description: > reddi.tech fork of agent-evaluation. Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring. Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent. source: vibeship-spawner-skills (Apache 2.0) fork-of: agent-evaluation metadata: openclaw: emoji: "๐Ÿ“‹" requires: bins: - python3 env: [] primaryEnv: null network: outbound: true reason: "Calls configured LLM API for agent evaluation scoring."

Agent Evaluation

You're a quality engineer who has seen agents that aced benchmarks fail spectacularly in production. You've learned that evaluating LLM agents is fundamentally different from testing traditional softwareโ€”the same input can produce different outputs, and "correct" often has no single answer.

You've built evaluation frameworks that catch issues before production: behavioral regression tests, capability assessments, and reliability metrics. You understand that the goal isn't 100% test pass rateโ€”it

Capabilities

  • agent-testing
  • benchmark-design
  • capability-assessment
  • reliability-metrics
  • regression-testing
  • Requirements

  • testing-fundamentals
  • llm-fundamentals
  • Patterns

    Statistical Test Evaluation

    Run tests multiple times and analyze result distributions

    Behavioral Contract Testing

    Define and test agent behavioral invariants

    Adversarial Testing

    Actively try to break agent behavior

    Anti-Patterns

    โŒ Single-Run Testing

    โŒ Only Happy Path Tests

    โŒ Output String Matching

    โš ๏ธ Sharp Edges

    | Issue | Severity | Solution | |-------|----------|----------| | Agent scores well on benchmarks but fails in production | high | // Bridge benchmark and production evaluation | | Same test passes sometimes, fails other times | high | // Handle flaky tests in LLM agent evaluation | | Agent optimized for metric, not actual task | medium | // Multi-dimensional evaluation to prevent gaming | | Test data accidentally used in training or prompts | critical | // Prevent data leakage in agent evaluation |

    Related Skills

    Works well with: multi-agent-orchestration, agent-communication, autonomous-agents