You cannot unit test creativity. You cannot assert on vibes.
Traditional testing assumes deterministic outputs. AI systems are probabilistic. The same input can produce different outputs, and "correct" is often subjective.
This post covers testing strategies that actually work for AI applications.
| Layer | Traditional | AI Equivalent |
|---|
| Unit tests | Function assertions | Component validation |
| Integration tests | API contracts | Pipeline tests |
| E2E tests | User flows | Scenario evaluation |
| Manual QA | Human review | Human evaluation |
| Layer | Coverage | Frequency | Cost |
|---|
| Deterministic tests | 20% | Every commit | Low |
| Evaluation suites | 50% | Every PR | Medium |
| Human evaluation | 20% | Weekly | High |
| Production monitoring | 10% | Continuous | Medium |
| Component | Description | Example |
|---|
| Inputs | Representative queries | "What is your return policy?" |
| Expected outputs | Ideal responses | Policy explanation |
| Metadata | Context, category | Category: FAQ, Difficulty: Easy |
| Annotations | Quality dimensions | Accuracy: 5, Helpfulness: 4 |
| Source | Quality | Volume | Cost |
|---|
| Production samples | High | High | Low |
| Expert creation | Very high | Low | High |
| Synthetic generation | Medium | Very high | Low |
| User feedback | High | Medium | Low |
| Use Case | Minimum Size | Recommended | Update Frequency |
|---|
| Smoke tests | 20-50 | 100 | Monthly |
| Regression | 100-200 | 500 | Bi-weekly |
| Full evaluation | 500-1000 | 2000+ | Weekly |
| Category | Percentage | Purpose |
|---|
| Happy path | 40% | Core functionality |
| Edge cases | 25% | Boundary conditions |
| Adversarial | 15% | Robustness |
| Regression | 20% | Known issues |
| Metric | What It Measures | Automation |
|---|
| Exact match | Identical output | Full |
| BLEU/ROUGE | Text similarity | Full |
| Semantic similarity | Meaning match | Full |
| Format compliance | Structure validity | Full |
| Factual accuracy | Correctness | Partial |
| Approach | Accuracy | Cost | Speed |
|---|
| Simple prompt | 70% | Low | Fast |
| Rubric-based | 85% | Medium | Medium |
| Multi-judge | 90% | High | Slow |
| Fine-tuned judge | 92% | Medium | Medium |
| Component | Purpose | Example |
|---|
| Task description | Context | "Evaluate customer service response" |
| Criteria | What to assess | "Accuracy, helpfulness, tone" |
| Scale | Scoring range | "1-5 for each dimension" |
| Examples | Calibration | "5 means completely accurate..." |
| Method | Accuracy | Speed | Cost |
|---|
| Binary (good/bad) | 85% | Fast | Low |
| Likert scale (1-5) | 80% | Medium | Medium |
| Comparative (A vs B) | 90% | Medium | Medium |
| Detailed rubric | 95% | Slow | High |
| Approach | Detection Rate | False Positives |
|---|
| Exact match diff | 60% | High |
| Semantic diff | 80% | Medium |
| Quality score diff | 90% | Low |
| Human review | 98% | Very low |
| Step | Action | Automation |
|---|
| 1 | Run evaluation suite | Automated |
| 2 | Compare to baseline | Automated |
| 3 | Flag significant changes | Automated |
| 4 | Review flagged items | Manual |
| 5 | Update baseline or fix | Manual |
| Metric | Warning | Block |
|---|
| Overall quality | -2% | -5% |
| Accuracy | -3% | -5% |
| Format compliance | -5% | -10% |
| Latency | +20% | +50% |
| Stage | Tests | Block Deploy |
|---|
| Pre-commit | Format, lint | No |
| PR | Smoke tests (50 cases) | Yes if over 10% drop |
| Merge | Full eval (500 cases) | Yes if over 5% drop |
| Deploy | Canary eval | Yes if over 3% drop |
| Test Type | Duration | When to Run |
|---|
| Smoke | Under 2 min | Every PR |
| Standard | 5-15 min | Pre-merge |
| Full | 30-60 min | Nightly |
| Comprehensive | 2-4 hours | Weekly |
| Strategy | Savings | Trade-off |
|---|
| Cache embeddings | 50% | Stale risk |
| Sample evaluation | 70% | Less coverage |
| Parallel execution | 0% (faster) | More compute |
| Staged rollout | 40% | Delayed feedback |
| Phase | Traffic | Duration | Rollback Trigger |
|---|
| Shadow | 0% (parallel) | 1 hour | Any regression |
| Canary | 5% | 2 hours | Over 5% quality drop |
| Gradual | 25% | 4 hours | Over 3% quality drop |
| Full | 100% | - | Over 2% quality drop |
| Element | Approach | Sample Size |
|---|
| Prompts | 50/50 split | 1000+ per variant |
| Models | Weighted routing | 500+ per variant |
| Parameters | Multi-arm bandit | 200+ per variant |
| Metric | Collection | Target |
|---|
| User ratings | Thumbs up/down | Over 85% positive |
| Regeneration rate | Track retries | Under 15% |
| Task completion | Flow tracking | Over 90% |
| Error rate | Logging | Under 2% |
| Category | Examples | Impact |
|---|
| Prompt injection | "Ignore instructions..." | Security |
| Jailbreaks | Role-play attacks | Safety |
| Edge inputs | Empty, very long | Reliability |
| Manipulation | Leading questions | Accuracy |
| Test Type | Frequency | Automation |
|---|
| Known attacks | Every PR | Full |
| Fuzzing | Nightly | Full |
| Red team | Monthly | Manual |
| Bug bounty | Ongoing | External |
| Defense | Test Method | Success Criteria |
|---|
| Input filtering | Known payloads | 99% blocked |
| Output filtering | Harmful content | 99.9% caught |
| Rate limiting | Burst traffic | No degradation |
| Fallbacks | Simulated failures | Graceful handling |
| Component | Size | Purpose |
|---|
| Smoke tests | 50 cases | Basic functionality |
| Format tests | 20 cases | Output structure |
| Edge cases | 30 cases | Boundary conditions |
| Component | Size | Purpose |
|---|
| Full evaluation | 500 cases | Comprehensive quality |
| Regression suite | 200 cases | Change detection |
| Adversarial | 100 cases | Security/safety |
| Performance | 50 cases | Latency benchmarks |
| Component | Size | Purpose |
|---|
| Golden dataset | 2000+ cases | Full coverage |
| Domain-specific | 500+ per domain | Specialized testing |
| User scenarios | 200+ flows | End-to-end |
| Continuous monitoring | Real-time | Production health |
- 1Golden datasets are essential - You need curated examples with expected outputs to measure quality.
- 2LLM-as-judge scales - Use LLMs to evaluate LLMs for cost-effective automated evaluation.
- 3Regression testing prevents disasters - Block deploys that drop quality below thresholds.
- 4Production is the final test - Canary deployments catch issues that lab tests miss.
- 5Adversarial testing is mandatory - If you do not test for attacks, attackers will find them.
- 6Start small, grow systematically - 50 well-chosen test cases beat 1000 random ones.
Testing AI is harder than testing traditional software, but it is not optional. Build your evaluation infrastructure early, or pay for it later in production incidents.