This document describes the comprehensive error handling and recovery system implemented in LibraVDB as part of the competitive features phase 3.
The advanced error handling system provides:
- Structured Error Types - Detailed error classification with context and metadata
- Automatic Recovery Mechanisms - Intelligent recovery strategies for different failure modes
- Graceful Degradation - System-wide degradation under resource pressure
- Circuit Breaker Pattern - Fault tolerance for failing components
- Health Monitoring - Continuous system health assessment
- Recovery Orchestration - Coordinated recovery across all system components
Enhanced error type with rich context and recovery information:
type VectorDBError struct {
Code ErrorCode `json:"code"`
Message string `json:"message"`
Details any `json:"details,omitempty"`
Retryable bool `json:"retryable"`
Severity ErrorSeverity `json:"severity"`
RecoveryAction RecoveryAction `json:"recovery_action"`
Context *ErrorContext `json:"context,omitempty"`
Cause error `json:"cause,omitempty"`
Timestamp time.Time `json:"timestamp"`
RetryCount int `json:"retry_count"`
MaxRetries int `json:"max_retries"`
}Features:
- Structured error codes for different failure types
- Severity levels (Info, Warning, Error, Critical, Fatal)
- Recovery action hints (None, Retry, Fallback, Graceful Degradation, Restart, Rebuild)
- Rich context with component, operation, stack trace, and metadata
- Automatic retry logic with backoff
Usage:
// Create a structured error
err := NewVectorDBErrorWithContext(
ErrCodeMemoryExhausted,
"memory limit exceeded during vector insertion",
true, // retryable
"collection",
"insert",
).WithSeverity(SeverityCritical).
WithRecoveryAction(RecoveryGracefulDegradation).
WithMetadata("memory_usage", "1.2GB").
WithRequestID("req-123")Coordinates automatic error recovery using pluggable strategies:
type ErrorRecoveryManager struct {
recoveryStrategies map[ErrorCode]RecoveryStrategy
circuitBreakers map[string]CircuitBreaker
degradationManager GracefulDegradationManager
healthMonitor SystemHealthMonitor
}Recovery Strategies:
- Memory Pressure Recovery - GC, cache eviction, memory mapping
- Quantization Recovery - Fallback to uncompressed storage
- Index Corruption Recovery - Index rebuild or simplification
- Batch Operation Recovery - Chunking, retry with backoff
Usage:
erm := NewErrorRecoveryManager()
// Register recovery strategies
memoryStrategy := NewMemoryPressureRecoveryStrategy(memoryManager)
erm.RegisterRecoveryStrategy(ErrCodeMemoryExhausted, memoryStrategy)
// Attempt recovery
err := NewVectorDBError(ErrCodeMemoryExhausted, "memory exhausted", true)
if recoveryErr := erm.TryRecover(ctx, err); recoveryErr != nil {
// Recovery failed
}Implements system-wide graceful degradation under various failure conditions:
type GracefulDegradationManager interface {
HandleMemoryPressure(ctx context.Context, pressureLevel int) error
HandleQuantizationFailure(ctx context.Context, fallbackEnabled bool) error
HandleIndexCorruption(ctx context.Context, rebuildEnabled bool) error
GetDegradationLevel() int
SetDegradationLevel(level int) error
}Degradation Levels:
- Level 1 - Garbage collection, lightweight cleanup
- Level 2 - Cache eviction, memory optimization
- Level 3 - Memory mapping activation
- Level 4 - Quantization fallback to uncompressed storage
- Level 5 - Index simplification (aggressive)
Configuration:
config := DegradationConfig{
EnableQuantizationFallback: true,
EnableMemoryMapping: true,
EnableCacheEviction: true,
EnableIndexSimplification: false, // More aggressive
MaxDegradationLevel: 5,
AutoRecoveryEnabled: true,
RecoveryCheckInterval: 30 * time.Second,
}
gdm := NewGracefulDegradationManager(config)Implements the circuit breaker pattern for fault tolerance:
type CircuitBreaker interface {
Execute(ctx context.Context, fn func() error) error
State() string // "CLOSED", "OPEN", "HALF_OPEN"
Reset()
}States:
- CLOSED - Normal operation, requests allowed
- OPEN - Circuit open, requests rejected immediately
- HALF_OPEN - Testing recovery, limited requests allowed
Configuration:
config := CircuitBreakerConfig{
Name: "quantization",
MaxFailures: 5,
Timeout: 30 * time.Second,
MaxRequests: 3,
FailureThreshold: 0.6, // 60% failure rate
MinRequests: 10,
}
breaker := NewCircuitBreaker(config)Continuously monitors system health and triggers recovery:
type SystemHealthMonitor interface {
RegisterHealthCheck(name string, check HealthCheck) error
GetHealthStatus() HealthStatus
Start(ctx context.Context) error
Stop() error
}Health Levels:
- HEALTHY - Component operating normally
- DEGRADED - Component experiencing minor issues
- UNHEALTHY - Component experiencing significant issues
- CRITICAL - Component in critical state
- UNKNOWN - Health status unknown
Usage:
monitor := NewSystemHealthMonitor(5 * time.Second)
// Register health checks
monitor.RegisterHealthCheck("memory", func(ctx context.Context) (HealthLevel, error) {
usage := memoryManager.GetUsage()
if usage.Limit > 0 && float64(usage.Total)/float64(usage.Limit) > 0.9 {
return HealthUnhealthy, fmt.Errorf("high memory usage")
}
return HealthHealthy, nil
})
monitor.Start(ctx)Coordinates recovery across all system components:
type AutomaticRecoveryOrchestrator struct {
errorRecoveryManager *ErrorRecoveryManager
memoryRecoveryManager MemoryRecoveryManager
quantizationRecoveryManager QuantizationRecoveryManager
batchRecoveryManager BatchRecoveryManager
}Features:
- Cross-component recovery coordination
- Recovery attempt tracking and statistics
- Failure pattern analysis
- Recovery success rate monitoring
| Error Code | Description | Recovery Actions |
|---|---|---|
ErrCodeMemoryExhausted |
Memory limit exceeded | GC → Cache eviction → Memory mapping → Quantization fallback |
ErrCodeMemoryPressure |
High memory usage | Proactive cleanup and optimization |
ErrCodeMemoryMappingFailure |
Memory mapping failed | Retry with different parameters, fallback to RAM |
ErrCodeCacheFailure |
Cache operation failed | Cache cleanup, size reduction |
| Error Code | Description | Recovery Actions |
|---|---|---|
ErrCodeQuantizationFailure |
Quantization operation failed | Retry with reduced complexity → Fallback to uncompressed |
ErrCodeQuantizationCorruption |
Quantization data corrupted | Retrain quantizer → Fallback to uncompressed |
ErrCodeQuantizationTraining |
Training failed | Adjust parameters → Reduce complexity → Fallback |
| Error Code | Description | Recovery Actions |
|---|---|---|
ErrCodeIndexCorruption |
Index data corrupted | Rebuild index → Simplify index → Fallback to flat search |
ErrCodeIndexFailure |
Index operation failed | Retry → Rebuild → Simplify |
ErrCodeIndexRebuild |
Index rebuild required | Automatic rebuild with progress tracking |
| Error Code | Description | Recovery Actions |
|---|---|---|
ErrCodeBatchFailure |
Batch operation failed | Retry failed items → Reduce batch size → Sequential processing |
ErrCodeBatchTimeout |
Batch operation timed out | Reduce batch size → Increase timeout → Chunking |
ErrCodeBatchSizeLimit |
Batch too large | Automatic chunking with optimal size |
// Create collection with error handling
collection, err := db.CreateCollection("vectors", CollectionConfig{
Dimension: 128,
Metric: DistanceEuclidean,
})
if err != nil {
if vectorErr, ok := err.(*VectorDBError); ok {
log.Printf("Error: %s (Code: %d, Severity: %s)",
vectorErr.Message, vectorErr.Code, vectorErr.Severity)
if vectorErr.IsRetryable() {
// Retry logic
}
}
return err
}// Setup comprehensive error recovery
erm := NewErrorRecoveryManager()
// Memory recovery
memoryManager := NewMemoryManager(MemoryConfig{MaxMemory: 2 * 1024 * 1024 * 1024}) // 2GB
memoryStrategy := NewMemoryPressureRecoveryStrategy(memoryManager)
erm.RegisterRecoveryStrategy(ErrCodeMemoryExhausted, memoryStrategy)
// Quantization recovery
quantStrategy := NewQuantizationRecoveryStrategy(true) // Enable fallback
erm.RegisterRecoveryStrategy(ErrCodeQuantizationFailure, quantStrategy)
// Graceful degradation
config := DefaultDegradationConfig()
gdm := NewGracefulDegradationManager(config)
gdm.SetMemoryManager(memoryManager)
erm.SetDegradationManager(gdm)
// Health monitoring
healthMonitor := NewSystemHealthMonitor(5 * time.Second)
healthMonitor.RegisterHealthCheck("memory", memoryHealthCheck)
healthMonitor.RegisterHealthCheck("quantization", quantizationHealthCheck)
erm.SetHealthMonitor(healthMonitor)
// Start monitoring
healthMonitor.Start(ctx)
gdm.StartAutoRecovery(ctx)// Setup circuit breakers for critical components
cbManager := NewCircuitBreakerManager()
quantizationBreaker := cbManager.GetOrCreate("quantization",
DefaultCircuitBreakerConfig("quantization"))
// Use circuit breaker for quantization operations
err := quantizationBreaker.Execute(ctx, func() error {
return quantizer.Train(ctx, vectors)
})
if err != nil {
log.Printf("Quantization failed: %v (Circuit state: %s)",
err, quantizationBreaker.State())
}// Health-aware error handler
handler := NewHealthAwareErrorHandler(healthMonitor, erm)
// Operations automatically adapt to system health
err := someOperation()
if err != nil {
recoveredErr := handler.HandleError(ctx, err)
if recoveredErr != nil {
// Recovery failed, handle appropriately
return recoveredErr
}
// Recovery succeeded, continue
}// Get recovery statistics
orchestrator := NewAutomaticRecoveryOrchestrator(erm)
stats := orchestrator.GetRecoveryStats()
fmt.Printf("Recovery Success Rate: %.2f%%\n", stats.SuccessRate*100)
fmt.Printf("Total Attempts: %d\n", stats.TotalAttempts)
fmt.Printf("Average Duration: %v\n", stats.AverageDuration)
// By error type
for errorCode, count := range stats.ByErrorCode {
fmt.Printf("Error %d: %d attempts\n", errorCode, count)
}// Monitor health status changes
healthMonitor.RegisterCallback(func(status HealthStatus) {
log.Printf("System health: %s", status.Overall)
for component, level := range status.Components {
if level != HealthHealthy {
log.Printf("Component %s: %s", component, level)
}
}
})// Track degradation events
gdm.RegisterCallback(func(event DegradationEvent) {
log.Printf("Degradation %s: %s (Level %d, Success: %t)",
event.Action, event.Type, event.Level, event.Success)
})
// Get degradation history
history := gdm.GetDegradationHistory()
for _, event := range history {
fmt.Printf("%s: %s - %s\n",
event.Timestamp.Format(time.RFC3339),
event.Type,
event.Reason)
}- Use appropriate error codes for different failure types
- Set correct severity levels (Critical for data loss, Warning for performance issues)
- Include relevant context and metadata
- Use structured errors consistently across the codebase
- Implement recovery strategies that are idempotent
- Use progressive recovery (try lightweight solutions first)
- Include circuit breakers for external dependencies
- Log recovery attempts and outcomes for debugging
- Design degradation levels that maintain core functionality
- Make degradation actions reversible when possible
- Monitor system health continuously
- Implement automatic recovery when conditions improve
- Register health checks for all critical components
- Use appropriate health check intervals (not too frequent)
- Include meaningful error messages in health check failures
- Monitor health trends, not just current status
- Test error conditions and recovery paths
- Simulate various failure scenarios
- Verify graceful degradation behavior
- Test circuit breaker state transitions
- Validate health monitoring accuracy
The error handling system is designed to be lightweight:
- Error objects are created only when needed
- Recovery strategies are stateless and reusable
- Health monitoring uses efficient polling
- Circuit breakers have minimal overhead
- Error creation: ~1-5μs overhead
- Recovery attempts: Variable (depends on strategy)
- Health checks: Configurable interval (default 5s)
- Circuit breaker checks: ~100ns overhead
All components are thread-safe and designed for concurrent use:
- Error recovery manager uses read-write locks
- Health monitor supports concurrent health checks
- Circuit breakers handle concurrent requests safely
- Degradation manager coordinates actions atomically
- Recovery Loops - Ensure recovery strategies don't cause the same error
- Circuit Breaker Stuck Open - Check failure thresholds and timeout settings
- Health Check Timeouts - Verify health check implementation efficiency
- Memory Leaks in Error Tracking - Recovery history is automatically pruned
Enable detailed error logging:
// Log all error recovery attempts
erm.RegisterRecoveryStrategy(errorCode, &LoggingRecoveryStrategy{
underlying: actualStrategy,
logger: log.New(os.Stdout, "RECOVERY: ", log.LstdFlags),
})Monitor system metrics:
// Export metrics for monitoring
metrics := map[string]interface{}{
"recovery_success_rate": stats.SuccessRate,
"circuit_breaker_states": cbManager.GetStates(),
"degradation_level": gdm.GetDegradationLevel(),
"health_status": healthMonitor.GetHealthStatus(),
}This advanced error handling system provides comprehensive fault tolerance and automatic recovery capabilities, making LibraVDB more robust and suitable for production deployments.