Abodi Telemetry & Predictive Self-Healing
Autonomous health monitoring, S.M.A.R.T. kernel telemetry, and pre-failure predictive data migration.
Abodi is Aarkam’s autonomous health and predictive diagnostics subsystem. Operating as a lightweight sidecar on every storage host (Abodi.Agent) alongside a central correlation engine (Abodi.Server), Abodi forecasts hardware failures and orchestrates data migration before drives catastrophically fail.
The Reactive Storage Problem
Traditional enterprise storage architectures (e.g., Ceph, MinIO, ZFS) are strictly reactive:
- A drive suffers bad sectors or mechanical death.
- The operating system encounters read/write I/O timeouts (often hanging for 30–60 seconds).
- The cluster flags the node offline and begins an emergency "recovery storm"—flooding the network with terabytes of reconstruction traffic during peak business hours.
Traditional Storage (Reactive):
Drive Degrades ──► Silent I/O Timeouts ──► Drive Dies ──► Emergency Rebuild Storm
Abodi (Predictive):
Wear Signals ──► ML Failure Forecast ──► Planned Quorum Drain ──► Zero-Downtime Swap
How Abodi Works: The Predictive Loop
Abodi eliminates emergency rebuild storms through a continuous 4-stage monitoring and prevention loop:
sequenceDiagram
autonumber
participant Node as Rokka.StorageNode
participant Agent as Abodi.Agent (Sidecar)
participant Server as Abodi.Server / ML Engine
participant Coord as Rokka.Coordinator
participant Portal as Aarkam.IO Portal
loop Every 10 Seconds
Agent->>Node: Query sysfs / WMI (SMART, IOPS, Temperature)
Agent->>Server: Stream Normalized Health Metrics
end
Note over Server: Correlate Multi-Signal Stress Facts & Run ML Model
alt Impending Drive Failure Detected (Confidence > 85%)
Server->>Coord: Emit Urgent Drive Quarantine Event
Coord->>Coord: Mark Node Read-Only & Evict from Write Ring
Coord->>Portal: Push Health Alert to Aarkam.IO Dashboard
Coord->>Node: Trigger Gentle Off-Peak Chunk Migration
end
1. Multi-Signal Metric Collection
Abodi.Agent runs directly on the storage host, reading kernel-level telemetry directly via WMI (Windows) or /sys/block and smartctl (Linux):
- Reallocated Sector Count (Attribute 5)
- Current Pending Sector Count (Attribute 197)
- Uncorrectable Sector Count (Attribute 198)
- NVMe Flash Wear Leveling & Available Spare Capacity
- IOPS Latency Jitter & Tail Spikes
2. Multi-Signal Correlation & ML Rules
Single noisy metrics do not trigger premature disk swaps. Abodi's rules engine evaluates composite degradation patterns (e.g., rising reallocated sector count accompanied by escalating 99th-percentile write latency).
3. Graceful Background Migration
When a drive's health score drops below threshold, Rokka.Coordinator proactively transitions the node to Quarantined:
- The node remains readable, serving existing data without interruption.
- New writes are routed to healthy cluster peers.
- Background healing gently copies the at-risk chunks to spare capacity at throttled rates, preventing network congestion.
Health Telemetry Ports
| Component | Port | Transport | Purpose |
|---|---|---|---|
| Abodi.Agent | 7776 (HTTP) / 57776 (HTTPS) | gRPC / REST | Host daemon reporting local disk counters |
| Abodi.Server | 7777 (HTTP) / 57777 (HTTPS) | gRPC / REST | Central telemetry aggregator & ML rules engine |