Files
peardata/docs/DATA-MODEL.md
T
Raven Scott c541c27ad8
CI / test (push) Successful in 1m17s
Release rolling / release (push) Successful in 8m15s
Prune work
2026-07-30 12:38:17 -04:00

215 lines
7.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Data model
Canonical shapes for metrics, anomalies, alerts, health, and jobs.
Implemented in `shared/data-model.js`, `shared/metrics.js`, and agent services.
## NodeInfo
```json
{
"nodeId": "a1b2c3…",
"hostname": "homelab-1",
"publicKeyHex": "…64 hex…",
"platform": "linux",
"arch": "x64",
"release": "6.x",
"cpus": 8,
"totalMemMiB": 32000,
"agentVersion": "0.1.0",
"protocolVersion": 1,
"startedAt": 1710000000000,
"charts": ["system.cpu", "system.ram", "…"],
"sampleIntervalMs": 1000,
"sampleCount": 42
}
```
## Metric contexts & charts
Canonical catalog lives in `shared/metrics.js` (`STATIC_CHART_DEFS` + runtime instance charts). Collector fills these from `/proc` + `/sys` on Linux (`os` / bare-os fallback elsewhere).
### Host-wide (static)
| Family | Contexts / charts | Units (typical) |
|--------|-------------------|-----------------|
| CPU / scheduler | `system.cpu`, `system.intr`, `system.ctxt`, `system.forks`, `system.processes`, `system.active_processes`, `system.load`, `system.uptime`, `system.entropy` | %, interrupts/s, load, seconds |
| Memory | `system.ram`, `mem.available`, `mem.swap`, `mem.swap_cached`, `mem.kernel`, `mem.slab`, `mem.writeback`, `mem.committed`, `mem.swapio`, `system.pgpgio`, `system.pgfaults` | MiB, KiB/s, faults/s |
| Disk aggregate | `system.io` | KiB/s (`in` / `out`) |
| Network aggregate | `system.net`, `system.ip`, `system.ipv6` | kilobits/s |
| TCP / IPv4 | `ip.tcppackets`, `ip.tcperrors`, `ip.tcpopens`, `ip.tcpsock`, `ipv4.packets`, `ipv4.errors`, `ipv4.udppackets`, `ipv4.udperrors` | packets/s, connections |
| PSI pressure | `system.*_{some,full}_pressure` | % (avg10/60/300) |
| Memory deep | `mem.zswap`, `mem.hugepages`, `mem.thp*`, `mem.reclaiming`, `mem.oom_kill`, `mem.numa`, `mem.ksm*` | MiB / events/s |
| TCP advanced | `ip.tcpsyncookies`, `ip.tcp_*_queue`, `ip.tcpconnaborts`, `ip.tcpofo`, `ip.tcpreorders`, `ip.tcphandshake`, `ip.tcpmemorypressures` | events/s |
| Softnet / ICMP | `system.softnet_stat`, `ipv4.icmp`, `ipv4.icmp_errors` | events/s |
See also `shared/metrics-deep.js` and gaps in [MISSING_METRICS.md](./MISSING_METRICS.md).
`system.cpu` dimensions: `guest_nice`, `guest`, `steal`, `softirq`, `irq`, `user`, `system`, `nice`, `iowait`, `idle`.
### Instance charts (registered at runtime)
| Context | Chart id pattern | Notes |
|---------|------------------|-------|
| `cpu.cpu` | `cpu.cpu{N}` | Per-core utilization (same dims as `system.cpu`) |
| `disk.io` / `ops` / `util` / `await` / `avgsz` / `qops` / `busy` / `iotime` / `discard` | `disk_*.{dev}` | Per-disk from `/proc/diskstats` |
| `disk.space` / `disk.inodes` | `disk_space.{mount}`, `disk_inodes.{mount}` | Per-mount via `statfs` |
| `net.net` / packets / errors / drops / speed | `net*.{iface}` | `/proc/net/dev` + sysfs speed |
| `cgroup.cpu` / `mem` / `io` | `cgroup.*.{id}` | cgroup v2 (default on Linux) |
| `sensors.temperature` / `thermal` | `sensors.*` | hwmon + thermal zones |
Chart summary objects use fields: `id`, `context`, `units`, `dimensions`, `update_every`, `first_entry`, `last_entry`, …
## MetricSample (live push)
```json
{
"chart": "system.cpu",
"context": "system.cpu",
"ts": 1710000000123,
"values": { "user": 12.4, "system": 3.1, "idle": 84.0 }
}
```
## QueryResult (history)
```json
{
"chart": "system.cpu",
"context": "system.cpu",
"labels": ["time", "user", "system", "nice", "iowait", "irq", "softirq", "idle"],
"data": [[1710000000, 12.4, 3.1, 0, 0, 0, 0, 84.0]],
"view_update_every": 1,
"after": 1709999940,
"before": 1710000000,
"points": 60,
"format": "json"
}
```
`time` is unix seconds. Nulls allowed for missing dimensions.
## Storage tiers
| Tier | Resolution | Default retention | Env / control |
|------|------------|-------------------|---------------|
| 0 (hot RAM) | 1s | 3600 points (~1h) | `PEARDATA_TIER0_POINTS` → seeds Data Manager |
| 1 (warm RAM) | avg every N samples (default 60 → ~1m) | 1440 points (~24h in RAM) | `PEARDATA_TIER1_POINTS`, `PEARDATA_TIER1_EVERY` |
| HyperDB warm | tier1 flush (static charts only) | **3 months** age rotate + **1 GiB** soft disk budget (defaults) | `PEARDATA_WARM_RETENTION_MS` / `PEARDATA_WARM_MAX_BYTES` seed; live policy in `retention.json` |
Live retention, soft disk budgets, and auto-prune are managed via desktop **Settings → Data** (RPC `getStorageInfo` / `setRetentionConfig` / `pruneHistory`). See [STORAGE-HYPERDB.md](./STORAGE-HYPERDB.md) and [user-guide/settings.md](../user-guide/settings.md).
## AnomalyEvent
```json
{
"id": "cpu_user_high:1710000000123",
"chart": "system.cpu",
"context": "system.cpu",
"dimension": "user",
"severity": "warning",
"score": 0.6,
"value": 88.2,
"threshold": 80,
"comparator": ">",
"message": "CPU user time high: user=88.2 > 80",
"ts": 1710000000123,
"cleared": false
}
```
## AlertConfig / AlertState
Config fields: `id`, `chart`, `dimension`, `warn`, `crit`, `comparator`, `lookbackSec`, `enabled`, `info`.
State adds: `status` (`CLEAR`|`WARNING`|`CRITICAL`|`UNDEFINED`), `value`, `lastStatusChange`.
## HealthSnapshot
```json
{
"status": "ok",
"score": 1,
"checks": [{ "id": "cpu_user_high", "ok": true, "detail": "…" }],
"ts": 1710000000123
}
```
Aggregate: any CRITICAL → `critical`; else any WARNING → `degraded`; else `ok`.
## JobRecord
```json
{
"id": "uuid",
"name": "snapshot",
"status": "done",
"startedAt": 1710000000000,
"finishedAt": 1710000000500,
"result": { "ok": true }
}
```
## Weights / Metric Correlations
Engine: `server/services/weights.js`. Not Pearson pairwise correlation — highlight-vs-baseline (or single-window) **scoring**.
| Method | Windows | Score meaning |
|--------|---------|---------------|
| `volume` | highlight + baseline | Average change × time above/below baseline |
| `ks2` | highlight + baseline | Two-sample KS on pairwise diffs |
| `anomaly-rate` | highlight | Anomaly / alert activity in window |
| `value` | highlight | Aggregated value (default `time_group=cv`) |
| `alerts` | none | Legacy active-alert ranking (Related panel boost) |
Default baseline when omitted: ends at highlight start, duration ≈ highlight × 4 (power-of-two snap). Reject if fewer than ~15 usable points.
Response shape:
```json
{
"method": "volume",
"view": {
"highlight": { "after": 1710000000, "before": 1710000060 },
"baseline": { "after": 1710000000, "before": 1710000000 }
},
"results": [
{ "id": "system.cpu", "weight": 0.91, "context": "system.cpu", "family": "cpu" }
],
"stats": { "scored": 100, "elapsed_ms": 35 }
}
```
Desktop UX: [user-guide/metric-correlations.md](../user-guide/metric-correlations.md). Related (Pearson/taxonomy) is separate: `shared/related-metrics.js`.
## Log entries
Normalized shape from `queryLogs` / `GET /api/v*/logs`:
```json
{
"id": "anomaly:1710000000000:system.cpu",
"ts": 1710000000000,
"source": "anomaly",
"severity": "warning",
"unit": "system.cpu",
"message": "…",
"fields": {}
}
```
Sources: `anomaly` | `audit` | `journal`. Desktop UI and omitted `source` default to **journal** (admin); viewers fall back to anomalies. See [user-guide/logs.md](../user-guide/logs.md).
## REST ↔ RPC parity
| Concept | RPC | REST |
|---------|-----|------|
| Charts | `listCharts` | `GET /api/v1/charts` |
| Data | `queryData` | `GET /api/v3/data` |
| Weights | `getWeights` | `GET /api/v1\|v2\|v3/weights` |
| Logs | `queryLogs` | `GET /api/v1\|v2\|v3/logs` |
| Contexts | `listContexts` | `GET /api/v3/contexts` |
| Nodes | `getNodeInfo` | `GET /api/v3/nodes` |
| Alerts | `listAlerts` | `GET /api/v3/alerts` |
| Export | `getAllMetrics` | `GET /api/v3/allmetrics` |
| Health | `getHealth` | `GET /api/v3/health` |