Add storage stack modules and track them in git

The modules workspace previously ignored every path matching `storage-*`,
which kept all Hypercore, Hyperbee, Hyperdrive, and Autobase packages out
of version control. Narrow .gitignore to test-artifact patterns only so
production category trees are committed.

Storage packages (23 total):
- storage-hypercore (7): replicator, seed-policy, fork-picker, merkle-sync,
  priority-fetch, bitfield-scheduler, audit-chain
- storage-hyperbee (5): batch-write, diff-follow, range-watch,
  secondary-index, tombstone-gc
- storage-hyperdrive (6): entry-catalog, gc-sweep, mirror-sync, mount-bridge,
  version-snapshot, watch-notify
- storage-autobase (5): fork-choice, view-sync, writer-lease, indexer-bus,
  light-writer

Implementation highlights:
- Shared attach/gossip helpers in _shared/storage-gossip-base.js
- P2P modules use Protomux gossip via p2p-bare; local planners omit swarm
- Recent deepen pass: priority queues, batch caps, audit export/import,
  watch/notify aliases on range-watch, fork/view lease helpers, etc.
- Per-package README, docs/api.md, docs/architecture.md, tests, examples

Also update modules/README.md doc hub links for the full storage stack
(hypercore, hyperbee, hyperdrive, autobase).

Co-authored-by: Cursor <[email protected]>
This commit is contained in:
Raven Scott
2026-05-21 00:17:51 -04:00
co-authored by Cursor
parent b10507b060
commit a11e22badc
214 changed files with 57132 additions and 5 deletions
@@ -0,0 +1,113 @@
# API: hyper-p2p-core-bitfield-scheduler
**Protocol:** `core-bitfield-scheduler/v1`
**Export:** `HyperP2PCoreBitfieldScheduler`
## Overview
Schedules Hypercore bitfield download ranges by priority so replication can respect bandwidth budgets when composed with `hyper-p2p-bandwidth-broker`.
## Constructor
```js
const mod = new HyperP2PCoreBitfieldScheduler(opts)
```
| Option | Type | Default | Description |
|--------|------|---------|-------------|
| `core` | object \| null | null | Attached core instance (`attach()` also supported) |
## Methods
### `attach(…)`
- **Returns:** module-specific (see implementation)
- **Throws:** — (none in method body)
### `scheduleRange(…)`
- **Returns:** module-specific (see implementation)
- **Throws:**
- `Error: start must be non-negative`
- `Error: len must be positive`
### `nextRange(…)`
- **Returns:** module-specific (see implementation)
- **Throws:** — (none in method body)
### `markRangeDone(…)`
- **Returns:** module-specific (see implementation)
- **Throws:** — (none in method body)
### `pendingCount(…)`
- **Returns:** module-specific (see implementation)
- **Throws:** — (none in method body)
### `getStats(…)`
- **Returns:** module-specific (see implementation)
- **Throws:** — (none in method body)
### `ready(…)`
- **Returns:** module-specific (see implementation)
- **Throws:** — (none in method body)
### `close(…)`
- **Returns:** module-specific (see implementation)
- **Throws:** — (none in method body)
## Events
| Event | Payload |
|-------|---------|
| `scheduled` | item `{ start, len, end, priority, at, key }` |
| `range` | scheduled range item |
| `closed` | no payload |
## getStats()
Returns `{ ...this._stats, protocol }` plus module-specific counters (pending queues, registry sizes, gossip in/out when P2P).
Local modules report hot-path counters only; P2P modules include gossip traffic when `topic` is set.
## Errors
Stable message substrings: see [`../../_shared/ERROR_CODES.md`](../../_shared/ERROR_CODES.md).
Validation helpers may throw `ValidationError` (e.g. `peer is required`, `path is required`).
### Documented `throw new Error(...)` strings
- `start must be non-negative`
- `len must be positive`
## P2P
Library-only: no swarm join. `ready()` resolves immediately.
`getStats().protocol` still reports the module protocol id for logging.
## Testing
```bash
npm install && npm test
```
## Common flows
1. `attach(core)` then `scheduleRange(start, len, priority)` to enqueue ranges.
2. `nextRange()` drains the priority queue and emits `range`.
3. `markRangeDone(start, len)` prevents re-serving the same range.
@@ -0,0 +1,59 @@
# Architecture: hyper-p2p-core-bitfield-scheduler
**Category:** Storage (Hypercore) · **Protocol:** `core-bitfield-scheduler/v1` · **P2P:** no (local planner)
## Role
Schedules **contiguous byte ranges** on a Hypercore for bitfield-driven download or upload work. Higher `priority` ranges dequeue first. Tracks completed ranges in `_served` so duplicate schedules are skipped.
Use with `hyper-p2p-core-priority-fetch` when you split work by **block index** vs **byte span**.
## State model
| Field | Type | Description |
|-------|------|-------------|
| `_queue` | `RangeItem[]` | Pending `{ start, len, end, priority, key, at }` |
| `_served` | `Set<string>` | Keys `start:len` already handed to replication |
| `core` | Hypercore? | Optional; `coreLength` in stats |
## Operations
1. `scheduleRange(start, len, priority)` — push + sort queue; emit `scheduled`
2. `peekNext()` — highest priority pending without dequeue
3. `nextRange()` — shift first unserved item; emit `range`
4. `markRangeDone(start, len)` — mark served, purge queue entries
5. `listPending()` — snapshot of unserved queue
## Wire messages
Local only — no Protomux channel. Pair with gossip modules on the same process when peers need policy agreement before scheduling.
## Events
| Event | Payload |
|-------|---------|
| `scheduled` | `RangeItem` |
| `range` | `RangeItem` dequeued |
| `closed` | — |
## Composition
- **Upstream:** Hypercore replication / custom fetcher consuming `nextRange()`
- **Peers:** `hyper-p2p-core-replicator`, `hyper-p2p-bandwidth-broker` (rate limits per range)
- **Downstream:** Holepunch `hypercore` bitfield updates (application applies ranges)
## Failure modes
- `start < 0` or `len <= 0` throws
- Re-scheduling a served key returns `null` and increments `skipped` stat
## Diagram
```mermaid
flowchart LR
App[Fetcher] --> Sched[BitfieldScheduler]
Sched --> Core[Hypercore]
Sched --> Q[Priority queue]
```
Shared: [`../../../_shared/storage-gossip-base.js`](../../../_shared/storage-gossip-base.js).