# Performance profile and reproducible workload

Run `moon run examples/benchmark --target native` once to build and verify the workload. Then measure repeated runs with an external wall-clock and memory tool appropriate to the host. The executable performs 20 encode/reconstruct cycles with `k=6`, `m=3`, 8192 bytes per source shard, and two fixed erasures. It checks every reconstructed byte, prints the source byte count, and reports 1 cache miss followed by 19 hits. The workload does not claim a particular throughput on other hardware or backends.

For a comparison, temporarily change the example to call `codec.reconstruct(erased)` instead of `session.reconstruct(erased)`, hold the same build mode, input, target and repetitions, and compare medians over several process runs. Restore the example afterward. The expected algorithmic difference is one dense `O(k^3)` inversion instead of twenty; field multiplication over shard bytes still dominates for long shards. For a single small stripe, session bookkeeping may outweigh the saved inversion.

The portable implementation uses table-backed GF(256) multiplication and dense matrices. Encoding is `O(m k L)` for shard length `L`; a cache miss adds `O(k^3)` inversion, while a hit avoids it. Working memory includes input and returned frames. The complete-object API holds all frames; incremental APIs limit internal pending data to one stripe but callers can still retain outputs. No SIMD or native-only optimization is assumed. Performance figures should be collected on the actual target and workload before making capacity claims.
