The research question
Modern AI accelerators derive extraordinary throughput from low-precision matrix arithmetic. Hopper made FP8 a first-class tensor format. Blackwell and CDNA4 push further into four-bit microscaled formats. Those formats are excellent for machine learning, but consensus imposes a different standard: every honest verifier must reach the same protocol-visible result.
That tension is Spinel's starting point. We are not assuming that a floating-point GEMM can simply be dropped into a blockchain. We are defining the arithmetic boundary first, then testing whether real hardware can execute it efficiently and reproducibly.
R0 workstreams
01. Canonical FP8
Specify allowed FP8 encodings, scaling, accumulation, rounding, overflow behavior, and all exceptional cases. The first target is E4M3-class arithmetic with an explicit accumulator contract.
02. MXFP4
Investigate OCP microscaling as a four-bit work class. The design goal is not vendor-specific peak FLOPS; it is a portable protocol primitive that can map efficiently onto multiple accelerator families.
03. Cross-vendor oracle parity
Every optimized implementation must be judged against an independent reference. A CPU implementation is the oracle. NVIDIA and AMD kernels are acceleration paths, not alternative definitions of correctness.
04. Transcript design
A candidate protocol needs a compact transcript bound to chain state and the performed tensor work. The transcript should be cheap enough to maintain in the mining kernel and structured enough to verify without re-running miner-scale computation.
Mining should be expensive. Verification should not.
A useful proof-of-work design fails if every full node needs the same accelerator that generated the proof. Spinel therefore treats verifier asymmetry as a first-order constraint, not a later optimization.
- Proof checking must be substantially cheaper than producing the underlying tensor work.
- Malformed proofs must be rejected before expensive verification whenever possible.
- Pool share verification must be engineered separately from full block verification.
- Verifier cost and denial-of-service resistance are protocol metrics.
Related work
Spinel is not emerging in a vacuum. Pearl Research demonstrated a Proof-of-Useful-Work protocol centered on matrix multiplication and publishes an open INT design. Their work is an important reference for useful compute, transcript construction, and the economics of matrix-compute mining. Spinel's specific research focus is different: deterministic consensus for the FP8/MXFP4 precision frontier and cross-vendor tensor execution.
We also draw on classical randomized verification of matrix multiplication, modern GPU kernel engineering, reproducible numerical computing, and open specifications for microscaled numerical formats.
Pearl Research →
Pearl INT whitepaper →
Planned outputs
TensorPoW arithmetic memo
Canonical FP8 semantics, challenge derivation, transcript, and target comparison.
CPU reference + vectors
Independent verifier and public randomized/adversarial test corpus.
Hopper miner
First native FP8 mining kernel with oracle parity receipts.
Multi-node chain
Two or more independent nodes mining and validating identical chain state.
MXFP4 work class
Blackwell/CDNA4 implementations behind one protocol definition.