Performance
Card figures are measured. Node figures are estimated.
Throughput
By configuration
Efficiency
Ops per watt
| N8 · at wall | ~230K ops/s/W at 1400 W (estimated) |
|---|---|
| Card · typical, excluding host | 533K ops/s/W at 75 W |
Baselines
Two axes
Vs lattice FHE
Conventional lattice overhead ~1000×. Umbra: no bootstrapping, no noise refresh by design.
Vs native
End-to-end system overhead ~50×2. Ciphertext throughout: no plaintext in memory during execution.
When the cost is payable
Not a cheaper GPU
End-to-end overhead versus native is the system cost. It is payable when the alternative is not a cheaper accelerator: it is not running the workload at all: data that cannot leave a key domain, declassification that never finishes, or a privacy review that blocks the lake.
Fraud scoring, coalition correlation, and citizen inference are classes where that arithmetic holds. A batch that could have run in plaintext on a GPU farm is the class where it does not.
Methodology
How figures are produced
Encrypted logic operations per second are measured on production silicon with a fixed harness to thermal steady state. Node figures multiply measured single-card throughput by card count and are labelled estimated until end-to-end node measurement completes.
Encrypted logic operations per second, measured on production silicon. Node figures estimated from measured single-card throughput.
End-to-end system overhead against equivalent native computation. Operation-level acceleration against lattice FHE is reported separately.
Is 40M measured?
Yes. Card figures are measured on production silicon.
Are node figures measured?
No. N4 and N8 are estimated from single-card throughput until end-to-end measurement completes.
Full benchmark harness
Test vectors, reproducibility instructions and raw measurement logs. Sent to your work address.