Select a logic file below to see how the VANA engine's sparse routing performs on real GPU silicon. All results are measured on an NVIDIA RTX 3060 (12GB VRAM) and verified on an NVIDIA RTX 5080 (16GB VRAM, Blackwell), both emulating the future VANA SoC architecture. The CPU processes every connection sequentially. The GPU skips entire regions of inactive neurons, doing less work and finishing faster.
80x is the measured per-tick speedup at 250K neurons, the largest payload the CPU completes in reasonable time. CPU time scales linearly with connection count. At 1M neurons the CPU needs ~164ms per tick. At 100M neurons: ~16s. At 365M neurons: over 60s, the CPU times out entirely. GPU sparse skip stays under 1ms regardless of scale because it only processes the active region. At production scale (millions of neurons), the real advantage is 10,000x or more.
The CPU processes 22.5 million connections sequentially at 41,017 microseconds per tick. The GPU runs the same computation natively in parallel, hitting 99% utilization at 2.3GB VRAM. With sparse routing, only 1 of 977 regions fires. The GPU does less work and finishes 80x faster per tick at this scale. At production scale (millions of neurons), the CPU times out entirely while the GPU stays under 1ms. On consumer RTX 3060 silicon that stands in for the eventual custom SoC.
| Payload File | CPU (Emulated) | GPU Dense Route | GPU Region Skip | GPU Sparse Skip |
|---|---|---|---|---|
| sparse_250k (250K Neurons, 22.5M Connections) | 41,017 us (24.4 Hz) | 7,554 us (5.4x) | 3,086 us (13.3x) | 511 us (80.3x) |
| neural_50k (50K Neurons, 4.5M Connections) | 9,051 us (110.5 Hz) | 1,408 us (6.4x) | 635 us (14.3x) | 141 us (64.2x) |
| neural_10k (10K Neurons, 900K Connections) | 1,677 us (596.4 Hz) | 336 us (5.0x) | 186 us (9.0x) | 133 us (12.6x) |
| circuit_1k (1K Neurons, 90K Connections) | 166 us (6.0K Hz) | 108 us (1.5x) | 106 us (1.6x) | 73 us (2.3x) |
Same source code, two GPU architectures. The v5_blood_bridge.cu benchmark compiled and ran on both
NVIDIA RTX 3060 (Ampere, sm_86) and RTX 5080 (Blackwell, sm_120) using -arch=native.
Sparse routing scales across hardware. The 5080 finishes the same 250K-neuron sparse sweep in 382 microseconds.
-arch=native. The GPU emulates the future VANA SoC architecture; it runs the same computation natively in parallel. The CPU processes every connection sequentially. GPU utilization was verified at 99% during dense runs. The benchmark tool (vana_bench.cu) compiles with nvcc and ships in the Sanguis SDK.