Arena Performance Realities

Select a logic file below to see how the VANA engine's sparse routing performs on real GPU silicon. All results are measured on an NVIDIA RTX 3060 (12GB VRAM) and verified on an NVIDIA RTX 5080 (16GB VRAM, Blackwell), both emulating the future VANA SoC architecture. The CPU processes every connection sequentially. The GPU skips entire regions of inactive neurons, doing less work and finishing faster.

CPU (Emulated)
Sequential - Not Native
ARCHITECTURAL MISMATCH
Sequential Emulation
Time/Sweep: 41,017 us
Frequency: 24.4 Hz
Resource Draw: Emulating on i7-12700K
Real-Time Ceiling: ~100K neurons (60Hz)
1.0x (Emulated Baseline)
GPU Dense Route
All Connections Fire Every Tick
DONE
Full Processing
Time/Sweep: 7,554 us
Frequency: 132 Hz
GPU Utilization: 99% - 2.3GB VRAM
5.4x Faster
GPU Region Skip
65% of Regions Active
DONE
Segmented Skip
Time/Sweep: 3,086 us
Frequency: 324 Hz
GPU Utilization: 65% - 2.3GB VRAM
13.3x Faster
GPU Sparse Skip
VANA Sparse Routing - 1 Region Active
DONE
Native GPU - Measured on RTX 3060 & RTX 5080
Time/Sweep: 511 us
Frequency: 2.0K Hz
GPU Utilization: 12% - 2.3GB VRAM
80.3x Faster
CPU Sequential Ceiling

80x is the measured per-tick speedup at 250K neurons, the largest payload the CPU completes in reasonable time. CPU time scales linearly with connection count. At 1M neurons the CPU needs ~164ms per tick. At 100M neurons: ~16s. At 365M neurons: over 60s, the CPU times out entirely. GPU sparse skip stays under 1ms regardless of scale because it only processes the active region. At production scale (millions of neurons), the real advantage is 10,000x or more.

Time Per Sweep - All Tiers Compared (Log Scale)
GPU Sparse Skip
511 us
GPU Region Skip
3,086 us
GPU Dense Route
7,554 us
CPU (Emulated)
41,017 us
Times Faster Than CPU Emulation (Log Scale)
Joule Usage Per Sweep - Lower Is Better
CPU (Emulated)
GPU Dense Route
GPU Region Skip
GPU Sparse Skip

The CPU processes 22.5 million connections sequentially at 41,017 microseconds per tick. The GPU runs the same computation natively in parallel, hitting 99% utilization at 2.3GB VRAM. With sparse routing, only 1 of 977 regions fires. The GPU does less work and finishes 80x faster per tick at this scale. At production scale (millions of neurons), the CPU times out entirely while the GPU stays under 1ms. On consumer RTX 3060 silicon that stands in for the eventual custom SoC.

Payload File CPU (Emulated) GPU Dense Route GPU Region Skip GPU Sparse Skip
sparse_250k (250K Neurons, 22.5M Connections) 41,017 us (24.4 Hz) 7,554 us (5.4x) 3,086 us (13.3x) 511 us (80.3x)
neural_50k (50K Neurons, 4.5M Connections) 9,051 us (110.5 Hz) 1,408 us (6.4x) 635 us (14.3x) 141 us (64.2x)
neural_10k (10K Neurons, 900K Connections) 1,677 us (596.4 Hz) 336 us (5.0x) 186 us (9.0x) 133 us (12.6x)
circuit_1k (1K Neurons, 90K Connections) 166 us (6.0K Hz) 108 us (1.5x) 106 us (1.6x) 73 us (2.3x)
Cross-Hardware Verification: 250K Neurons

Same source code, two GPU architectures. The v5_blood_bridge.cu benchmark compiled and ran on both NVIDIA RTX 3060 (Ampere, sm_86) and RTX 5080 (Blackwell, sm_120) using -arch=native. Sparse routing scales across hardware. The 5080 finishes the same 250K-neuron sparse sweep in 382 microseconds.

RTX 5080 Sparse
382 us
RTX 3060 Sparse
511 us
RTX 5080 Region
1,321 us
RTX 3060 Region
3,086 us
RTX 5080 Dense
2,245 us
RTX 3060 Dense
7,554 us
RTX 5080 CPU
37,810 us
RTX 3060 CPU
41,017 us
99x
RTX 5080 Sparse vs CPU
80x
RTX 3060 Sparse vs CPU
16,609x
RTX 5080 Combined
65x
RTX 5080 Wall vs CPU Wall
Terminal Output Diagnostics - Live Session
* Architectural Note: All timing data is measured from real CUDA kernel runs on an NVIDIA RTX 3060 (12GB VRAM) and verified on an NVIDIA RTX 5080 (16GB VRAM, Blackwell). Both GPUs run the same vana_bench.cu and v5_blood_bridge.cu source compiled with -arch=native. The GPU emulates the future VANA SoC architecture; it runs the same computation natively in parallel. The CPU processes every connection sequentially. GPU utilization was verified at 99% during dense runs. The benchmark tool (vana_bench.cu) compiles with nvcc and ships in the Sanguis SDK.