AMD Lucebox Outperforms NVIDIA DGX Spark by 3.63x in DeepSeek V4 Flash Benchmark
Key Takeaways
- ▸Lucebox achieves 3.63x faster token generation speed than single NVIDIA DGX Spark (51.1 vs 14.09 tok/s) on DeepSeek V4 Flash
- ▸AMD's dual-GPU architecture optimally distributes mixture-of-experts layers across heterogeneous hardware with different memory and speed profiles
- ▸Lucebox offers 2.6x better throughput-per-dollar than single DGX Spark, undercutting dual-Spark configurations on both cost and performance
Summary
AMD's new Lucebox inference system has achieved a significant performance advantage over NVIDIA's DGX Spark in benchmarks running DeepSeek V4 Flash, a 284-billion-parameter language model. Testing by GreenGames showed Lucebox delivering 51.1 tokens/second compared to 14.09 tok/s on a single DGX Spark—a 3.63x speedup—using an optimized load-balancing strategy across its AMD Radeon AI PRO R9700 and Strix Halo GPUs.
The performance gains stem from Lucebox's innovative approach to handling the mixture-of-experts architecture. Rather than traditional tensor parallelism, the system distributes individual expert blocks across the two GPUs based on their memory capacities (32 GB for R9700, 128 GB for Strix Halo) and processing speeds, allowing both to work in parallel on different expert selections for each token.
Priced at $6,499 for the complete system, Lucebox costs 38% more than a single DGX Spark ($4,699) but delivers 2.6x better throughput per dollar. It also undercuts two DGX Sparks ($9,398) by 31% while delivering superior performance, positioning AMD as a serious challenger to NVIDIA's near-monopoly in specialized AI inference hardware.
- Performance tested on 284B-parameter DeepSeek V4 Flash with speculative decoding, representing real-world inference workloads
Editorial Opinion
This benchmark marks a watershed moment for AMD's push into AI inference. A 3.63x performance advantage is substantial enough to merit serious consideration from enterprises, yet the higher upfront cost and architectural differences from NVIDIA's ecosystem could slow adoption. The real test won't be benchmarks—it will be whether software vendors and cloud providers integrate Lucebox into their stacks.



