BotBeat
...
← Back

> ▌

AMDAMD
RESEARCHAMD2026-07-30

AMD Lucebox Outperforms NVIDIA DGX Spark by 3.63x in DeepSeek V4 Flash Benchmark

Key Takeaways

  • ▸Lucebox achieves 3.63x faster token generation speed than single NVIDIA DGX Spark (51.1 vs 14.09 tok/s) on DeepSeek V4 Flash
  • ▸AMD's dual-GPU architecture optimally distributes mixture-of-experts layers across heterogeneous hardware with different memory and speed profiles
  • ▸Lucebox offers 2.6x better throughput-per-dollar than single DGX Spark, undercutting dual-Spark configurations on both cost and performance
Source:
Hacker Newshttps://www.lucebox.com/blog/deepseek-v4-asymmetric-parallelism↗

Summary

AMD's new Lucebox inference system has achieved a significant performance advantage over NVIDIA's DGX Spark in benchmarks running DeepSeek V4 Flash, a 284-billion-parameter language model. Testing by GreenGames showed Lucebox delivering 51.1 tokens/second compared to 14.09 tok/s on a single DGX Spark—a 3.63x speedup—using an optimized load-balancing strategy across its AMD Radeon AI PRO R9700 and Strix Halo GPUs.

The performance gains stem from Lucebox's innovative approach to handling the mixture-of-experts architecture. Rather than traditional tensor parallelism, the system distributes individual expert blocks across the two GPUs based on their memory capacities (32 GB for R9700, 128 GB for Strix Halo) and processing speeds, allowing both to work in parallel on different expert selections for each token.

Priced at $6,499 for the complete system, Lucebox costs 38% more than a single DGX Spark ($4,699) but delivers 2.6x better throughput per dollar. It also undercuts two DGX Sparks ($9,398) by 31% while delivering superior performance, positioning AMD as a serious challenger to NVIDIA's near-monopoly in specialized AI inference hardware.

  • Performance tested on 284B-parameter DeepSeek V4 Flash with speculative decoding, representing real-world inference workloads

Editorial Opinion

This benchmark marks a watershed moment for AMD's push into AI inference. A 3.63x performance advantage is substantial enough to merit serious consideration from enterprises, yet the higher upfront cost and architectural differences from NVIDIA's ecosystem could slow adoption. The real test won't be benchmarks—it will be whether software vendors and cloud providers integrate Lucebox into their stacks.

Large Language Models (LLMs)MLOps & InfrastructureAI HardwareMarket Trends

More from AMD

AMDAMD
PRODUCT LAUNCH

AMD Launches Ryzen AI Embedded X100 to Expand into Physical AI Market

2026-08-02
AMDAMD
INDUSTRY REPORT

AMD Gains Critical Momentum in AI Race with Anthropic and Microsoft Deployments

2026-07-31
AMDAMD
OPEN SOURCE

AMD Brings HDMI 2.1 Low-Latency Gaming Features to Linux Radeon Driver

2026-07-30

Comments

Suggested

Hugging FaceHugging Face
OPEN SOURCE

Strangers Pretrain 15M-Parameter Language Model Using GitHub Actions and Hugging Face PRs

2026-08-02
Alibaba (Cloud)Alibaba (Cloud)
INDUSTRY REPORT

Token Diplomacy: China Positions Open-Source AI as Global Strategic Resource

2026-08-02
Independent ResearchIndependent Research
RESEARCH

Novel Persistent State Machines Framework Achieves Ultra-Low-Power LLM Attention on FPGA

2026-08-02
← Back to news
© 2026 BotBeat
AboutPrivacy PolicyTerms of ServiceContact Us