South Korean AI chip startups develop specialized hardware for edge and data center inference

South Korean AI chip startups are developing specialized hardware to challenge established players in edge and data center inference. DeepX's DX-M1 offers 25 TOPS at 3-5W, targeting high-performance edge AI with significant power savings over traditional GPUs. HyperAccel is building Latency Processing Units (LPUs) like the Bertha ASIC, which aims for 90% memory bandwidth utilization in large language model (LLM) workloads. Mobilint provides versatile solutions with its high-performance Aries chip and the ultra-compact Regulus for power-constrained devices like drones and security cameras. FuriosaAI's RNGD (Renegade) processor, built on TSMC's 5nm process, targets data center LLM inference with 48GB of HBM3 and a custom Tensor Contraction Processor architecture. Rebellions, which recently merged with Sapeon to form a "unicorn" entity, is developing the Rebel chiplet-based processor featuring 144GB of HBM3e and one petaflop of FP16 performance. These companies emphasize energy efficiency, low latency, and specialized software stacks to optimize transformer-based workloads. Many are currently transitioning from R&D to mass production, with several flagship chips slated for release in 2025 and 2026.

South Korean AI chip startups are developing specialized hardware to challenge established players in edge and data center inference. DeepX's DX-M1 offers 25 TOPS at 3-5W, targeting high-performance edge AI with significant power savings over traditional GPUs. HyperAccel is building Latency Processing Units (LPUs) like the Bertha ASIC, which aims for 90% memory bandwidth utilization in large language model (LLM) workloads. Mobilint provides versatile solutions with its high-performance Aries chip and the ultra-compact Regulus for power-constrained devices like drones and security cameras. FuriosaAI's RNGD (Renegade) processor, built on TSMC's 5nm process, targets data center LLM inference with 48GB of HBM3 and a custom Tensor Contraction Processor architecture. Rebellions, which recently merged with Sapeon to form a "unicorn" entity, is developing the Rebel chiplet-based processor featuring 144GB of HBM3e and one petaflop of FP16 performance. These companies emphasize energy efficiency, low latency, and specialized software stacks to optimize transformer-based workloads. Many are currently transitioning from R&D to mass production, with several flagship chips slated for release in 2025 and 2026.

DeepX's DX-M1 chip delivers 25 TOPS of performance while consuming only 3 to 5 watts of power. HyperAccel's LPU architecture achieves 90% effective memory bandwidth utilization for large language model execution.

Mobilint's Aries chip provides 80 TOPS of performance for robotics and smart city infrastructure at 20-25 watts. FuriosaAI's RNGD processor uses a custom Tensor Contraction Processor to optimize transformer workloads in data centers.

Rebellions' upcoming Rebel chip uses a chiplet-based design with 144GB of HBM3e memory for high-scale LLM inference. The merger of Rebellions and Sapeon created South Korea's first AI semiconductor unicorn company.

Chapter guide

Worth noting

  • Many featured products, including the DX-M1M, Bertha, Rebel, and RNGD mass production units, are unreleased or slated for 2025-2026.
  • Performance claims, such as outperforming NVIDIA H100 or A100, are company-reported and have not been independently verified.
  • Reported revenue for some startups remains low as they transition from research and development to commercial deployment.
  • The merger between Rebellions and Sapeon is a recent development with long-term integration results yet to be seen.

Watch the original video ↗