AI Hardware platforms provide specialized computing infrastructure and devices optimized for artificial intelligence workloads. These solutions offer enhanced processing capabilities, accelerated computation, and optimized performance for AI applications.
34 verified AI-first sites in Hardware.
Last reviewed .
Latest changes
removed: Rain (ceased operations); Esperanto Technologies (ceased operations, IP sold to Nekko.ai)
LPU inference hardware and GroqCloud neocloud for low-latency LLM serving. After a major Nvidia licensing and talent deal, Groq remains independent and scales inference capacity with LPU and partner GPUs.
Pricing: API and cloud usage; enterprise capacity via GroqCloud.
AI processor and hardware platform led by Jim Keller. Features Tensix core RISC-V architecture in Wormhole and Blackhole cards, plus Galaxy servers for scale-out training and inference. Open-source TT-Metalium and TT-Buda software stacks with PyTorch and JAX integration. Also licenses AI and RISC-V IP to chip vendors. Backed by Samsung, Hyundai, LG, and Bezos Expeditions.
Pricing: Wormhole n150s from ~$999; Blackhole and Galaxy systems priced on request. IP licensing available.
Wafer-scale AI accelerators (WSE) for training and inference, plus Cerebras cloud access. Extreme on-chip memory and throughput for frontier LLM workloads; public-company AI silicon leader.
Pricing: Cloud and enterprise systems; see cerebras.ai for pricing.
Purpose-built AI inference platform with RDU (Reconfigurable Dataflow Unit) chips. Features 100x energy efficiency with dataflow architecture and three-tier memory. Powers sovereign AI deployments in UK (Argyll), EU (Infercom, OVHcloud), and Australia (SouthernCrossAI). Runs DeepSeek-R1, Llama 4, Qwen at 200+ tokens/second. Partners include Meta, Hugging Face, CrewAI. SambaCloud, SambaStack, and SambaManaged offerings.
Intel's Gaudi AI accelerators from its Habana Labs acquisition, for deep learning training and inference. Gaudi 3 is the last generation: Intel ended the line and is moving to Crescent Island inference GPUs and Jaguar Shores rack-scale systems. Still sold through server OEMs and cloud partners, with PyTorch support via the Gaudi software stack.
Pricing: Available via cloud partners. Enterprise: Contact Intel.
Innovative AI chip company specializing in transformer-optimized processors for AI inference. Features the Sohu chip, a custom ASIC built on TSMC's 4nm process that delivers superior performance for running large language models, text, image, and video transformers. Includes cloud-based developer platform for testing and deployment. Offers significant improvements in speed and energy efficiency compared to traditional GPUs.
Pricing: Hardware pricing based on volume. Developer Cloud access available for testing.
Neuromorphic processor company developing ultra-low power spiking neural processors for always-on edge AI applications. Features Spiking Neural Processor T1 microcontroller with integrated spiking neural network engine and 32-bit RISC-V core, Pulsar neuromorphic microcontroller for mass-market applications, sub-milliwatt power levels enabling always-on sensors, and support for smart building and home automation systems. Includes partnerships with Sony for integration into image sensors enabling low-latency object detection in smart cameras. Secured $21M Series A funding in 2025. Evaluation kits and development tools available.
Pricing: Contact for pricing. Evaluation kits available.
AI semiconductor company building inference accelerators for efficient large-model serving. Focuses on high-throughput, lower-power hardware for data-center and enterprise AI workloads.
Fabless chip company developing AI inference processors designed for performance-per-watt and low-latency execution. Targets production AI serving where efficiency and cost are critical.
Pricing: Commercial hardware engagements and partnerships; pricing on request.
AI hardware startup focused on inference processors for frontier model deployment. Emphasizes speed and cost improvements for large-scale, low-latency model serving.
Pricing: Early-access/commercial pricing available via enterprise inquiry.
AI compute hardware company building inference-focused accelerators for generative AI workloads. Designed for high-performance serving with improved efficiency in data-center environments.
Pricing: Enterprise platform pricing through commercial engagement.
South Korean fabless company building inference-focused AI accelerators and infrastructure. Features REBEL-Quad UCIe accelerator card, ATOM-Max NPUs, RebelServer, RebelRack, and RebelPOD systems for sovereign-scale LLM deployment. SDK supports PyTorch, vLLM, Triton, and Hugging Face with mixed-precision inference optimized for cost per watt. Strategic investors include Samsung, SK Hynix, and Aramco.
Pricing: Enterprise hardware and infrastructure; contact for deployment.
Generative inference hardware company building Transformer-optimized accelerators and appliances. Features Atlas inference system shipping now for models up to 500B parameters, upcoming Titan and Asimov chips, and OpenAI-compatible API endpoints. Maps Hugging Face Transformers models directly to hardware for maximum throughput and lower power than GPU baselines on Llama-class workloads.
Pricing: Contact for Atlas appliance and enterprise deployment pricing.
Photonic computing company building silicon photonics for AI data centers. Envise photonic AI accelerator uses light instead of electrons for matrix math, and Passage interconnect fabric provides high-bandwidth, low-power chip-to-chip and rack-scale communication. Targets hyperscale training and inference at radically lower energy per operation. Backed by GV, Fidelity, T. Rowe Price, and SIP Global Partners; valued over $4B.
Pricing: Enterprise hardware and interconnect systems; contact for pricing and deployment.
Deep-tech startup turning any deep-learning model into custom silicon ("the model is the computer"). Compiles transformer, diffusion, and other neural nets directly into ASICs, targeting orders-of-magnitude gains in tokens per second, per dollar, and per watt versus GPU inference. Founded by Tenstorrent alumni; its HC1 chip serves Llama 3.1 8B at roughly 17K tokens/sec. AMD agreed to acquire Taalas in August 2026, with closing expected in Q4.
Dominant AI accelerator platform for training and inference-HGX/DGX systems, Blackwell and Hopper GPUs, CUDA, and NVLink scale-out. The default silicon stack for hyperscalers, labs, and enterprises building frontier and production AI.
Pricing: Systems and cloud via NVIDIA and partners; enterprise quotes on site.
AMD Instinct GPU accelerators (MI300/MI350 class) for large-scale AI training and inference. High HBM capacity and the open ROCm software stack as the primary GPU alternative to NVIDIA in data-center AI.
Pricing: Through OEMs and cloud providers; see AMD Instinct product pages.
Google Tensor Processing Units purpose-built for training and serving large ML models on Google Cloud. Custom ASIC generations (including Trillium-class) power Google's own models and customer Cloud TPU workloads as a primary hyperscaler alternative to GPUs.
Pricing: Cloud TPU pricing via Google Cloud; see cloud.google.com/tpu.
AWS custom AI training accelerators (Trainium / Trainium2) for large-scale model training on Amazon EC2. Designed for high throughput and lower cost versus GPUs for many LLM and foundation-model training jobs inside AWS.
Pricing: EC2 Trainium instance pricing; see AWS Trainium pages.
AWS custom inference chips (Inferentia / Inferentia2) for high-throughput, cost-efficient model serving on Amazon EC2. Tuned for generative and classical ML inference as AWS's silicon alternative to GPU endpoints.
Pricing: EC2 Inferentia instance pricing; see AWS Inferentia pages.
AI chip startup building high-throughput accelerators for large language model training and inference. Focuses on dense matrix performance and efficiency for frontier LLM workloads versus general-purpose GPUs.
Pricing: Enterprise and partner engagements; contact MatX.
Chinese AI accelerator company (寒武纪) shipping MLU smart accelerator cards and edge modules. Purpose-built silicon for training and inference across data center and edge form factors.
Pricing: Hardware and OEM programs; contact Cambricon / distributors.
BirenTech AI GPUs for high-efficiency, high-performance general AI compute. Fabless Chinese accelerator vendor targeting training and inference with a full software stack.
Pricing: Enterprise and partner channels; see birentech.com.
Full-stack AI GPU company (摩尔线程) building GPUs, systems, and software for training and inference. Ships data-center and workstation AI compute products in China and partner markets.
Pricing: Systems and cards via Moore Threads channels; see mthreads.com.
Microsoft's custom Azure AI accelerators (Maia 100/200) for large-scale training and inference inside Azure. Hyperscaler silicon co-designed with Microsoft's model and cloud stack.
Pricing: Available via Azure AI infrastructure; see Microsoft Azure for access.
Analog in-memory AI accelerators for efficient edge-to-cloud inference. Targets much higher TOPS/W than conventional digital GPUs for generative and on-device models.
Pricing: Hardware and OEM partnerships; contact EnCharge AI.
Optoelectronic AI computing company shipping photonic accelerator cards and optical interconnects for AI data centers. Matrix compute and disaggregated memory over light.
Pricing: Products and evaluation platforms; contact Lightelligence.
Huawei Ascend AI processors and full-stack CANN software for training and inference at scale. Major non-NVIDIA AI silicon path across China and partner clouds.
Pricing: Hardware and developer community resources; see Ascend / HiAscend.
Baidu-affiliated AI chip company (昆仑芯) building Kunlun accelerators for training and inference. Domestic AI silicon used across Baidu and partner deployments.
Pricing: Enterprise AI accelerators; see Kunlunxin.
Meta's custom Training and Inference Accelerator family for ranking, recommendation, and GenAI workloads at hyperscale. Homegrown silicon co-developed with Broadcom, deployed in Meta data centers alongside GPUs for cost-efficient AI serving.
Pricing: Internal Meta infrastructure; not sold as merchant silicon-see Meta AI research posts.
Qualcomm's cloud AI inference cards (Cloud AI 100 / Ultra) for generative and classic AI serving. High on-die SRAM, PCIe accelerators, and Qualcomm AI Inference Suite for on-prem and cloud deployments without GPU racks.
Pricing: Hardware through Qualcomm and partners; see Qualcomm data-center AI products.
Intelligence Processing Units and systems for ML training and inference. SoftBank-backed IPU hardware and cloud access as a dedicated AI accelerator alternative to mainstream GPUs.
Pricing: Systems and cloud access; see Graphcore for current offerings.