AI News Today.

Artificial intelligence, professionally covered

Company profile

FuriosaAI

Designs and develops data center accelerators for advanced AI models.

furiosa.aiProfile compiled July 20263 source pages read
Category
Chips & hardware
Headquarters
Seoul
Sells to
Enterprise
Business model
Hardware
Deployment
On-premise, Cloud / SaaS
Pricing
Not published
Builds own models
Yes
Modalities
Text, Multimodal

FuriosaAI designs and develops data center accelerators for the most advanced AI models and applications. Their mission is to make AI computing sustainable so everyone on Earth has access to powerful AI. Founded in 2017 by three engineers from hardware, software, and algorithm fields, the company aims to build the world's best AI chips. They offer full-stack solutions for optimal programmability and have raised over $246 million. FuriosaAI focuses on efficient tensor contraction operations with their Tensor Contraction Processor architecture, breaking away from fixed-sized matrix multiplication instructions to unlock powerful performance and efficiency for AI inference.

  • RNGD PCIeGen 2 data center accelerator (TSMC 5nm) designed for efficient tensor contraction operations.
  • NXT RNGD ServerA 3 KW inference appliance for agentic systems, delivering exceptional performance with cost-efficient scalability for inference with advanced LLM and agentic AI applications. Designed for air-cooled data centers, it can be deployed on-premises, in managed environments, or colocation facilities. It features 8 RNGD cards, 4 petaFLOPS (512 TFLOPS x 8 cards), 384 GB HBM3 capacity, 12 TB/s memory bandwidth, and 3 kW power consumption.
  • Furiosa SDKA comprehensive toolchain for LLM inference and agentic workloads, from compilation and optimization to production deployment. It includes a new kernel framework and supports PyTorch 2.x integration.
  • Furiosa-LLMA high-performance inference engine for LLM models, part of the Furiosa SDK. It supports features like OpenAI-compatible server, tool calling, structured output, vision-language models, prefix caching, hybrid KV cache management, data-parallel routing, and model parallelism.
  • Gen 1 Vision NPUThe first generation AI chip (Samsung 14nm) for vision applications, which entered official volume production with Samsung Foundry and ASUS.
  • Furiosa Access ProgramA structured path for customers and partners to evaluate, integrate, qualify, and deploy Furiosa accelerators through both online and offline access, available worldwide.
  • Tensor Contraction Processor architecture (ISCA 2024)
  • High-throughput, low-latency execution for capable models
  • Reduced energy draw and fewer racks for lower TCO
  • Standard air cooling support
  • Compiler designed to optimize evolving AI workloads
  • Comprehensive toolchain for LLM inference and agentic workloads
  • Maximizing data center utilization with containerization, SR-IOV, Kubernetes, and cloud native components
  • Robust ecosystem support with PyTorch 2.x integration
  • Pre-optimized and pre-compiled models available on Hugging Face Hub
  • OpenAI-Compatible Server
  • Tool Calling with parsers and choice options
  • Structured Output generation
  • Vision-Language Models with image inputs
  • Prefix Caching for improved performance
  • Hybrid KV Cache Management
  • Data-Parallel Routing (scoring-based)
  • Model Parallelism (Tensor/Pipeline/Data parallelism)
  • Furiosa SMI CLI for managing NPUs
  • Furiosa SMI Library for managing NPUs
  • Host PCI Optimization Tuning
  • Data center accelerators for advanced AI models and applications
  • Inference with advanced LLM and agentic AI applications
  • Enterprise AI scale solutions
  • Deployment of capable models with high-throughput, low-latency execution
  • LLM inference and agentic workloads
  • Maximizing data center utilization
  • Transitioning models into production
  • Creating inference applications from PyTorch models
  • Model quantization
  • Model serving and deployment
  • Serving Vision-Language models

FuriosaAI designs and develops data center accelerators (NPUs) for advanced AI models and applications, focusing on efficient tensor contraction operations. They provide a full-stack solution including hardware and a software SDK for deep learning model inference, particularly for LLMs and agentic AI. They build their own proprietary NPU architecture and software stack.

Tech named: Tensor Contraction Processor Architecture, RNGD PCIe, NXT RNGD Server, Furiosa SDK, NPUaaS, LLM, Agentic AI, PyTorch 2.x integration, containerization, SR-IOV, Kubernetes, Furiosa-LLM (inference engine), OpenAI-Compatible Server, Tool Calling, Structured Output, Vision-Language Models, Prefix Caching, Hybrid KV Cache Management, Data-Parallel Routing, Model Parallelism (Tensor/Pipeline/Data), Furiosa SMI CLI, Furiosa SMI Library, Host PCI Optimization Tuning

  • Semiconductor Manufacturing
  • Data Centers
  • Enterprise AI
  • Cloud Computing
  • Proprietary Tensor Contraction Processor (TCP) architecture specifically designed for efficient tensor contraction operations, a higher dimensional generalization of matrix multiplication, unlike most commercial deep learning accelerators that use fixed-sized matmul instructions.
  • Outperforms RTX Pro 6000 with the latest SDK (benchmark claim).
  • Enables 4x more inference capacity per rack compared to 7.5 kW servers.
  • Offers a compelling combination of excellent real-world performance, dramatic reduction in total cost of ownership, and straightforward integration.
  • First AI chip startup to outperform Nvidia on MLPerf Inference (in 2021).
  • Full-stack solutions for optimal programmability.

Korea Development Bank, Industrial Bank of Korea, Keistone Partners, PI Partners, Kakao Investment

From the AI funding tracker — rounds as reported by the linked publications.

This profile was compiled from FuriosaAI's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.