Company profile
FuriosaAI
Designs and develops data center accelerators for advanced AI models.
- Category
- Chips & hardware
- Headquarters
- Seoul
- Sells to
- Enterprise
- Business model
- Hardware
- Deployment
- On-premise, Cloud / SaaS
- Pricing
- Not published
- Builds own models
- Yes
- Modalities
- Text, Multimodal
What FuriosaAI does
FuriosaAI designs and develops data center accelerators for the most advanced AI models and applications. Their mission is to make AI computing sustainable so everyone on Earth has access to powerful AI. Founded in 2017 by three engineers from hardware, software, and algorithm fields, the company aims to build the world's best AI chips. They offer full-stack solutions for optimal programmability and have raised over $246 million. FuriosaAI focuses on efficient tensor contraction operations with their Tensor Contraction Processor architecture, breaking away from fixed-sized matrix multiplication instructions to unlock powerful performance and efficiency for AI inference.
Products
- RNGD PCIeGen 2 data center accelerator (TSMC 5nm) designed for efficient tensor contraction operations.
- NXT RNGD ServerA 3 KW inference appliance for agentic systems, delivering exceptional performance with cost-efficient scalability for inference with advanced LLM and agentic AI applications. Designed for air-cooled data centers, it can be deployed on-premises, in managed environments, or colocation facilities. It features 8 RNGD cards, 4 petaFLOPS (512 TFLOPS x 8 cards), 384 GB HBM3 capacity, 12 TB/s memory bandwidth, and 3 kW power consumption.
- Furiosa SDKA comprehensive toolchain for LLM inference and agentic workloads, from compilation and optimization to production deployment. It includes a new kernel framework and supports PyTorch 2.x integration.
- Furiosa-LLMA high-performance inference engine for LLM models, part of the Furiosa SDK. It supports features like OpenAI-compatible server, tool calling, structured output, vision-language models, prefix caching, hybrid KV cache management, data-parallel routing, and model parallelism.
- Gen 1 Vision NPUThe first generation AI chip (Samsung 14nm) for vision applications, which entered official volume production with Samsung Foundry and ASUS.
- Furiosa Access ProgramA structured path for customers and partners to evaluate, integrate, qualify, and deploy Furiosa accelerators through both online and offline access, available worldwide.
Key capabilities
- Tensor Contraction Processor architecture (ISCA 2024)
- High-throughput, low-latency execution for capable models
- Reduced energy draw and fewer racks for lower TCO
- Standard air cooling support
- Compiler designed to optimize evolving AI workloads
- Comprehensive toolchain for LLM inference and agentic workloads
- Maximizing data center utilization with containerization, SR-IOV, Kubernetes, and cloud native components
- Robust ecosystem support with PyTorch 2.x integration
- Pre-optimized and pre-compiled models available on Hugging Face Hub
- OpenAI-Compatible Server
- Tool Calling with parsers and choice options
- Structured Output generation
- Vision-Language Models with image inputs
- Prefix Caching for improved performance
- Hybrid KV Cache Management
- Data-Parallel Routing (scoring-based)
- Model Parallelism (Tensor/Pipeline/Data parallelism)
- Furiosa SMI CLI for managing NPUs
- Furiosa SMI Library for managing NPUs
- Host PCI Optimization Tuning
Use cases
- Data center accelerators for advanced AI models and applications
- Inference with advanced LLM and agentic AI applications
- Enterprise AI scale solutions
- Deployment of capable models with high-throughput, low-latency execution
- LLM inference and agentic workloads
- Maximizing data center utilization
- Transitioning models into production
- Creating inference applications from PyTorch models
- Model quantization
- Model serving and deployment
- Serving Vision-Language models
AI approach
FuriosaAI designs and develops data center accelerators (NPUs) for advanced AI models and applications, focusing on efficient tensor contraction operations. They provide a full-stack solution including hardware and a software SDK for deep learning model inference, particularly for LLMs and agentic AI. They build their own proprietary NPU architecture and software stack.
Tech named: Tensor Contraction Processor Architecture, RNGD PCIe, NXT RNGD Server, Furiosa SDK, NPUaaS, LLM, Agentic AI, PyTorch 2.x integration, containerization, SR-IOV, Kubernetes, Furiosa-LLM (inference engine), OpenAI-Compatible Server, Tool Calling, Structured Output, Vision-Language Models, Prefix Caching, Hybrid KV Cache Management, Data-Parallel Routing, Model Parallelism (Tensor/Pipeline/Data), Furiosa SMI CLI, Furiosa SMI Library, Host PCI Optimization Tuning
Industries served
- Semiconductor Manufacturing
- Data Centers
- Enterprise AI
- Cloud Computing
What it says sets it apart
- Proprietary Tensor Contraction Processor (TCP) architecture specifically designed for efficient tensor contraction operations, a higher dimensional generalization of matrix multiplication, unlike most commercial deep learning accelerators that use fixed-sized matmul instructions.
- Outperforms RTX Pro 6000 with the latest SDK (benchmark claim).
- Enables 4x more inference capacity per rack compared to 7.5 kW servers.
- Offers a compelling combination of excellent real-world performance, dramatic reduction in total cost of ownership, and straightforward integration.
- First AI chip startup to outperform Nvidia on MLPerf Inference (in 2021).
- Full-stack solutions for optimal programmability.
Funding rounds we track
Korea Development Bank, Industrial Bank of Korea, Keistone Partners, PI Partners, Kakao Investment
From the AI funding tracker — rounds as reported by the linked publications.
This profile was compiled from FuriosaAI's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.