Company profile

Featherless AI

Serverless inference via GPU orchestration and model load-balancing.

featherless.aiProfile compiled July 20262 source pages read
Category
AI infrastructure
Headquarters
San Francisco, California
Sells to
Developers
Business model
Usage-based API
Deployment
API, Cloud / SaaS
Pricing
Published
Builds own models
No — builds on existing models
Modalities
Text, Code, Multimodal

Featherless AI provides serverless inference for large language models (LLMs) and other open-weight AI models through an API. The platform offers GPU orchestration and model load-balancing, enabling organizations to size their server fleet to throughput needs, not the number of models in their catalog. It supports fine-tuning and provides instant access to a catalog of over 40,000 open models, including those for role-playing, creative writing, coding assistance, and reasoning. The API is OpenAI compatible, allowing for easy integration into existing client programs.

  • Serverless inference
  • GPU orchestration
  • Model load-balancing
  • OpenAI compatible API
  • Access to 40,000+ open models
  • Fine-tuning enablement
  • Cost calculation for savings
  • Public cloud inference for 10k+ open weight models
  • Role-playing
  • Creative writing
  • Coding assistance
  • Agentic coding
  • Reasoning
  • Multimodal instruct
  • Math problem solving
  • Long-form roleplay
  • Prose-focused writing
  • Uncensored assistant
  • Efficient reasoning
  • Frontier reasoning
  • Vision-language model
  • Multilingual chat
  • Dual-mode chat model
  • Chat & coding model
  • Dual-mode reasoning
  • Compact reasoning
  • Tool use

Featherless AI provides a serverless inference platform for open-weight AI models, primarily large language models (LLMs). They offer GPU orchestration and model load-balancing to enable serverless inference and fine-tuning. Users access a catalog of 40,000+ models via an OpenAI-compatible API.

Tech named: GPU orchestration, model load-balancing, serverless inference, fine-tuning, OpenAI compatible API, large language models (LLMs), linear-transformer model (QRWKV)

  • Run GLM 5.2 on AMD exclusively
  • Freedom to reliably deploy any open model effortlessly
  • One API key for instant access
  • Serverless inference platform
  • Continually expanding library of open-weight models
  • OpenAI compatible API interface

AMD Ventures, Airbus Ventures, BMW i Ventures, Kickstart Ventures, Panache Ventures, Wavemaker Ventures

Airbus Ventures, 500 Global, Kickstart Ventures, HF0, Panache Ventures, Oakseed Ventures

From the AI funding tracker — rounds as reported by the linked publications.

This profile was compiled from Featherless AI's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.