Company profile
LightOn
Generative AI platform for document intelligence in critical environments.
- Category
- AI infrastructure
- Headquarters
- Paris, Île-de-France
- Sells to
- Mixed
- Business model
- Usage-based API, SaaS subscription, Services & consulting, Freemium
- Deployment
- On-premise, Cloud / SaaS, API, Self-hosted
- Pricing
- Starter (Free + PAYG), Business (149€/mo + PAYG), Enterprise (Custom) · from $160.7/mo · free tier
- Builds own models
- Yes
- Modalities
- Text, Multimodal
What LightOn does
LightOn is a leading European player in generative AI, delivering cutting-edge document intelligence for critical environments. The company deploys an on-premise RAG platform for unstructured data, accessible via API or interface. LightOn enables enterprises and public organizations to deploy state-of-the-art AI behind their firewall, safely leveraging their most sensitive data. Founded in 2016, LightOn has pioneered in the field of Large Language Models since 2020, developing over 12 models, including open-sourced foundation models with over 100 billion parameters, and commercializing its flagship model, Alfred. LightOn's mission is to help businesses seize the opportunities of Gen AI, by putting confidentiality and value creation at the heart of their solutions, with a commitment to user-centricity.
Products
- LightOn APIA production-ready Multimodal RAG API for developers to search, parse, and ingest documents at scale. It offers endpoints for parsing documents, extracting fields, and grounded retrieval with citations. It is designed to integrate secure reasoning into applications without managing complex AI stacks.
- LightOnOCR-2A state-of-the-art parsing engine that turns scans, tables, handwriting, and multi-column layouts into structured Markdown. It supports over 20 languages natively and is the parsing engine behind every retrieval workflow.
- Paradigm by LightOnA Generative AI platform for all Business Units, accessible via an intuitive Agent interface or dedicated APIs. It allows business units to tailor solutions to their specific needs, empowering them to drive productivity gains while maintaining data privacy and cost control. It can be deployed as an on-premise API engine or a ready-to-use white-label Chat & Search interface.
- Forge by LightOnProfessional services support for integrating Gen AI into enterprises. It offers custom AI creation, use case design, custom business case development, expert training, infrastructure management support, and personalized model customization.
- LateOnAn open-source ColBERT family retrieval model used for hybrid retrieval.
- NextPlaidAn open-source ColBERT family retrieval model used for hybrid retrieval.
- PyLateA model from LightOn's open-source collection.
- DenseOnA model from LightOn's open-source collection.
- AlfredLightOn's flagship commercialized Large Language Model.
- LightOn-rerankA state-of-the-art multimodal LLM reranker that reads text passages and scanned pages on the same relevance scale from a single adapter and deployment.
Key capabilities
- On-premise RAG platform for unstructured data
- API and interface access
- State-of-the-art parsing with LightOnOCR-2 (20+ languages)
- Field extraction using JSON schema
- Grounded retrieval with citations
- Hybrid retrieval (dense, sparse, late-interaction)
- LLM-agnostic infrastructure
- Model Context Protocol (MCP)-native
- Workspaces and ACLs at chunk level
- Unlimited workspaces (Starter plan)
- Native data source connectors (Starter plan)
- Data governance & access controls (RBAC) (Starter plan)
- SSO (SAML, OIDC, LDAP) & SCIM provisioning (Starter plan)
- Audit logs & usage analytics (Starter plan)
- Bring Your Own LLM or LightOn LLMaaS (Starter plan)
- EU sovereign hosting (Starter plan)
- Dedicated environment options (cloud, VPC, on-prem) (Enterprise plan)
- SecNumCloud / regional options (MENA, APAC, US) (Enterprise plan)
- Managed deployment & setup (Enterprise plan)
- Custom data source connectors (Enterprise plan)
- Dedicated CSM + Enterprise SLAs (Enterprise plan)
- Volume discounts (Enterprise plan)
- Optional custom AI engineering & integration (Enterprise plan)
- Build searchable knowledge bases
- Classify and organize documents with workspaces, tags, and facets
- Process documents on the fly (parse, extract)
- Multimodal retrieval API
- White-label Chat & Search interface
Use cases
- Deploying state-of-the-art AI behind firewalls
- Safely leveraging sensitive data
- Parsing documents
- Extracting any field from documents (e.g., invoice numbers, lease end dates, claim IDs, contract clauses)
- Grounded retrieval with citations
- Building knowledge-retrieval pipelines
- Building searchable knowledge bases
- Ingesting documents for search and Q&A
- Classifying and organizing documents
- Converting PDFs, Office files, and images to structured Markdown
- Enhancing productivity in business units
- Information retrieval for public agents (e.g., Île-de-France Regional Council)
- Extracting key insights from technical papers for engineers (e.g., Safran)
- Enhancing SEO strategy and content creation (e.g., Babbar)
- Integrating secure reasoning into CRM and ERP applications
AI approach
LightOn is a leading European generative AI company specializing in document intelligence for critical environments. They deploy an on-premise RAG platform for unstructured data, accessible via API or interface. They build their own large language models, including open-source foundation models with over 100 billion parameters and a flagship model named Alfred. Their technology includes LightOnOCR-2 for state-of-the-art parsing of various document types, and a hybrid retrieval system (dense, sparse, late-interaction) built on their open-source ColBERT family models (LateOn, NextPlaid). They offer an LLM-agnostic infrastructure, allowing users to bring their own models.
Tech named: VLM-4, LLM, Gen AI, RAG platform, API, LightOnOCR-2, LateOn, NextPlaid, PyLate, DenseOn, ColBERT family, Model Context Protocol, Alfred
Industries served
- Public Sector
- Aerospace
- Software Development
- Enterprise
What it says sets it apart
- On-premise RAG platform for critical environments
- Focus on confidentiality and value creation
- User-centric approach
- Developed over 12 Large Language Models, including open-sourced foundation models and a flagship commercial model (Alfred)
- Built on open research with models like LateOn and NextPlaid (open-source ColBERT family)
- Hybrid retrieval that picks the right signal (dense, lexical, late-interaction) automatically
- Grounded by default with auditable answers and source passages
- LLM-agnostic, allowing customers to bring their own models
- MCP-native for integration into various agent systems
- Workspaces and ACLs at chunk level for scoped corpora and permissions
- Transparent usage-based pricing with no seat licenses or commitments
- Offers both API for developers and ready-to-use interface for business teams
- Expert team for custom AI development, training, and infrastructure support (Forge)
- Ability to adapt OCR model to new languages through targeted training (e.g., Arabic)
This profile was compiled from LightOn's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.