AI solutions, for everyone
and everything

InferencePort has the full stack - from hosting, routing, and scaling AI models in the cloud through our unified API, and to running them locally on your own hardware. Lightning-fast, cost-effective, and open-source.

1,000
words / sec (cloud)
3
platforms supported
Free
to start

Two products. One AI backbone.

The InferencePort AI App for chat, hosting, and model management. The Developer Console for production APIs and fraud detection. Both powered by Lightning.

InferencePort AI App

Chat, host, and manage models.

The desktop app for end users and teams. Cloud chat powered by the Generation API, self-hosted AI servers with team access, local Ollama models, and a built-in marketplace.

💬
Generation API (Subscription)
Default chat path. Plan quotas, OpenAI-style endpoint, optimized for everyday chatting.
🌐
Self-Hosted AI Server
Run local LLMs, invite teams by email, authenticate via accounts or custom API keys.
📦
Model Marketplace
Browse, download, and run community models and tools.
💻
Local Models (Ollama)
100% on-device, fully offline, unlimited local chat.
{ } Developer Console

Production APIs & security.

The backend for developers. Credit-metered API for enterprise integrations, and real-time fraud detection for sign-ups and traffic.

P2G API
Credit-Billed Production
  • ✓ Wallet-based billing, separate ledger
  • ✓ Enterprise & production integrations
  • ✓ Managed through the console
AI Shield
Fraud Detection
  • ✓ Email, phone, IP, username, device signals
  • ✓ Risk score + confidence level
  • ✓ Duplicate & linked-account detection
Open Console →
Lightning

The model behind the platform.

Lightning is a model available over the API and the default model used by all InferencePort AI subscriptions. Up to 1,000 words/sec, with smart routing to the best provider for each task.

1,000
words / sec
Default model
API
Available over
OpenAI Llama Qwen Nemotron
Try It →

Cloud or Local — you decide.

Cloud Models
  • ✦ No GPU or hardware required
  • ✦ Instant access to frontier models
  • ✦ Up to 1,000 words/sec throughput
  • ✦ Auto-updated — always latest models
  • ✦ Image, video & audio generation
vs
Local Models
  • ✦ 100% private — nothing leaves your device
  • ✦ Works fully offline
  • ✦ Unlimited local chat
  • ✦ Full Ollama model compatibility
  • ✦ Remote server connection supported

Ship in three steps.

1
Sign in

Open the console to view your wallet, plan, usage, and API keys.

2
Choose your path

Cloud APIs for hosted workloads, or launch an AI Server for local models.

3
Ship it

Chat traffic on subscriptions, production on credit-based P2G. Scale when ready.

Scale when you're ready.

Start free with generous limits. Upgrade for unlimited cloud generation.

Free
$0
forever

Cloud
50 Cloud Chats / day
10 Images / day
3 Videos / day
1 Audio / week

Local
Unlimited Local Chat
Marketplace Access
HuggingFace Spaces
Remote Server
Get Started
AI Light
$9.99
per month

Cloud
Unlimited Cloud Chat
50 Images / day
10 Videos / day
5 Audio / week

Local
Unlimited Local Chat
Marketplace Access
HuggingFace Spaces
Remote Server
Subscribe
AI Professional
$99.99
per month

Cloud
Unlimited Cloud Chat
Unlimited Images
Unlimited Videos
75 Audio / week

Local
Unlimited Local Chat
Marketplace Access
HuggingFace Spaces
Remote Server
Subscribe

Join the Community

Collaborate, contribute, and explore new possibilities with developers worldwide.