Liquid Inference

LLM router where providers compete for every prompt

Visit website

An LLM router where providers bid for every prompt, giving you the lowest price that meets your requirements.

Liquid Inference homepage

Liquid Inference homepage

1 / 2

Overview

What is Liquid Inference?

LLM router where providers compete for every prompt

Liquid Inference is an LLM router where providers compete for every prompt: send your request, and providers bid to answer it at the lowest price that meets your requirements. Prices change with supply and demand, so you always get the cheapest qualifying provider.

One API key covers OpenAI Chat Completions and Anthropic Messages on a single base URL, so existing code works unchanged. Set routing rules for model, region, speed and minimum quality, put a price limit on every request, and pay only for the tokens produced. Hundreds of open- and closed-weight models are supported, including GPT, Claude, Gemini, DeepSeek, Qwen, Kimi and Grok.

Pricing is per token with output tokens priced separately — for example around $0.08 per million output tokens for GPT OSS 20B. An interactive calculator on the site shows what your workload would cost right now.

Platforms and languages

Liquid Inference Availability

Platforms

WebApi

Languages

English

Capabilities

Liquid Inference Key Features

Provider bidding

Providers compete to answer every prompt; you pay the lowest price that meets your requirements.

One key, two APIs

OpenAI Chat Completions and Anthropic Messages on one base URL; existing code works unchanged.

Routing you control

Choose model, region, speed and minimum quality; only qualifying providers compete.

Price limit on every request

You know the most a request can cost before the model starts; pay only for the tokens produced.

Independent benchmarks

Every provider is tested regularly with standard benchmarks and the results are published.

Best for

Who uses Liquid Inference?

Developers and teams building on LLMs

Cut inference costs by routing every prompt to the cheapest qualifying provider

Plans and access

Liquid Inference Pricing

Paid

Per-token pricing; providers bid - e.g. $0.08/M output tokens for GPT OSS 20B

Common questions

Liquid Inference FAQs

How does pricing work?

Per token, with output tokens priced separately. Providers bid for every prompt, so prices move with supply and demand.

Do I need to change my code?

No. One API key works for OpenAI Chat Completions and Anthropic Messages on a single base URL.

Can I control which models serve my requests?

Yes. Choose the model, region, speed and minimum quality, and only providers that meet them compete for your prompt.

Reviews

See what the community thinks and share your experience.

—

Based on 0 ratings

Rating distribution

0
0
0
0
0

Leave a review

Sign in to rate

Community reviews

No reviews

No written reviews yet.

Explore by category

AI Apis Sdks

View all AI Apis Sdks websites
Semwright preview
Semwright

Free and open source drivers that give AI agents structured, application-native access to real software.

AI Apis Sdks
judged.systems preview
judged.systems

An API that judges customer support tickets — returning choices, scores or probabilities so your thresholds decide.

AI Apis Sdks
Claude Haiku 5.5 preview
Claude Haiku 5.5

Anthropic's fastest, cheapest small AI model for high-volume tasks like summaries, classification and coding subagents.

AI Apis Sdks
fal preview
fal

Generative media AI inference platform: 4,000+ models via API plus serverless GPUs.

AI Apis Sdks
Mistral Studio preview
Mistral Studio

Mistral's developer console and API platform — keys, Playground, evaluations, agents, and SDKs to ship AI apps. Free mode; pay-as-you-go.

AI Apis Sdks
BlueDoAI preview
BlueDoAI

One OpenAI-compatible API for supported AI models, with live catalog pricing and usage-based billing.

AI Apis Sdks

Explore by category

AI Development Infrastructure

View all AI Development Infrastructure websites

Explore by category

AI Development Platforms

View all AI Development Platforms websites
Ollama preview
Ollama

Free, open-source runtime for running open LLMs locally on macOS, Windows, and Linux — CLI, OpenAI-compatible API, GGUF import. Local runs unlimited.

AI Development Platforms