v1.14.2 · Open source · Self-hosted · AGPL-3.0

xLyra — Open-source AI gateway

Bring official APIs, subscription accounts and relay stations behind a single endpoint. Your apps talk to one address; xLyra handles routing, protocol conversion and cost estimation.

xlyra.example.com / live flow
Live traffic
18+upstream integration types
4text protocols, any-to-any
1downstream endpoint
9+gateway endpoints
01ONE ENDPOINT

Change one address,
and you're done

Every upstream you added used to mean another SDK config, another key and another bill. Now your app only needs to know xLyra.

  • No longer maintained separately
  • OpenAI official keyapi.openai.com
  • Anthropic official keyapi.anthropic.com
  • Gemini keygoogleapis.com
  • Balances on two relay stationsrelay-*.example.com
  • Codex subscription token refreshOAuth
  • Claude Code subscription authorizationOAuth
client.py−2 / +2
client = OpenAI(-   base_url="https://api.openai.com/v1",-   api_key="sk-…",+   base_url="https://your-xlyra/v1",+   api_key="<downstream API key>",)# Same for the Anthropic SDK; model names stay unchanged
Endpoints to maintain6 → 1
Sites · OAuth

Official APIs, subscriptions, relays —
added in one click

18 integration types. Authorize Codex and Claude Code subscriptions directly. xLyra refreshes tokens, syncs quotas where the provider supports it, and picks among your accounts.

Add site dialog listing official and relay integration types
02ROUTING

When an upstream fails,
requests take another road

Every request scores all available upstreams. A failing one goes into cooldown and rejoins the queue once it recovers.

POST/v1/messagesclaude-sonnet-5-5from Cursor workstation
Claude officialpriority 3 · 410ms
—
Relay Novapriority 2 · 520ms
—
Relay Alphapriority 1 · 51% success
—

Switching happens before the first byte is written; once a stream has started, the upstream is never swapped.

  • 01Health
  • 02Latency
  • 03Cooldown
  • 04Priority
Request log showing model, downstream key, site, status and latency for each request

Every attempt stays in the request log: which site, which upstream key, and why it switched.

Protocols

Speak any protocol — upstreams understand

Chat Completions, Responses, Anthropic Messages and Gemini convert into one another. Pick a downstream protocol:

Downstream · CHAT COMPLETIONSPOST /v1/chat/completions
{
  "model": "claude-sonnet-5-5",
  "messages": [
    { "role": "system", "content": "You are a helpful assistant" },
    { "role": "user", "content": "Hello" }
  ],
  "stream": true
}
Upstream · Claude official · ANTHROPIC MESSAGESPOST https://api.anthropic.com/v1/messages
{
  "model": "claude-sonnet-5-5",
  "system": "You are a helpful assistant",
  "max_tokens": 1024,
  "messages": [
    { "role": "user", "content": "Hello" }
  ],
  "stream": true
}
03CONTROL

One key per app,
every cent accounted for

Upstream credentials stay inside the gateway; apps only get a downstream key. Who can use what, how much, and what it cost — all in one console.

API key list with quotas, rate limits and status
  1. 1One downstream address that every app points to
  2. 2Allowlists by model, site or site group
  3. 3Total, daily and weekly quotas: a key stops when one runs out; daily and weekly quotas reset on their own
  4. 4Per-key RPM and TPM limits
Billing

Every request shows exactly how its cost was calculated

Input, output and cache hits are priced separately from the price table, upstream key multipliers included. Estimated costs are split by key, model and site — no more guessing at month end.

Request detail with usage and the cost formula
04A DAY WITH XLYRA

A day with xLyra

09:30Morning

Cursor and Claude Code, one address

Some of the team use Cursor, others use Claude Code. Both point at xLyra with their own downstream keys; the gateway decides which upstream and which account serves each request.

Console view: the site and model that actually served a request
14:20Afternoon

The official API got rate-limited. Nobody noticed.

Claude official returned 429, and the request switched to Relay Nova before the first byte. The rate-limited keys cool down, then rejoin the queue.

Live flow map: downstream apps through the gateway to upstream sites
18:00End of day

The support bot ran out of quota

Its key only has a $50 quota. When it is spent that key stops — it never eats into another app's budget. Daily and weekly quotas also reset on their own.

Quota usage on the API keys page
23:30Before bed

One glance: what each key cost today

Cost trends stacked by model, a ranking by user, cache hit rate. Which app costs the most is obvious at a glance.

Console: usage and cost trends
05DETAILS

And a few more things we thought of

Model mappingHard maps, soft maps and wildcard fallbacks keep the model names you expose stable.
Traffic topologyLive request flow between downstream apps, the gateway and upstream sites.
PlaygroundChat and generate images in the console to check that routing does what you expect.
WebSocket ResponsesThe Responses API over WebSocket goes through the gateway too.
Images · Embeddings · SpeechImage generation and editing, embeddings and TTS all share the same entry point.
S3-compatible backupBack up configuration and data to any S3-compatible storage on a schedule.
TOTP two-factorConsole sign-in supports one-time passwords.
Audit logEvery change an admin makes is recorded.
06DEPLOY

Any machine that runs Docker
will do

One image contains the Go backend and the React console. Your data stays in your own PostgreSQL.

$ curl -O https://raw.githubusercontent.com/Yachiyo-5i/xLyra/main/docker-compose.yml
$ docker compose up -d
✔ Container xlyra-postgres Started
✔ Container xlyra Started
→ Open http://your-server:5801 and create the admin account
docker-compose.yml
1services:2  postgres:3    image: postgres:17-alpine4    container_name: xlyra-postgres5    restart: unless-stopped6    environment:7      POSTGRES_DB: xlyra8      POSTGRES_USER: xlyra9      POSTGRES_PASSWORD: postgres_password10    volumes:11      - ./postgres:/var/lib/postgresql/data12    networks:13      - xlyra1415  xlyra:16    image: yachiiiiyo/xlyra:latest17    container_name: xlyra18    restart: unless-stopped19    depends_on:20      - postgres21    environment:22      DB_HOST: postgres23      DB_NAME: xlyra24      DB_USER: xlyra25      DB_PASSWORD: postgres_password26    volumes:27      - ./data:/data28    ports:29      - "5801:5801"30    networks:31      - xlyra3233networks:34  xlyra:35    driver: bridge
Before exposing it publicly, replace the database password and put the console behind HTTPS.See the latest config in the repo →
  • One imageThe Go backend and the React console ship in the same image.
  • First visitOpen the console, create the admin account, then add your first site.
  • Before going liveChange the database password and put the console behind HTTPS.
Cloud serverPublic deployment — just add HTTPS
NASSynology, Unraid and others with built-in Docker
Company networkGive the team a single exit point
Your laptopSpin it up locally while developing
07FAQ

Frequently asked questions

What is xLyra?

xLyra is an open-source AI gateway and control plane. It sits between your applications and AI providers, exposes unified OpenAI- and Anthropic-compatible endpoints, and handles routing, protocol conversion, rate limiting, cost estimation and failover.

Do I need to change my code to use xLyra?

Usually not. Point the base URL of your existing OpenAI or Anthropic SDK at xLyra and use a downstream API key issued in the console.

Which upstreams does xLyra support?

18 integration types are available: OpenAI, Anthropic Claude, Google Gemini, OpenCode Go, DeepSeek, MiniMax, Xiaomi MiMo, Moonshot, Kimi Code, GLM, GLM Code and TypeSafe; Grok connects through xAI device authorization; NewAPI and xLyra can act as cascaded or aggregated sites; Codex OAuth, Antigravity OAuth and Claude Code OAuth accounts are supported as well.

Do requests switch automatically when an upstream fails?

Yes. The routing engine scores every upstream on health, latency, cooldown state and priority, and moves to the next one when a request fails. Switching only happens before the first byte is written; once a stream has started, the upstream is never swapped.

Where is my data stored?

You deploy xLyra yourself. Site credentials, request logs and usage data live in your own PostgreSQL, with S3-compatible backup support.

How do I deploy xLyra?

Start xLyra and PostgreSQL with Docker Compose, then open the console and create the admin account. Before exposing it publicly, change the database password and set up HTTPS, a reverse proxy and access control.

Why not call each provider's API directly?

Calling them directly means maintaining protocols, keys, rate limits, failure handling and billing separately. xLyra gathers all of that in one gateway: your app only needs to know one address, and the control plane does the rest.

What open-source license does xLyra use?

xLyra is released under AGPL-3.0, with the source hosted on GitHub.

Every upstream,
behind one endpoint

Open source, self-hosted, compatible with the SDKs you already use.

ONE ENDPOINT · EVERY MODEL