AI gateway for any team

One gateway.
Every AI model.
Your rules.

GateLLM gives your entire organisation — technical or not — safe, governed access to every AI model. Manage access, enforce standards, protect data, optimise costs, and track usage from one place.

Data stays in your networkWorks for any teamBring your own key or pay as you go

auth · DLP scan · standards · cache · routing · quota · log

Up to 80%

reduction in provider token costs via caching

0+ providers

OpenAI, Anthropic, Gemini, Groq, Mistral, self-hosted vLLM/Ollama & more

0+ features

Chat, KB, DLP, Guardrails, Standards, Posture, Residency & more

< 1 day

typical onboarding time for a whole team

Who it's for

Real value for every role
in your organisation.

You don't need to be a developer to benefit from GateLLM. Every persona gets concrete outcomes — not just access to a model.

Organisation Leadership

CEO · CTO · CFO

  • One consolidated AI cost — no scattered invoices from 10 different providers
  • DLP and incident tracking protect against regulatory fines before they happen
  • AI Standards enforce consistent quality and compliance across the whole company
  • Risk scoring surfaces problematic usage early, not after a breach
  • Prompt caching cuts provider token costs by up to 80% automatically

Developers & DevOps

Engineers · Platform teams

  • One API key and base URL — works with Cursor, Continue, Claude Code, any SDK
  • Coding standards injected automatically into every IDE request — no per-dev setup
  • Switch AI providers without changing a line of application code
  • Per-request logs with latency, tokens, and cost for debugging and optimisation
  • OpenAI-compatible format — no new SDKs, no migration effort

Business & Operations teams

Sales · HR · Finance · Marketing

  • AI chat in the browser — no API key, no install, no technical knowledge needed
  • Attach PDFs, Word docs, and images — OCR extracts text, AI answers questions about them
  • Knowledge Base: ask questions answered from your org's own documents, not the model's training data
  • Research Workspace: research a company or competitor from public URLs and GitHub, get a one-click brief grounded in real sources
  • Prompt Library: shared templates for reports, emails, analysis — reuse with one click
  • Compare GPT-4o, Claude, and Gemini side-by-side to find the best answer
  • Personal usage dashboard so you know what you've used before hitting a limit

IT, Security & Compliance

InfoSec · Legal · IT ops

  • DLP scans every message for PII, secrets, financial data — mask, block, or log
  • Content Guardrails block prompt injections and jailbreaks before they reach the model
  • Developer AI Posture: per-user risk scores, key hygiene, incident history in one view
  • Full audit trail: every request logged with user, model, latency, cost, and result
  • Automated offboarding: removing a member instantly revokes all their API keys
  • Policy Packs: apply jurisdiction bundles (EU AI Act incl. GDPR, UAE PDPL, India DPDP) with one click
  • Data residency: an active jurisdiction pack blocks routing to a provider key outside the required region — enforced, not just advisory
  • MCP Governance: tool calls from connected MCP servers get the same policy, DLP, and audit trail as chat traffic
  • Policy Tester: simulate prompts against your rules before activating them
  • Self-hostable — data never leaves your network, no third-party data sharing

The problem

AI adoption without governance
costs more than it saves.

When people connect to AI providers on their own, the organisation has no oversight — no visibility, no protection, no control over spend.

Everyone manages their own API key

Keys get shared, forgotten, or stay active after someone leaves. No central view of who uses what or what it costs.

Sensitive data reaches AI providers unscanned

Financial data, PII, credentials, and internal documents travel to external AI providers with no record and no way to stop it.

No shared standards, inconsistent quality

Every person prompts AI differently. Quality and results vary by individual with no way to set expectations at scale.

Everyone on your team

What your people get

Powerful AI tools.
No setup required.

Access is as simple as opening a browser. For technical users, the same gateway connects to their IDE, CLI, and code — without changing how they work.

AI Chat — built in, no install

Multi-turn conversations with any configured model. Works in any browser. No API key or technical knowledge required.

Document Reader — attach & analyse

Drag and drop PDFs, scanned images, Word docs, and spreadsheets. OCR extracts text automatically so you can ask the AI anything about the content.

Knowledge Base — ask your own docs

Query your org's uploaded documents with semantic search. Get answers grounded in your internal policies, runbooks, and specs — not the model's training data.

Research Workspace — briefs from real sources

NEW

Point it at a company's site or GitHub repo for a quick-check summary, or start a Project for persistent, RAG-grounded Q&A and one-click research briefs.

Compare models side by side

Send one prompt to GPT-4o, Claude, and Gemini simultaneously. Compare answers, speed, and cost in one view.

Team prompt library

Shared templates for code review, report drafting, and analysis. Insert any template with a / slash command.

Session management

View every device signed in to your account. Revoke any session instantly — takes effect immediately.

Personal usage dashboard

Token usage, cost, and model breakdown updated in real time. Know exactly what you've used before you hit a limit.

IDE & CLI integration (developers)

One API key and base URL connects Cursor, Windsurf, Continue, or any SDK. AI Standards applied automatically.

Admins & owners

What you stay in control of

Full visibility.
Precise control.

Every setting lives in one dashboard. Grant access, set limits, review incidents, optimise costs, and see usage — without touching anyone's machine.

Per-person access control

Grant or revoke chat, IDE, CLI, and API access per person. Set monthly token limits and rate limits — changes take effect immediately.

Data loss prevention (DLP)

Scan every message for secrets, PII, financial data, and sensitive content before it reaches any AI provider. Configure each severity level to allow, mask, block, or log.

AI Standards — system prompts at scale

Write your guidelines once. GateLLM injects them automatically into every request across chat, IDE, CLI, and SDK — invisible to users, enforced on every message.

Provider routing & failover

12+ providers including self-hosted vLLM and Ollama — keep traffic entirely on your own infrastructure. Route by round-robin, weight, or failover with zero client changes.

Smart Auto-Routing — best model per request

NEW

Point traffic at gatellm/auto and it picks the best real model for each request automatically, based on your configured routing rules — no per-request model choice needed.

Content Guardrails — block threats at the gateway

NEW

Detect and block prompt injections, jailbreak attempts, blocked topics, and harmful model outputs. Configurable by category with block or log-only actions.

Knowledge Base — org document management

NEW

Upload internal documents for RAG-powered search. Members query your docs from chat and get answers grounded in your org's own knowledge.

Prompt caching — cut token costs 60–80%

Automatically cache repeated system prompts and message prefixes at the provider level. Cache hits cost a fraction of full input tokens on supported providers.

Response Cache — zero tokens for identical requests

NEW

Cache complete responses to identical non-streaming requests. Repeat queries return instantly with no provider call — configurable TTL from 5 minutes to 24 hours.

Conversation summary — keep long chats useful

When enabled, long conversations are automatically summarised so context is never lost and requests stay within model limits — no user action required.

Risk scoring & incident management

Every member receives a live risk score based on DLP violations and policy breaches. Drill into incident history per user, filter by severity, and export for compliance.

Analytics, audit logs & notifications

Token and cost analytics by user, model, and provider. Full exportable audit trail. Alert rules via Slack, email, Teams, or webhook for any event.

Developer AI Posture — per-user security view

NEW

See every developer's risk level, DLP/guardrail incidents, active and dormant API keys, channel usage, and top models — in one governance dashboard with 7d/30d/90d range and live refresh.

Automated offboarding — keys revoked on removal

NEW

When an admin removes a member, all their active API keys are revoked instantly — no manual cleanup step. Key count is recorded in the audit log.

Policy Packs — governance rule bundles

NEW

Apply versioned jurisdiction bundles (EU AI Act incl. GDPR, UAE PDPL, India DPDP) that merge most-restrictive-wins into your DLP and guardrail config. Preview impact before activating.

Data Residency — enforced, not advisory

NEW

An org with an active jurisdiction pack can only route to a provider key tagged for a permitted region. No matching key, no route — a real routing-level control, not just a content policy.

MCP Governance Gateway

NEW

Register upstream MCP servers and govern their tool calls the same way LLM traffic is governed — policy, DLP scanning, and audit logging apply to tool calls too.

Policy Tester — simulate before you ship

NEW

Run any prompt through your current DLP, guardrail, and AI Standard config in a sandboxed simulator — see exactly what would be blocked, masked, or injected before your rules go live.

Cost optimisation

GateLLM can pay for itself.

Three built-in features reduce the token costs your org pays to AI providers — automatically, with no changes needed from users.

Prompt Caching

ENABLE_PROMPT_CACHING=true

System prompts and repeated message prefixes are cached at the provider level. When the next request reuses the same prefix, it reads from cache instead of re-processing all those tokens.

60–80%

typical savings on prefix tokens

Anthropic OpenAI

supported providers

~5 min

provider-side cache TTL

A team sending 1,000 requests per day with a 2,000-token system prompt saves roughly 1.6 million cached input tokens daily — on Anthropic that is over $2,000 per month at list pricing.

Conversation Summary

ENABLE_CONV_SUMMARY=true

Long conversations balloon in token count — each reply sends the full history. When enabled, GateLLM automatically summarises older turns and prepends the summary, keeping the context window efficient without losing meaning.

Without summary

Message 40 sends ~39,000 tokens of history. Cost and latency grow linearly.

With summary

Older turns are condensed. Context stays useful. Token count stays manageable.

Especially valuable for long coding sessions, research threads, and support conversations where context is critical but token budgets are not unlimited.

Headroom Compression

ENABLE_COMPRESSION=true

A self-hosted compression pass runs on every eligible request before it reaches the model — normalising whitespace and trimming redundant content — shrinking the token count without changing what you asked.

Runs entirely inside your own deployment — no request content is ever sent to a third party to compress it.

All three together

Teams that enable prompt caching, conversation summary, and compression typically see 40–70% total reduction in monthly provider spend versus ungoverned direct API access — often covering the GateLLM platform cost many times over.

How it works

One endpoint in. Every provider out.

Every AI request from your team flows through GateLLM. Your policies run on every message, automatically, before anything reaches an external provider.

WEBAI Chat (any browser)
IDECursor, Windsurf, Continue
CLIllm, opencommit, ai-shell
SDKPython, Node.js, Go, LangChain
GateLLMgateway
authDLPstandardscacheroutingquotarisklog
Anthropic ClaudeANT
OpenAI GPT-4oOAI
Google GeminiGEM
OpenRouter + moreOR
Your own vLLM / OllamaSELF

Competitive position

The only AI gateway built
for the whole organisation.

Other tools solve one problem well — routing, observability, or proxying. GateLLM handles all of them, plus the governance and non-technical access layers that enterprises actually need — deployed entirely on your own infrastructure.

CapabilityGateLLMself-hostedLiteLLMopen sourcePortKeySaaSHeliconeSaaSOpenRouterSaaS
Built-in chat UI for non-developers✓————
AI Standards (auto system prompts)✓————
Data loss prevention (DLP)✓~~——
Content Guardrails (prompt injection, jailbreak)✓~~——
Knowledge Base (RAG on org docs)✓————
Document Reader (OCR + AI analysis)✓————
Per-user access control & quotas✓————
Self-hostable / on-prem (full stack, data never leaves your infra)✓✓~——
Data residency (region-pinned routing enforcement)✓————
Prompt caching (prefix-level)✓~✓——
Response Cache (full response dedup)✓—~~—
Conversation summary / token optimisation✓————
Risk scoring per user✓————
Notifications (Slack / email / webhook)✓—~~—
Prompt Library for non-tech users✓————
Developer AI Posture (per-user risk dashboard)✓————
Automated API key revocation on offboarding✓————
Policy Packs (jurisdiction compliance bundles)✓————
Policy Tester (sandbox rule simulation)✓————
MCP server tool-call governance✓————
Multi-provider routing & failover✓✓✓~✓
Usage analytics by user & model✓~✓✓—
~ = available via a bolt-on/third-party integration or a paid tier, not native out of the box. Re-verified against public vendor docs as of August 2026 — check current docs before relying on this for a purchase decision. Portkey's gateway is open-source and self-hostable, but its control plane remains SaaS-hosted in most deployments. Helicone has been in maintenance mode since its March 2026 acquisition by Mintlify.

What only GateLLM does

Chat UI for every employee

Every other gateway is a developer proxy. GateLLM ships a full chat workspace so non-technical staff get governed AI access without a separate product or a per-seat SaaS subscription.

AI Standards injected at the gateway

No other self-hosted gateway supports injecting system prompts per channel at the proxy layer. Standards apply to every tool — chat, Cursor, llm CLI, SDK — without the user knowing or doing anything.

Knowledge Base + Document Reader

Upload your org's docs and let members query them from chat using RAG. Attach files for OCR-powered AI analysis inline — no separate tool, no data leaving your infrastructure.

Developer Posture & automated offboarding

GateLLM is the only self-hosted gateway with a per-developer security posture view — live risk scores, key hygiene, incident history, and channel breakdowns per person. Remove a member and all their API keys are revoked instantly — no manual step.

Policy Packs & sandbox testing

Jurisdiction governance bundles (EU AI Act incl. GDPR, UAE PDPL, India DPDP) merge most-restrictive-wins into your DLP and guardrail config with one click. The Policy Tester lets you simulate any prompt against your current rules before they go live — so you never deploy a config blind.

Data residency, actually enforced

Every other tool treats jurisdiction as a content-policy label. GateLLM ties it to routing: an active jurisdiction pack blocks a request from ever reaching a provider key outside the permitted region — a real routing-level control, not a checkbox.

✓Choose GateLLM if

  • Your whole org needs safe AI access, not just developers behind a proxy
  • Data residency or on-prem is a hard requirement — self-hosted, not just SaaS-with-a-VPC-option
  • You mix commercial APIs with self-hosted models (vLLM, Ollama) and want one governance layer over both
  • You need DLP, guardrails, and jurisdiction policy enforced natively, not through a third-party add-on

—Look elsewhere if

  • You're a solo developer or small team that just needs a thin proxy for cost/latency observability
  • You want a fully-managed SaaS with nothing to run yourself — GateLLM is self-hosted by design
  • All you need is routing and failover with no governance requirements — a lighter open-source proxy may suffice

Pricing model

Use your own keys or
pay only for what you use.

Two simple options — choose what fits your organisation. No token markup, no hidden fees, no surprise bills.

Bring Your Own Key

You supply the API keys.
We run the gateway.

Connect your existing OpenAI, Anthropic, or Google API keys. GateLLM routes all traffic through them — you pay your provider directly, and only pay GateLLM a flat platform fee.

  • Zero token markup — provider rates apply directly
  • Connect keys from OpenAI, Anthropic, Google, and more
  • All gateway features: DLP, routing, caching, quotas, audit logs
  • Swap or rotate keys without touching client config
  • Flat platform fee — predictable regardless of usage volume

Best for organisations that already have AI provider agreements or enterprise contracts, and want governance without a second vendor relationship for tokens.

Usage-Based

No provider accounts.
Pay only for what you use.

No existing API key needed. GateLLM provides model access and bills you in one consolidated invoice based on actual usage.

  • No direct provider account required to get started
  • Access GPT-4o, Claude, Gemini through a single relationship
  • Pay one invoice — we consolidate all provider costs
  • Per-user and per-team quotas prevent budget overruns
  • Full cost breakdown by model and user

Best for teams who want to start immediately, or who want a single consolidated AI spend they can charge back per department.

Works with

Plug in with no workflow change.

GateLLM speaks the OpenAI API format. Any tool that accepts a base URL and API key connects instantly — no workflow changes for your team.

IDEs & Editors
ContinueVS Code & JetBrainsRecommended
CursorPro plan required
Windsurfvia Continue extension
Claude Codevia base URL override
Continue is free and open source. Cursor requires a paid plan for a custom API endpoint.
CLI Tools
llm CLI
opencommit
ai-shell
aider
Set OPENAI_BASE_URL in your shell profile — all these tools pick it up automatically.
SDKs & Frameworks
Python openai
Node.js openai
LangChain
Vercel AI SDK
Go / REST / cURL
Pass base_url and a GateLLM API key to any OpenAI-compatible SDK.
Set once in your environment — works everywhere your team uses AI
export OPENAI_BASE_URL="https://api.gatellm.com/v1"
export OPENAI_API_KEY="gatellm_live_..."

Common questions

Frequently asked questions.

Does GateLLM see my prompts or store my data?

No. GateLLM is self-hosted — all traffic flows through your own infrastructure. Prompts pass through the gateway to your chosen AI provider and are never sent to any GateLLM server. Logs and conversation history are stored only in your own database.

Can I use my existing OpenAI or Anthropic API key?

Yes. You can bring keys from OpenAI, Anthropic, Google, Azure, and most other providers. GateLLM stores them AES-encrypted in your database and routes requests through them transparently. You pay your provider directly at their published rates.

How do non-technical users access AI through GateLLM?

Through the built-in chat workspace in the dashboard. Non-technical users get a full chat interface with file upload, conversation history, and AI Standards enforced automatically — no API key or setup needed.

What happens if my AI provider goes down?

GateLLM's failover routing detects failures and automatically retries with the next available key or provider in your group. You configure the fallback order in the admin dashboard — zero code changes required.

Is GateLLM compatible with tools like Cursor or Continue?

Yes. Any tool that accepts a custom base URL and API key works immediately. Set OPENAI_BASE_URL to your GateLLM endpoint and the tool routes through your gateway. Your org's DLP, quotas, and AI Standards apply to every request — from any tool.

How do we get started or request a demo?

Email hello@gatellm.io and we'll get back to you within one business day. We'll walk you through the setup, help you choose the right deployment option (self-hosted or managed), and make sure your team is up and running quickly.

Get started

Give your whole team governed AI access today.

Request access and we'll help you choose the right setup — bring your own keys or usage-based. Onboarding takes less than a day.