AI gateway for any team
GateLLM gives your entire organisation — technical or not — safe, governed access to every AI model. Manage access, enforce standards, protect data, optimise costs, and track usage from one place.
auth · DLP scan · standards · cache · routing · quota · log
Up to 80%
reduction in provider token costs via caching
0+ providers
OpenAI, Anthropic, Gemini, Groq, Mistral, self-hosted vLLM/Ollama & more
0+ features
Chat, KB, DLP, Guardrails, Standards, Posture, Residency & more
< 1 day
typical onboarding time for a whole team
Who it's for
You don't need to be a developer to benefit from GateLLM. Every persona gets concrete outcomes — not just access to a model.
CEO · CTO · CFO
Engineers · Platform teams
Sales · HR · Finance · Marketing
InfoSec · Legal · IT ops
The problem
When people connect to AI providers on their own, the organisation has no oversight — no visibility, no protection, no control over spend.
Keys get shared, forgotten, or stay active after someone leaves. No central view of who uses what or what it costs.
Financial data, PII, credentials, and internal documents travel to external AI providers with no record and no way to stop it.
Every person prompts AI differently. Quality and results vary by individual with no way to set expectations at scale.
What your people get
Access is as simple as opening a browser. For technical users, the same gateway connects to their IDE, CLI, and code — without changing how they work.
Multi-turn conversations with any configured model. Works in any browser. No API key or technical knowledge required.
Drag and drop PDFs, scanned images, Word docs, and spreadsheets. OCR extracts text automatically so you can ask the AI anything about the content.
Query your org's uploaded documents with semantic search. Get answers grounded in your internal policies, runbooks, and specs — not the model's training data.
Point it at a company's site or GitHub repo for a quick-check summary, or start a Project for persistent, RAG-grounded Q&A and one-click research briefs.
Send one prompt to GPT-4o, Claude, and Gemini simultaneously. Compare answers, speed, and cost in one view.
Shared templates for code review, report drafting, and analysis. Insert any template with a / slash command.
View every device signed in to your account. Revoke any session instantly — takes effect immediately.
Token usage, cost, and model breakdown updated in real time. Know exactly what you've used before you hit a limit.
One API key and base URL connects Cursor, Windsurf, Continue, or any SDK. AI Standards applied automatically.
What you stay in control of
Every setting lives in one dashboard. Grant access, set limits, review incidents, optimise costs, and see usage — without touching anyone's machine.
Grant or revoke chat, IDE, CLI, and API access per person. Set monthly token limits and rate limits — changes take effect immediately.
Scan every message for secrets, PII, financial data, and sensitive content before it reaches any AI provider. Configure each severity level to allow, mask, block, or log.
Write your guidelines once. GateLLM injects them automatically into every request across chat, IDE, CLI, and SDK — invisible to users, enforced on every message.
12+ providers including self-hosted vLLM and Ollama — keep traffic entirely on your own infrastructure. Route by round-robin, weight, or failover with zero client changes.
Point traffic at gatellm/auto and it picks the best real model for each request automatically, based on your configured routing rules — no per-request model choice needed.
Detect and block prompt injections, jailbreak attempts, blocked topics, and harmful model outputs. Configurable by category with block or log-only actions.
Upload internal documents for RAG-powered search. Members query your docs from chat and get answers grounded in your org's own knowledge.
Automatically cache repeated system prompts and message prefixes at the provider level. Cache hits cost a fraction of full input tokens on supported providers.
Cache complete responses to identical non-streaming requests. Repeat queries return instantly with no provider call — configurable TTL from 5 minutes to 24 hours.
When enabled, long conversations are automatically summarised so context is never lost and requests stay within model limits — no user action required.
Every member receives a live risk score based on DLP violations and policy breaches. Drill into incident history per user, filter by severity, and export for compliance.
Token and cost analytics by user, model, and provider. Full exportable audit trail. Alert rules via Slack, email, Teams, or webhook for any event.
See every developer's risk level, DLP/guardrail incidents, active and dormant API keys, channel usage, and top models — in one governance dashboard with 7d/30d/90d range and live refresh.
When an admin removes a member, all their active API keys are revoked instantly — no manual cleanup step. Key count is recorded in the audit log.
Apply versioned jurisdiction bundles (EU AI Act incl. GDPR, UAE PDPL, India DPDP) that merge most-restrictive-wins into your DLP and guardrail config. Preview impact before activating.
An org with an active jurisdiction pack can only route to a provider key tagged for a permitted region. No matching key, no route — a real routing-level control, not just a content policy.
Register upstream MCP servers and govern their tool calls the same way LLM traffic is governed — policy, DLP scanning, and audit logging apply to tool calls too.
Run any prompt through your current DLP, guardrail, and AI Standard config in a sandboxed simulator — see exactly what would be blocked, masked, or injected before your rules go live.
Cost optimisation
Three built-in features reduce the token costs your org pays to AI providers — automatically, with no changes needed from users.
ENABLE_PROMPT_CACHING=true
System prompts and repeated message prefixes are cached at the provider level. When the next request reuses the same prefix, it reads from cache instead of re-processing all those tokens.
60–80%
typical savings on prefix tokens
Anthropic OpenAI
supported providers
~5 min
provider-side cache TTL
A team sending 1,000 requests per day with a 2,000-token system prompt saves roughly 1.6 million cached input tokens daily — on Anthropic that is over $2,000 per month at list pricing.
ENABLE_CONV_SUMMARY=true
Long conversations balloon in token count — each reply sends the full history. When enabled, GateLLM automatically summarises older turns and prepends the summary, keeping the context window efficient without losing meaning.
Without summary
Message 40 sends ~39,000 tokens of history. Cost and latency grow linearly.
With summary
Older turns are condensed. Context stays useful. Token count stays manageable.
Especially valuable for long coding sessions, research threads, and support conversations where context is critical but token budgets are not unlimited.
ENABLE_COMPRESSION=true
A self-hosted compression pass runs on every eligible request before it reaches the model — normalising whitespace and trimming redundant content — shrinking the token count without changing what you asked.
Runs entirely inside your own deployment — no request content is ever sent to a third party to compress it.
All three together
Teams that enable prompt caching, conversation summary, and compression typically see 40–70% total reduction in monthly provider spend versus ungoverned direct API access — often covering the GateLLM platform cost many times over.
How it works
Every AI request from your team flows through GateLLM. Your policies run on every message, automatically, before anything reaches an external provider.
Competitive position
Other tools solve one problem well — routing, observability, or proxying. GateLLM handles all of them, plus the governance and non-technical access layers that enterprises actually need — deployed entirely on your own infrastructure.
| Capability | GateLLMself-hosted | LiteLLMopen source | PortKeySaaS | HeliconeSaaS | OpenRouterSaaS |
|---|---|---|---|---|---|
| Built-in chat UI for non-developers | ✓ | — | — | — | — |
| AI Standards (auto system prompts) | ✓ | — | — | — | — |
| Data loss prevention (DLP) | ✓ | ~ | ~ | — | — |
| Content Guardrails (prompt injection, jailbreak) | ✓ | ~ | ~ | — | — |
| Knowledge Base (RAG on org docs) | ✓ | — | — | — | — |
| Document Reader (OCR + AI analysis) | ✓ | — | — | — | — |
| Per-user access control & quotas | ✓ | — | — | — | — |
| Self-hostable / on-prem (full stack, data never leaves your infra) | ✓ | ✓ | ~ | — | — |
| Data residency (region-pinned routing enforcement) | ✓ | — | — | — | — |
| Prompt caching (prefix-level) | ✓ | ~ | ✓ | — | — |
| Response Cache (full response dedup) | ✓ | — | ~ | ~ | — |
| Conversation summary / token optimisation | ✓ | — | — | — | — |
| Risk scoring per user | ✓ | — | — | — | — |
| Notifications (Slack / email / webhook) | ✓ | — | ~ | ~ | — |
| Prompt Library for non-tech users | ✓ | — | — | — | — |
| Developer AI Posture (per-user risk dashboard) | ✓ | — | — | — | — |
| Automated API key revocation on offboarding | ✓ | — | — | — | — |
| Policy Packs (jurisdiction compliance bundles) | ✓ | — | — | — | — |
| Policy Tester (sandbox rule simulation) | ✓ | — | — | — | — |
| MCP server tool-call governance | ✓ | — | — | — | — |
| Multi-provider routing & failover | ✓ | ✓ | ✓ | ~ | ✓ |
| Usage analytics by user & model | ✓ | ~ | ✓ | ✓ | — |
| ~ = available via a bolt-on/third-party integration or a paid tier, not native out of the box. Re-verified against public vendor docs as of August 2026 — check current docs before relying on this for a purchase decision. Portkey's gateway is open-source and self-hostable, but its control plane remains SaaS-hosted in most deployments. Helicone has been in maintenance mode since its March 2026 acquisition by Mintlify. | |||||
Every other gateway is a developer proxy. GateLLM ships a full chat workspace so non-technical staff get governed AI access without a separate product or a per-seat SaaS subscription.
No other self-hosted gateway supports injecting system prompts per channel at the proxy layer. Standards apply to every tool — chat, Cursor, llm CLI, SDK — without the user knowing or doing anything.
Upload your org's docs and let members query them from chat using RAG. Attach files for OCR-powered AI analysis inline — no separate tool, no data leaving your infrastructure.
GateLLM is the only self-hosted gateway with a per-developer security posture view — live risk scores, key hygiene, incident history, and channel breakdowns per person. Remove a member and all their API keys are revoked instantly — no manual step.
Jurisdiction governance bundles (EU AI Act incl. GDPR, UAE PDPL, India DPDP) merge most-restrictive-wins into your DLP and guardrail config with one click. The Policy Tester lets you simulate any prompt against your current rules before they go live — so you never deploy a config blind.
Every other tool treats jurisdiction as a content-policy label. GateLLM ties it to routing: an active jurisdiction pack blocks a request from ever reaching a provider key outside the permitted region — a real routing-level control, not a checkbox.
Pricing model
Two simple options — choose what fits your organisation. No token markup, no hidden fees, no surprise bills.
Connect your existing OpenAI, Anthropic, or Google API keys. GateLLM routes all traffic through them — you pay your provider directly, and only pay GateLLM a flat platform fee.
Best for organisations that already have AI provider agreements or enterprise contracts, and want governance without a second vendor relationship for tokens.
No existing API key needed. GateLLM provides model access and bills you in one consolidated invoice based on actual usage.
Best for teams who want to start immediately, or who want a single consolidated AI spend they can charge back per department.
Works with
GateLLM speaks the OpenAI API format. Any tool that accepts a base URL and API key connects instantly — no workflow changes for your team.
Common questions
Does GateLLM see my prompts or store my data?
No. GateLLM is self-hosted — all traffic flows through your own infrastructure. Prompts pass through the gateway to your chosen AI provider and are never sent to any GateLLM server. Logs and conversation history are stored only in your own database.
Can I use my existing OpenAI or Anthropic API key?
Yes. You can bring keys from OpenAI, Anthropic, Google, Azure, and most other providers. GateLLM stores them AES-encrypted in your database and routes requests through them transparently. You pay your provider directly at their published rates.
How do non-technical users access AI through GateLLM?
Through the built-in chat workspace in the dashboard. Non-technical users get a full chat interface with file upload, conversation history, and AI Standards enforced automatically — no API key or setup needed.
What happens if my AI provider goes down?
GateLLM's failover routing detects failures and automatically retries with the next available key or provider in your group. You configure the fallback order in the admin dashboard — zero code changes required.
Is GateLLM compatible with tools like Cursor or Continue?
Yes. Any tool that accepts a custom base URL and API key works immediately. Set OPENAI_BASE_URL to your GateLLM endpoint and the tool routes through your gateway. Your org's DLP, quotas, and AI Standards apply to every request — from any tool.
How do we get started or request a demo?
Email hello@gatellm.io and we'll get back to you within one business day. We'll walk you through the setup, help you choose the right deployment option (self-hosted or managed), and make sure your team is up and running quickly.
Get started
Request access and we'll help you choose the right setup — bring your own keys or usage-based. Onboarding takes less than a day.