One gateway for every model you'll ever call.
Portkey AI sits between your application and every LLM provider you use — routing requests, enforcing policy, redacting sensitive data, and logging every call, without changing a line of your prompt logic.
Everything between your app and the model.
Six systems that work together so your team can ship LLM features without building infrastructure first.
AI Gateway
Unified API access to 1,600+ models from every major provider. Switch models or providers by changing a config value, not your application code.
Observability
Real-time dashboards for latency, error rate, token spend, and per-request traces across every model and provider you call.
Guardrails
Policy enforcement at the request and response layer — block disallowed content, redact PII automatically, and enforce output schemas.
Governance
Role-based access control, full audit logs, and per-team spend limits, with controls designed to map to HIPAA and SOC 2 requirements.
Prompt management
Version, test, and deploy prompts independently of application releases, with rollback and per-environment configuration.
Cost optimization
Semantic caching, batch routing, and provider-aware load balancing reduce token spend without touching response quality.
Every request passes through one gate.
Your application talks to a single Portkey endpoint. We handle everything else.
Request arrives at the gateway
Your app sends a standard chat or completion request to your Portkey endpoint instead of calling a provider directly.
Policy and guardrail checks run
Guardrails inspect the request for PII, disallowed content, and policy violations before it reaches any model.
Router selects the best model
Based on your rules — cost, latency, capability, or fallback order — the router picks a provider and forwards the request.
Response is logged and returned
The response is checked against output guardrails, logged for observability, and streamed back to your application.
Built for every team shipping LLM features.
The same gateway, configured differently depending on who's behind the request.
Engineering teams
Standardize on one API across every model your product uses, with automatic failover when a provider degrades.
Healthcare & finance
Route sensitive workloads through guardrails that redact PII and PHI before any payload reaches a third-party model.
Internal AI platforms
Give every team in your org a governed, observable way to call models, without each team managing their own keys.
Early-stage product teams
Start on a single model, then add fallbacks and cost controls as you scale, without rewriting your integration.
Fits into the stack you already run.
Portkey AI connects to the infrastructure and developer tools your team already uses.
Treated as production infrastructure, not a wrapper script.
Every request through Portkey AI is encrypted in transit, logged for audit, and subject to the access controls your team defines — whether you're a five-person startup or a regulated enterprise.
Compliance status reflects controls currently implemented and audits in progress. Contact us for current certification status and a signed DPA.
How a mid-size fintech consolidated four model providers into one gateway
A product team running separate integrations for OpenAI, Anthropic, and two regional providers moved all traffic behind Portkey to get a single point of observability and policy control — without changing their application code.
Start free. Scale when you need to.
Usage-based pricing with no markup on provider token costs.
Common questions
By default we retain request metadata (latency, model, status) for observability. Full payload logging is configurable per workspace and can be disabled entirely for sensitive workloads.
Yes — any provider with an OpenAI-compatible or custom REST interface can be added as a custom route through our configuration API.
You define a fallback order per route. If your primary model errors or times out, Portkey automatically retries against the next model in your list.
Self-hosted deployment is available on Enterprise plans. Contact sales to discuss your VPC and data residency requirements.
We charge a flat platform fee per workspace plus pass through provider token costs at no markup. Caching can reduce your effective token spend.
Route your first request in minutes.
Free to start. No credit card required.