Every team calls a different model. None of them should call it unguarded — Kong AI Gateway puts one governed path between your applications and the models, MCP tools and agents they call — PII redacted, prompts guarded, tokens capped and every call costed, at $100 a month per model on Konnect Plus.
Buy through TechBag
Same software. Better outcome — at a lower cost.
How it’s rated
Full scoreboard ↓Quick answer
This page covers Kong AI Gateway — the gateway for LLM, MCP and agent traffic. The rest:
Most product pages skip this. We start here — so you buy a capability, not a buzzword.
One controlled path between your applications and the models, MCP tools and agents they call.
What consolidation actually replaces, dimension by dimension.
| Dimension | Apps calling models directly | Kong AI Gateway |
|---|---|---|
| Provider keys | Pasted into every application | Held once per provider entity |
| Personal data in prompts | Trusted to each developer | Redacted by the sanitiser in flight |
| Switching models | A code change per application | A routing change at the gateway |
| Agent and MCP calls | Direct, outside any policy | Same auth, limits and logs as models |
| The monthly bill | One provider invoice, no owner | Token cost per model and consumer |
| What it is NOT | — | Not a model host or answer evaluator |
The cheapest test is one application: route it through the gateway for a fortnight and compare its token bill and redaction log with the month before.
Vendors love diagrams; buyers need to know what they’re actually operating. Here’s the whole platform, demystified.
A provider entity holds the connection and credentials for one upstream — OpenAI, Anthropic, Azure AI, Amazon Bedrock, Gemini and others — so applications never carry vendor keys.
Each model is a named, governed endpoint mapped to a provider. Kong bills Konnect Plus per unique model proxied, measured hourly, so the model list is also the price list.
MCP servers and agents are first-class entities in 2.0: Kong can expose APIs as MCP tools and route agent-to-agent calls through the same policies as model traffic.
Version 2.x runs on its own control plane in Konnect, replacing the plugin-by-plugin setup of V1. Guards, caching, limits and cost reporting attach to the entities above.
Providers behind one API, models as the unit of price and policy — with MCP servers and agents governed on the same path.
Kong AI Gateway puts one policy layer in front of every model, MCP tool and agent your applications call.
Callers use one interface while Kong translates to OpenAI, Anthropic, Azure AI, Bedrock, Gemini and more — swap a model without a code change.
Spread requests across providers and models, and fail over when one slows down or errors, with streaming responses supported throughout.
Expose existing APIs as MCP tools, register MCP servers, and route agent-to-agent calls under the same auth and limits as model calls.
The AI Sanitizer strips personal data from prompts before they reach an upstream model — useful when support tickets or KYC notes feed a prompt.
Prompt Guard blocks listed topics and keywords; Semantic Prompt Guard catches injection and jailbreak attempts by meaning, not exact wording.
Semantic Response Guard checks what comes back; AWS Guardrails, Azure Content Safety and GCP Model Armor can be called as external checks.
Similar prompts are answered from cache instead of a fresh model call, and the Prompt Compressor trims long prompts before tokens are spent.
AI Rate Limiting Advanced enforces token and spend limits per consumer, so one runaway agent cannot drain the month’s model budget.
Tracks tokens, latency and cost per model and consumer, with audit logs, OpenTelemetry spans and metrics, and Konnect dashboards.
A five-minute quickstart, connecting a provider and creating MCP servers in 2.0, and controlling token cost at scale.
First model proxied, start to finish.
The new provider entity in practice.
Registering MCP servers under 2.0.
Token limits and cost reporting.
Want a live, India-context walkthrough for your environment?
Book a guided demo →Here’s what genuinely sets it apart — and exactly where it stops.
In The Forrester Wave: API Management Software, Q3 2026, Forrester wrote that Kong’s gateway has “more policies for LLM governance than any other vendor in this evaluation.” In practice that means PII redaction, topic and semantic prompt guards, a response guard and hooks into AWS, Azure and Google safety services on one path.
Version 2.0 treats providers, models, MCP servers and agents as named entities on one control plane. Model calls, tool calls and agent-to-agent traffic pass the same authentication, limits and logging — rather than three separate proxies bolted on as teams adopt agents.
Konnect Plus lists $100 a month per unique model proxied, metered hourly, with a cap of five models. That makes a pilot easy to cost. The catch is that the enterprise AI plugins — semantic cache, sanitiser, advanced rate limiting — are an unpriced add-on on Plus and only included on Enterprise.
It governs traffic; it does not host models or evaluate answer quality. 2.x runs on a Konnect control plane, so a fully self-run setup means staying on the V1 plugins, which stay supported in Gateway 3.14 LTS and become opt-in from Gateway 3.18, as reported at the 2.0 launch. SSO and audit logs are Konnect Enterprise features.
Count the unique models you call and which apps and agents call them — that sets the Plus bill or an Enterprise quote.
Agree what gets redacted, which topics are blocked and each team’s token budget, and price the AI plugins you need.
Route one production app through a provider and model entity, with the sanitiser and a token limit switched on.
Register MCP servers and agents, move provider keys out of application code, and turn on semantic caching.
Hand finance per-team cost reports and plan any remaining V1 plugin setups onto 2.0 before Gateway 3.18 makes the AI plugins opt-in.
Modelled on Gartner Peer Insights structure. *Counts and breakdowns are illustrative pending verified review collection.
“Customer chat transcripts go to the model with account numbers and phone numbers already masked. Compliance signed off because of that.”
“We moved a summarisation workload from one provider to another in an afternoon. The apps kept calling the same endpoint throughout.”
“Per-team token caps ended the month-end surprise. Finance now gets a cost report per product line instead of one big invoice.”
“Plus was simple to cost at $100 a model, but the semantic cache and sanitiser needed the add-on conversation with sales.”
“Our agents call internal APIs as MCP tools now, behind the same auth as everything else. Security stopped asking for exceptions.”
“We ran the V1 plugins for a year. Moving to the 2.0 entity model was worth it, but budget real time for the migration.”
Analyst firms bury this view behind paywalls, and G2 retired its Grid. So here’s TechBag’s synthesis of the AI gateway market — tap any vendor to see why it sits where it does.
Execution strength vs product vision — the classic market map, minus the paywall.
Models, MCP and A2A; $100 per model on Plus.
The grid nobody publishes — how much control you get over where it runs and stores data vs how deep its LLM guardrails and limits go.
Konnect IN geo; widest LLM policy set per Forrester.
Positions are TechBag’s illustrative synthesis of public review-platform data and vendor documentation — not a reproduction of any analyst graphic. Verify before relying on it.
Against Cloudflare AI Gateway, Portkey, LiteLLM, Azure API Management and Google Apigee — on deployment, coverage, price, guardrails, limits and India.
| Dimension | Kong AI Gateway | Cloudflare AI Gateway | Portkey (Palo Alto Networks) | LiteLLM | Azure API Management (AI gateway) | Google Apigee (AI policies) |
|---|---|---|---|---|---|---|
| What it is | LLM, MCP and A2A gateway | Edge AI proxy | AI gateway, now PANW | Open-source LLM proxy | Policies inside APIM | Policies inside Apigee |
| Deployment | Konnect control plane | Cloudflare edge only | SaaS, OSS or VPC | Self-hosted, air-gap | Azure, some self-host | Google-hosted or hybrid |
| Coverage | Models, MCP, agents | Model providers | Universal model API | Many providers, one API | Models, MCP and A2A | Token and safety rules |
| Pricing model | Per model, hourly | Free core + usage | Per recorded logs | Free OSS or quote | Tier unit-hours | Per call + environment |
| Published entry price | $100/model/month | $0 on every plan | Free; $49/month Prod | Free OSS; Ent quote | 1M calls free | $100 per 1M calls |
| Included vs add-on | AI plugins: add-on | Core free, guards billed | Cache from Production | Governance is Enterprise | Cache needs Redis | Security add-on billed |
| Scale and limits | 5 models on Plus | Log caps by plan | Logs by plan | Your own capacity | Calls by tier | QPS by environment |
| Guardrails and security | PII, prompt, response | DLP and Llama Guard | Guardrails, Prisma AIRS | Mar 2026 PyPI incident | Content Safety policy | Model Armor, preview |
| Integrations | OTel, clouds, Konnect | Cloudflare stack | Fallbacks, private LLMs | SDK plus proxy | Token metrics, Foundry | Google Cloud native |
| Governance and SSO | SSO on Enterprise | Via Cloudflare account | RBAC; SSO on Ent | SSO on Enterprise | Token quotas per key | Quotas per API product |
| India storage region | Konnect IN geo | No region stated | VPC on Enterprise | Wherever you run it | Central India priced | Hybrid on subscription |
| Support | Email on Plus | By Cloudflare plan | Priority on Enterprise | SLA on Enterprise | Azure support plan | SLA by subscription |
| Lock-in and exit | Konnect-tied 2.x | Easy in, easy out | Roadmap now PANW’s | MIT, self-run | Tied to Azure | Tied to Google Cloud |
| Best fit | Governed AI at scale | Fast, free start | Palo Alto estates | Engineers who self-run | Microsoft-stack teams | Apigee customers |
Honest fit signals — because the fastest way to lose your trust is to pretend one product wins every scenario.
Kong AI Gateway is one of 35 developer tools products TechBag carries. The Developer Tools guide narrows them to a shortlist and shows the reasoning. →
Drag the sliders (developers building on LLM APIs; developer-hour cost). Estimates model developer time spent wiring provider keys, retries, rate-limit handling, redaction and cost reconciliation into each application — at an assumed 1.5 hours per developer a year, with 70% of it moved into the gateway. Both figures are assumptions. Illustrative.
Loaded cost = salary + overheads per productive hour. Illustrative only — your TechBag quote models your actual environment and modules.
Published: on Konnect Plus, $100 a month per unique model proxied, metered hourly, up to five models. The AI enterprise plugins are an add-on on Plus with no price shown, and included on Enterprise, which is custom-priced and billed annually. Prices are in USD and exclude taxes. TechBag counts your models and callers first, then quotes in INR with GST.
Best for a first governed AI workload
Best for a broader rollout
Best for regulated, multi-team AI traffic
Whatever the list prices above, TechBag negotiates a significantly better deal — with GST-compliant INR invoicing and local support. Ask us for your discounted quote.
Tell us your requirements and current tools — we’ll model it against what you spend today.
Take this into your next vendor call — including ours.
How many unique models will you proxy? Konnect Plus charges $100 a month each and stops at five.
Do you need semantic cache, sanitiser or advanced rate limiting? On Plus they are an add-on — get the price in writing.
Are you on the V1 plugins today? They stay supported in Gateway 3.14 LTS and become opt-in from 3.18, so plan the move to 2.0.
Is a Konnect-hosted control plane acceptable, given 2.x lives there rather than on your own servers?
Has Kong confirmed AI Gateway 2.x in the Konnect IN geo — and is that storage commitment in the contract?
Which fields must never reach a model — Aadhaar, PAN, account numbers — and has the sanitiser been tested on them?
Do you need SSO and audit logs? Both are Konnect Enterprise features, not part of Plus.
Which agents and MCP servers exist today, and who owns their authentication once they sit behind the gateway?
Count the models, callers and agents you would proxy first, or let a TechBag advisor scope a pilot that puts one application behind the sanitiser and a token limit.
Stats, ratings, review counts and pricing are illustrative and sourced from public materials; verify before purchase.