Every product here meters on volume: hosts monitored, gigabytes ingested, events captured, sessions recorded. A per-user mental model — the one that works everywhere else on this site — produces a forecast that is wrong by an order of magnitude.
Datadog publishes three different meters inside one platform: APM and Infrastructure per host, Log Management per GB ingested plus per million events indexed, RUM per thousand sessions. Model your volumes against the meter, not the rate.
Already decided — What this page decides
Still yours to weigh
Observability is the practice of keeping enough telemetry — logs, metrics, traces — that you can answer a question you did not anticipate. Monitoring answers questions you wrote a check for in advance. That distinction sounds academic until you see the bill: keeping the detail is exactly what you pay for.
The thirteen products below split by meter more usefully than by feature. Four are per-host or per-GB platform modules. Six are event-metered and materially cheaper for the same job at small volumes. Two are quote-only. And two — Mixpanel’s pair — are product analytics rather than infrastructure observability, included because the meter and the buying team overlap, and labelled so you do not buy one for the other.
The row to get exactly right
The pricing meter. This is the category’s equivalent of a firewall’s throughput figure: a per-host quote and a per-GB quote for the same estate can differ by an order of magnitude, and no comparison on the internet models it against your volumes. Ask every vendor to price your actual host count, ingest volume and event rate — then compare the totals, never the rates.
Often confused withDatabases & Data Tools — why the query plan changed →·Developer Tools — where the code and the pipeline live →·SIEM & Log Management — the same logs, asked a security question →
These are not tiers. Observability is not monitoring done better — it answers questions you never wrote a check for, from telemetry you pay to keep, at a materially different cost.
Monitoring vs observability
Did you know in advance what to check?
Observability vs APM
The whole estate, or the application request?
Logs vs metrics vs traces
Which one answers your question, and what does it cost to keep?
APM vs RUM vs error tracking
Server-side performance, device experience, or exceptions?
Seven variables move the shortlist. The first one moves it more than the other six combined.
The pricing meter
Per host, per GB ingested, per million events indexed, per session, per user — or quote-only. Same estate, different meters, order-of-magnitude difference.
Retention, and what keeping it costs
The licence covers the first tranche. Shortening retention to control cost is the standard reflex, and the incident then falls outside the window.
Language and runtime coverage
For the stack you actually run, including the older service nobody wants to touch.
Cardinality limits
The failure mode nobody forecasts: one well-meant custom tag turns a metric into millions of time series.
Alerting and on-call integration
Whether alerts reach the rota you already run, or need a second product.
Open-source viability
Prometheus, Grafana and OpenTelemetry are genuinely credible here. A page pretending otherwise is not worth reading.
India region and residency
Undocumented across every vendor on this page — logs frequently carry personal data, so confirm rather than assume.
Pick what the telemetry has to cover and how it has to be billed. Products drop out with the reason stated, never silently.
What it has to see
How it has to work
What the bill counts
India
India residency is annotated rather than used to eliminate: no vendor here documents an India telemetry region, so the product is flagged for you to confirm rather than ruled out.
per host / month, published list; distributed tracing across services with the span-level detail that names the slow hop — billed on top of Infrastructure Monitoring rather than instead of it
Estates running microservices that need to follow one request across many services, and that already accept a platform bill rather than a per-tool one.
The catch: The meter is per host, so the bill tracks your infrastructure rather than your traffic — and it compounds with every other Datadog module you enable. Cost management is the standing operational task, not a one-time negotiation.
per host / month, published list; the base layer of the platform — host, container and cloud-service metrics with the dashboards and alerting most estates start from
Teams that want one place for infrastructure health across cloud accounts, and are choosing the platform deliberately rather than by accident.
The catch: Metrics only: it will tell you the p99 got worse and not which span caused it. Custom metrics and high cardinality are separately metered, and cardinality is the line that multiplies a bill without traffic multiplying.
per GB ingested PLUS per million events indexed, published list; ingest and indexing are separate meters, so what you keep searchable costs more than what you merely collect
Estates whose debugging happens in logs and who are prepared to decide, deliberately, which logs are worth indexing.
The catch: Two meters on one product is where forecasts break: teams model ingest and forget indexing. Almost every runaway observability bill in this category is a log bill.
per 1,000 sessions / month for RUM, with synthetic tests metered separately by run; what the browser or mobile app actually experienced, including the network and third-party scripts the backend never sees
Teams whose users report slowness that server-side metrics insist is not happening — the gap between what the server did and what the device experienced.
The catch: Session-metered, so a traffic spike is a bill spike with no infrastructure change to explain it. Synthetics are a third meter again.
per month on the Team tier (about ₹2,158), Business about $80, both annual; metered on events captured rather than hosts run — exceptions grouped by fingerprint with the release, the commit and the users affected
Teams whose actual question is “why did that crash, in which release, for how many people” — answered at a fraction of full APM pricing.
The catch: Errors and performance for the application, not the infrastructure: no host metrics, no cloud-service integration, no log platform. It is deliberately one job done well rather than a platform.
metered on transactions/spans within the same published Sentry tiers; developer-first tracing joined to the errors and the release that introduced them
Engineering teams that want tracing tied to the commit that caused the regression, without adopting a per-host platform.
The catch: Sampled by design and scoped to the application — it will not replace infrastructure monitoring, and deep cloud-service telemetry is outside it.
metered per replay within the Sentry tiers; a recording of the session that produced the error, so the reproduction step is watching rather than guessing
Teams losing time to bugs that cannot be reproduced from a stack trace alone.
The catch: Replays are a separate meter and privacy masking must be configured deliberately — an unmasked replay of a checkout page is a data-protection problem, not a debugging win.
quoted; a full-fidelity platform built on OpenTelemetry with no-sample tracing. Splunk does not publish list rates for this SKU — the per-host and per-session figures circulating online are third-party, not vendor list
Estates that want every trace kept rather than sampled, and that are already in the Cisco/Splunk relationship — a distinct SKU from Splunk Enterprise Security, which is carded on the Security category.
The catch: No published list pricing, so it cannot be compared like-for-like on a spreadsheet and the quote is the only real number. Full-fidelity tracing is a genuine differentiator and a genuine volume commitment.
quoted on the Splunk platform; service-level health modelling and AIOps event correlation over data the platform already holds — the layer that turns thousands of alerts into a handful of service states
Large estates already ingesting into Splunk whose problem is alert volume rather than missing telemetry.
The catch: It assumes the Splunk platform underneath and its value is proportional to how much you already ingest there. Service modelling is configuration work measured in weeks, not a switch.
quoted per monitored resource; SaaS-delivered infrastructure and application monitoring from the vendor whose depth is databases — the cloud front end to a Foglight estate
Estates already running Foglight for databases that want the same console over cloud infrastructure rather than a second platform.
The catch: Narrower application-tracing depth than the observability specialists; its strength is the database line, and buying it as a general APM is buying the wrong half of the portfolio.
quoted per monitored host or VM; virtual and cloud infrastructure monitoring with capacity planning — built for VMware-shaped estates and their migrations
Estates with a large virtualised footprint that need capacity forecasting alongside health, particularly through a hypervisor migration.
The catch: Infrastructure-scoped: no application tracing and no log platform. Its natural buyer is the virtualisation team, not the engineering team this page is written for.
about $0.28 per 1,000 events over the free allowance, Enterprise from roughly $25–30k a year; product analytics — what users did in the product, funnels and retention, not whether the server was healthy
Product and engineering teams asking which features are used and where people drop out — a different question from whether the system is up.
The catch: This is product analytics, not infrastructure observability: it will not tell you a service is down or a query is slow. Included here because the event meter and the buying team overlap, not because it substitutes for APM.
metered within the Mixpanel event tiers; replays joined to the analytics events, so a funnel drop-off can be watched rather than inferred
Teams that have the funnel data and still cannot explain the drop-off at one step.
The catch: Tied to the Mixpanel platform and to product questions rather than engineering ones. Privacy masking is a deliberate configuration, not a default.
distributed tracesRules out Datadog Infrastructure Monitoring, Datadog Log Management, Datadog Digital Experience Monitoring, Sentry Error Monitoring, Sentry Session Replay, Splunk ITSI, Quest Foglight Cloud, Quest Foglight Evolve, Mixpanel Product Analytics and Mixpanel Session Replay — no distributed tracing. That leaves Datadog APM, Sentry Performance Monitoring and Splunk Observability Cloud.
log searchRules out Datadog APM, Datadog Infrastructure Monitoring, Datadog Digital Experience Monitoring, Sentry Error Monitoring, Sentry Performance Monitoring, Sentry Session Replay, Quest Foglight Cloud, Quest Foglight Evolve, Mixpanel Product Analytics and Mixpanel Session Replay — not a log platform. That leaves Datadog Log Management, Splunk Observability Cloud and Splunk ITSI.
real user monitoringRules out Datadog APM, Datadog Infrastructure Monitoring, Datadog Log Management, Sentry Error Monitoring, Sentry Performance Monitoring, Splunk ITSI, Quest Foglight Cloud, Quest Foglight Evolve and Mixpanel Product Analytics — server-side only: it cannot see what the device experienced. That leaves Datadog Digital Experience Monitoring, Sentry Session Replay, Splunk Observability Cloud and Mixpanel Session Replay.
error trackingRules out Datadog APM, Datadog Infrastructure Monitoring, Datadog Log Management, Datadog Digital Experience Monitoring, Sentry Performance Monitoring, Sentry Session Replay, Splunk Observability Cloud, Splunk ITSI, Quest Foglight Cloud, Quest Foglight Evolve, Mixpanel Product Analytics and Mixpanel Session Replay — no dedicated error grouping with release and commit context. That leaves Sentry Error Monitoring.
OpenTelemetry-nativeRules out Datadog APM, Datadog Infrastructure Monitoring, Datadog Log Management, Datadog Digital Experience Monitoring, Sentry Error Monitoring, Sentry Performance Monitoring and Splunk ITSI — OpenTelemetry is supported but the vendor agent is the primary path; Sentry Session Replay, Quest Foglight Cloud, Quest Foglight Evolve, Mixpanel Product Analytics and Mixpanel Session Replay — OpenTelemetry support is not documented. That leaves Splunk Observability Cloud.
one platformRules out Sentry Error Monitoring, Sentry Performance Monitoring, Sentry Session Replay, Quest Foglight Cloud, Quest Foglight Evolve, Mixpanel Product Analytics and Mixpanel Session Replay — focused on one job rather than covering the estate. That leaves Datadog APM, Datadog Infrastructure Monitoring, Datadog Log Management, Datadog Digital Experience Monitoring, Splunk Observability Cloud and Splunk ITSI.
on-premisesRules out Datadog APM, Datadog Infrastructure Monitoring, Datadog Log Management, Datadog Digital Experience Monitoring, Sentry Session Replay, Splunk Observability Cloud, Quest Foglight Cloud, Mixpanel Product Analytics and Mixpanel Session Replay — SaaS only. That leaves Sentry Error Monitoring, Sentry Performance Monitoring, Splunk ITSI and Quest Foglight Evolve.
published pricingRules out Splunk Observability Cloud, Splunk ITSI, Quest Foglight Cloud and Quest Foglight Evolve — quote-only: no published list, so the quote is the only real number. That leaves Datadog APM, Datadog Infrastructure Monitoring, Datadog Log Management, Datadog Digital Experience Monitoring, Sentry Error Monitoring, Sentry Performance Monitoring, Sentry Session Replay, Mixpanel Product Analytics and Mixpanel Session Replay.
not per hostRules out Datadog APM and Datadog Infrastructure Monitoring — per host, so the bill tracks infrastructure and autoscaling moves it. That leaves Datadog Log Management, Datadog Digital Experience Monitoring, Sentry Error Monitoring, Sentry Performance Monitoring, Sentry Session Replay, Splunk Observability Cloud, Splunk ITSI, Quest Foglight Cloud, Quest Foglight Evolve, Mixpanel Product Analytics and Mixpanel Session Replay.
event meteringRules out Datadog APM, Datadog Infrastructure Monitoring, Datadog Log Management, Splunk Observability Cloud, Splunk ITSI, Quest Foglight Cloud and Quest Foglight Evolve — not event-metered. That leaves Datadog Digital Experience Monitoring, Sentry Error Monitoring, Sentry Performance Monitoring, Sentry Session Replay, Mixpanel Product Analytics and Mixpanel Session Replay.
India regionRules nothing out on published terms. It flags Datadog APM — An India data region for telemetry is not documented on the vendor's pages, Datadog Infrastructure Monitoring — An India data region for telemetry is not documented on the vendor's pages, Datadog Log Management — An India data region for telemetry is not documented on the vendor's pages, Datadog Digital Experience Monitoring — An India data region for telemetry is not documented on the vendor's pages, Sentry Error Monitoring — An India data region for telemetry is not documented on the vendor's pages, Sentry Performance Monitoring — An India data region for telemetry is not documented on the vendor's pages, Sentry Session Replay — An India data region for telemetry is not documented on the vendor's pages, Splunk Observability Cloud — An India data region for telemetry is not documented on the vendor's pages, Splunk ITSI — An India data region for telemetry is not documented on the vendor's pages, Quest Foglight Cloud — An India data region for telemetry is not documented on the vendor's pages, Quest Foglight Evolve — An India data region for telemetry is not documented on the vendor's pages, Mixpanel Product Analytics — An India data region for telemetry is not documented on the vendor's pages and Mixpanel Session Replay — An India data region for telemetry is not documented on the vendor's pages — marked on the cards, not removed.
The meter decides the bill, not the rateA per-host quote and a per-GB quote for the same estate can differ by an order of magnitude. Datadog alone runs three meters — per host for APM and infrastructure, per GB ingested plus per million events indexed for logs, per session for RUM. Model your own volumes against each meter before comparing any two vendors.
Splunk does not publish list pricing hereThe per-host and per-session figures circulating online for Splunk Observability are third-party, not vendor list. We say so rather than repeating them. The quote is the only real number, and it is a volume commitment.
Cardinality is the line nobody forecastsAdd a user ID or request ID as a metric tag and one metric becomes millions of time series. It is the most common cause of a bill that multiplies while traffic does not, and no vendor stops you doing it.
Open source is genuinely viable herePrometheus for metrics, Grafana for dashboards, OpenTelemetry for instrumentation, Loki or Elastic for logs. Unlike most categories on this site, the open-source path is credible for real production estates. It costs engineering time instead of licence — typically an owner, not a side project — and that trade is worth making explicitly rather than by default.
If one of these is your sentence, the shortlist is short.
Why: Exceptions grouped by fingerprint with the release, the commit and the affected users — event-metered, published from about $26 a month.
The trade-off: No infrastructure monitoring and no log platform. If you also need host health, this is half the answer.
Why: Distributed tracing is the only signal that attributes time per hop across service boundaries.
The trade-off: Datadog meters per host and compounds with its other modules; Splunk keeps every trace but publishes no list price.
Why: Move from a host or ingest meter to an event meter, and cut what you index rather than what you collect.
The trade-off: Event-metered tools are narrower. The real fix is usually cardinality and log indexing discipline, not a new vendor.
Why: Only real user monitoring sees the device, the network and the third-party scripts the backend never touches.
The trade-off: Session-metered, so a traffic spike is a bill spike with no infrastructure change to explain it.
Why: OpenTelemetry-native ingest keeps the instrumentation portable, so changing vendor does not mean reinstrumenting.
The trade-off: Quote-only, and full-fidelity tracing is a real volume commitment.
Why: All three offer a self-hosted path where residency or policy forbids SaaS telemetry.
The trade-off: You own the upgrades, the storage and the scaling — the operational burden the SaaS price was covering.
Why: Service-level health modelling collapses alert volume into a handful of service states.
The trade-off: Assumes the Splunk platform underneath, and the service modelling is weeks of configuration.
Why: Genuinely credible for real production estates in this category, unlike most others on this site. Named here because pretending otherwise would waste your time.
The trade-off: It costs an owner rather than a licence. Budget the person, the storage and the upgrade path — and revisit when that person leaves.
Every other category on this site prices per user, per device or per instance — numbers that move when you hire, buy hardware or open an office. You can forecast those. Observability prices on how much your software is used, and that number moves when a marketing campaign works.
The same estate on three different meters:
Model all three at your current volume, at double, and at ten times. The vendor that wins at today’s volume frequently loses badly at ten times, and the contract you sign is usually multi-year.
Ask before signature
Scale here means telemetry volume, not team size — which is the whole point of the category.
One application, modest traffic
Put this in your PoC
Price the event meter before looking at any platform.
Several services, real traffic
Put this in your PoC
Decide what gets indexed versus merely collected.
Microservices at scale
Put this in your PoC
Model the bill at 2× and 10× before signing multi-year.
Very high volume
Put this in your PoC
Name the person who owns cardinality. Nobody does until the bill arrives.
Where a vendor does not publish list pricing, this page says so rather than repeating a third-party figure.
Instrumentation is the lock-in, not the data.
Instrumentation
Vendor agents mean reinstrumenting; OpenTelemetry means changing an endpoint
Dashboards
Rebuilt in the new tool every time — no interchange format exists
Historical telemetry
Exportable in principle, rarely worth the cost of moving
Alert rules and on-call routing
Re-authored, and the tuning that made them quiet is re-learned
The practical consequence: instrument with OpenTelemetry from the start if you expect to change vendor, even if you use a vendor agent today. It is the single cheapest insurance in this category.
By meter, in USD and INR, modelled at more than one volume.
Four checks, in the order most likely to return a yes.
This is the one category on the site where the open-source answer is genuinely competitive for production estates. Saying otherwise would cost you credibility with the engineer reading this.
Three meters, and the published rates that exist.
TechBag quotes every one of these in INR with GST, and models your actual volumes against each meter rather than comparing rates. Where a vendor publishes no list price, this page says so instead of repeating a third-party figure.
TechBag gives INR pricing, GST, PO cycle, minimums and tier-matched quotes. The INR above is conversion for scale at ≈₹83/$; the tier-matched INR quote is ours.
The largest hidden line. The licence covers a tranche; growth and retention are billed on top.
Frequently metered separately. One tag can multiply a metric into millions of series.
Not a licence, but real and recurring. Budget an owner, storage and an upgrade path.
Unless you instrumented with OpenTelemetry, changing vendor means changing every service.
Five ways this purchase goes wrong. Every one of them is a cost surprise rather than a capability gap.
The second-year bill after traffic grew
The forecast was built on today's volume against a meter that tracks usage. Model at 2× and 10× before signing a multi-year contract.
A cardinality explosion from one well-meant tag
Someone adds a user ID to a metric label. One metric becomes millions of time series and the bill multiplies with no traffic change.
Retention shortened to control cost
The reflex works until the incident that matters falls outside the window, and the post-mortem has no data.
Instrumenting everything, alerting on nothing actionable
Full telemetry and an on-call rota that ignores the pages. The tool is not the problem; nobody owns alert quality.
Buying APM when the requirement was error tracking
Ten times the price for a question that error tracking answers better. Read the boundary section above before shortlisting.
Vendor-neutral. No gated content. · Last reviewed