Any cloud · Any Kubernetes · Any host

Exact percentiles, without choosing buckets first.

AirMonitor takes the OpenTelemetry stream you already export and answers the Prometheus queries you already run, with exact percentiles instead of bucket estimates.

See the numbers →
p99 latency per minute · one routesame requests
true valueAirMonitorPrometheus classic bucketsPrometheus native histogram
0 ms250 ms500 ms750 ms1,000 ms0 min2 min4 min6 min8 min10 min
AirMonitor sits on the true value. Worst minute: classic buckets 38 % off, native histograms 10.2 % off.

0.01 % error

Worst p99 against the true value. Classic buckets: 38 %.

11× fewer series than classic buckets

7,306 against 78,567 for the same stream. A series is one tracked line, a metric plus its labels. A classic histogram keeps a line per bucket edge, plus a sum and a count: fourteen lines for each one here. AirMonitor keeps that distribution in one object. Native histograms already use one series; there the difference is accuracy, not the count.

6.6× faster

A p99-by-route panel: 77 ms against 507 ms.

One exporter

Added to your Collector, where you already export OpenTelemetry. No change to those services. A different instrumentation path, we can add.

The problem

A histogram makes you choose the buckets before you know the question.

A histogram's buckets are chosen when the code is written, and every answer after that is an estimate between two of them. AirMonitor keeps the measurements whole.

The usual histogram path
Estimate

Buckets first. An estimate later.

SDKbuckets chosen
→
Scrapecounts per bucket
→
Seriesone per bucket
→
Queryinterpolate
Edges guessedFourteen seriesBetween two edges
The AirMonitor path
Exact

Store the measurement. Ask when you know.

Spanas emitted
→
Collectorone exporter
→
AirTreeper minute
→
Queryread it off
No bucketsOne objectAny quantile
Beyond a histogram

Latency, size, and status — kept on the same request.

  • 2DLatency by response size: do bigger responses take longer?
  • 3DAdd the status: do failed requests take longer?
  • 4DAdd a fourth, such as database time.

See it sliced →

Every request, by response size and latency

1 KB10 KB100 KB10 ms100 ms1 sresponse size (log scale)latency (log scale)count · log scale1101001k15,988empty
An example service: 250,000 requests in one stored object.
What you get

The dashboards you have, with answers you can stand behind.

Percentiles

Exact quantiles

Any quantile, over any window.

What you can ask →

PromQL

Your queries, unchanged

Grafana adds a datasource and nothing else.

Alerting

Your rules, unchanged

Evaluated inside, pushed to your Alertmanager.

Joint distributions

Sliced after the fact

Latency by size by status, from the object already stored. No new metric.

AI assistants

Ask in plain English

A built-in, read-only MCP server.

View MCP →

Operations

Runs anywhere

In your account, or hosted by AirMettle.

Run anywhere →

Early access

See it on your telemetry.

We run AirMonitor beside your Prometheus and show you the same panels, side by side.

How it works

A drop-in beside the Collector and Grafana you have.

One daemon between the Collector you have and the Grafana you have. It files every value into an AirTree, a compact structure that holds a whole distribution exactly, and answers your queries from it.

Your services spans · counters · gauges unchanged OpenTelemetry Collector + one exporter queues and retries AirMonitor daemon one AirTree per series per minute exact distributions · counters · gauges rolled up for long retention Grafana a Prometheus datasource Alertmanager your alert rules
Services, Collector, Grafana and Alertmanager stay as they areOne binary · one directory · one port
The idea

Store the measurement. Choose the question later.

No extra instrumentation

The span you already emit is the measurement

It carries the latency, so your application records nothing more.

Decided at query time

Any quantile, any window

A p99.9 nobody planned for is answered from what is already stored.

Merging loses nothing

Long retention, same answers

Minutes roll up into hours and days and still answer the same questions.

What you can ask

The Prometheus queries you have, answered from the stored measurement.

Dashboards and alerting rules written for Prometheus run as they are.

histogram_quantile

Exact percentiles

The real quantile of the stored distribution, not an estimate between bucket edges.

rate · increase

Rates that add up

No extrapolation: 0.7 % off at worst, against 26 % for Prometheus.

gauges

Statistics over time

Averages, minimums, maximums, deltas and predictions.

rules

Alerting and recording rules

Your Prometheus rule files, pushed to your Alertmanager.

slices

Questions you did not plan

One measurement sliced by another, with no new metric.

lint

Checked before you switch

One command confirms a dashboard or a rule file will run.

PromQL coverage

The stable PromQL language is covered in full, subqueries and the @ modifier included.

Beyond a histogram

Latency, size, and status — kept on the same request.

A histogram holds one measurement. An AirTree holds up to four of the same request in one object, so any of them can be asked against any other, later.

  • 2DLatency by response size: do bigger responses take longer?
  • 3DAdd the status: do failed requests take longer?
  • 4DAdd a fourth: any number your spans carry, such as database time.

Every request, by response size and latency

1 KB10 KB100 KB10 ms100 ms1 sresponse size (log scale)latency (log scale)count · log scale1101001k15,988empty
An example service: 250,000 requests in one stored object. Bigger responses take longer, and one band is slow whatever the size.

The same object, by outcome

SucceededFailed1014.8223248681001522163204807041024153623043072latency, ms (log scale)count · log scale1101001k10k60,150empty
Sliced by status: the failed requests are the ones that took longest. Violet is fewest, red is most.
AirMonitor MCP

Ask in plain English. Read-only. The number comes from the store.

A built-in MCP server lets your AI assistant question what AirMonitor holds and get the exact number back.

Works with Claude Code, Cursor, Copilot and any other MCP-capable assistant.

Memory on the worker pods doubled since yesterday. Why?
what_moved · now, against a day earliermemory shifted ×2.1, and a new version arrived
shift · that metric, by versionthe new version carries the change
quantile_by · the new version's requestswhere its p99 latency sits
quantile_by · the old version, a day backthe same picture before the rollout
Four calls, every number exact.
You ask in plain English AI tool Claude Code, Cursor, Copilot… MCP airmonitord mcp read-only tools AirMonitor daemon Your store exactanswers
How it works

Your AI tool, your daemon, your data.

  • Local or remoteLaunched by your AI tool beside the daemon, or reached over HTTPS.
  • Read-onlyAn assistant can look, but cannot change anything.
  • Sized for a chatAnswers come back short enough for a conversation.
What the assistant can do

More than forty tools, built for the questions an incident asks.

Orient

What is here, and how busy

overviewlist_metricscoverage
Ask

Exact numbers

quantilesoutliersquery
Find what changed

What moved, and when

what_movedshiftwhen_did_it_change
Releases and SLOs

Is the rollout safe

release_gateregression_sinceslo_status
Alerts and rules

Check before you save

alertsbacktest_rulecheck_query
Capacity and reports

The Monday numbers

time_tocapacity_reportweekly_summary
Connect

One line in your AI tool.

Built-in investigations, from what went wrong to a capacity check, show up as commands.

# Claude Code, beside the daemon
claude mcp add airmonitor -- \
  airmonitord mcp --url http://localhost:4320
Measured against Prometheus

Same requests. Same stream. Both systems.

AirMonitor and Prometheus received identical telemetry from the same OpenTelemetry Collector. Every figure here is read from those runs.

0.01 %
worst p99 error. Prometheus classic buckets: 38 %
6.6×
faster on a p99-by-route panel: 77 ms against 507 ms
11×
fewer series than classic buckets for the same stream: 7,306 against 78,567
1.9×
smaller store on a compressing volume: 23.6 MB against 45.4 MB
1 · Accuracy

The reported percentile against the true one, minute by minute.

The same requests, read three ways: AirMonitor, classic buckets and native histograms.

Error of the reported percentile

AirMonitorPrometheus classicPrometheus native
0.0 %6.2 %12.5 %18.8 %25.0 %p50p95p990.0 %0.0 %0.0 %8.6 %20.7 %15.0 %0.5 %1.2 %1.7 %
Average error across the run, at p50, p95 and p99.

p99 per minute, one route

true valueAirMonitorPrometheus classicPrometheus native
0 ms250 ms500 ms750 ms1,000 ms0 min2 min4 min6 min8 min10 min
AirMonitor sits on the true value. Worst minute: classic buckets 38 % off, native histograms 10.2 % off.
2 · Footprint

At a thousand requests a second.

A service with 1,000 routes, both systems on the same stream.

7,306
series held, against 78,567 for classic buckets on the same stream. One line per bucket edge is what makes the larger number.
203 MB
of memory. Prometheus: 340 MB

What one observation costs your application, nanoseconds

OTel SDK, explicit-bucket histogram1,479 nsOTel SDK, exponential histogram2,220 nsPrometheus client library histogram419 nsAirTree on the daemon5 ns
A Prometheus histogram is computed inside your application on every observation. With AirMonitor the span you already emit carries the value.
Early access

Hosted, or in your account

AirMonitor is onboarding early customers.

Talk to us

See these panels on your telemetry

We run it beside your Prometheus and show you both, side by side.

Run anywhere

Your account, your datacenter, or ours.

The same AirMonitor runs in all three.

In your account

Any cloud, any Kubernetes

A Helm chart, a .deb or a container image. Your telemetry never leaves your account.

Hosted by AirMettle

One address for your Collector

Point your Collector at the endpoint and add a Grafana datasource.

Highly available

A pair for availability

Two daemons hold the same data, and queries fail over.

Operating it

One binary, one directory, one port.

01

Nothing beside it

No database to run alongside.

02

Health built in

It reports its own health and diagnoses itself.

03

Checked before the switch

Dashboards and rules are verified before you point anything at it.

Early access

Hosted, or in your account

AirMonitor is onboarding early customers.

Talk to us

See it on your telemetry

We run it beside your Prometheus and show you both, side by side.

Security & trust

In your account, nothing leaves. Hosted, no shared store.

ID

Authenticated access

Only callers you authorise can send or query, and credentials rotate without a restart.

TLS

Encrypted in transit

HTTPS throughout, with your own certificate when you run it.

ISO

Tenant isolation

No shared store and no cross-tenant queries.

AI

Read-only for assistants

An AI assistant can look, but cannot change anything.

FAQ

Questions, answered.

Is this a Prometheus replacement?

On the query side, yes. It serves the Prometheus API and the stable PromQL language in full, so Grafana, alert rules and Alertmanager work unchanged. It can also run beside Prometheus while you compare. Where you need a different exporter or query path, we can work on that with you.

Do I have to change my instrumentation?

Not if you already export OpenTelemetry to a Collector. Spans give it latency distributions; counters and gauges come from the metrics you already export. If the result is what you want and your stack is different, we can add that integration.

Can an AI assistant use it?

Yes. A built-in MCP server lets Claude Code, Cursor, Copilot or any MCP-capable assistant ask in plain English. It is read-only.

What does it cost on disk?

About half of what Prometheus needs for the same stream on a compressing volume: 23.6 MB against 45.4 MB in our thousand-route run.

What does fewer series mean?

A series is one line Prometheus stores and queries: the metric name plus labels such as route. Classic histograms store a separate line for each bucket edge, plus a sum and a count, so the same traffic fanned out to 78,567 lines against 7,306 here. Fewer lines means a smaller index and fewer lines for a query to touch. It does not mean the percentile is stored as less data — disk is a separate comparison. Native histograms already avoid that fan-out; compare those on accuracy.

Which clouds does it run on?

Any. A Helm chart for Kubernetes, a .deb or a container image for any host, or hosted by AirMettle.

How is availability handled?

Two daemons receive the same stream and hold the same data. Queries fail over, and a daemon that was down catches up from its peer.

How exact is exact?

Percentiles stayed within 0.01 % of the true value in the worst minute of our runs.

Is it open source?

The AirTree engine is source-available on GitHub. The daemon and the hosted service are in early access.

Your address is used for this request and never shared.