7 Prometheus Alternatives in 2026: An Honest Comparison

#Prometheus#PromQL#alternatives#comparison#observability#metrics#VictoriaMetrics#Grafana Mimir#Thanos#GreptimeDB#Grafana#monitoring#time-series#time-series database#Arc
Cover image for 7 Prometheus Alternatives in 2026: An Honest Comparison

Most people who search for a "Prometheus alternative" do not actually want to stop using Prometheus. They want to stop losing data after fifteen days. They want two replicas that do not disagree, one query that spans every cluster, and a TSDB that does not run out of memory because someone added a user_id label on a Friday afternoon.

Those are storage problems, and most of the products sold as Prometheus alternatives solve exactly those problems while keeping the parts of Prometheus you like: the scrape model, the exporters, PromQL, your Grafana dashboards, and your alerting rules. A smaller group of options asks you to give up PromQL in exchange for something else. That is a much bigger decision, and it should be made on purpose rather than by accident.

So this post is organized around that split. The first option is staying on Prometheus and fixing it in place. The next four keep PromQL. The last two do not.

I should be upfront about my bias. I worked at InfluxData. I now run Basekick Labs, where we build Arc, which is the seventh option on this list. Arc does not speak PromQL and does not accept Prometheus remote write today, so for most readers of this post it is not the answer, and I will say plainly when it is and when it is not. A comparison that never recommends a competitor is an ad.

Prometheus does four jobs, and most alternatives replace one

Before comparing anything, separate the jobs a Prometheus server does:

  1. Collection. It discovers targets and scrapes /metrics endpoints on an interval.
  2. Storage. It writes samples to a local TSDB on disk.
  3. Query. It evaluates PromQL for dashboards and ad hoc queries.
  4. Rules. It evaluates recording and alerting rules and sends alerts to Alertmanager.

Almost every product below replaces job 2 and keeps jobs 1, 3, and 4 intact or nearly intact. Prometheus (or an agent such as vmagent, Grafana Alloy, or the OpenTelemetry Collector) keeps scraping, then forwards samples with remote write to a backend that stores them durably and answers PromQL. That is why migrations to VictoriaMetrics or Mimir are measured in days, while migrations away from PromQL are measured in quarters.

Keep this in mind while reading. When a vendor says "drop-in Prometheus replacement", ask which of the four jobs they mean.

Why teams look past Prometheus

Prometheus is not a bad product. It is a CNCF graduated project, it is the default metrics system for Kubernetes, and it is shipping steadily: v3.15.0 landed on September 25, 2026, roughly six weeks after 3.14. The pressure comes from the edges of its design.

Local storage is a single node by design. The storage documentation is refreshingly direct about it: "A limitation of local storage is that it is not clustered or replicated. Thus, it is not arbitrarily scalable or durable in the face of drive or node outages and should be managed like any other single node database." Long retention on one disk works until the disk, the node, or the restore window becomes the problem.

High availability means duplicates. The standard HA pattern is two identical Prometheus servers scraping the same targets. Each has slightly different data, and neither knows about the other. Something has to deduplicate at query time, which is a job Prometheus itself does not do.

There is no global view. Each server sees its own targets. Federation exists, but the federation docs describe it for pulling selected series and keeping "only aggregated data" on global servers, not for merging every cluster's raw samples into one queryable place.

Cardinality has a ceiling, and the docs tell you to stay under it. Prometheus's own instrumentation guidance says "each labelset is an additional time series that has RAM, CPU, disk, and network costs," advises keeping metric cardinality below 10, and says that above 100 you should "investigate alternate solutions." The naming guide is blunter: "Do not use labels to store dimensions with high cardinality." If your actual questions are per customer, per device, or per request, Prometheus is telling you it is the wrong tool for that part of the workload.

Push is a second-class path. Prometheus 3 can receive OTLP and remote write, but the API docs warn that OTLP ingestion "is not considered an efficient way of ingesting samples," should be used "for specific low-volume use cases," and "is not suitable for replacing the ingestion via scraping." If your organization is standardizing on OpenTelemetry push, that warning matters.

To be fair, because this is the part most comparison posts skip: Prometheus 3 fixed a lot. If your mental model dates from 2.x, it is out of date.

What changed in Prometheus 3

Prometheus 3.0 shipped on November 14, 2024, the first major release in seven years. Since then:

  • Native histograms are stable. They became "stable, but optional" in 3.8.0 and the old feature flag became a no-op in 3.9.0 (January 2026). You enable scraping them with scrape_native_histograms.
  • UTF-8 metric and label names are on by default since 3.0, which removes a lot of friction with OpenTelemetry naming.
  • Agent mode is stable. Run prometheus --agent and you get a scrape-and-forward process with no local query or long-term storage, which is the right shape for the "collect locally, store centrally" architectures below.
  • OTLP ingestion is enabled with --web.enable-otlp-receiver at /api/v1/otlp/v1/metrics, with the efficiency caveat above.
  • Remote Write 2.0 is accepted by the receiver, but the 2.0 spec itself is still marked experimental, and support across backends is uneven. Mimir labels it experimental; VictoriaMetrics does not accept it yet.
  • 3.15 added scraping over Unix sockets, an OpenMetrics 2.0 scrape format, and a stable XOR2 chunk encoding (the release notes warn to check that tools reading the TSDB directly, such as the Thanos sidecar, support it before you upgrade).

None of this changes the single-node storage model. It does make Prometheus a much better collector for whichever backend you pick.

What to look for in a Prometheus alternative

Name the criteria before the list, and weight them for your own workload:

  1. PromQL compatibility. How much of your dashboard and alert corpus runs unmodified. Vendors publish percentages; your own queries are the only number that counts.
  2. Remote write support. Version 1 is universal. Version 2 is not, yet.
  3. Operational model. One binary, a handful of stateless and stateful services, or a cluster that also needs Kafka and object storage.
  4. Storage portability. Proprietary blocks, Prometheus TSDB blocks in a bucket, or open formats other tools can read.
  5. Cardinality behavior. What happens to memory and cost when series count grows 10x.
  6. License. Apache 2.0, AGPL-3.0, or proprietary managed service. Ask legal before the POC.
  7. Cost model. Infrastructure you run, or per-sample and per-series billing you do not control.

The alternatives

1. Stay on Prometheus 3

The option worth pricing first. If your problem is retention on one cluster or a few unruly metrics, you may not need a new system.

What it's best at. Zero migration. You keep everything: exporters, service discovery, PromQL, recording rules, Alertmanager, dashboards. Many teams can extend retention simply by giving Prometheus a bigger disk; the storage docs say samples average "only 1-2 bytes" and give the sizing formula (retention_time_seconds * ingested_samples_per_second * bytes_per_sample). Cardinality problems are often better solved with metric_relabel_configs that drop the offending label than with a new database. For HA, run a pair and put a deduplicating query layer in front only when you need it.

Where it falls short. Everything in the "why teams look past" section still applies: single-node durability, no global view, duplicate HA pairs, and a hard cardinality ceiling. Bigger disks also mean longer restarts while the WAL replays.

License. Apache 2.0.

Migration path. None.

Best for. Single-cluster setups whose pain is retention measured in weeks, not years, or a few high-cardinality metrics that should be relabeled away.

2. VictoriaMetrics

VictoriaMetrics is a Prometheus-compatible time-series database written in Go, available as a single binary or as a cluster.

What it's best at. Operational simplicity and resource efficiency. The single-node version is genuinely one binary, and the docs recommend it for ingestion "lower than a million data points per second," which covers most companies. It accepts Prometheus remote write, answers the Prometheus query API, and works with existing Grafana dashboards. Its query language, MetricsQL, is described as "backwards-compatible with PromQL" with a few documented, intentional differences in rate, increase, and NaN handling. The surrounding toolkit is complete: vmagent for scraping and forwarding, vmalert for Prometheus-format rules that send to Alertmanager, and vmctl for migration. The project claims up to 7x less RAM and storage than Prometheus, Thanos, or Cortex; that is a vendor number, so measure it on your own data. Releases are frequent (v1.153.0 shipped September 28, 2026) and there are two supported LTS lines at any time.

Where it falls short. It does not accept Remote Write 2.0 yet. Downsampling and multiple retention periods, two features people often assume are included, are Enterprise-only, along with several security and integration features. The cluster version (vminsert, vmselect, vmstorage) is still simpler than Mimir, but it is not a single binary. Logs and traces live in separate products: VictoriaLogs is GA at 1.x, while VictoriaTraces is still pre-1.0 (v0.12.0).

License. Apache 2.0, with a commercial Enterprise edition.

Migration path. Point Prometheus remote write at VictoriaMetrics, or replace Prometheus with vmagent. Backfill history with vmctl from a Prometheus snapshot or over remote read. Dashboards and alerts usually carry over unchanged; spot-check queries that depend on exact rate() edge behavior.

Best for. The default answer for most teams who want Prometheus with long retention and lower resource use, without running a large distributed system.

3. Grafana Mimir

Grafana Mimir is a horizontally scalable, multi-tenant Prometheus backend. Grafana forked it from Cortex in 2022.

What it's best at. Very large scale and multi-tenancy. Mimir 3.0 (October 2025) made the ingest storage architecture stable and preferred: writes go through Kafka or a Kafka-compatible system, which decouples the read path from the write path so heavy queries do not slow ingestion. The Mimir Query Engine is now the default, and Grafana claims up to 92% lower peak memory than the Prometheus engine. It accepts OTLP natively, supports native histograms, and accepts Remote Write 2.0 (labeled experimental). The current release is 3.2.1 (September 10, 2026). If you already run Grafana, Loki, and Tempo, Mimir fits the same operational vocabulary.

Where it falls short. This is the heaviest option on the list to self-host. The microservices mode involves distributors, ingesters, queriers, query-frontends, store-gateways, and compactors, and the preferred architecture now adds Kafka on top of object storage. The classic architecture still works but is slated for deprecation. A monolithic mode exists for smaller deployments, but if you are small enough for that, VictoriaMetrics single-node is usually less work.

License. AGPL-3.0. Fine for most internal use; check with legal if you modify it and offer it as a service.

Migration path. Remote write from existing Prometheus servers or Alloy. Historical TSDB blocks can be uploaded to Mimir's object storage. PromQL, dashboards, and rules carry over.

Best for. Platform teams running metrics as an internal multi-tenant service at very high series counts, with the people to operate Kafka and a microservices deployment.

4. Thanos

Thanos is a CNCF incubating project that adds a global query view, long-term object storage, and downsampling on top of existing Prometheus servers.

What it's best at. Incremental adoption. In sidecar mode, a Thanos sidecar runs next to each Prometheus server, uploads its TSDB blocks to object storage, and serves recent data to a central Querier. Nothing about how Prometheus scrapes changes. The Querier deduplicates HA pairs and merges every cluster into one PromQL endpoint. The Store Gateway serves historical blocks from the bucket, and the Compactor compacts, downsamples, and applies retention. If you would rather push, Thanos Receive accepts remote write instead. Data sits in object storage in Prometheus's own block format, which is as portable as metrics storage gets within the PromQL ecosystem.

Where it falls short. Many components, each with its own scaling and caching story. Query performance over long ranges depends heavily on compaction, downsampling, and caching being tuned well. The project is still pre-1.0 (v0.42.4, July 2026) and its minor releases come every four to eight months; it is actively maintained, but the cadence is slower than VictoriaMetrics or Mimir. Watch the interaction with new Prometheus storage features: the 3.15 notes specifically call out checking the sidecar's XOR2 support.

License. Apache 2.0.

Migration path. The gentlest PromQL-preserving path on this list. Add sidecars to the Prometheus servers you already run, point them at a bucket, and deploy a Querier.

Best for. Teams with several Prometheus servers who want a global view and cheap long-term storage without changing how they collect metrics.

5. Managed Prometheus

If the problem is that you do not want to operate any of the above, every major cloud and Grafana Labs will run a PromQL-compatible backend for you. You keep Prometheus or an agent for collection and remote write to the service.

The options.

  • Amazon Managed Service for Prometheus bills per sample ingested (US East: $0.90 per 10M samples for the first 2B, $0.35 per 10M for the next 250B), plus $0.03 per GB-month of storage and $0.10 per billion query samples processed.
  • Google Cloud Managed Service for Prometheus runs on Monarch, Google's internal monitoring system, bills $0.06 per million samples for the first 50B and $0.048 per million for the next 200B, and stores data for 24 months at no extra cost.
  • Azure Monitor managed service for Prometheus bills per sample ingested and processed and includes 18 months of retention.
  • Grafana Cloud bills per active series: the free tier covers 10,000 series with 14 days of retention, and Pro is $19 per month plus usage from $6.50 per 1,000 series with 13 months of retention.

What it's best at. No storage to operate, no upgrades, and retention measured in months or years by default.

Where it falls short. Cost scales with your series count and scrape interval, not with your hardware. As a rough illustration, 1 million active series scraped every 15 seconds is about 172.8 billion samples a month. At list prices that is roughly $6,200 a month for ingestion alone on Amazon's service, and roughly $8,900 on Google's, before queries and storage. Moving to a 60-second scrape interval cuts those numbers by roughly three to four times. The cost lever is always the same: fewer series and fewer samples.

License. Proprietary managed services.

Migration path. Remote write from what you run today. PromQL and dashboards carry over.

Best for. Teams that want to stop operating metrics storage and whose series count is predictable enough to budget for.

6. GreptimeDB

GreptimeDB is an open-source observability database that stores metrics, logs, and traces in one engine on object storage, and speaks both SQL and PromQL.

What it's best at. It is the only option on the main list that keeps PromQL and adds SQL over the same data. It accepts Prometheus remote write and remote read, OTLP, Loki push, Elasticsearch bulk, and InfluxDB line protocol. The docs claim over 90% PromQL support (a May 2026 blog post claims close to 100% of the official compliance suite; use the conservative number until you test). Version 1.0 went GA in April 2026, v1.2 added Remote Write 2.0 and native histogram ingestion, and it runs either standalone or as a distributed cluster with object storage as primary storage.

Where it falls short. It is younger than everything above it on this list, with a smaller operator community and fewer production war stories to learn from. Native histogram querying is still arriving in 1.3. Some PromQL features, including the @ modifier, are not supported. Enterprise features sit under a separate commercial license.

License. Apache 2.0, with a commercial Enterprise edition.

Migration path. Remote write from Prometheus, then point Grafana's Prometheus data source at GreptimeDB's PromQL endpoint. Verify your specific dashboards; then add SQL where PromQL was the wrong tool.

Best for. Teams who want to keep PromQL dashboards working while consolidating logs and traces into the same system and adding SQL.

7. Arc

Arc is an open, SQL-native time-series database for telemetry you need to keep. I built it after leaving InfluxData. It stores data as Apache Parquet on S3, Azure Blob, any S3-compatible store, or local disk, and you query it with PostgreSQL-compatible SQL.

Arc is not a drop-in Prometheus replacement, and that should be the first thing you know about it. Arc does not accept Prometheus remote write and does not speak PromQL. Your Grafana panels written in PromQL do not carry over, and neither do your Prometheus alerting rules. Arc's own README tells PromQL-dependent teams to evaluate VictoriaMetrics, and I agree with it.

I am including Arc because there is a specific kind of team that searches for "Prometheus alternative" and is really looking for something else.

What it's best at. Metrics that need to live next to everything else. If your real questions join request metrics with deploy events, customer tiers, or device registries, SQL with window functions, CTEs, and joins answers them directly, and Arc keeps metrics, logs, traces, and events in one place. Cardinality behaves differently: a label value is a value in a Parquet column, not a new in-memory series, so a per-customer or per-device dimension does not run into a series limit. It still costs something in storage and query time, but it does not take the database down. Retention is a storage bill rather than a RAM bill, because data sits as compressed Parquet in object storage you own, readable by Spark, Polars, ClickHouse, or any Parquet reader without an export step. Ingestion is fast: 34M records/sec sustained in our own benchmark over the MessagePack columnar protocol; expect lower end-to-end rates through a Telegraf scrape pipeline. Deployment is a single Go binary.

Where it falls short. PromQL is a better language than SQL for a lot of metrics work. rate() handles counter resets for you; in SQL you write a window function with LAG and handle resets yourself. Recording rules become continuous queries, and alert rules move to Grafana alerting with SQL queries. The exporter ecosystem still works, but through a collector (below) rather than Prometheus itself. Arc is newer than everything above it except GreptimeDB, so the community is smaller. Clustering, RBAC, and tiered storage are in Arc Enterprise rather than the open-source build.

License. AGPL-3.0, with a commercial license through Arc Enterprise.

Migration path. Keep your exporters. Replace the Prometheus scrape loop with Telegraf's inputs.prometheus plugin, which scrapes the same /metrics endpoints and supports Kubernetes pod discovery through annotations, and send the data to Arc with the Telegraf outputs.arc plugin. The Arc self-monitoring tutorial shows that exact config against Prometheus endpoints, and the Kubernetes monitoring walkthrough covers Telegraf, Arc, Grafana, and SQL alerting end to end. Dashboards are rewritten in SQL using the Grafana data source. Run both systems side by side until the SQL versions of your critical panels and alerts match.

Best for. Teams whose metrics are one part of a broader telemetry workload, who need high-cardinality dimensions or multi-year full-resolution history, and who would rather write SQL than PromQL. Not for teams whose main asset is a large PromQL dashboard and alert corpus.

Also worth knowing

  • Cortex. The CNCF project Mimir was forked from is very much alive: v1.21 shipped in April 2026 with 29 contributors and a new Parquet mode for its store gateway, and 1.22 is in release candidates. It is Apache 2.0, which matters if AGPL rules out Mimir for you.
  • ClickHouse. ClickHouse now has a TimeSeries table engine, Prometheus remote write and read endpoints, and a PromQL dialect that the company says covers over 85% of the language. As of ClickHouse 26.9 all of it is "private preview", and PromQL in ClickHouse Cloud is waitlist-only. Worth watching, not yet something to migrate a production Prometheus onto. ClickStack, its OpenTelemetry-based observability stack, takes metrics through OTel rather than PromQL.
  • Elasticsearch 9.5. Since August 2026 Elasticsearch accepts Prometheus remote write v1 and runs PromQL natively, with Elastic reporting that about 80% of a real-world query corpus runs unmodified. Remote Write 2.0 is not supported yet. I wrote a longer analysis of how Elastic built it. Most interesting if your logs already live in Elasticsearch; see also Arc vs Elasticsearch.
  • Cortex XCOR (formerly Chronosphere). Palo Alto Networks closed its Chronosphere acquisition in January 2026, and on October 1 relaunched the product as Cortex XCOR. Not to be confused with the CNCF Cortex project above. Its strength remains cost control, with aggregation and drop rules applied before data is persisted. Pricing is not public.
  • OpenObserve and SigNoz. Both are full observability platforms that accept Prometheus metrics. OpenObserve (AGPL-3.0, 1.0 in September 2026) takes remote write and supports PromQL; SigNoz supports PromQL and steers metrics ingestion through the OpenTelemetry Collector.
  • M3. Uber's metrics platform shipped v1.6.0 in September 2026, its first release since April 2022. Encouraging, but one release after a four-and-a-half-year gap is not yet a trend.

The comparison

Prometheus 3VictoriaMetricsMimirThanosManagedGreptimeDBArc
Query languagePromQLMetricsQLPromQLPromQLPromQLPromQL + SQLSQL
Remote write inv1, v2 (exp.)v1v1, v2 (exp.)v1 (Receive)v1v1, v2 (exp.)No
Long-term storageLocal diskLocal diskObject store + KafkaObject storeVendorObject storeParquet, S3 or disk
Global query viewNoClusterYesYesYesClusterEnterprise
High cardinalityLimitedGoodGoodAs PrometheusBilled by usageGoodNo series limit
Ops weightLowLowHighMedium-highNoneLow-mediumLow
LicenseApache 2.0Apache 2.0 + Ent.AGPL-3.0Apache 2.0ProprietaryApache 2.0 + Ent.AGPL-3.0 + comm.
Dashboards carry overYesYesYesYesYesMostlyNo (SQL)

MetricsQL is backwards-compatible with PromQL, with a few documented differences. "(exp.)" marks Remote Write 2.0, whose spec is still experimental. "Ent." and "comm." mark commercial editions.

How to choose

  1. Is your pain retention on one cluster? Stay on Prometheus 3, give it disk, and relabel away the worst cardinality offenders. It is the cheapest experiment.
  2. Do you want long retention and lower resource use with the least operational work? VictoriaMetrics.
  3. Do you have several Prometheus servers and want a global view without changing collection? Thanos.
  4. Are you running metrics as a multi-tenant platform at very large scale? Mimir.
  5. Do you want to stop operating storage entirely? A managed service, after you have modeled the bill at your real series count.
  6. Do you want to keep PromQL but add SQL and consolidate logs and traces? GreptimeDB.
  7. Are metrics one part of a SQL telemetry workload with high-cardinality dimensions and long history? Arc, knowing you are leaving PromQL behind.

And do not pick a backend from a blog post, including this one. Remote write makes it unusually cheap to test: point a second remote write target at your top two candidates for two weeks, replay your real dashboards and alert rules against each, and measure memory, storage, p99 query latency, and the number of queries that needed changes. Every vendor multiplier in this post, mine included, is marketing until you reproduce it.

FAQ

What is the best Prometheus alternative? It depends on which job you want to replace. For long-term storage with PromQL, VictoriaMetrics is the simplest choice for most teams, Thanos is the gentlest add-on to existing servers, and Mimir fits very large multi-tenant platforms. If you want to stop operating storage, use a managed Prometheus service. If you want SQL over metrics alongside other telemetry, look at GreptimeDB or Arc.

Is Prometheus still a good choice in 2026? Yes. As a collector and as short-retention storage for one cluster it remains the default, and Prometheus 3 made native histograms stable, made agent mode stable, and added OTLP ingestion. Most "alternatives" keep Prometheus or a compatible agent for collection and only replace the long-term storage.

What is the best long-term storage for Prometheus? For most teams, VictoriaMetrics, because it is the simplest to operate. Thanos if you want to keep your existing Prometheus servers and add object storage behind them. Mimir if you need multi-tenancy at very large scale and can operate Kafka.

Can I use a Prometheus replacement without rewriting my Grafana dashboards? Yes, with any PromQL-compatible backend: VictoriaMetrics, Mimir, Thanos, Cortex, the managed services, or GreptimeDB. Test your actual dashboards, because compatibility percentages are measured against someone else's queries. SQL-only databases, including Arc, require rewriting dashboards.

What is the difference between Thanos and Mimir? Thanos attaches to existing Prometheus servers with a sidecar and reads their blocks from object storage, so it is easy to adopt incrementally. Mimir is a standalone backend you remote write into, built for multi-tenant scale, and its preferred architecture uses Kafka. Thanos is Apache 2.0; Mimir is AGPL-3.0.

Does Arc support PromQL or Prometheus remote write? No, not today. Arc ingests Prometheus metrics through Telegraf's Prometheus input plugin and is queried with SQL. If PromQL compatibility is a hard requirement, choose one of the PromQL-compatible options above.

How do I reduce Prometheus cardinality without migrating? Find the top series by metric name, drop or rewrite high-cardinality labels with metric_relabel_configs at scrape time, and move per-user or per-request questions to a system built for them, such as logs, traces, or a SQL database.

This post will be updated quarterly. Last updated October 2026. If something here is wrong or out of date, open an issue on GitHub or in Discord.


Related reading:

Explore Arc:


About the author. I am Ignacio Van Droogenbroeck, founder of Basekick Labs and an ex-InfluxData engineer. I have been working with time-series data since 2018. Prometheus is one of the best pieces of infrastructure software of the last decade, and for most teams the right move is to keep it and give it a better backend. I built Arc for the telemetry that does not fit that shape: high-cardinality, long-lived, and queried with SQL alongside everything else. If that is your workload, I would like you to try Arc.

Ready to handle billion-record workloads?

Deploy Arc in minutes. Own your data in open files on your storage. Use for analytics, observability, AI, IoT, or data warehousing.

Get Started ->