Arc 26.09.3: Smoother Ingestion, Clearer Operations

#Arc#release#v26.09.3#ingestion#reliability#WAL#clustering#observability#security#DuckDB#Parquet#Kubernetes#Basekick Labs
Cover image for Arc 26.09.3: Smoother Ingestion, Clearer Operations

Arc 26.09.3 continues the operational work in 26.09.2: keeping ingestion moving when storage slows down, making recovery more precise, and giving operators a clearer view of what each cluster node is doing.

This patch improves how buffered writes are handled under load and during shutdown, adds durable WAL checkpoints, and keeps replicated files and tier metadata aligned. It also introduces replication-lag and backup metrics, lets DuckDB use its container-aware resource defaults, and corrects access-control checks across queries and listings.

As we begin Arc's second year, this is the work we're continuing to invest in: the behavior people depend on when a database runs every day. Below are the changes operators will notice, followed by the configuration and client updates to review before upgrading.

Ingestion handles backpressure more predictably

When a storage backend takes longer to write a file, the flush queue can fill. Previously, a size-triggered flush could remove a batch from its buffer before confirming that the queue had room, dropping the batch even though its writes had already been acknowledged. Arc now keeps those records in the buffer until a worker can accept them.

Deferred buffers are retried as workers free queue slots, with older buffers considered first. They no longer depend solely on another write to the same measurement or the next age-based sweep to make progress. In the release's test with 150 measurements, one flush worker, and a backend taking 1.5 seconds per write, this removed a 48-second idle stretch and kept the worker processing queued data.

Shutdown follows the same principle. Arc drains buffered and queued writes and waits for active flushes within its shutdown budget. A flush's storage timeout now starts when a worker picks it up, so time spent waiting in the queue does not consume its I/O allowance. A client disconnect also no longer cancels a schema-change flush containing other clients' acknowledged writes.

Two new metrics help distinguish a brief backlog from sustained storage pressure: arc_ingest_flush_deferred_total counts queue deferrals, and arc_buffer_deferred_buffers reports how many buffers are waiting. Buffered data still uses memory, so these changes make backpressure easier to handle and observe without introducing a global ingest-memory cap.

WAL recovery tracks what reached storage

For deployments with wal.enabled = true, 26.09.3 adds durable flush checkpoints. Arc records which tracked WAL entries reached Parquet, and recovery uses those checkpoints to avoid replaying batches that are already stored. Both synchronous and asynchronous flushes participate.

This is particularly useful for measurements without tags, where identical rows may represent separate events and compaction deliberately does not deduplicate them. In the release's crash-recovery test, 310 acknowledged records included 300 already written to Parquet. Recovery now replays the remaining 10 and returns 310 rows; the earlier behavior replayed all 310 and returned 610.

Periodic cleanup also retains rotated WAL files containing tracked, unflushed writes based on flush state rather than file age. A slow storage backend no longer makes those entries eligible for cleanup simply because time has passed. The WAL stats now expose pending_unflushed to show how much tracked work is waiting to become durable.

Recovery remains at-least-once: a crash between a successful Parquet write and its durable checkpoint can still replay that batch, and some legacy entry formats retain their existing replay behavior. The update does not remove historical duplicates from earlier recoveries. The full release notes explain the checkpoint coverage and remaining recovery limits.

Cluster nodes keep a better account of their files

On clusters combining peer file replication with a cold tier, a node's tier metadata previously tracked its own writes without consistently tracking files received from peers. That could prevent partition pruning or leave locally present hot files out of a query when the measurement also had cold data.

Replicated-file arrivals and local removals now update that node's tier metadata through a background writer. A startup scan also reconciles files already present when a node boots or upgrades. Queries and tiering statistics can therefore account for replicated files without waiting for the nightly migration schedule.

Several related changes improve replication and maintenance:

  • Compaction manifest updates are atomic. Registering the compacted output and removing its source entries happens under one lock, so a manifest reader cannot see a half-applied batch.
  • Rejoining nodes process removals from a received Raft snapshot. Arc compares the previous manifest with the restored one and schedules local removal of entries that disappeared. Files outside that comparison still depend on the existing orphan sweep; the notes describe that remaining case.
  • File pulls can try another peer after a checksum mismatch. Arc checks up to three candidates per attempt and keeps the existing committed copy while seeking the updated file. Every replacement still has to pass checksum verification.
  • Interrupted transfers are handled more precisely. A completed staging file is no longer mistaken for a committed Parquet file, and transfer resumes use staging bytes rather than an older committed copy.
  • Final shutdown files are registered before cluster coordination stops. The file registrar drains the final flush's registrations while Raft is still available.

More useful signals for everyday operations

The writer now exports per-peer arc_replication_lag_entries and arc_replication_lag_seconds gauges for connected WAL replication readers. They show outstanding entries and the age of the oldest outstanding entry, measured on the writer's clock.

These describe the reader's current connection, rather than its complete historical backlog. If the replication buffer saturates, the seconds gauge becomes a lower bound; pair it with the entry count and arc_replication_entries_dropped_total when setting alerts.

Backups also provide more detail. The manifest includes a sample of skipped file paths and a count of overlong keys, while backup listings expose skipped-file counts. The new arc_backup_skipped_files gauge makes the most recent backup's skipped files visible to monitoring. The existing skip-ratio check still determines when a backup must fail.

Compaction's deduplication row counts now use the correct Parquet metadata query, making the number of duplicate rows removed visible in logs. Audit events copy request values before background processing, keeping their recorded method, path, and database tied to the original request.

For people testing the experimental ArcX engine, Arrow IPC queries now participate in query tracking and cancellation, and interrupted streams caused by a writer panic carry an explicit incomplete-result signal.

Resource defaults follow the container

Arc now leaves an unset database.memory_limit or database.thread_count to DuckDB's own resource detection. Previously, Arc derived these defaults from the host's CPU count, which could give a small Kubernetes pod a query budget sized for the entire node.

With a container memory limit, DuckDB normally uses 80% of that limit. Compaction subprocesses also receive memory budgets derived from detected capacity, and their default thread counts account for the Go runtime's effective CPU allowance. Explicit configuration still takes precedence.

There is a practical upgrade detail here: without a container memory limit, DuckDB uses host memory, so its default budget can increase compared with Arc's previous CPU-derived value. Set an explicit database.memory_limit if you need a fixed budget, and leave room in database.temp_directory for spilling. These are engine budgets, rather than a cap on the whole Arc process and its concurrent compaction jobs.

Access controls and SQL handling

26.09.3 corrects several authorization paths. Non-admin tokens with team memberships now treat their RBAC grants as authoritative. Previously, an RBAC denial could fall back to the token's coarse permissions, allowing broader access than the grants specified. Denials are now final, and enforcement continues independently of license status. Tokens without team memberships continue using their coarse permissions, and admin tokens retain operator access. Database and measurement listings also apply the caller's permissions; write-only tokens can no longer enumerate them. See GHSA-mcqm-h7hj-99fg for the affected paths and behavior changes.

Query authorization now checks table references inside the measurement endpoint's where parameter, and SQL validation covers additional DuckDB relation forms. Permission extraction and query rewriting also agree on the multiline and CTE forms addressed by this release. The scope and technical details are in GHSA-qcf2-6hm7-62c5, GHSA-9rgq-j585-5fhq, and GHSA-h3rq-5r29-2wrh.

The SQL fixes also improve ordinary queries: comma-separated cross-joins, joins starting on a new line, and table functions under an x-arc-database header now resolve correctly.

Another storage-specific correction disables DuckDB's automatic Hive partition inference on Arc's Parquet reads. If a directory component is named key=value and that key matches a stored column, DuckDB could substitute the directory's value for the file's value. Arc now reads the stored columns without that inference. Deployments with such components in their storage or compaction input paths should follow the inspection guidance in the release notes: earlier compactions or deletes may have persisted altered data, and upgrading alone cannot reconstruct it. Recovery for affected files requires a backup from before the affected operation.

Before you upgrade

  • RBAC deployments: review team grants and wildcard patterns. Leading-wildcard grants now match as documented. Tokens scoped to specific databases must request those databases by name rather than enumerate through GET /api/v1/databases; clients need to handle 403 from that listing. Monitoring that reads compaction or retention endpoints now needs an admin token.
  • Continuous-query clients: send the complete definition on PUT, including name, database, both measurement names, query, and interval. Preserve optional fields too: omitting is_active sets it to false. Stored definitions are revalidated before each run; invalid names or SQL now produce a failed execution with a recorded reason. Check older definitions, including sources with dot-prefixed measurement names and edge-sync spoke names used as databases.
  • Containers: review memory limits and spill-directory capacity, especially if you previously relied on the automatic defaults without setting container limits.
  • WAL and shutdown: allow enough time for the configured shutdown budget. If that budget expires with writes outstanding, WAL-enabled deployments retain recovery data; with the WAL disabled, outstanding data cannot be recovered through replay.
  • Cold-tier configuration: the sample arc.toml now leaves S3 connection values unset so AWS defaults and the credential chain apply. If you relied on its bundled MinIO development values, explicitly configure the endpoint, credentials, addressing, and TLS settings together.
  • Storage paths and edge sync: review any key=value directory components using the notes' checklist. New spoke IDs and sync-path segments reject glob syntax and =; an existing invalid spoke registration needs re-registration and edge configuration updates. Its existing data is retained.

If you're upgrading from before 26.09.2, also read that release's upgrade guidance, including the coordinated Enterprise cluster restart and bundled-object-store migration.

Thank you to the Arc community

Thank you to @efegokdemir, @pujitha24, @lecodev-26, and @0utsights for contributions across replication, ingestion, recovery, backups, and audit handling. Thank you also to @drakeo338 for proposing the line-protocol WAL database fix, and to everyone who reported an issue, tested a change, or reviewed the behavior with us.

That work helps us turn production feedback into specific improvements, with explanations operators can use.

How to update

# Docker Hub
docker pull basekicklabs/arc:26.09.3
 
# or GitHub Container Registry
docker pull ghcr.io/basekick-labs/arc:26.09.3

On macOS, run brew upgrade basekick-labs/tap/arc. For binary and package installations, download 26.09.3 from the GitHub releases page. For Kubernetes, apply your reviewed values with the release chart:

helm upgrade --install arc \
  https://github.com/basekick-labs/arc/releases/download/v26.09.3/arc-26.09.3.tgz \
  -f your-values.yaml

The full 26.09.3 release notes contain the complete change list, deployment-specific checks, and recovery guidance.


Get started:

For questions or help planning an upgrade, join us on Discord or open a GitHub issue.

Ready to handle billion-record workloads?

Deploy Arc in minutes. Own your data in open files on your storage. Use for analytics, observability, AI, IoT, or data warehousing.

Get Started ->