NusaDB

Documentation / Limits and capacity

Limits and capacity

The constraints that change how you size and operate a deployment, stated plainly because finding them during a data load is far more expensive than reading them first.

Read this before loading a large dataset.

Table pages are cached and evicted, so a database can be larger than memory, but a few structures still cannot leave memory and one limit refuses writes once they reach it. The first section says which ones.

What still has to fit in memory

Table pages live in a page cache backed by the last checkpoint image. A page is loaded from the image's page segments on first use, a page unchanged since the last checkpoint can be evicted when memory is needed, and a changed page spills to a scratch file beside the log when there is nothing clean left to evict. The cache is bounded by --max-resident-bytes, or by a value derived from the detected memory budget when the flag is unset. A table scan reads a batch of rows at a time, and B-tree indexes, the primary key's included, live in pages too, so neither has to fit in memory.

What cannot leave memory is small but real: index entries too large for an index page (a key of roughly 2 KB or more), and a note per index entry an UPDATE moved to a new key, kept until the background purge removes the old entry. Once these reach the bound, the next insert or update is refused before it starts:

output
ERROR XX000: out of memory: the engine reached its resident-memory limit of 858993440 bytes
(859001088 bytes held by index entries); free rows (DELETE/TRUNCATE), drop indexes that are not
needed, raise the limit, or use a larger host

Three details are worth knowing:

  • A long-running transaction delays the purge, so heavy key-changing updates under one can grow the notes that count against the bound.
  • If the scratch file cannot be written (a full disk), changed pages stay in memory and the cache grows past the bound until the next checkpoint instead of failing the write.
  • Vector indexes (USING hnsw) are held in memory whatever the bound. A large vector index needs the memory for its whole graph.

DELETE, TRUNCATE, CREATE INDEX and the background purge are not refused at the bound, so space can always be freed. See configuration for the full picture.

The log grows until a checkpoint

The write-ahead log is the durable copy of the data. A checkpoint writes the pages changed since the last one into a new page segment, publishes a new image, and truncates the log, so the data directory holds live data plus the write history since the last checkpoint. A checkpoint is taken when a database is opened with a log past a few megabytes, on demand with the CHECKPOINT statement, and by a background worker once a database's log passes --checkpoint-threshold-bytes (64 MiB by default).

A checkpoint needs a moment with no transaction active. Under continuously overlapping transactions the worker briefly holds new ones (at most --checkpoint-max-pause seconds) so it can run. A transaction held open longer than that, such as an idle client inside BEGIN, defeats the pause, and the log keeps growing until that transaction ends. See configuration.

Restart time tracks the log tail

Recovery opens the last checkpoint image, loads pages from it only as they are used, and replays only the log after it, so start-up time grows with the writes since the last checkpoint, not with the size of the data or its whole history. The background checkpoint worker keeps that tail near its threshold; a transaction left open for a long time lets it grow, and so does a server run with the worker turned off.

Backup, point-in-time recovery and standby

A checkpoint image together with the page segments it names is a consistent copy of one database, even while the server keeps writing, and a hard-link snapshot of them takes milliseconds. With --wal-archive-dir, every checkpoint also archives its log segment and image, and a database can be rebuilt offline as of a moment or a log position with --restore-database. A second server started with --standby-from follows a primary through that archive and serves reads.

What you still arrange yourself: there is no built-in scheduled backup, so snapshots and archive pruning run from your own scripts. The standby lags the primary by its checkpoint cadence, because a segment reaches the archive only when the primary checkpoints; it is read-only, and promotion is a manual restart without --standby-from. See configuration for the recipes.

Single node

The engine runs on one machine. There is no clustering and no automatic failover: the availability of the deployment is the availability of that host and its disk, plus a standby you promote by hand.

One server per data directory

A running server holds an exclusive lock on global/cluster.lock in its data directory, and each open database one on its own lock file. A second server started on the same directory exits at once with a message naming the lock. The operating system releases the locks when the holder ends, even when it is killed, so there is no stale lock to remove by hand. Keep the data directory on a local file system: on some network file systems these locks are not enforced across machines.

A storage error stops that database

If reading or writing a page fails while a change is being made (a damaged page on disk, an unreadable page of the scratch file, a full disk under a ROLLBACK TO SAVEPOINT), that database stops rather than persist a half-made change. Every later statement against it fails with XX000 and a message saying the database stopped after a storage error, and the metrics endpoint reports nusadb_database_stopped{database="..."} 1; the stopped database never checkpoints, so its image stays as the last good checkpoint left it. Other databases on the same server keep serving. After fixing the cause, restart the server: the database is rebuilt from its last image and its log, with every committed transaction and nothing of the interrupted one. An error while a statement only reads a page fails that statement alone.

Query memory: spill or fail, never swap

The executor materialises each stage. With --work-mem set, a sort or hash join whose input exceeds the budget streams the overflow to --spill-dir when one is configured (a Linux host with a detected memory budget gets one by default); without a spill directory the stage fails with an error naming the limit and the flag, and the server stays responsive. Aggregation, DISTINCT and window functions do not spill yet; they fail at the budget. Without any budget, a large enough query can exhaust host memory, so setting one is the safer configuration.

Two more per-client bounds exist: --max-txn-write-bytes caps one transaction's uncommitted writes and --copy-max-bytes caps one bulk load. Both derive from the memory budget and both fail the offending client with a clear error rather than the whole server.

Vector index build cost

Vector search is available with a VECTOR(n) type, four distance operators, and an HNSW index whose recall reaches exact-search results at a high enough hnsw_ef_search. Building the index is expensive, substantially more so than a specialised vector extension, so plan index builds as scheduled work rather than something to do during a load. The graph is saved as it changes and reloaded at open rather than rebuilt, and it is held in memory whatever the resident bound.

Not in this release

AreaStatus
Scheduled backupNot built. Take hard-link snapshots of the image and page segments from your own scheduler, or archive with --wal-archive-dir.
Clustering and automatic failoverNot built. A read-only standby follows through the checkpoint archive; promotion is a manual restart.
Vector indexes larger than memoryNot built. USING hnsw graphs are held in memory whatever the resident bound.
Spill for aggregation, DISTINCT, window functionsNot built. They fail at the budget.
INSERT ... ON CONFLICT into a partitioned tableRefused; upsert into the partitions directly.
Row movement on UPDATE of a partition keyRefused loudly (23514); change the key by delete-and-insert.
INSTEAD OF triggersRefused.
Named time zonesSET TIME ZONE takes UTC or a fixed offset; an IANA region name is refused (no time-zone database, so no daylight-saving rules).
Locale collationsRefused; byte ordering only.
Index methodsEvery scalar index is served by the B-tree (hash / brin are recorded, gin / gist / spgist fall back); hnsw is the vector index.
Full-text configurations beyond simple, english and indonesian; prefix matchingRefused.
Password changes without a restartPasswords live in --auth-user; change one by restarting.
Latency and storage metricsFour server counters and a per-database stopped gauge only.
The .NET driver on NuGetIn the source tree, not yet published (Java is on Maven Central as com.nusadb:nusadb-jdbc).

Where NusaDB fits today

It is a reasonable choice when a single node with a read-only standby you promote by hand is acceptable, and when you are willing to schedule your own snapshots or archive pruning. Data can outgrow memory, since table and B-tree index pages are evicted and spilled, while vector indexes still need memory for their whole graph. It is not the right choice yet where failover must happen without an operator, where a standby must lag the primary by less than its checkpoint cadence, or where a backup must run without any scheduling on your side.

Everything on this page is a property of the current release rather than a permanent limitation. Check the release notes when a new version appears.