Reference

Advanced Internals

How Jazz works under the hood: raw tables, row histories, sync, the query pipeline, and the browser architecture.

This page describes Jazz's internal architecture. You do not need any of this to use Jazz, but it is helpful if you are debugging, reasoning about performance, or understanding why the system behaves the way it does.

Data model

Raw tables plus engine-managed fields

Jazz stays table-first all the way down.

Your schema defines normal application columns such as title, done, and projectId. Under the hood, the engine also tracks a small set of reserved _jazz_* columns that explain how each row behaves over time, such as:

  • a stable row id
  • the branch view the row belongs to
  • the current row-version id
  • ancestry pointers to earlier row versions
  • visibility state
  • confirmed durability tier
  • delete markers
  • engine/user metadata

The important physical fact is that Jazz stores one flat row_format row containing both the user columns and the reserved engine columns. Some Rust types still expose the user-column slice separately for convenience, but that is just a decoded view rather than a different storage model.

Visible entries and row histories

Each logical row has two important storage shapes behind it:

  • a visible entry for current reads
  • a row history containing every stored row version

Ordinary queries read the visible entry first. History is what makes replay, reconnect, branching, and future historical queries possible.

The simplest picture is:

todos
  visible: (branch, row_id) -> current winner for that branch view
  history: (row_id, version_id) -> row versions over time

This is why Jazz can feel like "just tables" at the app layer while still keeping rich local-first history underneath.

Both storage shapes are flat rows:

  • history rows use reserved _jazz_* columns plus the user columns
  • visible rows use a slightly larger _jazz_* prefix plus the same user columns

Indexing

By default, every column on every table is indexed. This keeps where, orderBy, and join lookups fast on any column without you having to think about it, at the cost of one index entry per column per row. The _id index for each table doubles as the authoritative row manifest, so discovering all rows in a table is just an _id index scan.

Sometimes you'll want to optimise for write performance or storage cost instead. Use indexOnly() to specify which columns in the table need indexes; only the columns you specify will be indexed.

Avoid overuse

table.indexOnly(["title", "done"]) does not mean "add indexes on title and done". It means "drop the indexes on every other column on this table". Read it as "index only these". indexOnly is an optimisation that you're unlikely to need to begin with, and if not used correctly, can significantly impact the performance of your app, especially on reads.

schema.ts
import { schema as s } from "jazz-tools";

const schema = {
  todos: s
    .table(
      {
        title: s.string(),
        done: s.boolean(),
        description: s.string().optional(),
        activityLog: s.string().optional(),
      },
      {},
    )
    .indexOnly(["title", "done"]),
};

In this example, where({ title: ... }) and where({ done: ... }) are still index-backed. A query that filters on description or activityLog works, but falls back to scanning every row in the table.

Use indexOnly when you're seeing slow writes and:

  • The column holds large or rarely-queried data (long text, serialised metadata, audit logs).
  • The table is write-heavy and the per-column index cost is showing up in your performance data.

Deletion

The user-facing delete API performs a soft delete. The row is preserved in history, but it disappears from ordinary live queries.

Internally, the current visible state leaves the live _id index and can still be addressed through deleted-row paths such as _id_deleted.

A hard delete mode also exists at the storage layer, but it is not currently exposed as the normal app-facing API.

Row history and truncation

Row history is append-only by default. Every write creates a new row version and keeps older versions available for replay and reconciliation.

There is a low-level truncation path that can drop older ancestry while preserving the current visible state, but it is not a normal application-facing feature yet.

Monotonic direct-write ordering

Each runtime instance maintains a small monotonic clock for direct writes. New row versions created by that runtime get strictly increasing local timestamps, which makes deterministic last-writer-wins ordering straightforward within a single device or process.

Merge strategies

By default, Jazz adopts a per-column last-writer-wins strategy to resolve concurrent edits. If two clients update the same column of the same row simultaneously, Jazz's deterministic hybrid logical clock ordering chooses the winner.

This is a sensible default, but can cause some unexpected behaviour with certain types of data. For example, imagine a voting app. Alice and Bob both read the current voteCount value as 2. They each want to increment the value. If they simultaneously write 3, then the new value will be 3, even though they actually each wanted to increment by one.

Counters

Use .merge("counter") on an integer or bigint column to keep deltas additive instead:

schema.ts
import { schema as s } from "jazz-tools";

const schema = {
  votes: s.table(
    {
      proposal: s.string(),
      voteCount: s.int().merge("counter"),
      totalVotingPower: s.bigint().merge("counter"),
    },
    {},
  ),
};

With merge("counter"), every update(...) on the column is recorded as a delta from the value the writer was looking at. When concurrent edits meet, the deltas are summed: Alice's +1 and Bob's +1 both apply, and the voteCount correctly lands on 4.

Counter merges are useful for things like:

  • A shared score or vote tally.
  • An inventory level being incremented and decremented from multiple devices.
  • Any other counter where you care about preserving every increment rather than which device wrote last.

merge("counter") is only valid on non-nullable integer or bigint columns. Calling it on a string, a nullable integer (s.int().optional()), a nullable bigint (s.bigint().optional()), or any other type throws at schema construction time.

Grow-only sets

Counters solve concurrent numbers; arrays have the same problem with concurrent membership. Under last-writer-wins, if Alice adds "urgent" to a tags array while Bob concurrently adds "blocked", one write clobbers the whole array and a tag is silently lost.

Use .merge("g-set") to make an array column a grow-only set instead — concurrent writes converge to the union of every replica's elements:

schema.ts
import { schema as s } from "jazz-tools";

const schema = {
  documents: s.table(
    {
      title: s.string(),
      tags: s.array(s.string()).merge("g-set"),
    },
    {},
  ),
};

When concurrent edits meet, the merged array is the union of all contributed elements, deduplicated and sorted into a canonical order so every replica converges on a byte-identical result. Alice's "urgent" and Bob's "blocked" both survive. An element written by one replica is never dropped by a concurrent write from another that never saw it.

Grow-only sets are useful for things like:

  • Accumulating tags, labels, or category membership.
  • Append-only logs of participants, contributors, or seen IDs.
  • Any collection where you care about keeping every element rather than which device wrote last.

merge("g-set") is grow-only: there is no element removal. It is only valid on non-nullable array columns; calling it on any other type, or on a nullable array (s.array(...).optional()), throws at schema construction time.

Cold start

On startup, Jazz loads indices first rather than eagerly decoding every row. Row content is then loaded on demand as queries reference it.

The result is that cold-start cost is much closer to "index size" than "total stored data size."

Browser architecture

Core browser runtime

In the browser, Jazz uses the direct Rust/WASM core for reads, writes, subscriptions, and sync. With driver: { type: "memory" }, the database runs in the main thread and syncs to the configured server over the core WebSocket protocol.

With driver: { type: "persistent" }, Jazz opens one directly durable SharedWorker runtime per database namespace. The worker owns the IndexedDB page store and the server connection; each tab communicates with that runtime through a message port. There is no tab-leader election, follower handoff, or per-tab durable database. Jazz does use one origin-wide Web Lock per physical database namespace to prevent retry generations or separately loaded worker assets from concurrently opening the same IndexedDB root. If that lock is unavailable or already held by another worker realm, the new realm fails closed instead of risking two durable owners.

driver.dbName (and dbName) selects a logical base, not a physical IndexedDB namespace. Jazz derives the physical IndexedDB and SharedWorker namespace from that base plus app, environment, and canonical authentication scope. Thus accounts can coexist on one browser: reopening the same base as Alice selects Alice's prior cache, while Bob selects Bob's. The derived names and durable metadata contain no credential, token, secret, or arbitrary claims.

On its first open Jazz also pins the exact non-secret logical owner (app, environment, and authentication scope) next to the page-store manifest. This is defense in depth against a derivation bug or low-level attempt to open the wrong physical root: that attempt fails before pages change. db.deleteClientStorage() / logout with wipeData destroys only the current scoped namespace; it does not erase another account's cache under the same logical base. The durable owner is neither the foreground replica/node ID nor a credential; those remain separate identities.

Calling db.disconnect() in one tab explicitly takes that whole persistent namespace offline: the worker has one upstream connection, so every attached tab uses local-first fallback for "remote-if-possible" reads and remote reads wait until any tab calls db.reconnect(). This is intentional; disconnecting only one tab while continuing to sync its writes through the shared worker would give that tab an incoherent offline contract.

The main thread keeps a responsive client-side peer while the SharedWorker commits durable state in the background.

With driver: { type: "memory" }, the worker and IndexedDB are skipped entirely, and the main-thread runtime syncs directly with the server.

IndexedDB crash safety

The IndexedDB page store commits a complete page generation and its metadata in one IndexedDB transaction. A failed or interrupted write therefore leaves the previous committed generation available on reopen.

React Native

React Native uses its native adapter rather than browser workers or IndexedDB.

Query engine

Execution pipeline

Queries compile into a graph of processing nodes:

IndexScan → [Union] → Materialize → [PolicyFilter]
  → [ArraySubquery] → [Filter] → [Sort] → [LimitOffset]
  → [Project] → Output

Nodes in brackets are only present when the query requires them. The graph processes deltas incrementally, which means that when data changes, only dirty nodes re-evaluate. That is what makes live subscriptions efficient: a single row change does not require re-running the whole query.

Materialization

Materialize is where candidate row ids turn back into rows.

It typically:

  1. looks up the visible entry for the relevant branch
  2. falls back to row history only when the query needs an older settled winner
  3. decodes or reprojects the flat row, dropping the reserved engine columns before returning app-facing values
  4. emits row-level deltas to the downstream graph

This is why the visible region matters so much: most current reads never need to reconstruct a row from full history.

One-shot queries

db.all() and db.one() are implemented as "create a temporary subscription, wait for the first durability-qualified snapshot, then auto-unsubscribe." They share the same reactive machinery as live subscriptions, which is why they participate in durability-tier gating and lens transforms.

Include performance

Initial setup and materialisation still scale with the size of the relation returned by an include() or array subquery. Once the query is maintained, a child-row change is routed incrementally to the affected result. A canonical scale canary checks that the allocation cost of one maintained relation change does not grow with the accumulated relation size.

Sync protocol

Transport

Jazz uses a single WebSocket sync transport plus a small HTTP surface for health and admin reads.

  • Sync: GET /apps/<appId>/ws upgrades to a WebSocket carrying the typed sync protocol.
  • Admin: GET /apps/<appId>/schemas, GET /apps/<appId>/schema/:hash, and POST /apps/<appId>/admin/... handle schema and permissions publication/read flows.
  • Health: GET /health.

Client identity

Each client generates and persists a stable ClientId. On reconnect with the same id, the server can treat it as the same logical peer rather than as a brand-new client with no prior state.

Reconnection

The TypeScript client uses exponential backoff with jitter. On reconnect, active query subscriptions are replayed as anti-entropy: the server re-evaluates them and resends any rows the client still needs.

Trust model and client roles

Sync is asymmetric:

  • Upward (client -> server): row versions, row-state changes, and catalogue updates are pushed toward trusted servers
  • Downward (server -> client): only rows matching the client's active query subscriptions are sent

Each client connection has a role that determines how writes are routed:

RoleWrite handling
UserWrites queued for permission policy evaluation before apply
AdminWrites applied directly, no permission check
PeerWrites applied directly, used for trusted runtime-to-runtime sync

Frontend clients usually authenticate as User. Backend services with a backend secret authenticate as Admin or Peer.

Schema evolution

Lenses

Migrations in Jazz produce lenses — bidirectional transformations between schema versions. When jazz-tools migrations create diffs two schemas, it generates a lens with declarative operations such as adding, removing, or renaming columns and tables.

At query time, Jazz can use lens paths to read older stored data through the current schema. At write time, it projects updates through the lens path while retaining the row's physical schema identity.

Catalogue sync

Schemas and lenses travel through a separate catalogue lane, not through the normal user-row history path. Clients publish catalogue entries, servers discover them lazily, and query execution uses that catalogue state to resolve schema context on demand.

Durability signals

Jazz separates two durability questions:

SignalGatesQuestion it answers
QuerySettledFirst read delivery"Has the query result settled at tier T?"
Write tier confirmation.wait({ tier }) promise completion"Has this write been confirmed at tier T?"

Both use the same tier lattice (local < edge < global), but they answer different questions. A query's first callback is held until QuerySettled reaches the requested tier. A .wait({ tier }) promise resolves when the requested tier confirms the write.

The read durability tier only gates the first delivery of a subscription. After the initial snapshot arrives at the requested tier, later updates are delivered as they reach the local node. That means tier: "global" gives you a globally settled first snapshot, not globally gated delivery forever after.

On this page