Nuthatch
The book
Version 4.12.0 · October 20261. The data problem
An Ethereum node can tell you what happened in a block. It is rather less interested in helping you answer a product question such as “which addresses delegated to this indexer last week?” That answer requires turning a stream of logs into a durable model, keeping it current as new blocks arrive, and making it cheap enough to ask repeatedly. This is the ordinary work of an indexer.
The usual answer is a hosted subgraph. That is often an entirely sensible answer. A schema and its mappings turn chain activity into a GraphQL endpoint, somebody else runs the infrastructure, and the application obtains a pleasant query surface. The awkward moment is when that endpoint is unreachable, unserved, behind the chain, or contains a capability that cannot be reproduced from the chain alone.
Nuthatch starts from a smaller promise. A nest describes a chosen set of contracts, pinned ABIs, events and SQL views. It reads logs from the chain, decodes those logs deterministically, and serves the resulting event data through SQL and HTTP. There is no allocation, indexing market or hosted control plane between the RPC endpoint and the query. A nest is something a team can keep beside the application that depends on it.
The useful boundary
A transaction log is public chain history. Given the same block range, contract addresses and ABI, two honest implementations should decode the same rows. This makes event indexing a good substrate for a portable local index. It also makes it possible to verify the result independently rather than taking an API’s word for it.
Not all subgraph data has that property. A mapping may make eth_calls whose result depends on the
block, fetch IPFS content, call an external service, or maintain bespoke off-chain state. Those are
not mistakes. They are simply dependencies that lie outside the event stream. Pretending that a
log index automatically offers parity with every such mapping is the beginning of a fairly tedious
outage.
Two of those dependencies Nuthatch now handles, and it is worth being precise about how, because the
manner matters more than the fact. A [[calls]] declaration reads a contract at a fixed block,
and the result is addressed by the question that produced it: chain, block, contract, calldata. An
[[ipfs]] declaration resolves a document and verifies the bytes against the address that names
them. Neither is the ambient, unrecorded fetch that makes a mapping hard to re-execute. The
property this chapter cares about - that two honest implementations agree, and that you may check
rather than trust - is preserved in both cases, because a pinned question has one answer and a
content address describes exactly one document.
What remains outside is what always was: an external service, a value that depends on when you asked, bespoke off-chain state. Those are still dependencies rather than data, and naming them is still the first honest step in a port.
The first question in a port is therefore not “can Nuthatch replace this subgraph?” It is “which reads are event-derived, and what does each remaining read depend on?” Many useful reads turn out to be straightforward: transfers, swaps, delegation changes, registration records and cumulative balances. Some current-looking values can be derived from ordered events, such as ERC-20 supply from mints and burns. Others genuinely require a contract read or external input. Name the boundary before promising a fallback.
A modest but strong promise
This bounded scope gives Nuthatch several useful properties:
- The input is inspectable. A nest vendors its ABI and says which events it accepts.
- The output is repeatable. Decoding is a pure transformation of log plus authored decode rules.
- The operator owns availability. An application may query its own machine or an API it controls.
- Historical data can be sealed and retained without keeping a full node or a bespoke database cluster alive forever.
- Sealed history travels. Since 4.9.0 an operator can publish a nest’s sealed segments as a mirror, and another operator can seed a new nest from it instead of backfilling over RPC, with every file checked against the content address the catalogue names. What the hashes do not prove is that the catalogue is true to the chain; that remains the reader’s question to ask of whoever published it.
The machine does not know the business meaning of an event. It knows how to preserve and present it without quietly changing its mind. The semantic layer belongs in views and application code, where it can be reviewed as such.
A running example
Imagine a protocol dashboard that normally asks a GraphQL subgraph for recent delegations. Its critical screen needs the delegator, indexer, timestamp and amount. Those values originate in a specific event. A nest can watch the protocol contracts, decode that event into a table and expose a view with precisely those columns. The dashboard retains its normal endpoint, but gains a small and independently operated fallback for the read it cannot afford to lose.
This is not an ideological replacement programme. It is a good operational shape: use the rich
surface you have, and keep the irreplaceable event data somewhere you control. The practical
subgraph fallback guide walks through that exercise. Where the
dashboard’s queries ask only for event-shaped fields, the separate nuthatch-graph download
(attached to every release since 4.11.0) can answer the subgraph’s own GraphQL for them and refuse
anything else by name; the default binary serves no GraphQL at all.
The next question is how that small description remains trustworthy after it has left your laptop. That is the work of the authored nest.
2. The authored nest
Before Nuthatch indexes a single block, somebody must make several decisions: which chain matters, which contracts matter, which ABI is authoritative, which event signatures should be decoded, and which derived reads readers may need. These choices are the nest. The database files, generated schema and decoded rows are consequences of that description, not the nest’s essence.
That distinction is easy to wave away until an ABI changes, a directory is copied, or a deployment must be reproduced on another machine. Then the authoring inputs are precisely the thing you need to preserve.
The authored inputs
A normal nest directory contains nuthatch.toml, vendored files under abis/, and optionally SQL
views, incremental entities, semantic descriptions and checks. The TOML names contracts and their
start blocks. The ABI pins the event layout used to decode their logs. A view may turn raw rows into
a consumer-shaped read without modifying the underlying historical facts. The TOML can also ask for
a little more than logs: on a chain that reports its L1 block in every header, [extract] l1_blocks = true (4.6.0) adds a table recording, for each block that carried a decoded row, the L1
block it settled against. That is an authored input like any other, so turning it on for a nest
that has already indexed is refused rather than quietly applied from that point on.
The ABI is deliberately vendored. An explorer API is useful when scaffolding a nest, but it is not an adequate long-term dependency for its definition. An explorer can change, a proxy can mislead, and a fetched ABI may not be the ABI that was intended at the time. Keeping the source artefact in the nest makes review and reproduction possible.
From those files Nuthatch generates a decode registry and a schema. The registry maps the address and event signature it receives from a log to the columns it will write. The generated artefacts are checked rather than treated as private magic. A packaged nest carries the registry hash its inputs produced, and installing one regenerates the registry and refuses the package if the two differ. A nest run from a directory carries no such claim, so the check runs the other way round: at startup the identity its configuration produces is compared with the one recorded in its store, and a mismatch refuses to start. The schema file is treated as what it is, a derived artefact, and is regenerated when it has gone stale rather than argued with. A stale decoder producing plausible rows is not a success condition.
Events first, then views
The raw event table is the durable base layer. It records chain coordinates such as block number, transaction hash and log index alongside decoded event fields. Those coordinates are not clutter. They establish order, support audit, and give a reader a route back to the originating chain fact.
Views sit above this base layer. They are ordinary SQL declarations authored with the nest. A view can make transfers friendly to query, calculate a balance from the ordered transfer stream, or present a protocol-specific activity table. Because it is a view rather than rewritten history, a consumer can inspect the derivation and the raw events remain available when the definition needs to change.
This is an important division of labour. A decoder answers “what did this log say according to this pinned ABI?” A view answers “what result do we want from those rows?” Conflating the two makes schema upgrades unnecessarily dangerous.
Content identity
Nuthatch packages authored inputs into a canonical manifest. Hashed as it stands, with SHA-256, that manifest gives the bundle hash, which also covers the version of nuthatch that wrote it. The nest identity, or NID, is hashed from the same manifest with that version replaced by a fixed placeholder and a domain prefix in front, so upgrading the binary alone never moves a nest’s identity. The NID identifies what was authored, not which directory happens to contain it and not who mounted it. Two copies of the same inputs have the same identity. A one-byte change in an ABI, configuration or view creates a new identity.
That is deliberately strict. The hash is not a version label chosen at a meeting. It is a statement that this exact package is what the runtime verified. Human names and versions remain useful for navigation, but the NID is the thing a machine can use to decide whether two claimed nests are identical.
Identity does not mean every NID requires a fresh backfill. Sometimes the package has changed in a way that leaves its event data byte-identical. How Nuthatch recognises and safely adopts that data is a later chapter. For now the point is simpler: the system begins by making the definition of a nest small, explicit and reproducible.
For the exact file layout and configuration, see authoring modes and the configuration reference. Next, the nest meets the chain.
3. The cursor
An indexer has two jobs which look similar from a distance and behave very differently in practice. It must collect old history, often millions of blocks, and it must then follow a live chain whose latest blocks are not yet final. Nuthatch calls the durable position that coordinates this work a cursor.
There is one cursor per chain in a runtime. Not per HTTP route, not per tenant and not per nest. That is a useful constraint, not a missing feature. One chain has one ordered tip, one finality boundary and one reorganisation story. Giving every nest an independent opinion about those things would multiply RPC work and make recovery needlessly inconsistent.
Backfill is a controlled walk through history
For a newly mounted nest, Nuthatch begins at the earliest start block its contracts declare and asks
the RPC for logs in bounded windows. (A nest that declares none starts 5,000 blocks behind the tip,
and --backfill N overrides either.) Providers impose limits on ranges, result counts and
concurrency, so the process does not assume that a heroic getLogs call will be welcome. It splits
work into windows, retries within its policy and records progress only after the rows have been
accepted by the hot store: a window’s rows, its block-hash checkpoint and the new last block land in
one transaction, so there is no moment at which the store claims a block it does not hold.
The details matter because RPCs are prone to giving an answer that is technically valid and
operationally useless. A provider may time out, cap a response, or make a 10,000-block request feel
like a personal insult. The window therefore adapts: it starts from the chain’s measured default,
shrinks when a request is refused and grows back when requests succeed, and --window is the
ceiling an operator places on that. A provider can also object to the filter rather than the
range: publicnode on BNB Chain and Polygon refuses a request naming ten or more addresses, and since
4.11.0 that is recognised as a refused filter rather than misread as a credentials failure, and the
address list is split in halves until a request is accepted. Concurrent fetching is reserved for the direct-seal backfill of
history already past finality; the ordinary walk is one window at a time, because its results must
enter the hot store in order. Faster is useful, but only if every accepted block remains attributable
and repeatable.
A window can carry more than logs. Where the nest declares [extract] l1_blocks, the cursor also
fetches the header of every block that produced a decoded row and records the L1 block it reports.
A header without that field refuses the whole window rather than storing a zero, because a row
that says “settled against L1 block 0” is a lie with a plausible shape.
Backfill catches a nest up to the present. It does not establish that the present is permanent.
The tip is provisional
At the chain tip, a block can be replaced by a competing block. This is a reorganisation. A reader who saw an event in the first branch must not be left with that event after the chain selects the other branch. Nuthatch therefore keeps unfinalised data in its hot store, with enough block-hash checkpoints to notice when the chain no longer agrees with the path it had followed.
When a mismatch is detected, the cursor finds the fork point and each affected hot store rolls back to that point. It then indexes forward along the winning branch. The invariant is not “we never briefly served a provisional result”. No honest near-tip system can promise that. The invariant is “provisional rows are marked by their place in a reversible part of the pipeline, and the store converges to the canonical chain.”
Finality is what ends that reversible period. Once a block sits sufficiently behind the head under the configured chain policy, Nuthatch can seal it. Sealed history is no longer subject to ordinary tip rollback and becomes the cold, durable half of the query surface.
One chain, one source of ordering
In a multi-nest runtime the cursor obtains the union of needed logs and routes them to the nests
that own their address and event signature. A log may be relevant to more than one nest, in which
case each gets its own decoded rows. Fetching is shared; the nests’ datasets are not silently
merged. This distinction keeps ownership and rollback manageable while avoiding N copies of the
same RPC polling. One consequence is worth knowing when sizing a runtime: a factory nest cannot
name its children’s addresses in advance. Since 4.11.0 a cursor hosting one asks by the factory’s
address and every child discovered so far, rediscovering as it goes, until the union passes 500
addresses; past that it fetches by topic alone, and the whole union loses its address filter. Its
neighbours then pay, in logs fetched and discarded, for the factory’s open-endedness. An endpoint
that refuses an address-less eth_getLogs keeps the cursor asking by address whatever the count.
Different chains need different cursors. They have different heads, different finality rules and different failure domains. A runtime can host them, but it does not pretend that Arbitrum and Ethereum form one sequence merely because they have both inconvenienced the same operator.
The cursor is the reason an index can say where it is. The next chapter is about where the data sits once it has passed through that cursor, and why it lives in two forms.
4. Storage and sealing
Chain data has an awkward temporal property. The newest blocks must remain reversible because a reorganisation may replace them. Old blocks ought to be cheap to retain and scan for years. Trying to satisfy both needs with one storage shape usually produces a database that is rather busy doing neither particularly well.
Nuthatch separates the two deliberately. The hot store holds the near-tip, reversible working set. Finalised history is sealed into immutable Parquet segments. The query layer joins the two so a reader asks one question without having to know whether the answer happened ten minutes or ten months ago.
Hot data is for change
The hot store receives decoded rows while the cursor backfills and follows the tip. It supports the operations a live index needs: inserting new event rows, recording checkpoints, rolling back a reorganisation and deleting a range that belonged to the discarded branch. It is per dataset, which gives a nest a clear write boundary and prevents an error in one package from rewriting another package’s working set.
The hot store is not a second-class cache. Before a block has crossed the finality boundary, it is the authoritative record of what the cursor currently believes the chain says. Its mutability is a feature. Treating tip data as immutable merely moves the eventual correction into an application bug, where it is more expensive and less visible.
Sealing is a commitment
When rows are final enough, Nuthatch writes them to Parquet segments and records them in the
segment catalogue, segments/manifest.json. A segment is immutable content with a bounded block
range, table identity and hash; the hash is the SHA-256 of the file’s bytes, and it is also the
file’s name. The catalogue describes the set of segments that make up the sealed historical view
and is itself part of the evidence an operator can inspect.
“Final enough” is decided per chain, in the binary’s chain table rather than in the nest’s
configuration. Ethereum mainnet seals 64 blocks behind the tip; the L2s that serve a finalized
tag use it, with a fixed depth to fall back on when an endpoint will not answer; a chain the table
does not know gets 64 blocks. Finality says when a row may be sealed, not when it is. The
sealer cuts a segment once the finalised rows reach 20,000 or 64 MiB, or once they span the
chain’s seal span (1,800 blocks on mainnet), whichever comes first, because a segment per poll
would be a catalogue of crumbs. Until a cut, finalised rows stay in the hot store with the
watermark pinned at a checkpoint, and a finalised range with no rows at all simply advances the
watermark. A cut that leaves a table with fewer than 1,000 rows writes a segment marked
provisional, which the table’s next seal folds in and replaces; a quiet table is not condemned to a
thousand tiny files.
The operation is not “move some old rows to a different folder and hope for the best”. Each
segment is written to a temporary name, fsynced, renamed into place and its directory fsynced;
only then is the catalogue rewritten the same way and swapped in with one rename. If a process
dies at an inconvenient moment, the previous committed catalogue remains coherent and the orphaned
file is just a file. What this path does not do is read its output back: the hash is computed
from the bytes in memory before they are written. Verification is a separate job, done at every
startup, when each catalogued segment is re-hashed and a damaged one quarantined, and on demand by
nuthatch doctor. Recovery may be boring, which is the highest compliment available to storage
machinery.
Because segments are immutable and named by their content, a runtime keeps them in one shared
store, segments/<hash>.parquet beside its datasets, and each dataset’s catalogue points into it.
A second nest sealing byte-identical rows finds the file already there. This does not mean every
nest shares every table. It means the runtime can avoid retaining identical sealed history twice
while preserving each dataset’s package and query surface. The distinction becomes important
during upgrades.
Content addressing is also what lets sealed history leave the machine. nuthatch publish mirrors a
nest’s non-provisional segments, schema and catalogue to an object store under its data identity,
and nuthatch seed (4.9.0) fills a nest that has never indexed from such a mirror, checking every
file against the hash the catalogue names before installing it. The hashes prove the files are
the ones the catalogue describes. They prove nothing about who wrote the catalogue.
One query surface
Burrmill, the embedded query engine on DataFusion, reads sealed Parquet and the current hot rows
together: each table is the sealed segments unioned with the hot rows above the sealed watermark.
Nuthatch builds a read-only SQL surface over that union; only SELECT and WITH are admitted. A
query for a block range that straddles the finality boundary does not require the caller to issue
two requests or reconcile duplicate rows. The storage boundary is an implementation detail, albeit
one worth understanding when diagnosing performance. (Burrmill has been the engine since 4.1.0,
released 2026-10-02. DuckDB served from the first release until then, and appears in this book
only as that history.)
There are guards around this freedom. Queries have concurrency, time, row and unsealed-row bounds: two statements at a time per cursor unless the operator raises it, thirty seconds, 50,000 rows, and two million hot rows or 64 MiB of them, past which a query is refused rather than quietly answered from sealed history alone. Since 4.10.0 the engine also counts the memory a scan holds against a pool the operator sizes, and startup refuses a configuration whose pool, ingestion reservation and headroom do not fit the cursor’s budget, naming the largest value that would. Those guards protect the node from an enthusiastic analytical query becoming a denial-of-service tool. They do not provide customer identity, quotas or billing. Those belong at the gateway, where there is actually an authenticated caller to reason about.
Because an answer is a function of its inputs, a repeated statement can be remembered without ever
being stale. The answer cache keys an entry on what the answer depends on: the statement, the sealed
segments it is served, the generation of the nest’s hot rows, each entity’s watermark and the
authored files. A committed row or an admitted segment changes the key, and the next request
computes. Up to 4.10.1 the key followed the store’s write counter and the seal watermark, both of
which move on every cursor poll; since 4.11.0 it follows the rows and segments themselves, so a nest
that is quiet at the tip answers a repeated statement from the cache. It is bounded by
NUTHATCH_SQL_MEMO_BYTES (64 MiB by default, 0 turns it off), and a degraded answer is never
remembered.
The full life of an event
Take one Transfer log. The cursor sees it in a block, selects the ABI decoder and writes a row
with its chain coordinates into the hot store. A reader can now query it, knowing it remains inside
the reorg window. After finality, and once enough finalised rows have gathered for a cut, the
sealer writes the row into a Parquet segment, commits the updated catalogue, and then in one hot
store transaction prunes the corresponding hot copy and advances the sealed watermark. The SQL view
continues to return it, because it reads hot plus sealed history as one logical table.
The row changes physical home exactly once under normal operation. Its meaning does not change at all. That is the point of keeping decoding separate from storage and of treating finality as a first-class boundary.
For operational details, see storage and sealing and reorgs and finality. We can now turn to the thing readers actually use: the index’s query surface.
5. Reading the index
An index is useful only when somebody can ask it a question. Nuthatch exposes decoded event data as tables and makes that data available through read-only HTTP and SQL surfaces. The system is intentionally more ordinary than a bespoke query language. SQL is well understood, inspectable and capable of expressing both a simple point lookup and a careful event-derived calculation.
Start from the raw event
Every selected event has a table, including an event which the contract has not emitted yet. In
that case the table is present and empty, rather than appearing only on the day the first log
arrives. The table contains decoded event fields and chain coordinates: block number, block hash,
block timestamp (on unless the nest was scaffolded with --no-timestamps), transaction hash, log
index, emitting address and a sequence number. There is no transaction index; the log index is
block-wide, so it alone orders a block. These columns let a consumer order simultaneous events,
trace a result back to a transaction and decide how to display the distance from finality.
Raw tables are the audit trail. They may not be the API a product wants to hand to a browser, and that is quite all right. They are the stable base upon which a product-specific interface can be built. When a view looks suspicious, the raw log-derived rows provide the way to check it.
Views give event data a useful shape
An authored SQL view is part of the nest package. It can rename columns, select a narrow consumer
surface, join related event tables and derive incremental-looking answers from the event history.
For example, an ERC-20 total supply can be calculated from mints and burns, and a latest Uniswap V2
reserve can be selected from the latest Sync event per pair. These are not opaque mapping code.
They are reviewable SQL over a known, pinned input.
This is a strong capability, but it has limits. A view is only as sound as the event model beneath
it. A historical event stream cannot reproduce a state variable that was changed without an event,
or content that a subgraph fetched from IPFS. An eth_call at a particular block may be necessary
for some questions.
Where such a read is necessary, it is declared rather than performed quietly: a [[calls]] or
[[ipfs]] block in the nest’s configuration, pinned to a block or checked against a content address,
and stored in a table of its own that a reader can see and interrogate like any other. The
distinction this section is drawing survives intact. Nuthatch makes the event-derived part
dependable, and it does not perform other forms of data acquisition behind the reader’s back - it
performs them in front of the reader, or not at all.
An answer that is already computed
For a question that is expensive, repeated, and naturally keyed, Nuthatch 3.0 adds a third surface:
an authored incremental entity. Its SELECT is deliberately restricted, declared with a key and a
row bound in entities.toml, then maintained as each block arrives. The relationship is not a cached
view. A reorganisation submits the removed facts as retractions, and the maintained answer retracts
with them.
This changes where the cost falls. A normal view is free until a reader asks, and is the correct default. An entity reserves memory whether it is queried or not, and makes every incoming block pay a small update cost. The payoff is that a reader need not fold all sealed history again. On Lodestar, the same 82-row rewards result moved from 2.15 seconds at p50 as a view to 87.7 ms as an entity.
That comes with an operator’s obligations. The declared maximum rows is a hard admission bound, a
fault quarantines rather than serving a frozen answer, and a restart seeds the entity from locally
stored facts before serving it. The maintained relation is read by key under /derived and by
name from /sql, beside the decoded tables. The entities guide spells out
those limits.
Serving safely
The SQL endpoint is read-only and bounded. It has a concurrency semaphore, a request timeout, a row cap and an upper limit on the unsealed rows a scan may include. A public deployment should use named queries or an allowlist when it does not intend to offer general SQL. The runtime enforces these local resource bounds, while a gateway in front supplies authentication, rate limits and tenant policy.
The HTTP API offers conventional endpoints for discovery and tables as well as /sql. Admin routes
are a distinct set of routes under /_admin, and they may mutate runtime state, such as mounting or
unmounting a dataset. They are not a separate listener: they sit on the same port as the data API,
and the binary mounts them only when that port is loopback or an admin token has been set in the
environment, with --no-admin to leave them out altogether. It is important not to describe the
whole server as “read-only” merely because its data API is. The administrator can make changes, so
the port that carries the admin routes needs the corresponding protection.
MCP and semantic descriptions provide more guided access for agents and tools. They are not a second data model. They describe and query the same underlying tables, views and entities.
A good consumer contract
For a dashboard or service, define a small query contract: the rows, ordering, finality behaviour and error policy it actually needs. Put the contract in an authored view or named query. Test it at a fixed watermark. Then an application can switch from a degraded primary endpoint to the nest without improvising SQL during the incident, which is a period when even ordinary punctuation can become a tactical challenge.
The SQL reference, HTTP API and authored views guide provide the exact surfaces. The next chapter explains how this data can be retained and reused even as its surrounding nest package evolves.
6. Maintaining an answer
An indexer has two chances to pay for a calculation. It can wait until somebody asks, scan the facts then, and return the result as a view. Or it can do a little work for each arriving block, retain a bounded answer, and make the later read cheap. Neither is universally better. The first is flexible; the second is useful precisely because it is constrained.
Nuthatch 3.0 makes that choice available to a nest author through an authored incremental entity. It is the first release in which the author, rather than only the built-in balance circuits, can say that a relation should be maintained during indexing.
Two kinds of authored SQL
An ordinary views/*.sql file contains CREATE VIEW. It runs on the hot and sealed query surface when
a reader asks for it. It can use the broad read-only SQL surface. It has no permanent state and does
not affect ingest. This is the sensible default.
An entity has one SELECT in entities/ and a declaration in entities.toml. That declaration gives
the result a key and a maximum row count. Nuthatch compiles the admitted subset to an incremental
circuit and feeds it facts as they arrive. The relation can be queried through SQL or read by key.
The distinction stops a pleasing but false sales pitch. Naming a view does not make it incremental. An entity does, but it gains that property by accepting resource limits and less SQL, not by attaching an adjective to a SQL query.
Why the restriction earns its keep
Incremental maintenance must be able to add a block and retract a block without rescanning all earlier
ones. Projection, filtering, exact arithmetic, grouping, supported aggregates and inner equijoins have
that shape. ORDER BY, LIMIT, window functions, outer joins, DISTINCT, percentiles, recursive
queries, HAVING, common table expressions and count(x) as opposed to count(*) do not belong in
the first safe subset. They remain valid questions for views.
The row bound is equally important. A maintained answer is state retained at the cursor, alongside
the chain’s hot data. max_rows participates in admission rather than being a hopeful comment, and
it bounds more than the answer: it caps the live input rows the circuit retains, summed across its
input relations, including while it seeds on restart. A count(*) grouped into a handful of rows can
still retain thousands of inputs, so size max_rows for the inputs, not only the result. Past it, the
entity faults and the nest is quarantined. An answer that has stopped
being maintained must not continue to be served as current.
The lifecycle is still one circuit
During backfill, tip following and a reorganisation, the same relation receives facts. A reorg feeds the removed facts with negative weight, so no bespoke author rollback can drift from the forward path. After a restart, the relation is seeded from sealed segments and the hot tail already held locally. No historical RPC replay is involved.
This is observable. Each entity has applied-through, current, row-count, state-bytes, faulted,
unavailable and seconds-since-progress metrics, labelled by nest and entity. The row count is the
answer’s rows, not the input rows max_rows bounds, so the two are not the same number. The
reader’s provenance and the operator’s metrics answer the same question from opposite ends: is this
result current, and if not, why are we still looking at it?
One ordering detail was wrong until 4.10.1 and is worth knowing because it is the kind of thing a
circuit makes easy to get wrong: a window’s facts were applied to the entity before the hot store
had committed them, so for a moment the entity’s watermark ran ahead of what /sql could read,
and a window whose commit then failed had been folded all the same. The entity now applies a window
only after the store has it.
There is one deliberately blunt limitation, introduced in 3.0 and still standing. --seal-direct
bypasses the ingest route used to maintain entities, so it is refused for a nest declaring one. It
is better to reject an unsupported fast path than to finish a backfill with an elegant, empty
answer.
The practical authoring details are in the entities guide. The next chapter returns to identity, where changing an entity definition has different consequences from changing the facts beneath it.
7. Identity, upgrades and reuse
Software versioning is often an agreement among humans: this release is called 2.5.0, and these are the changes we say it contains. That remains valuable, but it is not sufficient for deciding whether two indexer packages decode the same data. Nuthatch uses content identity for that question.
The nest identity, or NID, is the SHA-256 hash of the canonical authored manifest, taken with the
manifest’s generator version replaced by a fixed placeholder and behind a nuthatch-nid-v1 domain
prefix, so a new nuthatch binary does not by itself make a new nest; the plain manifest hash, version
included, is the separate bundle hash nuthatch nest bundle prints. Change a
contract selection, pinned ABI, event configuration or authored view, and the NID changes. The
runtime stores data by that identity under data/<nid>/; a mount maps an operator-facing alias and
tenant to it. An alias may change without changing data. Two tenants may mount the same NID without
starting two indexers. The name is for people. The NID is for evidence.
Package identity is not data identity
The useful subtlety is that two different nest packages can produce exactly the same decoded data. Perhaps the only change is a view, a semantic description, a comment-like metadata field or some other authored input that does not affect event selection or decoding. The package has honestly changed, so it must have a new NID. But forcing a complete backfill merely because a view changed would be an expensive ceremony with no new information in it.
Nuthatch calculates a data identity from the inputs that actually determine decoded event rows: the
schema version, the registry hash, and the path and digest of every authored file that can change a
stored byte. What it leaves out is as deliberate as what it takes in. Views, entities and their
declarations, the named-query allowlist, the semantic description, llms.txt, the README, the
scaffolded agent skill and, since 4.9.0, any hidden file are all part of the NID and none of the
data identity, because none of them is read on the indexing path. (That last one was learnt the
hard way: a stray .cache/ made two copies of one nest compute different identities, so neither
could find the other’s mirror.) Vendored ABIs and address labels stay in, because the decode and
the exposure annotations do read them.
On mounting a new identity, or on migration, Nuthatch compares that data identity with the datasets
it already holds. If one matches, and its store records the same registry hash, it adopts that data
into the new identity instead of indexing the chain again. On the mount path the adoption is staged:
the existing dataset’s hot store and catalogue are copied into a sibling directory, the new package
material is laid over them, and the result is renamed into place, with the staging directory removed
if anything fails before that rename. nuthatch migrate does the same copy without the staging
step. In either case the existing source remains, and in a runtime the copied catalogue points into
the shared segment store, so the sealed bytes are not duplicated; only the catalogue and the hot
store are.
This is not a claim that every update is free. If an ABI, event selection, contract range or decode
rule changes, the decoded dataset may differ. That is a real data change. nuthatch migrate names
the breaking change and refuses it unless the operator passes --allow-breaking, rather than
presenting an old table with a new label and hoping no one notices. An adoption needs no such
acknowledgement, because the identity match is the proof that nothing broke.
The guarantee has to hold on the ordinary path too
That last sentence describes an intention, and for a while it was only true where somebody had gone looking for it. A packaged nest records its expected registry hash, and mounting one regenerates the registry and verifies it. A nest run the ordinary way, from a directory on your own machine, did not compare anything.
So adding an event to a running nest’s configuration and restarting produced this: the process started, observed it was already at the chain tip, indexed nothing, and served - stamping the new registry hash onto rows the old configuration had produced. Every query then reported provenance under a decode registry that had never run. The only visible symptom was an unrelated view failing to load, and only because that view happened to reference one of the new tables; a change touching tables no view mentioned said nothing at all.
It is worth being precise about why that is worse than missing rows. A content address is a claim about what produced this data. A wrong one is not an inconvenience, it is a lie in the one field whose entire job is to be checkable. And it was a lie the system told about itself, in the direction that looks healthy.
A nest now compares the decode identity recorded in its store against the one its configuration
produces, and refuses to start when they differ, naming both hashes and the remedy. The identity is
the registry hash folded with the nest’s [[calls]] and [[ipfs]] declarations, because those
also decide what gets stored. A store written before the check existed has no recorded hash;
refusing those would break every running deployment for a fault it may not have, so it adopts the
hash and logs that it was recorded, not verified - which is a different claim, and says so.
Reuse at two levels
There are two distinct wins:
- Exact package equality: two mounts with the same NID share one dataset immediately. This is what happens when two tenants mount the same nest.
- Data equality across package versions: distinct NIDs can adopt data that has the same data identity. The packages remain distinct, but the indexer avoids useless re-ingestion.
Sealed segments are immutable content, so they are the natural part of history to share, and in a runtime they already live in one content-addressed store that every dataset’s catalogue points into. The hot store remains dataset-local because it must participate in live writes and reversible reorg handling; an adoption copies it rather than sharing it. This keeps a new package from accidentally inheriting a mutable working store another mount is still writing.
Grafting, honestly
“Grafting” is sometimes used loosely to mean any form of reuse. In the current runtime, adoption is whole-dataset reuse where data identity proves compatibility. Views are query-time definitions, not independently materialised histories that can each be grafted in a different way. There is no promise today that a changed view receives a special partial backfill or that arbitrary schemas can be spliced together. Those might be useful future capabilities, but a book should say what the machine does now.
The nuthatch migrate command plans the move, can dry-run it, and refuses named breaking changes
unless the operator explicitly permits them. It is designed to be idempotent. A migration that
discovers it needs to re-index data has failed its central job.
The detailed, operational companion is nest identity and reuse and upgrading a nest. Once packages and datasets have these clean boundaries, a single process can host many of them without becoming a muddle of names.
8. The runtime
A nest is a portable description plus its dataset. A runtime is the process that makes one or more such datasets live. This is where Nuthatch separates the questions people often tie together by habit: what a package is, what an operator calls it, who has mounted it, which chain is followed, and what query surface is exposed.
The runtime reads mounts.toml. A mount records an alias, tenant, NID and SQL access policy. The
alias is the route a caller sees. The tenant is an opaque operator label. The NID selects the
identity-keyed dataset. Because those are separate, two tenants can mount one NID and share its
data while exposing different aliases or named-query allowlists.
This is multi-tenancy in the useful, modest sense: multiple independently named datasets in one runtime. It is not an identity service, quota system or billing platform. Those require knowing who the caller is, and the runtime does not. Put a gateway in front when the deployment needs that policy. The node protects its own memory and query capacity regardless of who is asking.
One cursor per chain
Within one chain, a runtime uses a shared cursor. The cursor sees the chain once, obtains the union of relevant logs and routes each log to the mounted nests that need it. This reduces duplicate RPC polling and means that finality and reorg detection have one coherent boundary. Each dataset still has isolated hot storage, package files and query context.
The rule is per chain. A multichain runtime owns a separate cursor for each chain because chains do not share a head, a finality rule or a failure mode. The runtime groups mounts by chain before spawning the cursor work. It does not make a multi-chain deployment magically one ordered stream.
The effect is worth being precise about. A runaway factory or failed decoder should quarantine the affected nest rather than corrupt its neighbour’s data. A dead cursor must not hide behind healthy HTTP routes. And an operator needs metrics that distinguish a process-level number from per-nest and per-chain health.
Capacity is a physical constraint
Nuthatch has a default 2 GB resident-set budget per active-chain cursor, max_rss_mb in
mounts.toml. It estimates the impact of a mount before accepting it, at boot and again for every
live mount, and refuses one that would push its cursor over. The actual footprint is exposed as
nuthatch_rss_bytes, which is the whole process rather than one cursor, so a two-chain runtime
reads it against the sum of its two budgets. The number is not a marketing density claim. A runtime
with two chains has two cursors and therefore two separately bounded workloads. A high-rate nest can
still be expensive even if it has very few neighbours.
Since 4.10.0 the budget is also something the process can be held to rather than hoped about. The query engine counts the memory a scan holds against a pool the operator sizes, and startup adds that pool to an ingestion reservation and a runtime headroom and refuses the configuration if the sum does not fit the wall, naming the term that counted and the largest value that would. The two new terms are measured, not guessed: the reservation is the high-water mark of the process indexing with no queries, and the headroom is the most the process has held outside the pool under the real query set. On the nest that forced the work, a copy of kittiwake’s QoS nest under its production queries, those came out at 372 MiB and up to 931 MiB, so the budget is written as a 704 MB pool, a 384 MB reservation and 960 MB of headroom, 2,048 in all; four runs under that budget inside a 2 GB cgroup peaked between 1,316 and 1,349 MiB. Before the engine counted scan memory the same nest peaked at 1,746 MiB, which is the difference between a budget and a wish.
The shared cursor pays for chain polling once. It does not make storage, decode work or analytical queries free. The runtime keeps those costs visible so density does not become a pleasant-sounding route to an out-of-memory kill.
The runtime owns the lifecycle
A runtime that could only be changed by editing a file and restarting would make every change as wide as a fault: restarting to add one nest stops all the others. So the set of mounts is changed while the process runs, through an admin API, and since 3.13.0 the runtime does the whole of it. A caller hands it a NID and a name. The runtime fetches the bundle from a registry if it does not hold it, checks that the bundle’s own manifest computes to that NID, installs it, lets it catch up and then serves it. Nothing outside the process has to unpack, verify or arrange anything.
A mount that fetches and backfills can take far longer than a caller will hold an HTTP request open, so
it is a job rather than a call. The job moves through accepted, fetching, joining and then live or
failed (or reads suspended while an operator has paused it). Every job short of live, and every
failed one, is written to mount-jobs.json beside mounts.toml, so a restart in the middle picks
it up again instead of forgetting it; a live job needs no record of its own, because mounts.toml
is that record. The job list is kept beside the runtime’s lock rather than behind it: asking how a
mount is going must never wait for the mount it is asking about.
Joining is the delicate step. A cursor advances from the slowest of its live nests, so a nest spliced in far behind would drag every neighbour back through history. The newcomer therefore catches up on its own, beside the cursor, and joins at a window boundary once it is level. The first nest on a chain has no neighbours to spare, so it starts that chain’s cursor exactly as boot would. The last nest to leave does not stop the cursor; an empty cursor idles, because the next mount onto that chain should find it there.
Pausing is a different thing from leaving. A suspended mount comes off its cursor and gives up its store, but keeps its data, its record and its name, and its routes answer that it is suspended rather than that it has vanished. Resuming is an ordinary mount of the recorded identity, which catches up from where it stopped.
Changing a name’s version is where a naive design shows a gap. A new version of a nest is a new NID, so the name has to be repointed, and a reader must see the old nest and then the new one with no error in between. The runtime mounts the new identity under a staging name and lets it catch up. It then takes the old nest off the cursor, tells the cursor to know the staged nest by the real name, and swaps the whole composition of routes in one step. Until that swap a reader is served by the old nest’s stored data; after it, by the new. The swap is a single atomic replacement of the routing table, so there is no moment at which the name is served by neither.
When one process is not enough
Scaled mode moves the hot store into Postgres and puts a control plane beside it. The control plane
holds what the fleet should run and which workers exist; it does not hold ownership. Ownership is
a lease kept in the chain’s own hot-store schema, next to the data it protects, and a worker claims
it and fences its writes with a monotonically increasing value. A worker may only write while it
holds the current fence. If it dies, another worker can take the lease and continue, and since
4.11.0 a worker that finds its lease taken stops that cursor’s nests on the same tick. The fence
prevents a late former owner from writing as though nothing happened, which is the small but
crucial detail separating failover from two machines cheerfully scribbling over the same state.
None of this is in the default binary: scaled mode is a separate build, --features postgres-store, published as the nuthatch-scaled Linux tarball, because the embedded binary
carries no database driver and that is non-negotiable.
The data model remains recognisable: mounts identify datasets, cursors follow chains, hot data is reversible and sealed data is durable. Scaled mode changes who is allowed to perform the cursor work, not what the chain data means.
For configuration and practical commands, see run many nests, host nests for others and scaled mode. The final chapter is about operating the whole arrangement without taking a green health endpoint as proof of truth.
9. Operating a truthful index
An indexer can be available and wrong. It can answer HTTP requests while following the wrong chain, serve a stale cursor, decode against an unsuitable ABI, or retain a view whose meaning nobody has checked since the first enthusiastic afternoon. Availability matters. It is not the whole definition of health.
Nuthatch is designed to make its claims checkable. The remaining work is operational discipline: choose what to verify, give failures a route to a human, and avoid calling a system healthy merely because it remains capable of returning JSON.
Verify the inputs and the result
Start with the nest package. Inspect the contract addresses, deployment ranges, vendored ABIs and event selections. Rebuild the generated schema and decode registry from those inputs. Confirm the NID when loading or deploying a bundle. This establishes that the machine is indexing the package you intended, rather than an equally well-formatted stranger.
Then verify a result against the chain at a fixed watermark. A useful check has a known block range, a concrete expected answer and a way to replay the SQL. For an event-derived view, compare its rows with the events and, where appropriate, with an independent chain query. For an application fallback, exercise the exact named query and response shape the application will use. A test that only proves that an endpoint returned 200 has the emotional comfort of a fire alarm with its batteries removed.
Nuthatch’s own releases are held to the same rule, and the way they are is worth copying. Since
4.1.1 reached a production nest and refused its dashboard’s join-heavy views within minutes, with
every CI gate green, a release candidate is served over a copy of that nest, under the budget the
nest runs on, and sent the 74 statements its consumers actually issue. Each answer is compared with
the last release’s by digest, not only by status, because 4.3.0 once passed every statement and
served NULL for a column that had been populated the day before. A refusal, an error, an
out-of-memory or a time past a stated bound turns the candidate red, and red stops the roll to
production. The script is scripts/release-gate.sh in the repository, and the query set lives with
the consumer that sends it. The habit generalises: gate on the answers your readers depend on, at
the budget they are served under, and compare them with what was served yesterday.
A nest seeded from someone else’s mirror deserves one more question. The seed checks every file against the hash the catalogue names, and refuses a catalogue that overlaps itself or carries a provisional segment, so what arrived is what was published. Whether what was published is true to the chain is not something a hash can say. Spot-check a seeded range against an independent RPC before the nest serves anything a reader will act on.
Verify the endpoint, and then verify it again later
The package is not the only input. The RPC endpoint is one too, and it is the one that changes without telling you.
nuthatch doctor --rpc <url> asks an endpoint three questions before a backfill trusts it: the
widest eth_getLogs range it will serve, the largest JSON-RPC batch it accepts, and whether it has
archive depth. Each of those limits otherwise surfaces mid-backfill as a retry loop that looks
exactly like slowness, which is the worst way to learn it. Run it as nuthatch doctor --dir <nest>,
without --rpc, and it probes the nest’s own endpoints filtered to every address the nest declares,
rather than an empty filter or the first contract in the file. With an explicit --rpc and no
--address, doctor samples the endpoint for its busiest address and re-probes with that, so the
figure it recommends reflects a result-count limit as well as a range limit; pass --address when
you would rather it probed with yours.
It prints two window figures, and they mean different things. up to N blocks is what the endpoint
served. recommend --window M is what to set: half the measured limit from an address-filtered
probe. Only when doctor could find no address to probe with does it fall back to a range-only
figure, capped at 320 and labelled as such, because a probe without an address cannot see the
result-count limit a real nest meets. Read the recommendation as a starting point that holds, and
re-probe with --address for a figure that reflects both limits.
The part worth building a habit around is the second probe. Nuthatch ships measured endpoints for its built-in chains, and one of them was measured on a Tuesday and had silently lost archive depth by the Wednesday - a from-deployment backfill could no longer use it at all, while the recorded figure in the source still said otherwise. A recorded measurement is a snapshot presented as a property. Nothing about an endpoint’s past behaviour is a promise, including ours, so probe before a long backfill rather than trusting a number somebody wrote down once.
Observe the pipeline, not just the server
Metrics make the stages visible. For a single nest, nuthatch_tip_lag_blocks tells you whether the
cursor keeps up and nuthatch_last_poll_unixtime reveals a poller that has frozen. A runtime hosting
several datasets reports per nest instead: nuthatch_nest_tip_lag_blocks{nest="…"}, with each nest’s
poll time and stall state on its own /<name>/ready. Per-nest health tells you whether a
part of a shared runtime has been quarantined. nuthatch_cursor_live identifies the chain cursor
that has died even if another chain in the same process remains busy. RSS and query rejection
counters show whether the node is protecting itself as intended.
Alert on sustained lag, a stalled poller, quarantined nests and approaching memory limits. Treat the first three as a service issue and the last as an opportunity to reduce query concurrency or reconsider the mounting budget before the machine makes the decision rather more abruptly. The metrics guide contains runnable Prometheus examples.
Secure the separate surfaces
The data API is read-only, but a runtime may also expose administrative routes that mount or remove
nests, on the same port. The binary refuses to mount them off loopback unless NUTHATCH_ADMIN_TOKEN
is set, and then checks the token on every admin request; --no-admin leaves them out entirely.
Do not undo that care by putting a token-bearing port on the public internet and then feel
surprised when somebody experiments with it. Restrict the port, use a gateway where public access is
needed, and keep database credentials and RPC URLs in deployment configuration rather than the nest
package.
SQL needs its own care. General SQL is powerful enough to consume resources even when it cannot write. Use named queries or an allowlist for public services. Keep the node’s built-in timeout, concurrency and result bounds on. Per-caller rates and quotas require authenticated identity, so place them at the gateway. Calling a local concurrency semaphore a rate limiter would be a category error with a pleasingly dangerous outcome.
Recover without inventing history
When something fails, first establish which layer is unhealthy: RPC reachability, cursor progress, one dataset’s decoder, storage, query load or the gateway in front. The runtime is built so a quarantined nest, a lease handover and a failed query are observable conditions rather than silent reasons to return old data forever.
Do not repair a suspected data error by editing sealed rows or retroactively decoding history under a new ABI. Preserve the evidence, identify the package and range involved, build the corrected dataset under its own identity and migrate consumers deliberately. The cost of this restraint is small compared with explaining an untraceable historical rewrite later.
That is Nuthatch’s central ethic. It is not merely a fast way of turning logs into tables. It is a way of retaining a clear chain of custody from authored definition to chain event to query result. Once that chain of custody exists, a team can operate its own critical data path without asking a remote endpoint to be both available and believed.
Use production operation, security and verification as the practical checklist alongside this chapter.
Appendix A. One log, end to end
The architecture becomes less mysterious when reduced to one ordinary event. Take a real one: USDC
emitted a Transfer log at index 663 of block 26,128,681 on Ethereum mainnet, 135 USDC from
0x58df…47af to 0x64cd…9a53, at 20:52:47 UTC on 2026-10-05. The transaction succeeded, the
receipt holds the log, and the chain’s RPC can return it through eth_getLogs. What must happen
before a reader can safely ask Nuthatch for that transfer?
The short answer is fetch, route, decode, order, persist, then eventually seal. The longer answer
is worth knowing because each verb protects a different invariant. The figures below come from
running the released 4.10.1 binary on a laptop against two public mainnet endpoints, with a nest
scaffolded by nuthatch init for the USDC contract from 120 blocks behind the tip.
1. Fetch only a defined universe
The cursor has a registry built from the authored package. It knows the contract addresses and the event topic hashes that it is prepared to decode. For a solo nest it asks the source for logs inside a bounded block window. In a runtime with several mounts on the same chain, the cursor asks once for the union of those filters and routes a returned log to every live nest whose registry matches.
The RPC response is not yet application data. It is untrusted transport input. It may arrive out of order, it may include logs useful to another mounted nest, and an endpoint may reject a window or return a response too large for its local limits. The cursor’s retry, window-sizing and concurrency policy exist here. They deal with the ordinary weather of RPC infrastructure before a row reaches durable storage.
An empty result still matters. The cursor must advance over a block range with no selected events, otherwise it would return to the same quiet range forever. Progress is about blocks successfully accounted for, not merely rows received.
Where the nest declares [extract] l1_blocks, the window also fetches the header of every block
that produced a row, so the L1 block it settled against is recorded beside the logs. That is the
one case in which the fetch step reads more than eth_getLogs.
As run: init resolved USDC’s ABI through its proxy via Sourcify and wrote the nest in 4.4 s. dev
reported finality Depth(64), window 20, had the API live within 0.2 s of starting, and announced
cold start: backfilling from block 26128556 (tip 26128677). Eight seconds later it was caught up to tip at block 26128678 - 17881 events in 8s, 2321 ev/s, seventeen tables’ worth of rows across
122 blocks, with the adaptive window reporting 22 blocks by the end. Eighteen seconds from an empty
directory to a live, caught-up index is what the first-run promise looks like on a quiet afternoon.
2. Route and decode with a pinned registry
For the transfer log, the registry looks at the emitting address and topic zero, identifies the
Transfer(address,address,uint256) decoder from the vendored ABI, and decodes indexed and data
fields into a typed row. It also supplies the implicit chain columns: block number, block hash,
block timestamp (on unless the nest was scaffolded with --no-timestamps), transaction hash, log
index, address and a sequence number. There is no transaction index column; order within a block is
the log index, which is block-wide. The row for our transfer, read back from
/table/fiat_token_v2_2__transfer, carries exactly those: block_number 26128681, block_hash
0x627b…5b50, block_timestamp 1791233567, tx_hash 0xaca1…10c8, log_index 663, address
the USDC contract, and then from, to and value as the ABI names them.
No schema inference happens at this point. The table layout was generated from the nest’s inputs. An unknown event is not invited to become a new column because it happened to look interesting. Likewise, a malformed log does not earn a creative interpretation. The point of a pinned registry is that the mapping from byte sequence to row is reviewable and repeatable.
Factories add one wrinkle. A factory event can discover child contracts which should be indexed thereafter. On the ordinary path, discovery happens inside the one decode pass, log by log in chain order, so a child created earlier in a window is known by the time its first log is reached. The direct-seal backfill, which fetches a window by topic alone, runs a discovery pass over the window first and discards its rows, then performs the authoritative decode with the children known, so a child discovered in the window is already known when its own logs are decoded. The final row path remains the same either way. Discovery is not allowed to become a second, differently behaving decoder.
3. Establish a canonical local order
Providers are not entitled to return equivalent log sets in the same order. Nuthatch therefore sorts decoded rows by chain coordinates before writing or sealing. The ordering is block number, then log index, and the sealer breaks a tie on the row’s canonical bytes; there is no transaction key, because the log index already orders a whole block. This is not cosmetic. A view calculating a balance, a factory registry or a content-addressed segment must behave the same when two RPC endpoints return the same chain facts in different sequence.
This is also why concurrent backfill needs care. Fetches may happen in parallel, but their results must enter the sealing path in deterministic block order. Otherwise speed has changed the bytes of historical storage, and a supposedly content-addressed result has become dependent on timing. That would be a rather expensive way of saving a few seconds.
4. Commit to the hot store
For an unfinalised block, the sorted row is written to the dataset’s hot store alongside the cursor checkpoint. The write establishes two things together, in one transaction: the window’s rows, the hash of its last block and the new last-block mark all land or none do. The transfer is then visible to live reads, and the cursor has evidence of which chain block it believes it processed. A later reorg check compares that evidence with the chain’s current answer.
Derived views and incremental state receive the same event in this phase. Their update is part of the reversible hot path. If the log later belongs to a discarded branch, the rollback replays its effect with the opposite weight or rebuilds the affected hot state. A derived answer is therefore not permitted to outlive the raw event it depends upon.
5. Seal after finality
Once block 26,128,681 sits 64 blocks behind the tip it may be sealed, but it is not sealed yet.
The sealer cuts a segment when the finalised rows held reach 20,000 or 64 MiB, or when they span
the chain’s seal span of 1,800 blocks; until then they stay in the hot store, with a checkpoint
pinned at the finalised ceiling so a later reorg walk cannot step past the watermark. In the run,
the nest had 14,403 transfer rows across 126 blocks and had sealed nothing, which /ready reported
as sealed_through: 0 beside last_block: 26128681.
When the cut comes, Nuthatch serialises the rows into a content-addressed Parquet segment, written
to a temporary name, fsynced and renamed into place, and only then rewrites the catalogue and swaps
it in with one rename. The hash is computed from the bytes before they are written; re-hashing
the file is the job of the startup integrity pass, not of the seal. The same hot-store transaction
that advances the sealed watermark prunes the matching hot rows. The event is no longer subject to
ordinary reorg rollback. On a fork of mainnet where the clock could be advanced, the same nest
configuration reached its span cut at block 26,130,493 and wrote
fiat_token_v2_2__transfer-2d2d37dd….parquet, 1,800 blocks of transfers in 32,940 bytes, then
advanced the watermark on through the empty finalised tail to 26,130,693.
--seal-direct uses the same sealed representation for old, already-final history, bypassing the
hot store during an initial bulk backfill. It still resolves every declared [[calls]] input at its
pinned block before it commits the segment. It is an optimisation with an important precondition:
the range must be past finality. It does not use a fast path to omit declared inputs or declare
recent, reversible blocks permanent.
6. Answer a query with provenance
When a reader asks for the transfer or runs SQL over its table, the serving layer reads sealed
segments and the hot tail as one logical surface. The response carries watermarks and source
information so a caller can tell how current the answer is and whether the hot contribution was
available. A count over the table in the run came back with provenance naming as_of 26128681,
sealed_through 0, source hot+sealed, the registry hash 0xce63…f128 and the nest’s NID, and
with tip_unavailable: false. A hot-store failure must not quietly make a query look complete
while returning only sealed history. A hot scan that fails answers from sealed rows alone with
tip_unavailable: true, which is the field to read, and since 4.11.0 source agrees with it and
says sealed; up to 4.10.1 it still said hot+sealed (nuthatch #1935). degraded stays false in
that case, because it reports incomplete sealed history, not a missing tip.
That final detail is representative of the design. A Nuthatch answer is useful not only because it contains a number, but because it can say which package decoded it, how far the cursor had reached and what portion of history was sealed when the answer was formed.
Appendix B. A reorganisation, walked through
Reorganisations are where an indexer either demonstrates that it understands a blockchain or quietly begins preserving fiction. The normal case is recoverable precisely because Nuthatch keeps the recent part of history hot and reversible.
Consider a cursor that has processed blocks 100 through 110. Its finality policy has sealed through block 104. Blocks 105 through 110 remain in the hot store, along with block-hash checkpoints. A transfer in block 108 has incremented a derived balance view. At this instant that is a valid, useful result, but it is not yet final.
The chain changes its mind
On the next poll, the cursor obtains a head which does not agree with its recorded checkpoint for block 110. It walks backwards through the known checkpoints and the source’s current block hashes until it finds the deepest common ancestor. In this example, block 106 still agrees; blocks 107 to 110 were replaced. The ancestor is 106.
The cursor does not try to patch individual logs based on a hunch. It performs a rollback to the ancestor. The hot store deletes rows above 106, resets its last-block metadata to 106 and removes the checkpoints that belonged to the old branch. The derived balance view retracts the effect of the former transfer in block 108. Factory child discovery that occurred only on the discarded branch is rolled back as well. The cursor then starts again at 107 and indexes the canonical replacement blocks in the normal path.
The application may have observed a provisional balance before the rollback. That is inherent to serving near-tip data. What Nuthatch promises is convergence: once it notices the reorg, neither the raw table nor a maintained derivation may retain the old branch.
That is the example. Here is the same thing run on the released 4.10.1 binary, on 2026-10-05,
against a local fork of mainnet with a one-second poll. Three USDC transfers to the burn address
were sent in blocks 26,128,700, 26,128,703 and 26,128,706, and the built-in balance circuit showed
the burn address at 8,000,000 units. The fork was then reorganised five blocks deep, replacing
26,128,705 onward with empty blocks, which discarded the third transfer. The tip came back at the
same height, so no poll saw it move, and the rollback arrived on the idle re-check 9.8 seconds
later: reorg detected: rolled back to block 26128702 (removed 2 entities). Two things in that line
are worth a second look. The ancestor is 26,128,702, not the true fork point of 26,128,704, because
the hot store keeps one checkpoint per committed window and the walk stops at the deepest stored
checkpoint that still matches; the cost is re-fetching two blocks that had not changed. And the
rollback removed two rows, the discarded transfer and the surviving one in 26,128,703, then
re-indexed forward and put the survivor back. The balance read 3,000,000 afterwards,
nuthatch_reorgs_total read 1, and the discarded row was gone from the raw table.
Shared cursor, many datasets
Now place three nests on the same chain cursor. The cursor detects the hash disagreement once at its shared boundary, then fans the rollback out to every live dataset. Each dataset may have a different sealed watermark because it may have been mounted at a different time or progressed differently. A nest already at or below the ancestor does nothing. A nest whose hot range contains the discarded blocks retracts them. If one nest cannot roll back, it is quarantined rather than making the other datasets lie about their state.
This is why a shared cursor does not require shared mutable tables. Fetching and reorg detection are shared chain work. Row ownership, store mutation and local failure handling remain per dataset. The structure is slightly more machinery than one giant database, but much less machinery than trying to explain which tenant’s rows were inadvertently removed by a global repair.
The finality line
Return to the example. A reorg to ancestor 106 is repairable because the seal watermark was 104. Every affected block is still in the reversible hot layer. But imagine the source reports an ancestor of 102. Blocks 103 and 104 have already been sealed as immutable history. Deleting only the hot rows above 104 would leave sealed data from the discarded branch in the query surface.
Nuthatch refuses this condition. It reports a finality violation and halts the affected index rather
than silently presenting a half-correct history. A single nest run with nuthatch dev exits; in a
runtime the nest is quarantined as a terminal fault, named on /nests with its reason, and its
siblings on the same cursor carry on. This is not a graceful recovery in the marketing sense. It is
the only honest behaviour once the external finality assumption has been violated. The operator
must investigate the chain source, finality configuration and recovery procedure instead of
allowing a plausible but inconsistent index to keep serving.
On the same fork, the same nest was left to seal: with the clock advanced 2,048 blocks it cut a
segment spanning 1,800 blocks and advanced its watermark to 26,130,695. The fork was then
reorganised 80 blocks deep, to an ancestor of 26,130,679, with one replacement transaction so the
new branch genuinely differed, and one block was mined on top so the next poll would see the tip
move. 0.65 seconds after the reorganisation was issued the log read no checkpoint at or below block 26130759 is canonical, because each window commit prunes the checkpoints below the newest one
at or below the sealed watermark, and then: a fork deeper than every checkpoint this nest holds is below the sealed/finalized watermark 26130695 - a finality violation this indexer cannot repair; halting. Raise the chain's finality depth. The process exited with that line. Nothing was deleted, nothing
was rewritten, and the catalogue still names the segment the abandoned branch produced, which is
what the operator will need when deciding what to do next.
Detection is not instantaneous. The cursor learns that the chain changed when a reorg check runs, and until then it answers from what it last indexed, as it does for an ordinary reorg near the tip. A check runs when a poll sees the tip move; a fork that keeps the tip height is re-checked at most every twelve seconds while idle, and a failing RPC delays it further, so there is no fixed bound. The two runs above show both ends of that: 0.65 seconds when the tip moved, 9.8 seconds when it did not. Measured once before, on 3.13.3 against a forked chain at a one-second poll interval, the sealed rows of the abandoned branch were served for about half a second before the halt. No read-time setting removes the window, because below the seal the discarded rows are the sealed ones. A finality depth the chain honours is the only protection.
The same line applies to direct sealing. That fast backfill path only processes a range already behind the finality boundary. Its performance comes from avoiding hot writes, not from relaxing the definition of permanent history.
What to look for in practice
nuthatch_reorgs_total records ordinary detected reorgs. A nest’s /ready exposes the last indexed
and sealed watermarks (last_block, sealed_through); in a runtime that is /<name>/ready, while
the root /ready says only whether the whole is ready and which nests are quarantined or stalled.
/health is bare liveness and says only ok. A sudden tip lag accompanied by reorg growth calls
for a look at the source and chain conditions. A finality violation is a hard incident, not a
counter to wave away.
The important operational habit is to distinguish “we are behind” from “we are wrong”. Being behind can often be fixed by an RPC change or smaller windows. A reorg below seal means the system has lost the ability to correct historical facts automatically, and it should be treated accordingly.
Appendix C. A lease handover, walked through
Running a cursor on one machine is simple because there is one process that can write its hot state. Running it across machines introduces a more awkward possibility: the first worker is slow or partitioned, the control plane gives the work to a second worker, and then the first worker comes back convinced it still owns the chain. Without a fencing rule, both can write. This is how a failover exercise becomes a data-corruption exercise with better branding.
Scaled mode assigns cursor work through a Postgres-backed control plane, which holds what the fleet
should run and which workers exist. It does not hold ownership. The lease for a chain lives in that
chain’s own hot-store schema, as three rows of its meta table: lease_owner, lease_expires_at
and owner_fence, a monotonically increasing number that identifies one particular period of
ownership. Putting the lease next to the data it protects is what lets every write check it in the
same transaction.
This appendix is read against the 4.11.0 source and dated rather than re-run. Scaled mode is a
separate build, --features postgres-store, shipped as the nuthatch-scaled Linux tarball; the
default binary answers nuthatch worker with a refusal that says so. The two-machine run it
describes is the one that accompanied RFC-0022.
The normal sequence
Worker A starts, registers with the control plane and claims the lease for Arbitrum. The claim locks the fence row, reads 40, writes 41 and records A as the holder, all in one transaction. A runs the Nuthatch cursor, advances through blocks and renews its lease every five-second tick, for a thirty-second term. Every write it makes, a window commit, a rollback, a seal watermark, opens a transaction that first re-reads the fence and refuses to continue unless it still reads 41.
At this point there is one writer. Nobody infers that from a worker’s optimism. The chain’s store has a current lease record which says so.
A fails, B takes over
Suppose Worker A is killed, loses network access, or otherwise stops renewing. Once its lease expires, Worker B can claim the same chain. The claim records B as holder and increments the fence to 42. B now starts the cursor under fence 42. (The fence counts acquisitions, not changes of owner: a worker that lets its own lease lapse and claims it again also moves the number on, which is correct, since its old transactions are just as stale.)
The increment is not an ornament for logs. It distinguishes this ownership epoch from A’s old one. Any write that still arrives from A carrying fence 41 is refused at the fence check with a lost-ownership error, including A’s attempt to renew. Worker A cannot resume and extend its former claim merely because its process remained alive long enough to regain connectivity. The lease has moved on.
A worker should also stop its local ingestion task when it learns that it has lost the lease, to save the work and narrow the time during which it attempts stale operations. Since 4.11.0 it does: a renewal the store refuses because the fence has moved is recorded as a lost cursor rather than a failed tick, the worker logs that another holder has it, and stops that cursor’s nests on the same tick while its other cursors carry on. Up to 4.10.1 the ingestion task ran on, polling the RPC and decoding windows whose every commit the fence then refused, until the worker was restarted (nuthatch #1934). The data was safe even then, which is the point of the next sentence: the fence is necessary because process shutdown and network delivery are not atomic events. It is the backstop that makes delayed messages, slow death and a worker that has not yet noticed harmless.
Control plane outage is not automatic eviction
There is a separate failure worth naming. If the control plane is unavailable but the worker’s data store is, the cursor keeps indexing. There is no named policy for this; it is what the code does. A tick heartbeats the control plane before it renews any lease, so while the control plane is down no lease is renewed and the term runs out after thirty seconds, but an expired lease is not a taken one. Nobody else can claim it either, the fence is unchanged, and every write still passes its check. A control-plane outage is not evidence that a second worker has taken ownership. In the two-machine test, workers continued processing during a deliberate control-plane outage while Postgres remained available, then reconciled when the service returned.
The boundary is always ownership, not mere connectivity. If the worker can no longer establish that its lease is current, it must not keep acting as the sole authority by force of habit, and the fence sees to it that it cannot. Conversely, if the control-plane service is briefly unavailable but the lease record in the store still protects the writer, stopping the cursor needlessly would turn a control-plane incident into a data availability incident.
What this buys the operator
With leases and fences, failover is observable and testable. The control plane’s worker roster
shows registrations. The chain store’s meta rows show the holder and fence. A deliberate handover
should move the holder and increment owner_fence. A healthy test is not “two workers exist”. It is
“the old holder stopped being allowed to write, the new holder took over, and indexing continued
without two authorities claiming the same cursor.”
Scaled mode does not alter the event model, decode registry or sealing rules described elsewhere in this book. It gives those same rules a single writer across machines. The real achievement is not that a worker can restart. It is that the system can tell the difference between a restart and two writers, which is where the difficult bits live.