A backfill is the expensive part of running a nest: a million RPC credits for the larger Graph
Protocol nests, every one of them buying history somebody has already indexed. The public mirror
holds that history. nuthatch seed downloads it, checks every file against its
content address, and leaves your nest following the chain from the block where the mirror ends.
1 · the nest
Clone the nest and check out the commit it was published from. The mirror is found by the nest's data identity, so any other version finds nothing.
2 · the history
Seed downloads the sealed segments, refuses any file that does not hash to its address, and resumes if interrupted.
3 · the tip
Dev follows the chain from the next block, over your own RPC. Catching up a day costs a few hundred calls.
That every file is the one the mirror's catalogue names, so nothing changed on the way or at
rest. They do not prove who wrote that catalogue, or that it is true to the chain. Seed from a mirror run by someone you
would trust to run the nest for you, or re-index a sample of it yourself and compare.
Publish your own
Any nest can be a mirror. Run it with --publish-target s3://your-bucket/prefix, or
run nuthatch publish sync on a stopped one, and point people at the bucket's public
address. The mirror is append-only and costs nothing to read from a bucket that does not charge
for downloads. The operator guide has the bucket policy.
Read it without Nuthatch
The mirror is plain Parquet, so DuckDB, Trino or anything else that reads Parquet over HTTPS can
query it directly. A public bucket cannot be listed: each dataset's
manifest.json names its tables and, for every segment, a hash, and the
file is at <dataset>/<table>/<hash>.parquet. The manifest's
file field is the name inside a nest directory, not the address here.