For AI agents: the documentation index is at /llms.txt. Markdown versions of pages are available by appending .md to the URL.
Skip to main content

Production Indexer Reliability and What Happens When an Indexer Fails

Author:Jordyn LaurierJordyn Laurier··12 min read
Reviewed by:Dmitry Zakharov
Production Indexer Reliability with HyperIndex
TL;DR
  • When the node or data source behind an indexer fails, a well-built indexer switches to another source and keeps going. In HyperIndex, a fallback RPC takes over when the primary source stops returning new blocks or keeps failing, and the indexer tries the primary again 60 seconds later.
  • When the indexer process itself fails, nothing already saved is lost. HyperIndex tracks its progress in the database, so a restart resumes from the last saved point, not the start block. Each batch is written in a single database transaction, so there's never half-written data.
  • An unhandled error in a handler stops the indexer. It doesn't skip the event or write half a batch. You fix the handler and restart, and it resumes.
  • Reorgs and missing data are separate failures with separate protections. Reorg rollback is on by default for all EVM chains and Solana, and HyperSync checks blocks against their headers as it ingests them.
  • Securing indexed data means making sure it matches the chain and controlling who can query it. Envio Cloud paid plans add alerts, API keys and IP whitelisting.
  • HyperIndex passes every check in the reliability scenarios of the open indexer benchmark, from database restarts to deep reorgs and failing RPC nodes.
  • Want your coding agent to harden an existing indexer? Jump to the agent prompt.

A blockchain indexer depends on three things staying healthy. The node or data source it reads from, its own process, and the chain itself. Any of them can fail, and what matters is what happens to your data when they do. Each section below covers one kind of failure, what HyperIndex does about it, and what's left to you.

What Happens When the Node Behind an Indexer Fails​

Every indexer reads blocks and events from somewhere. For HyperIndex on supported EVM chains that's HyperSync by default, and on other chains it's an RPC endpoint you configure. Nodes go down, providers rate limit, and endpoints stall.

HyperIndex retries failed requests with backoff, and you can tune the timeouts and backoff for each RPC entry in the RPC reference. Retries only help if the source comes back, though. For production, the docs recommend adding a fallback RPC.

chains:
- id: 137 # Polygon
start_block: 0
rpc:
- url: https://polygon.your-rpc-provider.com
for: fallback
# contracts: ...

The fallback takes over when the primary source stops returning new blocks, after 20 seconds once the indexer is caught up to the chain head, or when requests to it keep failing. After switching, HyperIndex tries the primary again 60 seconds later, so you don't need to restart anything to get back onto your preferred source. One trade-off to know is that data read through a fallback RPC isn't checked against block headers the way HyperSync data is. The setup is covered in improving resilience with RPC fallback, and the for: sync, for: realtime and for: fallback roles are in the configuration file guide.

What Happens When the Indexer Process Fails​

Hosts restart, containers get rescheduled, and handlers throw. What matters is the state the indexer is in afterwards.

HyperIndex tracks its progress in the database. When the process restarts without the -r flag, it resumes from the last saved point instead of re-indexing from the start block. Each batch of entity changes is written in a single database transaction, so a crash mid-write never leaves half a batch behind. Dynamically registered contracts are stored in the database too, so the indexer doesn't forget contracts it discovered at runtime.

An unhandled exception in a handler stops the indexer. That's deliberate. Skipping the event would leave your data quietly wrong, and stopping leaves it correct up to the last saved point. You fix the handler, restart, and indexing continues from there.

When you self-host, restarting is your process manager's job. The indexer serves a /healthz endpoint, on port 9898 by default, that returns 200 once it's up, which works as a Kubernetes liveness probe (health checks).

Where Reorg Rollback Fits​

A reorg isn't a failure of your indexer. The chain replaces some recent blocks, and anything you indexed from the old ones is no longer true. HyperIndex keeps a history of entity changes for blocks that could still be reorged, and rolls back to the point where the chains diverged when it detects one. It's on by default for all EVM chains and Solana, and needs no code in your handlers. On EVM chains, how far back it tracks is set per chain by max_reorg_depth. On Arbitrum One and Nova, OP Mainnet and some testnets the default depth is 0, so set it yourself if you want rollback there.

We cover the settings, the depth limits and where rollback stops in Indexing and Reorgs, so we won't repeat them here.

Missing Data Is a Different Failure​

There's a quieter failure that neither retries nor reorg handling will catch. A node can return an eth_getLogs response with logs missing and no error. Nothing about the chain changed, and the response looks fine, but one missing transfer is enough to make a balance wrong.

HyperSync ingests whole blocks, so it can recompute each block's transactions and receipts roots and compare them with the block header as it ingests them. When a check fails, it refetches from a different source. On chains where a transaction type can't be rebuilt yet, the root checks wait until it can. The checks are built to catch faulty sources, not a source deliberately building a false but consistent block. The documented cases and how the checks work are in Nodes Silently Miss Events.

Testing It Against Real Failures​

Claims about reliability are easy to make and hard to check, so we test them. The open indexer benchmark has a reliability suite that breaks things on purpose and then checks what ended up in the database. It runs on a generated chain, so every failure happens the same way every run, and a check only passes if it passes every run.

Its scenarios include these.

  • The database restarts or freezes mid-sync.
  • The indexer is killed while a batch is being written, then started again.
  • The chain rewrites itself, including a reorg that happens while the indexer is offline, several reorgs in a row, and one deeper than the indexer can undo.
  • The RPC node errors, stalls, rate limits, caps block ranges or briefly reports an older head.
  • Legal but awkward values, like a token with an empty symbol or a transfer of the largest possible amount.

HyperIndex passes every check. In the database restart scenario it exits and needs restarting, which is the job of your process manager as covered above, and it then finishes with the data correct. The suite measures the RPC path, because the benchmark can't make HyperSync fail on demand. The results and what each check means are in the reliability README.

Knowing Something Went Wrong Before Your Users Do​

Self-hosted, HyperIndex exposes Prometheus metrics at /metrics, on port 9898 by default. The metric names are covered by semver, so dashboards built on them stay stable within a major release. A few worth alerting on are envio_progress_ready for whether the indexer is caught up to the chain head, envio_processing_stalled_on_fetch_seconds for time spent waiting on fetches during sync, and envio_reorg_detected_total for reorgs. The full list is in the observability docs.

On Envio Cloud, paid plans include built-in alerts. You're told when your production endpoint goes down, when the indexer stops processing blocks, and when it restarts or logs errors. Alerts go to Discord, Slack, Telegram, email or a webhook you can point at tools like incident.io or PagerDuty.

Updating an Indexer Without Downtime​

Deploying a new version is its own risk. On Envio Cloud, every push to your deployment branch starts a new deployment that re-indexes from the start block, while the previous one keeps serving queries until the new one is fully synced.

Production plans give each indexer a static endpoint. When the new deployment is ready, you promote it to production and the endpoint switches to it straight away. If something looks off, you can switch back to a previous deployment the same way. Your app keeps calling the same URL throughout. The details are in zero-downtime deployments.

Keeping Indexed Data Correct and Private​

Data on the chain is public and protected by the network's consensus. Nobody can quietly change a finalised transaction. An indexer takes that data and stores a copy in a database, shaped for your app. So the security question has two parts. Is the copy correct, and who can read it?

Correctness comes from everything above. A data source whose data is checked, reorg rollback and a process that stops rather than skipping bad events. The copy can also be rebuilt from the chain at any time, because re-indexing from the start block produces it again, as long as your handlers only use chain data.

Access is about the endpoint your app queries. On Envio Cloud paid plans you can require an API key, restrict requests to known IP addresses, or both (security). Requests without a valid key are rejected.

The practices we'd follow for any indexer are these.

  • Read from a source you can verify, and add a fallback so one provider's outage doesn't stop you.
  • Leave reorg rollback on, and raise max_reorg_depth on chains with a history of deep reorgs.
  • Keep side effects out of handlers. API calls and files written outside the database aren't rolled back on a reorg.
  • Let the indexer stop on errors instead of catching and ignoring them.
  • Alert on progress and restarts, not just on the endpoint being up.
  • Restrict who can query the endpoint if the data or your query budget matters.

Prompt for Your Coding Agent​

If you already have a HyperIndex indexer, you can hand this to Claude Code, Cursor or any coding agent to check it against this blog.

Copy the prompt
Review my Envio HyperIndex indexer for production
reliability. Use https://docs.envio.dev as the source
of truth and don't invent config keys.

1. In config.yaml, check each chain has a fallback
RPC (an entry with `for: fallback`, or with no
`for` on HyperSync chains). Add one where missing,
reading the URL from an environment variable.
Docs: /docs/HyperIndex/rpc-sync
2. Check `rollback_on_reorg` isn't set to false, and
suggest a `max_reorg_depth` for chains known for
deep reorgs, and for chains where it defaults
to 0 (Arbitrum, OP Mainnet).
Docs: /docs/HyperIndex/reorgs-support
3. Flag handler code that calls external APIs or
writes outside the database, since reorg rollback
won't undo it.
4. Flag any try/catch in handlers that swallows
errors.
5. List the metrics I should alert on and show a
Prometheus scrape config for port 9898.
Docs: /docs/HyperIndex/observability

Show me the diff before changing anything.

Frequently Asked Questions​

What Happens If the Node Behind an Indexer Fails?​

The indexer loses its data source until it retries successfully or switches to another one. In Envio HyperIndex, failed requests are retried with backoff. If you configure a fallback RPC, the indexer switches to it when the primary stops returning new blocks or keeps failing, then tries the primary again 60 seconds later. Data already indexed is kept, and indexing resumes from where it stopped.

How Secure Is Data Stored Using Blockchain Indexing Methods?​

Securing indexed data depends on two things. The data has to match the chain, which means handling reorgs, reading from a source whose data can be verified, and stopping on errors rather than skipping events. And access to the query endpoint has to be controlled. The data can be rebuilt from the chain by re-indexing, as long as your handlers only use chain data. On Envio Cloud paid plans you can protect an indexer's endpoint with API keys, IP whitelisting or both.

What Are Best Practices for Securing Onchain Data?​

For data you index, read from a verified source with a fallback, keep reorg rollback on, keep side effects out of event handlers, let the indexer stop on unhandled errors, alert on progress and restarts, and restrict who can query your endpoint. For data stored onchain, remember it's public. Anything written to a public chain can be read by anyone, so sensitive data shouldn't go there unencrypted.

How Secure Is the Data in Onchain Storage?​

Onchain data is protected against tampering by the network's consensus. Once a block is finalised, its transactions can't be changed without the network agreeing. It's not private, though. Anyone can read it, which is also what lets an indexer copy and organise it. Recent blocks can still be replaced by a reorg until they're finalised, which is why indexers need rollback.

What Common Challenges Arise in Blockchain Indexing Technologies?​

The ones covered in this blog are data source outages, reorgs, incomplete data from nodes, and indexers stopping without anyone noticing. For the wider list, including slow syncs and multichain complexity, see Blockchain Indexing Challenges and How to Solve Them.

Does HyperIndex Lose Data When It Restarts?​

No. HyperIndex tracks its progress in the database and restores it on restart, including dynamically registered contracts. Without the -r flag, it resumes from the last saved point instead of the start block. Each batch is written in a single database transaction, so there's no half-written data to clean up.

Build With Envio​

Envio is the fastest independently benchmarked EVM blockchain indexer for querying real-time and historical data. If you need an indexer that keeps going when its sources don't, start with the HyperIndex quickstart, deploy it on Envio Cloud for alerts and zero-downtime updates, and come talk to us about your data needs.

Stay tuned for more updates by subscribing to our newsletter, following us on X, or hopping into our Discord.

Subscribe to our newsletter

Website | X | Discord | Telegram | GitHub | YouTube | Reddit