Skip to main content

Command Palette

Search for a command to run...

Scaling Event Ingestion Service

Updated
•2 min read•View as Markdown

Trying to write something raw in the AI era 😄

Recently, I’ve been building an event ingestion system on my own machine: Intel i5 12th Gen, 8GB RAM, Debian, Node.js and PostgreSQL.

The goal: handle 8 million events/day reliably—roughly 93 events/sec on average, with room for bursts.

Here’s what the investigation taught me 👇

1. Database round trips add up

The earlier pipeline used Bull for background processing, but each event still needed database writes, identity lookups and contact updates.

With a three-connection pool, processing was around 14 events/sec. One staging run sent 400,000 events in five minutes—and took roughly eight hours to land them.

Accepting an event isn’t the same as finishing its processing.

2. Batching wasn’t the whole answer

I added multi-row inserts, buffered eligible contact updates and shared concurrent identity lookups for the same customer.

In a controlled local processing test at 500 events/sec, sharing those lookups brought p95 latency down from about 3.3 seconds to 174ms. All 10,000 events landed.

Useful evidence, but not a production throughput or durability guarantee.

3. More workers can make things worse 😅

I tried PM2 clustering to use multiple CPU cores.

At 30–50 events/sec, four workers increased CPU usage without improving throughput.

At a 3,000 events/sec target, three workers increased steady throughput from roughly 1,050 to 2,100 events/sec.

But startup connection failures and gateway errors remained. More capacity didn’t automatically mean reliability.

4. Separate ingestion from processing

The architecture became:

Client → Ingestion API → Redis Stream → Consumer workers → Batch writes → PostgreSQL

The API now acknowledges after appending to Redis, without waiting for database processing.

The stream buffers bursts—but if consumers stay slower than arrivals, the backlog keeps growing. It has to drain.

One embarrassing discovery: some earlier telemetry comparisons were invalid because the exporter wasn’t enabled. Verify your measuring tools too. 😂

The architecture is deployed, but sustained end-to-end verification is still pending. I’m not claiming “8 million/day proven” or “zero loss” yet.

Next: test sustained load, retries, duplicates, backlog recovery and failures.

That’s the part of backend engineering I enjoy most:

Measure → find the bottleneck → change something → test again.

What surprised you most when you first load-tested a system?

#BackendEngineering #SystemDesign #NodeJS #PostgreSQL #Redis