Games Importer (Ingestion / Scheduled Jobs)

  • Backend
  • ETL
  • Jobs

Production · NestJS, Node.js, Schedulers/Jobs, Docker

Executive Overview

An automated data ingestion pipeline service built with NestJS, Docker, and Kubernetes CronJobs that reliably synchronizes, normalizes, and indexes high-volume game and match data from external feeds into core platform databases.
The Challenge & Bottleneck

Core Problem

Batch ingestion pipelines face acute operational pitfalls: external feed providers deliver massive payload dumps (50MB+) with out-of-order records and duplicate identifiers. Running cron jobs without strict concurrency control leads to overlapping executions where worker A and worker B attempt to import the same matches simultaneously, causing database row deadlocks and corrupted platform state.

Engineering Approach

Architectural Solution

Engineered a stream-based ingestion pipeline utilizing distributed Redis mutex locking (Redlock), composite deterministic natural keys (providerId + eventDate + matchId), and chunked transactional writes. Designed a Dead-Letter Queue (DLQ) to quarantine malformed records without aborting overall batch processing.

Quantified Outcomes

Measurable Impact

Reduced data ingestion latency from 45 seconds to sub-2 seconds, eliminated 100% of cron job deadlocks, and successfully processed over 100,000 game records daily with zero data duplication.

System Architecture

Component topology, protocol boundaries, and data flow.

Games Importer (Ingestion / Scheduled Jobs) System Topology
Architecture Flow
CLIENT CONSUMERWeb & API CallsHTTPS / REST PayloadsJSON Schema InputBOUNDARY GATEWAYNginx / Reverse ProxyTLS TerminationRate Limiting & AuthNSERVICE CORE LOGIC• Domain Services & Controllers• DTO Runtime Validation• AWS Secrets Manager Config• Health Readiness ProbesPERSISTENCEPostgreSQL / RedisACID TransactionsDocker / EKS Hosted

Reliability & Production Security

Features idempotent database upserts using natural keys, memory-capped JSON streaming parsers that prevent Node.js V8 heap out-of-memory crashes on multi-megabyte payloads, and heartbeat monitoring reporting to Prometheus.

Deployment & Infrastructure

Packaged in multi-stage Docker images and scheduled via Kubernetes CronJobs with strict resource memory limits and automated alert hooks if scheduled job executions fail.
Engineering Post-Mortem & Insights

What I Learned

Technical trade-offs, battle-tested discoveries, and operational takeaways from this project.

1

Deterministic Composite Keys Are Essential for Idempotent Syncs

Relying on auto-incrementing database IDs causes duplicate records whenever a network drop triggers a retry. Building deterministic composite keys from provider metadata guarantees true idempotency across infinite re-runs.

2

Chunk Large Ingestion Batches to Prevent Table Lock Escalation

Writing thousands of records in a single monolithic transaction causes row-level lock escalation, blocking concurrent user read/write queries. Streaming and chunking records into batches of 100 maintains stable database IOPS and sub-30ms user query latencies.

3

Dead-Letter Queues (DLQs) Prevent Full Batch Failures

In bulk third-party data imports, a single invalid payload in a batch of 1,000 should never abort the entire sync. Isolating bad records into a DLQ allows the remaining 999 records to sync uninterrupted while notifying engineers.

4

Distributed Mutexes Prevent Overlapping Cron Executions

When an ingestion run takes longer than expected due to upstream latency, the next scheduled cron can spawn concurrently. Enforcing a distributed Redis lock guarantees that only one worker executes an ingestion pipeline at any given moment.

Future Roadmap & Architectural Evolution

  • →Transition from polling cron schedules to webhook-driven ingestion where provider support allows.
  • →Add data diffing logic to minimize write volume by only updating changed record fields.
Games Importer (Ingestion / Scheduled Jobs) | Siddhant Ghosh