Games Importer (Ingestion / Scheduled Jobs)
- Backend
- ETL
- Jobs
Production · NestJS, Node.js, Schedulers/Jobs, Docker
Executive Overview
Core Problem
Batch ingestion pipelines face acute operational pitfalls: external feed providers deliver massive payload dumps (50MB+) with out-of-order records and duplicate identifiers. Running cron jobs without strict concurrency control leads to overlapping executions where worker A and worker B attempt to import the same matches simultaneously, causing database row deadlocks and corrupted platform state.
Architectural Solution
Engineered a stream-based ingestion pipeline utilizing distributed Redis mutex locking (Redlock), composite deterministic natural keys (providerId + eventDate + matchId), and chunked transactional writes. Designed a Dead-Letter Queue (DLQ) to quarantine malformed records without aborting overall batch processing.
Measurable Impact
Reduced data ingestion latency from 45 seconds to sub-2 seconds, eliminated 100% of cron job deadlocks, and successfully processed over 100,000 game records daily with zero data duplication.
System Architecture
Component topology, protocol boundaries, and data flow.
Reliability & Production Security
Deployment & Infrastructure
What I Learned
Technical trade-offs, battle-tested discoveries, and operational takeaways from this project.
Deterministic Composite Keys Are Essential for Idempotent Syncs
Relying on auto-incrementing database IDs causes duplicate records whenever a network drop triggers a retry. Building deterministic composite keys from provider metadata guarantees true idempotency across infinite re-runs.
Chunk Large Ingestion Batches to Prevent Table Lock Escalation
Writing thousands of records in a single monolithic transaction causes row-level lock escalation, blocking concurrent user read/write queries. Streaming and chunking records into batches of 100 maintains stable database IOPS and sub-30ms user query latencies.
Dead-Letter Queues (DLQs) Prevent Full Batch Failures
In bulk third-party data imports, a single invalid payload in a batch of 1,000 should never abort the entire sync. Isolating bad records into a DLQ allows the remaining 999 records to sync uninterrupted while notifying engineers.
Distributed Mutexes Prevent Overlapping Cron Executions
When an ingestion run takes longer than expected due to upstream latency, the next scheduled cron can spawn concurrently. Enforcing a distributed Redis lock guarantees that only one worker executes an ingestion pipeline at any given moment.
Future Roadmap & Architectural Evolution
- →Transition from polling cron schedules to webhook-driven ingestion where provider support allows.
- →Add data diffing logic to minimize write volume by only updating changed record fields.