Production EKS & GitOps Infrastructure Platform
- DevOps
- Kubernetes
- Terraform
- AWS
- Cloud
Production · Terraform, AWS EKS, Kubernetes, Helm …
Executive Overview
Core Problem
Managing cloud microservices across multiple environments (dev, staging, production) via manual AWS console clicks creates configuration drift, security vulnerabilities from baked-in secrets, and brittle deployments prone to human error and downtime during service rollouts.
Architectural Solution
Architected a declarative Infrastructure-as-Code control plane using modular Terraform to provision AWS VPCs, EKS clusters, and node groups. Configured Ingress-Nginx behind an AWS Network Load Balancer (NLB) with TLS termination, Cert-Manager for automatic certificate renewal, External Secrets Operator to securely sync AWS Secrets Manager credentials directly into Kubernetes secrets, and ArgoCD for automated GitOps releases.
Measurable Impact
Reduced service deployment lead times from hours to under 3 minutes, eliminated 100% of environment configuration drift, and enabled zero-downtime rolling updates with instant rollback capabilities across all microservices.
System Architecture
Component topology, protocol boundaries, and data flow.
Reliability & Production Security
Deployment & Infrastructure
What I Learned
Technical trade-offs, battle-tested discoveries, and operational takeaways from this project.
GitOps Eliminates Configuration Drift
Allowing direct kubectl changes in production inevitably causes discrepancies between source code and running infrastructure. Enforcing ArgoCD as the sole delivery agent ensures that Git remains the single, auditable source of truth.
External Secrets Operator Solves Secret Injection Without Baking
Committing secrets to Git or baking them into container images is a severe security violation. Using External Secrets Operator to pull encrypted credentials dynamically from AWS Secrets Manager into ephemeral in-memory Kubernetes secrets ensures zero plaintext credential leakage.
Network Load Balancers Outperform ALBs for High-Throughput Ingress
While ALBs operate at Layer 7, AWS NLBs operate at Layer 4 and handle sudden traffic spikes without pre-warming delays. Placing Ingress-Nginx behind an NLB with TLS termination delivers ultra-low latency and seamless WebSocket connection upgrades.
Cluster Autoscaler Requires Precise Pod Resource Requests
Cluster Autoscaler makes scaling decisions solely based on pod CPU/memory requests, not live usage. Setting realistic, tuned resource requests on container specs is mandatory to prevent node starvation or wasteful over-provisioning.
Future Roadmap & Architectural Evolution
- →Implement Karpenter for faster, just-in-time node provisioning compared to classic Cluster Autoscaler.
- →Integrate Kyverno or Open Policy Agent (OPA) for automated Kubernetes admission control and security policy enforcement.