Sunil Kumar
Redhat Openshift | Azure, AWS, Kubernetes, Infrastructure, Terraform Engineer | Azure DevOps | Ansible | Certified Azure Administrator, Azure DevOps Expert, Kubernetes Administrator, AWS Solution Architect Associate
- Role
- Site Reliability Engineer at Bank of America
- Location
- Phoenix, AZ, US
- LinkedIn followers
- 500 followers
About Sunil Kumar
Cloud & DevOps Engineer skilled in Azure, AWS, Red Hat OpenShift, Kubernetes, and Linux/UNIX, delivering scalable, secure, automated infrastructure for enterprise applications. Experienced in designing and managing Azure virtual networks, subnets, security policies, Azure Firewall, and hybrid connectivity via ExpressRoute and BGP.Proficient in multi-cloud automation using Terraform and CloudFormation for AWS (VPC, EC2, RDS, EKS, S3, IAM) and Azure (VMs, storage, networking) focusing on scalability, security, and cost control. Managed AKS and on-prem Kubernetes via Rancher, handling cluster setup, RBAC, workload deployment, and autoscaling across Dev, Test, and Prod.Built CI/CD pipelines with Azure DevOps, GitHub Actions, Terraform, Ansible, and ARM templates to automate containerized app and infrastructure deployments. Experienced in Kubernetes orchestration, deploying highly available clusters with autoscaling, ingress controllers, and secure Nginx reverse proxies.Automated continuous deployment using Ansible playbooks for cloud services and configuration management, proficient in YAML. Implemented monitoring and observability with Datadog, Dynatrace, Prometheus, Grafana, Azure Monitor, Log Analytics, and Splunk, creating dashboards and alerts for cluster health and security compliance.Hands-on with Kubernetes tools like Argo CD (GitOps), Harbor registry, Longhorn storage, and Rancher control planes automated via Ansible. Expert in Azure Databricks data pipelines building Lakehouse architectures with PySpark ETL, integrating Power BI and Synapse for BI.Experienced in Azure AD Connect, ADFS configuration with PowerShell for identity management. Managed microservices platforms with Azure Service Fabric and Application Gateway for scalable app delivery.Release engineering background with UNIX/Linux administration, software builds, patching, source control, and CI/CD automation using Jenkins, Puppet, Git, Maven, Ant, Bash, Python, Shell, and PowerShell. Developed Puppet modules integrated with Jenkins for secure automated deployments.Managed source control and branching with Git, SVN, Bamboo, Bitbucket, ensuring robust versioning and release management. Automated infrastructure and app lifecycle tasks with Python and Bash scripts, including Kubernetes YAML, audits, and logs.Automated Linux system administration including cron jobs, log parsing, kernel tuning, SAN/LVM management, network optimization, and troubleshooting. Automated Windows and SQL tasks via PowerShell to improve provisioning and process automation.
Experience
Site Reliability Engineer
Mar 2024 — Present · Chandler, AZ, US
Maintained and operated multiple enterprise-grade Red Hat OpenShift clusters supporting 100+ application teams across dev, staging, and production — managing version upgrades, TLS/Route configs, ingress, autoscaling, and lifecycle operations.• Executed staged OpenShift cluster version upgrades with readiness validation, rollback plans, and minimal downtime; troubleshot pod failures from OOMKilled, CreateContainerConfigError, and ExceededQuota events.• Integrated Ceph-based persistent storage with OpenShift; resolved PVC binding issues and monitored storage performance. Managed node cordoning/draining for hardware maintenance with zero workload disruption.• Led quarantine image approval process via Bitbucket PRs, mapping external images to internal Artifactory registries with Aqua security scans and automated Ansible Tower mirroring post-merge.• Onboarded 100+ dev teams to OpenShift: provisioned namespaces, enforced RBAC, quotas, and network policies; supported GitOps deployments via Argo CD and Bitbucket.• Integrated HashiCorp Vault for dynamic secret injection via sidecar agents; centralized log aggregation with Splunk — built dashboards for cluster health, pod restarts, API latency, and error rates.• Wrote Python and Bash automation for namespace cleanup, resource audits, YAML manifest management, and PVC utilization tracking; scheduled cron jobs for production config snapshots.• Supported 24x7 on-call rotations resolving node pressure, pod crashes, image pull failures, and CI/CD disruptions; led resiliency audits and SLA compliance reviews.
Education
University of Central Missouri
Master's degree, Information Technology
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.