Ashutosh Chandra
Site Reliability Engineer Lead |SRE|Azure|GCP|AIOps|Grafana|Observability|ECommerce|OmniChannel|OMS|
- Role
- Site Reliability Engineer Lead at Albertsons Companies
- Location
- San Francisco, CA, US
- LinkedIn followers
- 500 followers
About Ashutosh Chandra
Senior Site Reliability Engineering leader with 18+ years (9+ years under H‑1B visa status) driving high-availability solutions for large-scale e-commerce and retail platforms. Expertise in cloud-native infrastructure (Azure, GCP), observability (Grafana, Prometheus, LOKI, OTEL, Log analytics), and AI-Ops automation for (Incident prediction and remediation using AI (Ollama) and Llama 3.0 model).Involved in RCA Hub and Summarization intelligence module with LLM-powered summarization and investigation assistant using Open AI. Proven track record mentoring SRE teams, implementing reliability best practices, and optimizing performance and scalability of mission-critical systems. Good experience in WebSphere Commerce Server (WCS), Sterling OMS & Hybris. Throughout his career, he has effectively developed internet applications, evaluated alternative solutions, and created multi-channel software solutions. He has hands-on experience with service integration, SOA integration, WebSphere Commerce Server, Sterling Order Management, Spring, and REST, utilizing Microsoft Azure Cloud platform & GCP.Ashutosh is proficient in analyzing and translating business requirements into technical requirements and architecture. He has been involved in understanding client needs through communication with clients and cross-functional teams, as well as in estimating e-commerce design, high-level and low-level design (both functional and technical), and testing documents.Worked for different E-commerce/Retail clients like Albertons/Safeway, Universal Studio, Staples,Macys.com, Walt Disney, Warner Brothers, Argos UK, David Jones Australia,Star CJ, F&P, Home Shop 18. • Programming & Web Frameworks: Java 8.0, JSF, Spring Boot, Hibernate, Maven, AJAX, Microservice, Apache Spark,Python. • Monitoring & Observability: Grafana, Prometheus, LOKI, Azure Log Analytics, OTEL,Synthetic Monitorng,Dynatrace D,ataDog, AppDynamics and Splunk. • Database & Messaging: Kafka, Confluent, Yarn, Spark UI, Ambari, Oracle, DB2, Cassandra, Cosmos /Mongo DB. • DevOps & Automation: RAD, WAS, CVS, SVN, Soap UI, Git, Stash, Yarn UI, POSTMAN, Version One, Jira, PIM, Automation Platforms, CI/CD Pipeline Management, Kubernetes (AKS), Container Platform, Docker, Helm,AIOps (Ollama),Llama/Open AI models,Azure /GCP Cloud, Databricks, Incident Management (Service Now) and Terraform. • Functional Modules: Payment System, Distribution/Order Management system, Fullfilment system, Call Center Management (CRN), Inventory System, Member and Catalog Merchandising system, MDM, OMS & Hybris.
Experience
Site Reliability Engineer Lead
Apr 2020 — Present · US
Implement and maintain system monitoring, alerting, and logging to ensure high availability and reliability using Grafana, Azure Log Analytics, and GCP log monitoring, following site reliability engineering best practices.•Involved in Incident Prediction outage – AI-Ops Implementation using AI (Ollama) and Llama models. •Involved in RCA Hub and Summarization (observability and intelligence module with LLM-powered summarization and investigation assistant) using Open AI. •Participated in setting up alerts for various microservices and data bricks logs using Grafana, GCP and ALA.•Implemented end-to-end Open Telemetry (OTel) across our observability stack, enabling unified LTM.•Involved in migration of Synthetic Monitoring to Prometheus Black Box Exporter.•Respond to and resolve incidents using established incident management processes, conducting root cause analysis and postmortm to prevent future occurrences.•Automate manual tasks using automation platforms and scripting (e.g. Unix Shell scripting, Scala Notebook)•Involved in resolving Databricks/Spark-related issues promptly, ensuring no impact on customers.•Define and track Service Level Objectives (SLOs) and Service Level Agreements (SLAs) to ensure performance.•Optimize system performance by identifying and resolving bottlenecks, latency issues, and resourc overhead.•Plan and execute capacity management strategies to handle increasing traffic and system growth.•Implement disaster recovery and business continuity strategies across cloud platforms to minimize downtime.•Provide on-call support, ensuring systems are monitored 24/7 and responding to critical incidents using SNOW.•Manage service outages by coordinating with stakeholders and making efforts to bring systems back online.•Evaluate new technologies and tools that can improve system reliability, automation to support our systems.•Lead and participate in on-call rotations to ensure systems are supported and issues are addressed promptly.
Education
Bharati Vidyapeeth
BE, Information Technology
2001 — 2005
Skills
- Tomcat
- Junit
- Software Development Life Cycle (Sdlc)
- Integration
- Jsp
- Hybris
- Cassandra
- E-Commerce
- Hibernate
- Microsoft Azure
- Web Development
- Ibm Db2
- Web Applications
- Javascript
- Struts
- Richfaces
- Play Framework
- Web Services
- Servlets
- Xml
- Jboss Application Server
- Kibana
- Eclipse
- Spring
- Java
- Ant
- J2ee Application Development
- Team Management
- Ibm Websphere Commerce
- Websphere Application Server
- Ajax
- Java Enterprise Edition
- Oracle
- Requirements Analysis
- Core Java
- Jsf
- Subversion
- Websphere
- Mysql
- Jdbc
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.