Pinjari Kullayiswamy
Hpc Ai Infrastructure Engineer @Tata Consultancy Services
Signup · Get unlimited contacts
WORK HISTORY
Hpc Ai Infrastructure Engineer @Tata Consultancy Services
Bengaluru, IN
Key Contributions:Provisioned and managed full-stack GPU clusters (AMD MI210, MI250, MI300, MI325, MI355) from scratch, including system imaging, Slurm setup, kernel tuning, and ROCm stack integration.Executed AI/ML benchmarking workloads using SGLang, DeepSeek, Megatron, Torch-Triton, PyTorch, JAX, and TensorFlow for inference performance validation analyzing throughput, latency, and GPU utilization.Performed GPU and interconnect benchmarking using TransferBench, RCCL, AGHFC, rocHPL, BabelStream, MLPerf, and RVS to evaluate compute and communication efficiency.Automated cluster provisioning and validation using Ansible, GitHub Actions, and shell scripts to streamline system setup and reduce manual efforts.Deployed and maintained Prometheus federation and Grafana dashboards for observability; integrated FluentBit logging and NGINX reverse proxy for system monitoring.Worked on InfiniBand and RoCEv2 networks, optimizing RDMA for high-performance GPU communication.Supported internal teams and users with Kubernetes and Slurm job submissions, sbatch scripting, and MPI-based workload debugging.Worked on AMD’s platforms (AACA, Plexus) for user onboarding, workload submissions, and bare-metal container provisioning.Provisioned AWS instances for deploying internal applications (Prometheus, AACA, Sentry).Collaborated with Platform, Product, and Performance teams to identify bottlenecks, validate ROCm kernels, and ensure AI workload stability across GPU generations.Worked on InfiniBand and RoCEv2-based clusters, used BMC/IPMI for remote power control, and fetched/upgraded BKC versions via Redfish API during node validation in coordination with the SWAT team.
EDUCATION
GATES Institute of Technology, Gooty
Bachelor of Technology - BTech, Mechanical Engineering
Government Polytechnic College Rayadurg
Diploma of Education, Mechanical Engineering
ABOUT PINJARI KULLAYISWAMY
Passionate HPC / AI Infrastructure Engineer with around 4 years of experience in managing GPU-based high-performance computing clusters, with deep expertise in the AMD ROCm ecosystem, Slurm workload management, and AI/ML framework benchmarking. Proven track record of provisioning and maintaining on-prem and cloud-based clusters (bare metal & VM) using Vultr and OpenStack.Skilled in deploying and validating AI/ML workloads using ROCm-enabled containers with PyTorch, JAX, TensorFlow, SGLang, DeepSeek, Megatron, and Torch-Triton for model inference benchmarking, analyzing throughput, latency, and GPU utilization across multi-GPU systems.Experienced in writing and managing sbatch scripts, executing benchmarking workflows with TransferBench, RCCL, rocHPL, BabelStream, RVS, MLPerf, and AGHFC, and optimizing GPU performance across AMD GPUs (MI210, MI250, MI300, MI325, MI355).Adept in DevOps and automation using GitHub Actions, Ansible, and shell scripting, with Kubernetes-based orchestration for containerized workloads. Set up Prometheus federation and Grafana dashboards for observability; managed FluentBit-based logging, NFS, and kernel-level tuning for system stability.Worked on AMD’s internal platforms like Plexus and AACA to support external customer deployments, workload submissions, and bare-metal container provisioning.Strong in cross-functional collaboration with Platform, Product, and Client teams to resolve issues, onboard users, and automate full-stack AI/HPC operations. Seeking challenging opportunities in AI Infrastructure, GPU Computing, and DevOps-driven cluster engineering roles.
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.