Deepak Kumar
Research Engineer, Cuda, Inference Optimization @Genbio Ai
Signup · Get unlimited contacts
WORK HISTORY
Research Engineer, Cuda, Inference Optimization @Genbio Ai
Palo Alto, CA, US
Building and optimising large-scale GPU inference pipelines for GenBio’s AIDO biological foundation models (DNA, RNA, Cell, Tissue, Protein) using H100/A100/L4 clusters, Ray Serve, and Kubernetes, delivering significant latency and throughput improvements. Designed and deployed a high-performance Ray Serve microservices architecture with multi-GPU worker groups, autoscaling, MIG/MPS partitioning, and optimised containerised model runtimes on GKE.Building production-grade inference engines with ONNX/TensorRT, Torch-TensorRT, CUDA graphs, memory-efficient attention, and end-to-end performance instrumentation, including stall analysis, warp scheduling, and roofline-based optimisation.
EDUCATION
Illinois Institute of Technology
Master's degree, Computer Science
Dr. A.P.J. Abdul Kalam Technical University
Bachelor of Technology - BTech
ABOUT DEEPAK KUMAR
I’m a CUDA-focused Software Engineer with a Master’s in Computer Science from Illinois Tech and over 6 years of experience building scalable backend systems and accelerating machine learning inference using NVIDIA’s ecosystem.Currently, I develop high-performance CUDA kernels, profile workloads using Nsight Systems/Compute, and optimize real-time inference pipelines with PyTorch and TensorRT. My work involves H100/A100 GPU tuning, multi-GPU scheduling, and memory-efficient deployment of large language models (LLMs) for low-latency, large-scale inference.In parallel, I’ve built robust backend systems during my experience at Oracle, Finoit, and Fluper, designing scalable monolithic and microservices architectures across domains like E-commerce, Healthcare, Automation, and Social Media. I specialize in Java, Spring Boot, Python (FastAPI, Django, Flask), and modern web technologies like React, along with SQL/NoSQL databases, caching, and event-driven systems.I’m passionate about engineering high-availability, distributed architectures involving sharding, replication, and consistent hashing, and deploying them via Docker/Kubernetes with CI/CD pipelines like Jenkins.I’m actively seeking opportunities as a Backend Engineer, Distributed Systems Engineer, or ML Systems Engineer, where I can push the limits of performance, scale intelligent systems, and work with a forward-thinking team.Skills: CUDA | Back-End Development | Distributed Machine Learning | Federated Learning | Natural Language Processing | TensorFlow | PyTorch | Keras | NumPy | Pandas | Scikit-learn | Kubernetes | Kubeflow | Kserve | Apache Kafka | Apache Lucene | Docker | Redis | MongoDB | MySQL | Cosmos DB | Node.js | Java | Python | Real-Time Communication/Multimedia Streaming
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.