Akhilesh K
AI/ML Engineer | LLMs & Scalable NLP Systems | PyTorch, DeepSpeed, Spark, AWS
- Role
- Devops Engineer at Meta
- Location
- San Jose, CA, US
- LinkedIn followers
- 500 followers
About Akhilesh K
I\'m an AI/ML Engineer with 4+ years of experience building and deploying large-scale AI systems, from foundation models (LLaMA 3/4) to production-ready ML pipelines in enterprise environments. I specialize in NLP, LLM fine-tuning, and scalable ML infrastructure using tools like PyTorch, DeepSpeed, Hugging Face, and Spark. 🤖 LLMs & Applied NLP – Fine-tuned and deployed large models (70B–405B) using LoRA, QAT, and RLHF, improving multilingual NLP, summarization, and RAG benchmarks. Delivered LLM-driven solutions in virtual assistants, fraud detection, and personalized financial advice. MLOps & Deployment – Built robust MLOps workflows with MLflow, Airflow, DVC, and GitHub Actions. Automated retraining, CI/CD, and canary deployments across AWS SageMaker and Kubernetes (EKS), reducing deployment overhead and improving reliability. Real-Time Inference & APIs – Deployed scalable, low-latency inference endpoints using FastAPI, TorchServe, ONNX Runtime, and Triton on AWS (EC2, Lambda, EKS). Integrated RESTful microservices in production apps with Docker and Flask. Data Engineering & Visualization – Engineered real-time ETL pipelines with Spark, PySpark, and Pandas, boosting processing throughput. Built dashboards using Tableau, Power BI, and Matplotlib to monitor drift, bias, and KPIs. Responsible AI & Explainability – Embedded fairness and explainability using SHAP, LIME, and Fairlearn, ensuring GDPR compliance and ethical model behavior in financial and customer-facing systems.
Experience
Devops Engineer
Dec 2023 — Present · TX, US
Designed and implemented scalable microservices architectures, enhancing system reliability and reducing downtime by 20%.• Developed distributed systems for large-scale data processing, handling millions of user requests per second.• Built and deployed CI/CD pipelines using Jenkins and GitHub Actions, accelerating release cycles by 35%.• Automated infrastructure provisioning using Infrastructure as Code (IaC) tools such as Terraform, enabling consistent and repeatable environment setups across development, staging, and production.• Managed containerized applications using Docker and orchestrated deployments on Kubernetes, ensuring seamless scalability.• Collaborated with cross-functional teams to integrate machine learning models into production pipelines, improving recommendation accuracy by 15%.• Optimized cloud infrastructure in AWS environments, achieving 10% cost reduction while maintaining high availability and performance.• Developed robust monitoring and logging solutions using Prometheus and Grafana, enabling proactive incident management and quick issue resolution.• Conducted root cause analyses and implemented fixes for critical system outages, reducing mean time to resolution (MTTR) by 25%.
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.