Huasong Shan
Engineer/Researcher (Infra, Cloud, MLOps)
- Role
- Senior Staff Engineer at XPENG
- Location
- San Jose, CA, US
- LinkedIn followers
- 500 followers
About Huasong Shan
15+ years’ experience in the internet (e-commercial), autonomous vehicles, and chip companies, good command of various technical domains: large-scale distributed systems, cloud computing, AIOps, MLOps for various AI applications.2. Open source contributions: ContainerDNS on Kubenetes:https://github.com/tiglabs/containerdnsFederated Learning: https://github.com/cyqclark/fedlearn-algo3. 10+ technical patents worldwide and 10+ top-tier publications in computer system/security/AI domains; regularly serve as TPC for ACM & computer conference (e.g, EuroSys, SoCC, USENIX Security, ICDCS, & hl=en)My focus engineering/research area:1. Build highly available, scalable, and performance system(micro-services, parallelized pipeline)2. Cloud computing: resource management, performance analysis, abnormal detection, root-cause diagnosis, etc.3. MLOps for various AI applications: e-commercial advertisement text generation, news recommendation, perception model validation for autonomous vehicles4. CI/CD: build and test automationTechnical Skills: AWS/Azure, Docker/Kubernetes, MongoDB, RabbitMQ/Kafka, Redis, Jenkins, Git, Redash, Python, Java, C/C++, Make, SQL, NumPy, Pandas, Django, Flink/Spark, Airflow, MLflow, Pytorch.
Experience
Senior Staff Engineer
Jul 2022 — Present · Santa Clara County, CA, US
CI/CD and MLOps, ML Platform: On-target Test for Autonomous Vehicle (MongoDB, RabbitMQ/Kafka, Jenkins, Redash)(1) Large-scale dataset on-target model evaluation/serving/inference for Perception model (100s of thousands of jobs per day)(2) Dataset ingestion pipeline: transcoding, validation, DDS replay/record (TB per day)(3) Model formate transition to TensorRT (10s of thousands of jobs per day)(4) Autonomous Vehicle Full Stack Performance Evaluation Test2. Cloud Resource Management: reliability and availability On-premise ARM/GPU cluster management and reliability, node anomaly detection, and auto-healing on K8S for NVIDIA Orin cluster (Kubernetes, Docker)3. Workload Orchestrator: performance and scalabilityEfficiency optimization, reducing the overhead of the on-target test, via batch scheduling algorithm and parallelized pipeline
Education
Huazhong University of Science and Technology
Master's degree, Computer Science
2003 — 2006
Huazhong University of Science and Technology
Bachelor’s Degree, Computer Science
1999 — 2003
Louisiana State University
Doctor of Philosophy - PhD, Computer Science
2015
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.