Salina Khadka
Senior Data Engineer @Orion Capital Group
Signup · Get unlimited contacts
WORK HISTORY
Senior Data Engineer @Orion Capital Group
Santa Clara, CA, US
Developed high-performance Spark applications using PySpark, Scala, and Spark-SQL/Streaming for large-scale fintech data processing and risk analytics.• Structured data into Bronze, Silver, and Gold layers using Azure Data Lake, ensuring scalability and quality.• Built scalable ETL pipelines with PySpark, Talend, FiveTran, Matillion, AWS Glue, and GCP Dataflow for multi-cloud integration.• Containerized ETL jobs using Docker, orchestrated with Kubernetes for fault-tolerant execution.• Automated infrastructure provisioning via Terraform across cloud environments.• Used Shell scripting (Bash) on Linux/Unix for ETL support and automation.• Migrated data from Oracle, SQL Server, and MongoDB to Azure Data Lake and GCP BigQuery using ADF and Google Cloud Storage.• Integrated Sqoop and Flume for ingesting relational and streaming data into HDFS.• Managed Kafka clusters and integrated with Spark Streaming for real-time analytics with performance monitoring via KPIs.• Performed data transformation using Spark Core, Spark SQL, Java, Apache Beam, and GCP Dataproc for scalable analytics.• Created MapReduce jobs, Hive UDFs, and HBase tables for optimized query processing.• Designed BI solutions with Azure Synapse, Databricks, BigQuery, and Cloud Pub/Sub for enterprise-scale analytics.• Modeled Star/Snowflake schemas using Erwin for Fact/Dimension table creation.• Built dashboards in Power BI and Tableau for automated reporting and business insights.• Orchestrated data workflows using Apache Airflow DAGs for end-to-end pipeline automation.• Designed efficient MongoDB schemas for performance and indexing.• Conducted data analysis using pandas and numpy to support ML models and predictive analytics.• Utilized AWS Glue and S3 for ETL and data storage in a multi-cloud setup.• Employed Git, Jenkins, GitLab for CI/CD, and used Jira for task and issue tracking.
EDUCATION
University of North Texas
Master's degree
ABOUT SALINA KHADKA
I’m a Data Engineer with 6+ years of experience designing and delivering high-performance data pipelines, real-time streaming systems, and cloud-native architectures across AWS, Azure, and GCP. My expertise spans Apache Spark (PySpark/Scala), Hadoop, Kafka, Airflow, and modern ETL tools like Talend, FiveTran, and dbt. I specialize in building robust data lakes and Medallion architecture in Databricks, optimizing performance through data modeling (Star, Snowflake), and enabling analytics through Snowflake, BigQuery, and Synapse. Skilled in both batch and real-time processing, I’ve engineered end-to-end solutions supporting fintech, healthcare, and enterprise-scale analytics. With hands-on experience in Terraform, Docker, and Kubernetes, I bring DevOps discipline to data engineering, streamlining CI/CD pipelines, automating infrastructure, and ensuring scalable, secure deployments. I thrive at the intersection of data, cloud, and AI, often collaborating cross-functionally to drive insights, governance, and business value. Let’s connect if you’re building next-gen data platforms or need someone to turn complex data into strategic intelligence
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.