Sushanta Kumar Pani

Senior Machine Learning Scientist | Conversational AI • Speech-to-Speech • Whisper ASR • TTS • LLM Fine-Tuning

Role
Senior Machine Learning Scientist at Stealth Startup
Location
Redmond, WA, US
LinkedIn followers
500 followers

About Sushanta Kumar Pani

Senior Machine Learning Scientist specializing in speech and language model development, with 14+ years of experience building advanced systems across ASR, TTS, speech-to-speech translation, LLM fine-tuning, and Generative AI. I focus on designing and training transformer-based models that enable multilingual, voice-preserving, real-time AI applications.In my current role as a Senior ML Scientist at a speech AI startup, I architect and train end-to-end speech-to-speech translation models that preserve speaker identity and prosody across languages. My work includes fine-tuning Whisper, Wav2Vec2, Tacotron2/FastSpeech2/YourTTS, and LLM decoders using multi-GPU Azure systems, improving WER, translation accuracy, and naturalness metrics through custom alignment, prosody, and embedding pipelines.Previously, as a Senior ML Scientist at Innoveren, I led the development of Generative AI–powered educational systems built on Whisper ASR, LLaMA/Mistral fine-tuning (QLoRA), and optimized RAG architectures. I helped scale AI solutions over hours of classroom content, improving transcript quality, summarization, and question-answering performance across multi-course learning workflows.Before industry, I conducted research at the iPRoBe Lab, Michigan State University, where I developed voice morphing algorithms and exposed vulnerabilities in speaker verification systems via adversarial morphing attacks. My academic background includes an MS in Computer Science from MSU, focused on transformer-based coreference resolution using BERT/SpanBERT.Earlier in my career at CDAC, I built foundational ASR and TTS systems for low-resource South Asian languages—spanning rule-based TTS to DNN-based synthesis and Kaldi ASR—contributing to national-scale multilingual speech initiatives.I’m deeply passionate about the intersection of speech, LLMs, and semantic representation, including speech embeddings, multimodal generative models, and scalable, accessible AI that enables inclusive global communication.Open to connecting with teams building cutting-edge speech, NLP, and generative AI systems with real-world impact.

Experience

  1. Senior Machine Learning Scientist

    Stealth Startup

    Jun 2024 — Present · Redmond, WA, US

    Architected and trained end-to-end speech-to-speech translation models, integrating semantic, acoustic, and speaker-identity embeddings to preserve voice characteristics across 6+ languages.Fine-tuned Whisper, Wav2Vec2, and transformer-based LLM decoders on 300+ hours of domain audio using multi-GPU training on Azure, reducing WER by 28% and improving translation fidelity by 22%.Enhanced Tacotron2 / FastSpeech2 and YourTTS models with custom alignment, duration-prediction, and prosody-stabilization modules, delivering +0.4 MOS improvement and more consistent voice preservation.Developed multilingual speech embedding pipelines (speaker + semantic + prosody) that increased cross-lingual speaker-similarity scores by 35% and improved clarity in voice-preserved output.Optimized inference on Azure GPU endpoints using mixed precision, pruning, and quantization, achieving <2s latency while maintaining output quality for long-form lecture content.

Education

  • Godbhaga M.E. School, Bargarh, Odisha

    Pre High School

  • Centre for Development of Advanced Computing (C-DAC)

    FPGDST, Software Technology

  • Govt High School Burla, Sambalpur, Odisha

    High School

  • Biju Patnaik University of Technology, Odisha

    Bachelor of Technology - BTech, Information Technology

  • SLB Farm Upper Primary School Goshala, Sambalpur, Odisha

    Primary

  • Gangadhar Meher University (GMU), Sambalpur

    HSE/ (+2 Science), Science

  • Michigan State University

    Master of Science - MS, Computer Science

Skills

  • Research
  • R&D
  • Speech Technology
  • Text-to-Speech
  • Speech Synthesis
  • Speech Signal Processing
  • Prosody
  • Acoustic Modeling
  • Speech Recognition
  • Natural Language Processing
  • Artificial Intelligence
  • Artificial Neural Networks
  • Machine Learning
  • Pattern Recognition
  • Feature Extraction
  • Deep Learning
  • Python
  • Matlab
  • Perl
  • C
  • Shell Scripting
  • Data Structures
  • Algorithms
  • Unix
  • Htk
  • Audacity
  • Linux
  • Ubuntu

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Sushanta Kumar Pani — Senior Machine Learning Scientist at Stealth Startup in Redmond, WA, US | Unifers