Mani A.
Data Scientist @ UCSF | MS in Data Science @ USFCA
- Role
- Data Scientist at Ucsf Radiation Oncology
- Location
- San Francisco, CA, US
- LinkedIn followers
- 500 followers
About Mani A.
I\'m a Data Scientist with experience in NLP, predictive modeling, and large-scale data engineering, currently completing my M.S. in Data Science at the University of San Francisco while interning at UCSF\'s Cancer Radiation Oncology team. Once upon a time, I was managing a Starbucks team during the morning rush. Now I\'m building transformer pipelines over brain tumor records. The throughline, I think, is that I\'ve always cared about doing the job properly, whether that\'s getting an order right or getting a model right. A few highlights- Built and benchmarked transformer embedding pipelines over 70K+ longitudinal EHR notes to predict glioma survival outcomes and tumor grade (AUC ~0.79)- Engineered features from 250K+ CDC health records and trained ensemble models achieving +17% accuracy over baseline- Built a distributed ETL pipeline across three data sources, producing a Random Forest Regressor with R² of 0.93 I\'m proficient in Python, SQL, PyTorch, PySpark, Hugging Face Transformers, XGBoost, Airflow, and GCP. Graduating Summer 2026 and actively looking for full-time Data Scientist roles. Feel free to reach out at m••••••••@gmail.com.
Experience
Data Scientist
Oct 2025 — Present · San Francisco, CA, US
Built a scalable EHR embedding pipeline to generate patient-level representations from ~70K clinical notes, enabling downstream prediction of survival, tumor grade, and clinical outcomes• Designed and implemented advanced text processing workflows including adaptive chunking, sentence-boundary detection, and embedding caching (SQLite + vector stores) to improve efficiency and reproducibility• Evaluated and benchmarked state-of-the-art embedding models (e.g, BERT variants, long-context models like Stella and Jina) against classical baselines (TF-IDF + feature engineering), achieving strong predictive signal (AUC ~0.77)• Developed parallelized multi-GPU processing pipelines to embed large-scale clinical datasets, reducing runtime and supporting ongoing large experiments across ~70K+ documents• Conducted embedding analysis and visualization (UMAP, PCA) to assess representation quality and robustness across chunking and pooling strategies
Education
University of California, Berkeley
Bachelor of Arts - BA, Data Science - Business Emphasis
University of San Francisco
Master of Science - MS, Data Science
Find verified contacts for anyone on LinkedIn
Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.
Free plan included · No credit card required
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.