
Tanya Walia
Looking for a Job
0 followersยท0 connections
Interested Industries
Skills
Bayesian Inference
expert
PySpark
expert
Tableau
expert
Random Forests
expert
Marketing
intermediate
NLTK
expert
Work Experience
Data Scientist Intern
Athena Enzyme Systems
2025-09 - Present
Designed and deployed AI-enabled analytics workflows using Python, SQL, and Airflow, embedded into internal client-facing systems, increasing task success rates by 15% and reducing manual review effort by 30%. Designed and deployed an AI-driven personalized messaging system that increased email conversion rates by 15%, directly supporting marketing strategy optimization. Built an agent evaluation harness with prompt/policy regression tests, failure-mode taxonomy, and MLflow dashboards for accuracy, latency, and safety, enabling rapid iteration without quality regressions.
Data Scientist and Machine Learning Intern
Nostopharma
2025-04 - 2025-08
Built a retrieval-augmented LLM pipeline in Python to parse, index, and ground generation on 1,000+ structured documents, reducing hallucinations and cutting manual review 75%. Fine-tuned a 3B-parameter LLAMA model (LoRA / PEFT) for domain-specific agent behavior, achieving a 28% improvement in generation consistency via blinded expert evaluation. Designed constraint-aware LLM agents with tool-based decision rules and human-in-the-loop review, enforcing safety and policy constraints upfront and reducing downstream rework 30%. Implemented multi-source context pipelines to support same-day decision-making, delivering grounded outputs to users through real-time Tableau dashboards.
Data Scientist
Allstate Insurance
2022-02 - 2024-07
Led behavioral segmentation across 5M users using K-Means and RFM, informing retention strategies that contributed to a 20% lift in renewals. Owned production deployment of multi-class ticket-routing models handling 1M+ annual interactions, contributing to a 30% reduction in resolution time through monitored inference pipelines. Implemented weakly supervised labeling pipelines combining heuristic and model-driven labels, reducing manual annotation cost and improving recommender accuracy by 15%. Refactored core SQL and Pandas ETL pipelines to enforce automated data quality checks, reducing reporting latency by 22%. Built NLP pipelines over large-scale customer feedback, improving sentiment classification accuracy by 20% and directly informing product roadmap prioritization. Developed and optimized Spark / PySpark data pipelines on cloud platforms (Databricks-compatible architectures) to process multi-million-record datasets for downstream modeling and reporting.
Education
Master of Professional Studies in Data Science
University of Maryland, Baltimore County
2024-08 - Present
Bachelor of Technology in Computer Science specialization in Big data
University of Petroleum and Energy Studies
2018-08 - 2022-05