
Sai Karthik Gutha
Looking for a Job and investors
1 followerยท1 connection
Interested Industries
Skills
AWS Glue
intermediate
credit risk
intermediate
Cloud Storage
intermediate
PyCharm
intermediate
Azure Kubernetes Service (AKS)
intermediate
model evaluation
expert
Work Experience
AI Engineer / Data Scientist
JP Morgan Chase
2024-08 - 2025-06
Refactored a legacy risk-scoring pipeline inherited from earlier teams, tightening feature constraints to stay within Basel III and Dodd-Frank limits while cleaning up brittle preprocessing paths for transaction and market feeds. Built a training stack in Python with TensorFlow and scikit-learn, wiring reproducible feature engineering, experiment tracking, and idempotent long-running jobs so model refreshes could resume cleanly after partial failures across 12 recurring risk datasets. Calibrated fraud and credit-risk thresholds after drift checks started flagging unstable score distributions, then grounded internal finance answers through LangChain, LlamaIndex, and vector retrieval so Gen AI outputs stayed traceable and audit-friendly.
Junior Data Scientist
Value Labs
2022-08 - 2023-04
Refactored brittle feature-prep scripts into maintainable Python modules, stripping legacy data-cleaning logic out of the training path and standardizing inputs through Azure Data Factory and Azure Databricks for repeatable model runs. Designed validation gates around refreshed datasets so train-test leakage stayed out of the evaluation flow, enforced schema checks from Azure Data Lake Storage Gen2 into Azure Machine Learning, and kept preprocessing identical across runs under a strict client delivery constraint. Documented experiment lineage and tuned scikit-learn classifiers and regressors with early stopping, then tracked runs in Azure DevOps and monitored production scoring jobs through Azure Monitor while the team iterated on a forecast model that lifted segmentation accuracy by 18 points on a 42,000-row client dataset.
Data Science Intern
IHub Data
2021-10 - 2022-07
Cancer risk scoring in Python with RandomForestClassifier, XGBoost, and logistic regression, recall was the main thing because false negatives were not acceptable, thresholds kept aligned with oncologists so the output could actually be used in clinic. Python ETL pulling from MongoDB and hospital SQL systems, then stitched into one training set. Automated lab value imputation, plus a weekly sync job since stale records kept showing up. A lot of EDA and data checks in Pandas, with Matplotlib heatmaps and statistical profiling. Clinical fields were inconsistent all over the place, fixed those before training, less noise in the downstream ML pipeline.
Education
Master of Science in Computer Science
2023-05 - 2025-05