Rohan Dawkhar

Rohan Dawkhar

Looking for a Job

0 followersยท0 connections

Activity Score: 42

Interested Industries

Artificial IntelligenceMachine LearningData ScienceSoftware Development

Skills

Graph Database Query

intermediate

Benchmarking

intermediate

Retrieval-Augmented Generation (RAG)

expert

Recommendation Systems

expert

Data Structures & Algorithms

intermediate

Event Stream Schema Design

intermediate

Work Experience

Research Assistant

MIRAGE Lab

2025-02 - 2025-04

Conducted end-to-end OCR benchmarking for 3D-printed mechanical parts, transitioning from classical ML models (SVM, Random Forest, 45% accuracy) to CNN-based vision models (78% accuracy) and finally Vision-Language Models (95%+ accuracy); engineered preprocessing pipelines, performed dataset augmentation and error analysis. Proposed and implemented dataset preprocessing and annotation refinements that enhanced data consistency, reduced labeling errors by 30%, and contributed to improved model generalization, leading to more robust OCR performance. Led collaborative experiments with Dr. McGregor and the MIRAGE Lab team, driving iterative improvements across models and boosting overall OCR reliability for serial number recognition, leading to a 20% reduction in misreads.

Data Scientist Intern

Kampd

2022-06 - 2022-12

Designed, developed, and tested Recommendation Engine's Content Mod Abstraction, implementing utility endpoints to optimize content handling and improve DEV environment workflows, reducing data preprocessing time by 40%. Built and deployed the Taxonomy Engine, automatically labeling content using titles, descriptions, and user tags; integrated live learning to adapt to changing data distributions, reducing manual and inconsistent label updates by 80%. Conducted data annotation and preprocessing on large user datasets, producing high-quality testing sets for ML model evaluation; improved dataset consistency, reduced labeling errors by 25%, and enabled accurate performance benchmarking. Performed rigorous graph database query testing (Gremlin) across modules such as kamps, bytes, and insights, ensuring product logic correctness and supporting personalization features; reducing query-related inconsistent errors by 30%. Led proof-of-concept (PoC) experiments for Personalization Phase 1, including event stream schema design (experimented JSON, Avro, Protobuf), event nomenclature, and catalog creation, enabling centralized event tracking for engineering and data teams and resulting in more reliable event-driven pipelines and faster iteration for personalization features. Developed and tested processor blocks in the graph layer, and created graph traversals and vector database PoCs (Milvus, Qdrant) in the model layer, leading to advanced personalization and recommendation capabilities of the engine. Actively collaborated with engineering and data teams to deploy, validate, and test ML pipelines, ensuring end-to-end reliability and performance improvements across the platform, thereby streamlining the software development lifecycle.

Education

University of Maryland College Park logo

Master of Science in Data Science

University of Maryland College Park

2024-08 - Present

Coursework: NLP, Deep Learning, Big Data, Cloud, ML, Algorithms

Rohan Dawkhar | Appli Network