Skip to content
View BIRJUNG's full-sized avatar

Block or report BIRJUNG

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
BIRJUNG/README.md

Hi πŸ‘‹, I'm Birjung Thapa

Data Scientist | Analytics Professional | Turning Business Problems into Data-Driven Solutions


πŸš€ About Me

I am a Data Scientist and Analytics Professional with 15+ years of experience driving business impact through data.

I combine:

  • πŸ“Š Business strategy + analytics
  • πŸ€– Machine learning + statistics
  • πŸ“ˆ Data visualization + storytelling

πŸ’‘ My strength is translating complex data into actionable insights for real-world decision-making.


🧠 What Makes Me Different

βœ” Strong mix of business + data science
βœ” Experience in real-world decision systems (not just models)
βœ” Hands-on with end-to-end ML pipelines
βœ” Ability to communicate insights to executives and stakeholders


πŸ› οΈ Tech Stack

πŸ’» Languages

Python β€’ SQL β€’ R β€’ SAS β€’ Excel β€’ Google Sheets

πŸ“Š Data Science & ML

Pandas β€’ NumPy β€’ SciPy β€’ Scikit-learn β€’ XGBoost

πŸ“ˆ Visualization & BI

Matplotlib β€’ Seaborn β€’ Tableau β€’ Power BI

🧠 NLP & CV

TF-IDF β€’ Feature Engineering β€’ Semantic Similarity
Sentence Transformers β€’ OpenCV β€’ YOLOv8

πŸ“ Statistics & Experimentation

Hypothesis Testing β€’ A/B Testing β€’ Probability β€’ Inferential Statistics


πŸ“Œ Featured Projects

πŸ”Ž Quora Duplicate Question Detection

NLP | Sentence Transformers | XGBoost | Streamlit

  • Built end-to-end semantic similarity pipeline (~450K dataset)
  • Used MiniLM embeddings (384-dim)
  • Engineered:
    • cosine similarity
    • embedding interactions
    • lexical features
  • Compared multiple models β†’ XGBoost selected
  • Optimized using threshold tuning
  • Deployed with Streamlit

πŸ‘‰ Repo: (Add your link here)


πŸš— Real-Time Object Detection & Vehicle Analysis

YOLOv8 | OpenCV | Computer Vision

  • Built real-time detection pipeline
  • Extracted structured insights from video
  • Applied for:
    • traffic monitoring
    • surveillance
    • smart city analytics

πŸ“ˆ GitHub Stats (Live)


πŸ“Š Contribution Graph

🎯 Data Science Strength Meter

Problem Solving           β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ  95%
EDA & Insights            β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ   92%
Business Thinking         β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ  96%
Machine Learning          β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ      84%
Statistics                β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ     86%
NLP                       β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ       82%
Computer Vision           β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ         78%
Visualization             β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ   90%
SQL                       β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ     88%
Deployment                β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ          70%

πŸ“š Currently Leveling Up

  • πŸš€ MLOps & Deployment
  • βš™οΈ Production ML Systems
  • ☁️ Cloud (AWS / GCP)
  • πŸ“‘ Model Monitoring
  • 🧩 Feature Engineering at Scale

πŸŽ“ Education

  • πŸŽ“ MS in Data Science β€” University of Colorado Boulder (In Progress)
  • πŸŽ“ BSc β€” Tribhuvan University

πŸ“œ Certifications

  • πŸ“œ Google Advanced Data Analytics
  • πŸ“œ Machine Learning with Python
  • πŸ“œ SQL for Data Science
  • πŸ“œ Probability for Data Science

πŸ’Ό Open To Roles

  • πŸ’Ό Data Scientist
  • πŸ’Ό Machine Learning Engineer
  • πŸ’Ό Applied AI / NLP
  • πŸ’Ό Business Data Scientist

🀝 Let’s Connect


β€œTurning data into insight β€” and insight into business impact.”

Pinned Loading

  1. -Project--Quora-Duplicate-Question-Detection-with-streamlit -Project--Quora-Duplicate-Question-Detection-with-streamlit Public

    Production-grade NLP system for semantic duplicate question detection combining transformer-based embeddings, advanced feature engineering, and supervised ML models with threshold tuning, deployed …

    Jupyter Notebook

  2. Customer-Analytics-Case-Study Customer-Analytics-Case-Study Public

    End-to-end customer analytics ML system leveraging KMeans segmentation, Random Forest churn prediction, and regression-based spend forecasting to drive data-driven marketing, retention, and revenue…

    Jupyter Notebook

  3. Medical-Insurance-Cost-Prediction Medical-Insurance-Cost-Prediction Public

    Medical Insurance Cost Prediction is a data science project that predicts individual insurance costs based on personal attributes such as age, sex, BMI, number of children, smoking status, and regi…

    Jupyter Notebook