01Work

Four projects, from question to result.

Chosen from 16 on GitHub. Each one shows the question, the result and the code behind it.

  1. Fig. 2Illustration: behavioural clusters projected to two dimensions.

    Clustering · 2025

    Hotel customer segmentation

    Which hotel guests behave differently, and how could the hotel tell them apart?

    records cleaned and analysed
    83,590
    clustering methods compared
    3
    How I did it

    Problem

    The hotel had 83,590 customer records, and a lot of them needed work before they were useful. Some columns were hashed IDs, revenue was split across two channels, and a field called DaysSinceCreation was really measuring how long someone had been a customer.

    Method

    1. I dropped the ID, NameHash and DocIDHash columns. They say nothing about how a guest behaves, and keeping them would have grouped people by identity instead of behaviour.
    2. I explored revenue, lead time, cancellations, no-shows and special requests before deciding which features to use.
    3. I clustered the data three ways. K-Means gave me compact groups, and DBSCAN and HDBSCAN picked up the unusual stay lengths and big spenders that averages hide.
    4. I used PCA and t-SNE to plot the groups in two dimensions, so I could actually see them and sanity-check them.
    5. I wrote up each group as a customer profile, with a suggestion for how the hotel could look after it.
  2. Fig. 3Illustration: train booking schema, three related tables.

    SQL · Database design · 2025

    Booking and retail databases in SQL

    Two MySQL databases that keep their own data clean, plus the queries I used to analyse them.

    databases designed from scratch
    2
    triggers and stored procedures
    5+
    How I did it

    Problem

    If a booking database has mistakes in it, every report built on top of it has the same mistakes. Ages go out of date, a changed passenger ID leaves other tables pointing at nothing, and one duplicated employee can double a headcount.

    Method

    1. For the train booking system I built three linked tables, with triggers that work out each passenger's age from their date of birth and update IDs everywhere when they change.
    2. I added a five-character rule for Passenger_id and proper foreign keys, so bad rows can't get in at all.
    3. I made backup tables and an EMPTY_DATA() procedure, so I could reset a test run with one call.
    4. For the retail store data I ranked stores and employees, averaged sales by product, and matched up staff records using joins and DISTINCT.
    5. I removed duplicate employees directly in the table instead of copying the clean rows into a second one.
  3. Fig. 4Illustration: churn trend, and the same rate cut three ways.

    Business intelligence · 2025

    Customer churn dashboard

    A Power BI report that shows not just how many customers left, but who they were and where they were.

    views in one report
    6
    breakdowns: demographics, plan, geography
    3
    How I did it

    Problem

    A single churn number doesn't tell a manager much. They need to know which customers left, where they were and what they had in common. Without that, the report gets a nod in a meeting and is never opened again.

    Method

    1. I used a public churn dataset from Kaggle, so anyone can check my numbers against the source.
    2. I showed churn over time rather than as one figure, so there's always something to compare against.
    3. I broke the rate down by demographics, subscription plan and geography to see where it was highest.
    4. I kept every chart cross-filterable and exportable, so people can dig into it themselves.
  4. Fig. 5Illustration: one scan in, one of four ordered stages out.

    Computer vision · CNN · 2025

    Alzheimer's stage classifier

    It reads an MRI scan and places it in one of four stages of Alzheimer's, instead of just saying yes or no.

    validation accuracy
    99.21%
    test accuracy
    99%
    How I did it

    Problem

    Most image classifiers only say whether something is there or not. With dementia that isn't very helpful, because catching it early matters most, and a plain yes or no can't show what stage someone is at.

    Method

    1. I trained the model to tell four stages apart: non-demented, very mild, mild and moderate.
    2. I built a convolutional neural network and trained it on the Well-Documented Alzheimer's Dataset from Kaggle.
    3. I published the model under an MIT licence and put it online, so anyone can try it on a real scan.

Sources

  1. 83,590 hotel customers grouped by behaviour. Row count of HotelCustomersDataset.xlsx, the dataset in the segmentation repository.

  2. 6 views in one Power BI churn report. Six views in the professional-power-bi-dashboard report.

  3. 99.21% validation accuracy on Alzheimer's MRI stages. Validation accuracy of the CNN, as reported in the project README. Test accuracy was 99%.

All 16 projects on GitHub (opens in a new tab)

02About

Business first. Data to back it up.

I studied marketing at Seneca Polytechnic in Toronto. Marketing taught me to ask what a decision needs before I reach for a chart.

I taught myself SQL, Python, Power BI and machine learning alongside it. Now I'm looking for my first data role, ideally somewhere people actually use the analysis to decide.

Looking for
Data analyst and marketing analytics roles
Based in
Ahmedabad, India
Tools
SQL, Python, Power BI, Excel, scikit-learn, PyTorch
Languages
English, Hindi, GujaratiEnglish: C1 · IELTS 8.0
Experience
Team Member, Tim HortonsJul 2025 to Aug 2026Ecommerce Analyst (Part-time), Canadian Outlet StoreMay 2024 to Oct 2024

How I work

  1. I think about the business first

    My diploma is in marketing, so before I touch the data I want to know who the customer is and what decision the numbers are meant to help with.

  2. I can do the whole job

    I can take a messy spreadsheet all the way to a clean dataset, a SQL database, a Power BI report or a trained model. All of my code is on GitHub if you'd like to look.

  3. I'm easy to work with

    I speak and write English fluently, and I spent over a year working shifts in a busy team at Tim Hortons in Canada. I turn up, I get things done, and I can explain what I did.

Education & certifications

  1. Business, Marketing

    Seneca Polytechnic, Toronto, Canada · Ontario College Diploma

    2025
  2. Google Data Analytics

    Google · Coursera

    Sep 2025View programme for Google Data Analytics (opens in a new tab)
  3. Data Science, ML, DL & NLP Bootcamp

    Udemy

    2025Verify for Data Science, ML, DL & NLP Bootcamp (opens in a new tab)
  4. Inbound Marketing

    HubSpot Academy

    Sep 2024View programme for Inbound Marketing (opens in a new tab)
  5. Data Analytics

    TOPS Technologies

    Oct 2023View programme for Data Analytics (opens in a new tab)

03Contact

Got data? Let's talk.

Send me the role. I'll come back with questions, and a time to talk.

Email is quickest. I usually reply within two days.