Data Analyst & Statistician
Turning data into decisions
Statistics student at the University of Manitoba. I build end-to-end analytical pipelines in R and Python regression, time series, classification and spend my free time building PCs, chasing bugs, and listening to jazz.
Featured Work
01 / 06
Binary classification on 7,043 telecom customers to predict churn and quantify its drivers. Odds ratios reveal month-to-month contracts and fiber optic service as the strongest risk factors backed by Type II likelihood-ratio tests.
AUC ≈ 0.84 in both R and Python
github.com/Jrudani21/telco-churn-logistic-regression ↗
02 / 06
Modelled 731 days of Capital Bikeshare data using Poisson and Negative Binomial GLMs. Detected severe overdispersion (variance/mean = 357×) and used IRRs to quantify weather and seasonal effects on daily rentals.
NB model AIC ~1,200 better than Poisson
github.com/Jrudani21/bike-share-demand-count-regression ↗
03 / 06
Full Box-Jenkins SARIMA pipeline log transform, stationarity testing (ADF/KPSS), ACF/PACF analysis, 36-month forecast with prediction intervals, and head-to-head comparison against Facebook Prophet.
SARIMA MAPE ~3–5% vs Prophet ~4–6%
github.com/Jrudani21/air-passengers-timeseries ↗
04 / 06
A comprehensive linear modeling showcase simple linear regression, one-way ANOVA with Tukey post-hoc, two-way ANOVA with interaction, and ANCOVA with adjusted means. R² improves from 0.759 to 0.869 by adding species after flipper length.
ANCOVA adjusted R² = 0.869
github.com/Jrudani21/palmer-penguins-linear-models ↗
05 / 06
Detects providers (NPIs) submitting Medicare claims from a US address and a foreign country simultaneously a documented CMS OIG phantom billing scheme. Five SQL queries (self-join date-overlap detection, risk scoring, country aggregation, specialty benchmarking) plus an interactive Streamlit dashboard with 5 analysis tabs. Built on the CMS Medicare Physician & Other Practitioners PUF schema with 9,976 synthetic claims across 2,000 NPIs.
18 flagged NPIs · $46K+ Medicare payments at risk · HIGH/MEDIUM/LOW risk tiers
github.com/Jrudani21/medicare-fraud-detector ↗
06 / 06
Real-time pandemic statistics application connecting to the disease.sh public API across five interactive views: global KPIs (cases, recoveries, deaths, CFR), 180-day trend analysis, top-15 country rankings, a choropleth map normalised by cases per million, and a searchable data explorer. Data refreshes every 10 minutes via a built-in cache to balance responsiveness with minimal network overhead.
Live data from 200+ countries · sub-second rendering · 180-day historical window
github.com/Jrudani21/covid-dashboard ↗
Capabilities
Languages
R
Python
SQL
Libraries & Frameworks
tidyverse · ggplot2 · forecast
pandas · numpy · matplotlib · seaborn
scikit-learn · statsmodels · pmdarima
Streamlit · Plotly · SQLite
Tools & Workflow
Git & GitHub
R Markdown · Jupyter Notebooks
VS Code
Statistical Methods
Linear & Logistic Regression
One-way / Two-way ANOVA & ANCOVA
Poisson & Negative Binomial GLMs
Time Series (ARIMA / SARIMA)
Hypothesis Testing & Inference
Model Evaluation
AUC / ROC Curves
Confusion Matrix & Classification Report
AIC / BIC Model Comparison
MAPE & Hold-out Validation
Residual Diagnostics & Ljung-Box
Beyond the Data
Designing and assembling custom rigs — from component selection to cable management and overclocking. The hardware side of computing is as satisfying as the software.
Hands-on work with breadboards, microcontrollers, and circuit design. There's something deeply satisfying about making physical things respond to logic.
From classic Miles Davis to modern fusion — the improvisational structure of jazz feels a lot like exploratory data analysis.
Strategy, RPGs, and competitive titles. Gaming honed my instinct for systems thinking and pattern recognition long before stats did.
Finding edge cases and breaking things before others do. QA thinking comes naturally — if a model or program can fail, I want to find out how.
Background
I'm Janak Rudani, a statistics student at the University of Manitoba with a focus on applied regression, statistical modelling, and data analysis. My work spans linear models, generalized linear models, and time series all implemented hands-on in both R and Python.
Every project in this portfolio shows the full analytical workflow: data cleaning, exploratory analysis, model selection, assumption checking, and actionable interpretation the same pipeline expected in a professional data analyst role.
Outside of data, I build PCs, tinker with circuits, play and listen to jazz, game competitively, and enjoy breaking software through bug testing. The same curiosity that drives those hobbies drives my work with data.
I'm actively looking for data analyst internships and entry-level roles where I can apply statistical reasoning to real business problems customer behaviour, demand forecasting, A/B testing, or operational analytics.
Full Name
Janak Rudani
Education
University of Manitoba Statistics and Mathematics
Location
Winnipeg, Manitoba, Canada
Primary Languages
R · Python
GitHub
Get In Touch
Open to data analyst internships and entry-level roles. Feel free to reach out.
Janak25rudani@icloud.com