Open to opportunities

Syed Yahiya

Data Scientist | Data Analyst | AI/ML

Building End-to-End Analytical Pipelines, GenAI-powered Analytics Agentic AI Projects, and Statistical Business Case Studies. Python · SQL · Tableau · LLMs · Hyderabad, India.

scroll
01

About

I'm an aspiring Data Analyst based in Hyderabad building a portfolio that spans data engineering, data & statistical analysis, hypothetical testing and AI-powered automation.

My work covers the full analytical loop, raw ingestion and cleaning through inferential testing to business recommendations. I believe the value of analysis lives in the decision it enables, not the chart it produces.

Targeting analytics roles across MBB consulting, Big 4, analytics pure-plays, and BFSI GCCs. Open to full-time roles

5
Portfolio Projects
8+
Tools in Stack
02

Projects

Open Source · Data Engineering · Community Hub
Data Engineering Resource Hub
Curation → GitHub Repo → GitHub Pages Web App + XLSX Tracker

Built and maintaining an open-source, community-driven resource hub cataloging 60+ verified data engineering assets across 10 categories, including 32 core infrastructure tools, roadmaps, and production design patterns. Building an ongoing real-time searchable, filterable web application deployed via GitHub Pages alongside a structured offline tracking template.

60+ Curated Resources · 32 Tools Cataloged · 5-Phase Production Roadmap
Open SourceData Engineering GitHub PagesHTML/CSS RoadmapsCommunity
AI · Automation · Data Analysis · Business Intelligence
Augmented Analytics
Raw CSV → Cleaning → Profiling → Analysis → Visualization

Built an interactive web app using a multi-agent AI system to automate the entire data analysis workflow. Upload a raw CSV and get actionable insights, visualizations, and what-if simulations in seconds—without writing a single line of code. Features include automated data cleaning, statistical profiling, insight generation, and conversational chart building.

Automated End-to-End Analysis · Conversational Visualization · What-If Simulation · Zero Code
PythonPandas StreamlitSeaborn MatplotlibGemini API
Declarative AI · RAG Pipeline · Containers
RICE: A Declarative LLM/RAG Configuration Engine
TOML Config → Parser → Orchestrator → Adapters → Podman Containers → Answer

Open-source declarative engine to define AI pipelines (ingestion, chunking, embeddings, vector DB, RAG, LLM) in a single TOML file. Eliminates manual setup and dependency conflicts by running tools in isolated Podman containers with sequential, RAM-optimized execution.

One TOML file → Full AI pipeline · Zero manual setup · RAM-optimized containerized execution
PythonTOML LLMPodman RAGChromaDB Ollama
Agentic AI · AutoML · On-Premise MLOps
Sovereign AutoML Agent
Raw data → DeepSeek R1 preprocessing → Feature eng → Model selection → Eval report

Agentic ML orchestration framework running entirely on local infrastructure via Ollama. DeepSeek R1 drives each stage autonomously: schema inference, missing-value strategy, feature engineering, algorithm selection (regression/classification/clustering), cross-validation, and a final model card with metrics. Purpose-built for BFSI and healthcare environments with PHI or MNPI data sovereignty constraints.

Zero cloud cost · Zero data egress · 60% faster experiments
PythonOllama DeepSeek R1Agentic Workflow AutoMLData Sovereignty
03

Skills

Languages
Python SQL
Libraries
PandasNumPy SciPySeaborn Matplotlib
BI & Cloud
TableauSnowflake AWS S3Neon Postgres Airbyte
AI & Tools
Gemini APIOllama DeepSeek R1Streamlit GitJupyter
Analytical Concepts
Exploratory Data Analysis Hypothesis Testing RFM Segmentation A/B Testing ETL Pipeline Design Statistical Inference Descriptive Statistics Data Cleaning
terminal — yahiya@analyst
$ echo $CURRENT_FOCUS
Building statistical case studies & cloud pipelines
$ cat learning_queue.txt
dbt · Apache Airflow · Advanced SQL window functions · MLOps
$ cat availability.txt
Open to full-time Data Analyst and Analytical Engineer roles.
04

Case Studies

Business analytics case studies structured as real product interviews — Problem → Hypothesis → Analysis → Recommendation.

Food Delivery · Revenue Analysis
Zomato — Revenue Flat Despite 18% Order Growth
Volume ↑ 18% → AOV decomposition → Discount audit → Retention forecast

Diagnosed why Zomato's Core Food Delivery revenue stayed flat despite an 18% surge in order volume. Identified a ₹200 flat-discount new-user campaign driving near-zero AOV orders. Recommended minimum order threshold and Q2 retention tracking framework.

Root cause: CAC campaign inflating volume, suppressing AOV by 20%
AOV Analysis CAC vs LTV Retention Rate Discount Audit Three-sided Marketplace
E-commerce · Supply Chain Analytics
Myntra — Net Profit Drop Despite Record Gross Sales
GMV at Record Highs → Return Rate Segmented (35%) → "Bracketing" Behavior Identified → Policy & Friction Fix

Diagnosed why Myntra's logistics costs surged despite record 'End of Reason Sale' Gross Sales. Identified existing VIP users exploiting a ₹3,000 cart threshold by "bracketing" multiple sizes of apparel, driving returns to 35%. Recommended pro-rated refunds and a targeted reverse pickup fee.

Root cause: Exploitative "bracketing" to hit discount thresholds driving massive reverse logistics costs.
Reverse Logistics Cohort Analysis GMV vs Net Sales Bracketing Policy Friction
05

Contact

Open to full-time Data Analyst roles in Hyderabad. Reach out through any channel below — I typically respond within 24 hours.