Open to Data Analyst / Scientist / AI Engineer roles

HI, I'M
HEERA SHANKER

Data Science & Machine Learning Engineer passionate about crafting intelligent systems and predictive models — from raw data to something people can actually use.

✉ Contact Me
3+
Production Projects
250+
DSA Problems
79M
Rows Handled
0.90
ROC-AUC Score
✦ Python
✦ SQL
✦ Scikit-learn
✦ Pandas
✦ NumPy
✦ TensorFlow
✦ Streamlit
✦ Power BI
✦ PyTorch
✦ Tableau
✦ MongoDB
✦ Hadoop
✦ Git
✦ VS Code
✦ Python
✦ SQL
✦ Scikit-learn
✦ Pandas
✦ NumPy
✦ TensorFlow
✦ Streamlit
✦ Power BI
✦ PyTorch
✦ Tableau
✦ MongoDB
✦ Hadoop
✦ Git
✦ VS Code

ABOUT ME

I'm Malathkar Heera Shanker — a final-year B.Tech student in Computer Science & Engineering with a Data Science specialization at JNTU Hyderabad (CGPA: 7.68/10).

Currently an AI/ML Intern at FlyRank AI, where I built an ML pipeline on real Google Search Console data (~79M rows), lifting page-refresh prioritization accuracy 3× (Precision@50: 0.24 → 0.74).

I hold a HackerRank Gold Badge in Python and a LeetCode 50-day badge, with 250+ DSA problems solved. I love turning messy, production-scale data into models people can actually act on.

PythonML End-to-EndNLP / RAG Power BISQLStreamlit TensorFlowScikit-learn
7.68
CGPA
3+
Live Projects
2
Internships
Heera Shanker

EXPERTISE

01 MACHINE LEARNING & PREDICTIVE MODELING XGBoost, Random Forest, Logistic Regression — from feature engineering to deployment. ↗
02 DATA ANALYTICS & DASHBOARDS Power BI, Tableau, Excel — translating data into decisions. ↗
03 NATURAL LANGUAGE PROCESSING (NLP) RAG pipelines, LLM integration, entity extraction, and semantic search. ↗
04 SQL & DATABASE ARCHITECTURE Relational databases, DuckDB, query optimization, and data pipeline design. ↗
05 AI VOICE & AUTOMATION SYSTEMS Speech recognition, REST API integration, modular Python architecture. ↗

PROJECTS

01 PharmaRAG — Clinical Decision Support System
PythonRAGLlama 3
+
✦ Multi-stage async RAG pipeline with source-grounded LLM outputs

A Retrieval-Augmented Generation pipeline integrating RxNorm, OpenFDA, and RxNav APIs for real-time drug interaction and contraindication alerts. Built a multi-stage architecture — entity extraction → normalization → parallel async retrieval → LLM synthesis — using Llama 3 with deterministic outputs via a Streamlit UI.

02 AI-Powered Customer Churn Prediction
XGBoostStreamlitScikit-learn
+
✦ 85%+ accuracy · 0.90 ROC-AUC

Binary classification pipeline on 10,000+ customer records. Evaluated Logistic Regression, Random Forest, SVM, and XGBoost. Designed 12+ engineered features via EDA with Pandas & Matplotlib. Deployed a real-time Streamlit dashboard for churn probability predictions.

03 House Price Prediction Pipeline
Ridge/LassoGradient BoostingStreamlit
+
✦ R² > 0.90 on Kaggle Ames dataset

Trained regression models (Linear, Ridge, Lasso, Gradient Boosting) on the Kaggle Ames dataset (1,460 samples, 81 features). Engineered 20+ features with missing-value imputation, categorical encoding, and outlier handling. Deployed an interactive Streamlit app.

04 Trackademy — Student Performance Analytics
AnalyticsDashboard
+

A student performance analytics dashboard. See the repository list on GitHub for the code and current status.

05 JARVIS — AI Voice Assistant
PythonNLPREST APIs
+
✦ 15+ user intents · REST API integrations

Full-featured AI Voice Assistant built in Python with speech recognition (SpeechRecognition), text-to-speech (pyttsx3), and multi-command NLP processing. Integrated OpenWeatherMap, Wikipedia, Gmail SMTP, YouTube, and Chrome automation. Modular architecture with CPU monitoring, screenshot capture, and music playback.

06 Product Recommender — NLP Engine
PythonNLP
+
✦ Text-based product recommendations

Recommends products from text descriptions using natural language processing techniques. Built end-to-end in Python with NLP preprocessing, similarity scoring, and ranking.

EXPERIENCE

AI/ML Intern
FlyRank AI — Applied Search Intelligence
2026 – Present · Remote
  • K-Means clustering + PCA on ~30K web pages for content-refresh prioritization
  • Built supervised classification pipeline on 44-column production dataset
  • Lifted Precision@50 from 0.24 (baseline) → 0.74 (model) — a 3× improvement
  • Extending pipeline to FlyRank's ~79M row production dataset via Hugging Face + DuckDB
Python Programming Intern
Oasis Infobyte
2026 · Remote
  • Built JARVIS — a full-featured AI Voice Assistant with 15+ user intents
  • Integrated OpenWeatherMap, Wikipedia, Gmail SMTP, YouTube APIs
  • Modular architecture with CPU monitoring, screenshot capture, and browser automation
Data Analytics Job Simulation
Tata iQ (Forage)
2025 · Virtual
  • Gen AI-powered data analytics simulation
  • Completed data wrangling, visualization, and insight generation tasks
  • Certificate of completion issued
B.Tech — CS Engineering, Data Science
JNTU Hyderabad
Aug 2023 – Jun 2027
  • CGPA: 7.68 / 10
  • Coursework: ML, AI, DSA, Big Data, Python, OOP, SQL, Data Visualization, Power BI
  • 250+ DSA problems across LeetCode, HackerRank, Striver's A2Z Sheet
AI/ML Intern
FlyRank AI — Applied Search Intelligence
2026 – Present · Remote
  • K-Means clustering + PCA on ~30K web pages for content-refresh prioritization
  • Built supervised classification pipeline on 44-column production dataset
  • Lifted Precision@50 from 0.24 → 0.74 — a 3× improvement
Python Programming Intern
Oasis Infobyte
2026 · Remote
  • Built JARVIS — a full-featured AI Voice Assistant with 15+ user intents
  • Integrated OpenWeatherMap, Wikipedia, Gmail SMTP, YouTube APIs

CERTIFICATIONS

📄

Resume

Full details of my experience, education, projects, and skills — ready to download.

⬇ Download Resume (PDF)

LET'S GET
IN TOUCH

Got messy data and a question worth answering? Tell me what the data is, what you want to know, and what you've tried — I'll reply with how I'd approach it.

heerashanker0214@gmail.com →