Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🏠 House Price (House Rate) Predictor

A complete end-to-end Machine Learning project that predicts house/property prices based on structured data such as location, area, rooms, amenities, distance to city, and other real‑world features.

This project is built for learning + production readiness, covering:

  • Data cleaning
  • Feature engineering
  • Model training
  • Evaluation
  • Inference
  • Web deployment (Streamlit)

⚠️ Educational & research project. Predictions are estimates, not financial or legal advice.


🎯 Project Objective

To build an AI system that can:

  • Learn patterns from historical property data
  • Predict house prices accurately
  • Explain predictions using model reasoning
  • Provide a usable web interface for real users

🧠 Problem Type

  • Machine Learning Type: Supervised Learning
  • Task Type: Regression
  • Input: Property features (area, rooms, location, etc.)
  • Output: Predicted price (₹)

🏗️ Project Architecture

House-Rate-Predictor/
│
├── data/
│   ├── raw/                # original raw dataset
│   ├── interim/            # cleaned/intermediate data
│   └── processed/          # final training-ready data
│
├── notebooks/
│   └── 01_eda.py           # exploratory data analysis
│
├── src/
│   ├── config.py           # project config
│   ├── features.py         # feature engineering
│   ├── eda.py              # EDA utilities
│   ├── train.py            # model training
│   └── infer.py            # prediction script
│
├── app/
│   └── streamlit_app.py    # web application
│
├── models/                 # trained models
├── requirements.txt
└── README.md

📊 Features Used

🔢 Numerical Features

  • area_sqft
  • bedrooms
  • bathrooms
  • age_yrs
  • floor
  • total_floors
  • amenities_count
  • distance_cbd_km
  • distance_metro_km
  • latitude
  • longitude

🧮 Engineered Features

  • bath_per_bed = bathrooms / bedrooms
  • floor_ratio = floor / total_floors
  • amenity_density = amenities / area

🏷️ Categorical Features

  • city
  • furnishing
  • builder
  • property_type

🧪 ML Pipeline

Step 1 — Data Ingestion

  • Load CSV dataset
  • Validate schema
  • Type conversion

Step 2 — EDA (Exploratory Data Analysis)

  • Missing value detection
  • Outlier detection
  • Distribution analysis
  • Correlation analysis

Step 3 — Data Cleaning

  • Missing value imputation
  • Outlier winsorization
  • Data normalization

Step 4 — Feature Engineering

  • Ratio features
  • Density features
  • Robust transformations

Step 5 — Model Training

  • Algorithm: XGBoost Regressor
  • Target transformation: log(price)
  • Pipeline-based training
  • Train/Validation split

Step 6 — Evaluation

  • RMSE (Root Mean Squared Error)
  • MAPE (Mean Absolute Percentage Error)

Step 7 — Deployment

  • Model saved using joblib
  • Web UI using Streamlit
  • Real-time predictions

🤖 Model Details

  • Model: XGBoost Regressor

  • Loss Function: Mean Squared Error

  • Target Transformation: log(price)

  • Why log(price)?

    • Reduces skew
    • Stabilizes training
    • Improves prediction reliability

📈 Evaluation Metrics

Metric Purpose
RMSE Penalizes large prediction errors
MAPE Business-friendly % error metric

🌐 Web Application

Built using Streamlit, allows users to:

  • Input property details
  • Upload features
  • Predict price
  • See prediction confidence band

⚙️ Installation

# create virtual env
python -m venv .venv

# activate
.venv\Scripts\activate   # windows
source .venv/bin/activate # mac/linux

# install dependencies
pip install -r requirements.txt

▶️ How to Run

1. Train Model

python -m src.train

2. Run Web App

streamlit run app/streamlit_app.py

📁 Dataset Format

id,city,lat,lon,area_sqft,bedrooms,bathrooms,age_yrs,floor,total_floors,
amenities_count,distance_cbd_km,distance_metro_km,furnishing,builder,property_type,price,listed_date

🛡️ Safety & Ethics

  • No personal data stored
  • No identity tracking
  • No financial decision automation
  • Predictions are estimates only

🚀 Future Improvements

  • CatBoost model integration
  • SHAP explainability
  • Satellite image integration
  • NLP-based property description parsing
  • Location clustering
  • RAG-based property advisory
  • API deployment (FastAPI)
  • Mobile app version

📚 Learning Outcomes

This project teaches:

  • Real-world ML system design
  • Production ML pipelines
  • Data engineering concepts
  • Model deployment
  • AI system architecture
  • End-to-end AI product building

👨‍💻 Author

Built as part of an AI learning journey to master:

  • Machine Learning
  • Generative AI
  • AI Engineering
  • Production AI systems

📜 License

MIT License — Free to use, modify, and learn from.


⭐ If you find this useful

Give the project a ⭐ and use it to build your AI portfolio.


This project is not just a model — it is a full AI system blueprint.

About

A supervised regression pipeline to estimate house prices from structured features. Tech: Python, scikit-learn, XGBoost, Streamlit.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages