A complete end-to-end Machine Learning project that predicts house/property prices based on structured data such as location, area, rooms, amenities, distance to city, and other real‑world features.
This project is built for learning + production readiness, covering:
- Data cleaning
- Feature engineering
- Model training
- Evaluation
- Inference
- Web deployment (Streamlit)
⚠️ Educational & research project. Predictions are estimates, not financial or legal advice.
To build an AI system that can:
- Learn patterns from historical property data
- Predict house prices accurately
- Explain predictions using model reasoning
- Provide a usable web interface for real users
- Machine Learning Type: Supervised Learning
- Task Type: Regression
- Input: Property features (area, rooms, location, etc.)
- Output: Predicted price (₹)
House-Rate-Predictor/
│
├── data/
│ ├── raw/ # original raw dataset
│ ├── interim/ # cleaned/intermediate data
│ └── processed/ # final training-ready data
│
├── notebooks/
│ └── 01_eda.py # exploratory data analysis
│
├── src/
│ ├── config.py # project config
│ ├── features.py # feature engineering
│ ├── eda.py # EDA utilities
│ ├── train.py # model training
│ └── infer.py # prediction script
│
├── app/
│ └── streamlit_app.py # web application
│
├── models/ # trained models
├── requirements.txt
└── README.md
- area_sqft
- bedrooms
- bathrooms
- age_yrs
- floor
- total_floors
- amenities_count
- distance_cbd_km
- distance_metro_km
- latitude
- longitude
- bath_per_bed = bathrooms / bedrooms
- floor_ratio = floor / total_floors
- amenity_density = amenities / area
- city
- furnishing
- builder
- property_type
- Load CSV dataset
- Validate schema
- Type conversion
- Missing value detection
- Outlier detection
- Distribution analysis
- Correlation analysis
- Missing value imputation
- Outlier winsorization
- Data normalization
- Ratio features
- Density features
- Robust transformations
- Algorithm: XGBoost Regressor
- Target transformation: log(price)
- Pipeline-based training
- Train/Validation split
- RMSE (Root Mean Squared Error)
- MAPE (Mean Absolute Percentage Error)
- Model saved using joblib
- Web UI using Streamlit
- Real-time predictions
-
Model: XGBoost Regressor
-
Loss Function: Mean Squared Error
-
Target Transformation: log(price)
-
Why log(price)?
- Reduces skew
- Stabilizes training
- Improves prediction reliability
| Metric | Purpose |
|---|---|
| RMSE | Penalizes large prediction errors |
| MAPE | Business-friendly % error metric |
Built using Streamlit, allows users to:
- Input property details
- Upload features
- Predict price
- See prediction confidence band
# create virtual env
python -m venv .venv
# activate
.venv\Scripts\activate # windows
source .venv/bin/activate # mac/linux
# install dependencies
pip install -r requirements.txtpython -m src.trainstreamlit run app/streamlit_app.pyid,city,lat,lon,area_sqft,bedrooms,bathrooms,age_yrs,floor,total_floors,
amenities_count,distance_cbd_km,distance_metro_km,furnishing,builder,property_type,price,listed_date
- No personal data stored
- No identity tracking
- No financial decision automation
- Predictions are estimates only
- CatBoost model integration
- SHAP explainability
- Satellite image integration
- NLP-based property description parsing
- Location clustering
- RAG-based property advisory
- API deployment (FastAPI)
- Mobile app version
This project teaches:
- Real-world ML system design
- Production ML pipelines
- Data engineering concepts
- Model deployment
- AI system architecture
- End-to-end AI product building
Built as part of an AI learning journey to master:
- Machine Learning
- Generative AI
- AI Engineering
- Production AI systems
MIT License — Free to use, modify, and learn from.
Give the project a ⭐ and use it to build your AI portfolio.
This project is not just a model — it is a full AI system blueprint.