Things I've built around real-world data.
A selection of machine learning, AI, and data engineering work from academic projects and professional experience.
Property Price Prediction
Can spatial and neighborhood-derived features improve property price prediction?
Problem
Predicting property prices requires more than basic property attributes. Location and neighborhood effects can strongly influence the final value.
Approach
Built a predictive modeling workflow using XGBoost and CatBoost, with Scikit-learn for preprocessing and evaluation.
Features included
- Spatial clustering
- KNN average price
- Naive Bayes ZIP-code probability
- Property attributes
- Location-based features
AI Property Description & Summary Generation
How can structured property data be transformed into consistent, useful language?
Built an LLM-powered content pipeline that aggregates building and listing information from Elasticsearch and generates structured building descriptions and buyer-oriented AI summaries.
MLS Data Pipeline & Entity Resolution
How can large volumes of listing data be cleaned, matched, and attributed reliably?
Built scalable Python and Elasticsearch workflows for extracting, cleaning, matching, and updating MLS listings across the US.
LISTING ATTRIBUTION
What these projects taught me.
DATA
Working with real-world data means dealing with quality, ambiguity, and scale.
MODELS
Good models depend on meaningful features, sound evaluation, and understanding the problem.
SYSTEMS
A useful solution needs reliable engineering around the model.